Kernel Engineering Workshop · Lecture 1
What happens inside a GPU when you ask a chatbot a question, and the speed limit on every GPU program.
Dr. Raj Dandekar · Vizuara AI Labs
So what?
36 blocks × (42 M + 151 M) weights
+ the output layer 622 M
= 7.57 billion multiply-adds for every token
Qwen3-8B, from its configuration: hidden size 4,096, feed-forward size 12,288, 36 blocks, 151,936-word vocabulary.
How fast can a GPU go?
132 SMs × 4 tensor cores × 1,024 operations per tick
× 1.83 billion ticks a second
= 989 trillion operations a second
Dense BF16 on the H100 SXM5 (NVIDIA datasheet). We call it 989 TFLOP/s.
If we could use all of it
15.2 billion operations ÷ 989 trillion a second
= 15 µs a token = 65,000 tokens a second
A 300-page book is about 130,000 tokens: at this speed, 2 seconds.
Reality
149 tokens a second, measured
the same book: 14 minutes
we are using 0.23% of the GPU
Measured on a Modal H100, 28 September 2026: Qwen3-8B in BF16 served by vLLM, one user.
Physics sets the limit
| what | value | source |
|---|---|---|
| fast memory on the H100 chip | 116 MB | 50 MB L2 + registers + shared memory |
| the model's weights read per word | 15.1 GB | Qwen3-8B, BF16 |
| SRAM area for those weights | 3,000 mm² | TSMC 2 nm SRAM, 38.1 Mb/mm² |
| largest printable chip | 858 mm² | lithography reticle, 26 × 33 mm |
| energy: read 32 bits from DRAM | 640 pJ | Horowitz, ISSCC 2014 |
| energy: one 32-bit multiply | 3.7 pJ | Horowitz, ISSCC 2014 |
The one equation
operations per second = bandwidth × operations per byte
(while memory is the limit)
Bandwidth is fixed by the hardware: 3.35 TB/s on the H100. Operations per byte is ours: it depends on how the work is organised, which is what kernels, batching and precision change.
One step, one user
memory time = 15.1 GB ÷ 3.35 TB/s = 4.5 ms
compute time = 15.2 G ÷ 989 T = 0.015 ms
step = the longer = 4.5 ms measured: 6.7 ms
The gap between 4.5 ms on paper and 6.7 ms measured is what kernel engineering closes.
One user in numbers
operations per byte = 15.2 G ÷ 15.1 GB ≈ 1
3.35 TB/s × 1 = 3.35 TFLOP/s at best: 0.34% of 989
measured: 2.3 TFLOP/s = 0.23% of the GPU
Same bytes, more math
| users | operations per byte | GPU share on paper | measured, tokens/s | GPU share measured |
|---|---|---|---|---|
| 1 | 1 | 0.34% | 148 | 0.23% |
| 64 | 60 | 20% | 8,006 | 12% |
| 256 | 195 | 66% | 20,818 | 32% |
| 1,024 | 452 | 100% | 26,303 | 40% |
Measured on a Modal H100 through the live server, 28 September 2026 (128 tokens of conversation per user). Past 295 operations per byte the chip is compute-bound.
Before the project
Hour 2 · the project
Qwen3-8B on one H100 on Modal, served with vLLM. One H100 costs about $3.95 an hour. How many of you can it serve at once, and what does each answer cost?
phones and laptops
→
on Modal, a class code
→
Qwen3-8B, BF16
→
users, tokens a second, cost
probe.py · chat.py · dashboard.py · loadgen.py · loadwatch.py
Predict first
| users | memory time | compute time | limit | each user, tok/s | together, tok/s | your guess |
|---|---|---|---|---|---|---|
| 1 | 4.52 ms | 0.02 ms | memory | 221 | 221 | ? |
| 8 | 4.56 ms | 0.12 ms | memory | 219 | 1,753 | ? |
| 64 | 4.88 ms | 0.98 ms | memory | 205 | 13,117 | ? |
| 256 | 5.96 ms | 3.94 ms | memory | 168 | 42,947 | ? |
| 512 | 7.40 ms | 7.87 ms | compute | 127 | 65,042 | ? |
| 1,024 | 10.29 ms | 15.74 ms | compute | 64 | 65,042 | ? |
Per step on paper: memory time = (15.1 GB + users × 128 tokens × 144 KiB) ÷ 3.35 TB/s; compute time = users × 15.2 billion operations ÷ 989 TFLOP/s. The step takes the longer one.
Your turn
Open the class page, type the class code, and ask the chatbot anything. Then watch the dot move to the right.
vizuaraai--l1-dashboard-web.modal.run
class code
roofline
Rehearsal run
| users | step | each user, tok/s | together, tok/s | per million tokens |
|---|---|---|---|---|
| 1 | 6.69 ms | 149 | 148 | $7.40 |
| 64 | 7.79 ms | 128 | 8,006 | $0.137 |
| 256 | 11.92 ms | 84 | 20,818 | $0.053 |
| 512 | 20.29 ms | 49 | 24,631 | $0.045 |
| 1,024 | 37.77 ms | 26 | 26,303 | $0.042 |
Measured on a Modal H100 through the live server and simulated users, 28 September 2026: Qwen3-8B in BF16, vLLM, 64-token prompts and 128-token answers.
What an answer costs
cost per million tokens = $1,097 ÷ tokens per second
because the H100 bills $0.001097 a second on Modal.
1 user, 148 tokens a second
$7.40
1,024 users, 26,303 tokens a second
$0.042
Before Lecture 2
Next: Lecture 2, the CUDA programming model. How a GPU actually runs your code.
→ or Page Down next · ← back · Space or a click on the video pauses and plays · R replay · 1 2 3 speed 0.5×, 0.75×, 1× · N notes · S speaker view · T restart the clock · F full screen. Videos stop on their last frame.