How Fast Can This Go?

Kernel Engineering Workshop · Lecture 1

How fast can this go?

What happens inside a GPU when you ask a chatbot a question, and the speed limit on every GPU program.

Dr. Raj Dandekar · Vizuara AI Labs

So what?

Almost all of it is matrix multiplication

36 blocks × (42 M + 151 M) weights

+ the output layer 622 M

= 7.57 billion multiply-adds for every token

Qwen3-8B, from its configuration: hidden size 4,096, feed-forward size 12,288, 36 blocks, 151,936-word vocabulary.

How fast can a GPU go?

The H100's ceiling, from its parts

132 SMs × 4 tensor cores × 1,024 operations per tick

× 1.83 billion ticks a second

= 989 trillion operations a second

Dense BF16 on the H100 SXM5 (NVIDIA datasheet). We call it 989 TFLOP/s.

If we could use all of it

One token would take 15 microseconds

15.2 billion operations ÷ 989 trillion a second

= 15 µs a token = 65,000 tokens a second

A 300-page book is about 130,000 tokens: at this speed, 2 seconds.

Reality

Measured on our H100: 149 tokens a second

149 tokens a second, measured

the same book: 14 minutes

we are using 0.23% of the GPU

Measured on a Modal H100, 28 September 2026: Qwen3-8B in BF16 served by vLLM, one user.

Physics sets the limit

The weights cannot fit next to the math

whatvaluesource
fast memory on the H100 chip116 MB50 MB L2 + registers + shared memory
the model's weights read per word15.1 GBQwen3-8B, BF16
SRAM area for those weights3,000 mm²TSMC 2 nm SRAM, 38.1 Mb/mm²
largest printable chip858 mm²lithography reticle, 26 × 33 mm
energy: read 32 bits from DRAM640 pJHorowitz, ISSCC 2014
energy: one 32-bit multiply3.7 pJHorowitz, ISSCC 2014

The one equation

Bandwidth times math per byte

operations per second = bandwidth × operations per byte

(while memory is the limit)

Bandwidth is fixed by the hardware: 3.35 TB/s on the H100. Operations per byte is ours: it depends on how the work is organised, which is what kernels, batching and precision change.

One step, one user

The step takes the longer of the two

memory time = 15.1 GB ÷ 3.35 TB/s = 4.5 ms

compute time = 15.2 G ÷ 989 T = 0.015 ms

step = the longer = 4.5 ms measured: 6.7 ms

The gap between 4.5 ms on paper and 6.7 ms measured is what kernel engineering closes.

One user in numbers

About one operation per byte

operations per byte = 15.2 G ÷ 15.1 GB ≈ 1

3.35 TB/s × 1 = 3.35 TFLOP/s at best: 0.34% of 989

measured: 2.3 TFLOP/s = 0.23% of the GPU

Same bytes, more math

Operations per byte grow with users

usersoperations per byteGPU share on papermeasured, tokens/sGPU share measured
110.34%1480.23%
646020%8,00612%
25619566%20,81832%
1,024452100%26,30340%

Measured on a Modal H100 through the live server, 28 September 2026 (128 tokens of conversation per user). Past 295 operations per byte the chip is compute-bound.

Before the project

The recipe

  1. Count the operations.
  2. Count the bytes.
  3. Divide: operations per byte.
  4. Compare with the ridge, 295: which roof holds you down?
  5. Predict the time: the longer of memory and compute.
  6. Measure, and explain the gap.

Hour 2 · the project

The class is the batch

Qwen3-8B on one H100 on Modal, served with vLLM. One H100 costs about $3.95 an hour. How many of you can it serve at once, and what does each answer cost?

Students

phones and laptops

→

Class page

on Modal, a class code

→

vLLM on one H100

Qwen3-8B, BF16

→

Live roofline

users, tokens a second, cost

probe.py · chat.py · dashboard.py · loadgen.py · loadwatch.py

Predict first

Write your numbers before we run anything

usersmemory timecompute timelimiteach user, tok/stogether, tok/syour guess
14.52 ms0.02 msmemory221221?
84.56 ms0.12 msmemory2191,753?
644.88 ms0.98 msmemory20513,117?
2565.96 ms3.94 msmemory16842,947?
5127.40 ms7.87 mscompute12765,042?
1,02410.29 ms15.74 mscompute6465,042?

Per step on paper: memory time = (15.1 GB + users × 128 tokens × 144 KiB) ÷ 3.35 TB/s; compute time = users × 15.2 billion operations ÷ 989 TFLOP/s. The step takes the longer one.

Your turn

Everyone, send a prompt on 3, 2, 1

Open the class page, type the class code, and ask the chatbot anything. Then watch the dot move to the right.

vizuaraai--l1-dashboard-web.modal.run

class code

roofline

QR code for the class page

Rehearsal run

What the H100 actually did

usersstepeach user, tok/stogether, tok/sper million tokens
16.69 ms149148$7.40
647.79 ms1288,006$0.137
25611.92 ms8420,818$0.053
51220.29 ms4924,631$0.045
1,02437.77 ms2626,303$0.042

Measured on a Modal H100 through the live server and simulated users, 28 September 2026: Qwen3-8B in BF16, vLLM, 64-token prompts and 128-token answers.

What an answer costs

Same GPU, 177 times cheaper per token

cost per million tokens = $1,097 ÷ tokens per second

because the H100 bills $0.001097 a second on Modal.

1 user, 148 tokens a second

$7.40

1,024 users, 26,303 tokens a second

$0.042

Before Lecture 2

Your assignment: the napkin audit

  1. Pick a 7B LLM. For its four op families (QKV projections, attention matmuls, MLP matmuls, elementwise and norm ops), compute the operations per byte at decode (batch 1, one token) and at prefill (2,048 tokens).
  2. Place all eight dots on one hand-drawn H100 roofline and label each memory-bound or compute-bound.
  3. Predict two families' speed, then measure them with the lab's timing harness. One page: the drawing, the dots, the two measurements, and your biggest surprise.

Next: Lecture 2, the CUDA programming model. How a GPU actually runs your code.