All Insights
#DiffusionModels
July 2026

Sunday Coffee & Code: - DiffusionGemma on a 24GB GPU (yup, still gemma)

Grab a big coffee - last week I said '๐˜ช๐˜ต ๐˜ธ๐˜ฐ๐˜ถ๐˜ญ๐˜ฅ ๐˜ฃ๐˜ฆ ๐˜ช๐˜ฏ๐˜ต๐˜ฆ๐˜ณ๐˜ฆ๐˜ด๐˜ต๐˜ช๐˜ฏ๐˜จ ๐˜ต๐˜ฐ ๐˜ต๐˜ณ๐˜บ ๐˜ต๐˜ฉ๐˜ฆ 26๐˜‰ ๐˜”๐˜ฐ๐˜Œ ๐˜ฎ๐˜ฐ๐˜ฅ๐˜ฆ๐˜ญ.' Google ๐——๐—ถ๐—ณ๐—ณ๐˜‚๐˜€๐—ถ๐—ผ๐—ป๐—š๐—ฒ๐—บ๐—บ๐—ฎ enters the room - a 26B MoE that doesn't generate text one token at a time, it ๐˜ฅ๐˜ฆ๐˜ฏ๐˜ฐ๐˜ช๐˜ด๐˜ฆ๐˜ด blocks of 256 tokens in parallel, like a photo coming into focus. Mission: get a new diffusion LLM running on my single L4 (24GB) via vLLM (๐˜ฎ๐˜ข๐˜บ๐˜ฃ๐˜ฆ ๐˜ต๐˜ช๐˜ฎ๐˜ฆ ๐˜ต๐˜ฐ ๐˜ถ๐˜ฑ๐˜จ๐˜ณ๐˜ข๐˜ฅ๐˜ฆ ๐˜ฎ๐˜บ ๐˜ˆ๐˜ž๐˜š ๐˜Œ๐˜Š2 ๐˜ด๐˜ฆ๐˜ณ๐˜ท๐˜ฆ๐˜ณ). Small problem: the weights are 49GB. My card is 24GB (Claude to the rescue).

By Steve Harris

๐—ฆ๐˜‚๐—ป๐—ฑ๐—ฎ๐˜† ๐—–๐—ผ๐—ณ๐—ณ๐—ฒ๐—ฒ & ๐—–๐—ผ๐—ฑ๐—ฒ - ๐——๐—ถ๐—ณ๐—ณ๐˜‚๐˜€๐—ถ๐—ผ๐—ป๐—š๐—ฒ๐—บ๐—บ๐—ฎ ๐—ผ๐—ป ๐—ฎ ๐Ÿฎ๐Ÿฐ๐—š๐—• ๐—š๐—ฃ๐—จ (๐˜†๐˜‚๐—ฝ, ๐˜€๐˜๐—ถ๐—น๐—น ๐—š๐—ฒ๐—บ๐—บ๐—ฎ)

Grab a big coffee - last week I said โ€๐˜ช๐˜ต ๐˜ธ๐˜ฐ๐˜ถ๐˜ญ๐˜ฅ ๐˜ฃ๐˜ฆ ๐˜ช๐˜ฏ๐˜ต๐˜ฆ๐˜ณ๐˜ฆ๐˜ด๐˜ต๐˜ช๐˜ฏ๐˜จ ๐˜ต๐˜ฐ ๐˜ต๐˜ณ๐˜บ ๐˜ต๐˜ฉ๐˜ฆ 26๐˜‰ ๐˜”๐˜ฐ๐˜Œ ๐˜ฎ๐˜ฐ๐˜ฅ๐˜ฆ๐˜ญ.โ€ Google ๐——๐—ถ๐—ณ๐—ณ๐˜‚๐˜€๐—ถ๐—ผ๐—ป๐—š๐—ฒ๐—บ๐—บ๐—ฎ enters the room - a 26B MoE that doesnโ€™t generate text one token at a time, it ๐˜ฅ๐˜ฆ๐˜ฏ๐˜ฐ๐˜ช๐˜ด๐˜ฆ๐˜ด blocks of 256 tokens in parallel, like a photo coming into focus. Mission: get a new diffusion LLM running on my single L4 (24GB) via vLLM (๐˜ฎ๐˜ข๐˜บ๐˜ฃ๐˜ฆ ๐˜ต๐˜ช๐˜ฎ๐˜ฆ ๐˜ต๐˜ฐ ๐˜ถ๐˜ฑ๐˜จ๐˜ณ๐˜ข๐˜ฅ๐˜ฆ ๐˜ฎ๐˜บ ๐˜ˆ๐˜ž๐˜š ๐˜Œ๐˜Š2 ๐˜ด๐˜ฆ๐˜ณ๐˜ท๐˜ฆ๐˜ณ).

Small problem: the weights are 49GB. My card is 24GB (Claude to the rescue).

๐—ง๐—ต๐—ฒ ๐—ผ๐—ฏ๐˜€๐˜๐—ฎ๐—ฐ๐—น๐—ฒ ๐—ฐ๐—ผ๐˜‚๐—ฟ๐˜€๐—ฒ:

  • CUDA mismatch: PyTorch wheels built for CUDA 13, my driver stuck at 12.8. Fixed with vLLMโ€™s cu129 wheel.
  • 49GB of weights vs 24GB of VRAM: Red Hat published a 4-bit quant version - 17GB, fits.
  • A vLLM bug where an unrelated modelโ€™s warmup kernel crashed startup. Patched locally, bug report heading upstream.
  • Memory Tetris: 0.90 GPU utilization OOMโ€™d, 0.80 starved the KV cache. 0.86 + 16k context: ๐˜ˆ๐˜ฑ๐˜ฑ๐˜ญ๐˜ช๐˜ค๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ ๐˜ด๐˜ต๐˜ข๐˜ณ๐˜ต๐˜ถ๐˜ฑ ๐˜ค๐˜ฐ๐˜ฎ๐˜ฑ๐˜ญ๐˜ฆ๐˜ต๐˜ฆ.

๐—ง๐—ต๐—ฒ ๐—ฝ๐—ฎ๐˜†๐—ผ๐—ณ๐—ณ:

  • Asked it about the Chamber of Secrets. Coherent, well-structured answer - you can literally read the modelโ€™s thinking block planning the outline before it writes.
  • The metrics tell the diffusion story: ๐Ÿญ๐Ÿฏ.๐Ÿฐ๐Ÿณ ๐˜๐—ผ๐—ธ๐—ฒ๐—ป๐˜€ ๐—ฐ๐—ผ๐—บ๐—บ๐—ถ๐˜๐˜๐—ฒ๐—ฑ ๐—ฝ๐—ฒ๐—ฟ ๐—ณ๐—ผ๐—ฟ๐˜„๐—ฎ๐—ฟ๐—ฑ ๐—ฝ๐—ฎ๐˜€๐˜€. An autoregressive model commits exactly 1. It even stopped denoising early when confident.
  • ~51 tokens/s on a $0.80/hr GPU - and thatโ€™s the floor, not the ceiling.

๐—ช๐—ต๐—ฎ๐˜ ๐—œ ๐—น๐—ฒ๐—ฎ๐—ฟ๐—ป๐—ฒ๐—ฑ:

  • The memory squeeze is three levers: quantization, context window (hello again, num_ctx), and GPU memory utilization. vLLM makes you set all three; Ollama quietly decides for you.
  • Model released โ†’ community quants โ†’ running on commodity hardware happens very quickly. The gap between โ€œannouncedโ€ and โ€œtest it yourselfโ€ has collapsed.

Anyone else poking at diffusion LLMs yet,?

A good Sunday, and the coffee went cold around the third OOM.

Want to Discuss This Topic?

Steve is always happy to have a direct conversation.