cpu-decode
How close can a hand-written int8 decoder get to a laptop’s memory-read limit?
Question
Generating text one token at a time is usually limited by memory: every new token reads all of the model’s weights. I wanted to know how close a small engine written from scratch could get to that limit on a laptop CPU, and how it would compare with llama.cpp, the standard tool for the job.
Method
The engine is a complete C++ forward pass for Qwen2.5-0.5B-Instruct: memory-mapped weights quantized to int8 with one scale per output row, a 32-bit key-value cache, and matrix-vector kernels in scalar, 256-bit and 512-bit SIMD code. Without quantization it matches Hugging Face Transformers to within 0.00044 on every logit checked, and all 32 greedy tokens agree. The ceiling divides the measured read bandwidth, 39 to 43 GB/s depending on thread count, by the bytes each token has to read: about 500 MB of weights plus the cache. llama.cpp ran at a pinned commit with Q8_0 weights, its closest 8-bit format, at the same thread counts. Each point is the median of three runs of 16 tokens.
Result
At short context the engine gets close to the ceiling. With 2 threads and 128 tokens of context it decodes 69.0 tokens per second, 83% of the ceiling, against 61.8 for llama.cpp. That is its best case: at 4 threads the two are level, and at 1 and 6 threads llama.cpp is faster. Long context is where the engine loses. At 4,096 tokens llama.cpp decodes 44.1 tokens per second and the engine 31.7, because its attention is still scalar code, which takes 17.2 of the 31.6 milliseconds per token.
Speed isn’t the whole comparison. The engine’s quantizer reads 5.5% fewer bytes than Q8_0, 8.03 against 8.50 bits per weight, but it distorts the output more: its mean KL divergence from the full-precision model is 0.029 nats against Q8_0’s 0.0097, although it matched the full-precision top token slightly more often, at 57 of 60 positions against 56. llama.cpp also ran with its default 16-bit cache and fused attention, both of which favor it.
Most of the speed came from vectorizing. The 256-bit SIMD kernels were 5.2 times faster than scalar code, and going to 512 bits added 10%. Switching from 16-bit to 8-bit weights changed scalar speed by only 1%. Every method slowed sharply at 12 threads, which is twice the laptop’s six physical cores.
Limits
One model on one laptop, with clock speeds and temperatures not fixed. The ceiling is an estimate from a read benchmark, not a measured DRAM utilization. Quality was checked on four short prompts; agreement at long context wasn’t tested.
Links
Code, data and results GitHub
llama.cpp Gerganov et al.
llama2.c Karpathy, 2023
Qwen2.5-0.5B-Instruct Qwen team, 2024