One level down: #001 a compiler for an LLM engine

Herbert is the engine I write to run language models in mAIstrow, our AI system. Its first public version did what everyone does: kernels written in advance, one per operation. Developing that code takes an enormous amount of time. And those kernels end up too specific to one particular model.
Looking for where a program spends its cycles, how many registers it occupies, which unit of the processor is really working, how the caches are being hit: that kind of question has drawn me since my first lines of assembly on a Z80. So when Rodrigue and I decided that I would write our own engine, I knew I would end up asking them. We did not write it to redo what exists. We wrote it to know exactly what the machine does at every cycle. And to be able to change it.
And yet, the version I present here does not ship a single line of hand-written assembly. Herbert itself writes the machine code, at startup, on your machine.
This number says how. And what it gives against three other engines. It stops just short of the details, which will get their own articles.
A compiler, not a library of kernels
A classic inference engine ships kernels written in advance. One for matrix multiplication, one for normalization, one for attention, with loops that read their bounds while running, because nobody knew the shape of the model when the kernel was written.
Herbert ships nothing of the kind. It ships emitters. One emitter per family of computation: the projection, the norm, the attention, the recurrence of the Gated DeltaNet mixer, the output head. An emitter does not compute. It writes the code that will compute, for one precise shape.
When the model is loaded, everything is known. The width of each layer, the number of heads, the size of a head, the size of the vocabulary. And at startup on a machine, the rest is known too: the width of the registers, the instructions this processor accepts, the number of physical cores. The emitters receive all of it and put constants where a classic kernel would have a variable. The number of loop iterations is an immediate. The offset to the next block of weights is an immediate. There is no tail code for the remainder of a division, because the division was done once, at emission.
A single dimension is still decided at run time: the number of rows in a block, that is, how many tokens go through at once. At prefill, it is the prompt in blocks. At decode, it is one row. For everything else, the code is specialized.
That is why I call it a compiler for an LLM engine. It takes a model and a machine. It produces the program that runs this model on this machine. That program has nothing left to read while it runs.
These emitters also serve me as a measuring instrument. The ideas about how to compute a matrix multiplication, how to lay out the weights in memory, which quantization to apply to which weights, everybody has them. The only thing that counts is what reality is willing to say about them. A processor that promises an instruction does not promise that it is worth using: that depends on its micro-architecture, on the way it was designed. Remember that, it is what will keep coming back throughout the series. It gets measured.
Then these kernels are assembled. For each regime, prefill and decode, the engine builds a plan: a list of nodes, one kernel each, with its task count and its dependencies. Then the plan is written into a single page of machine code. Every thread of the pool runs this page from the first block to the last. Between two nodes, a barrier, emitted in the code itself. There is no scheduler handing work to executors. There is a program and participants walking through it together.
Reading the same emitter filled in twice
An option writes to disk the code emitted for both regimes. I ran it on the Ryzen 9 9900X, in AVX-512, with two sizes of the same model, Qwen3.5-0.8B and Qwen3.5-4B. Here is the same projection, the one of the first layer at decode, in both cases.
target/release/herbert-bf16-chat ~/models/Qwen3.5-0.8B --dump-asm listings-0.8B target/release/herbert-bf16-chat ~/models/Qwen3.5-4B --dump-asm listings-4B
On the 0.8B, the loop that accumulates a block of sixteen columns of weights:
vxorps zmm0,zmm0,zmm0 mov rcx,0x200 prefetcht0 [rsi+0x40] vmovdqu16 zmm1,[rsi] vpbroadcastd zmm2,[rdi] vdpbf16ps zmm0,zmm2,zmm1 add rdi,4 add rsi,0x40 dec rcx jne short 0x12f
On the 4B, the same projection:
vxorps zmm0,zmm0,zmm0 vxorps zmm1,zmm1,zmm1 mov rcx,0x500 prefetcht0 [rsi+0x40] prefetcht0 [rsi+0x14040] vmovdqu16 zmm2,[rsi] vmovdqu16 zmm3,[rsi+0x14000] vpbroadcastd zmm4,[rdi] vdpbf16ps zmm0,zmm4,zmm2 vdpbf16ps zmm1,zmm4,zmm3 add rdi,4 add rsi,0x40 dec rcx jne short 0x135
Two things changed. Nobody wrote either. The iteration count, loaded into rcx: 0x200, that is 512, on one side, 0x500, that is 1,280, on the other. It is the width of the layer divided by two, since each vdpbf16ps consumes two values per lane: 1,024 for the 0.8B, 2,560 for the 4B. And the shape of the kernel: on the 4B, the emitter chose to accumulate two blocks of sixteen columns at once, in zmm0 and zmm1, the second block read 0x14000 bytes further, that is sixteen columns of 2,560 bf16 values. That distance is an immediate too.
Same emitter, two shapes, two codes. You can produce them again at home with the same command and compare them with mine.
What a token costs
An engine has two regimes. Prefill processes the question: those tokens are already known, we have them all at hand, which opens optimizations impossible otherwise. Decode moves forward token by token, together with the sampler, the little piece of code that picks the next one. Same modules, two methods. For whoever writes the kernels, that is twice as much work, at least. And the two regimes do not have the same physical constraints.
At decode, to write a single token, the model rereads all its weights. For Qwen3.5 in bf16, the language part weighs 1.505 GB on the 0.8B, 3.764 GB on the 2B, 8.412 GB on the 4B. I count the embeddings in full, since they are tied to the output head and read again by it. These figures come from the headers of the safetensors files.
No processor cache holds that. At every token, the weights go through memory. The maximum speed of an engine at decode is therefore the memory bandwidth divided by these bytes. Nobody goes above it, whatever the talent of the code.
Counting an engine's tokens per second is counting the machine's memory as much as the engine. The useful measure is the share of the bandwidth the engine manages to pull.
All of that holds for one request at a time. When a server serves several people in parallel, each read of the weights serves several tokens. Serving several people with Herbert requires specific optimizations. That will be the subject of another article.
And all of this holds for weights in bf16. Quantization, which cuts the bytes to reread at every token, already runs in production in mAIstrow. It will be the subject of one or more articles, with the method of the series: measurement first. That is where the JIT earns its keep: choosing, machine by machine, the weight format and the code that reads it, according to what the processor's instruction set offers.
Measuring the limit rather than assuming it
A machine's bandwidth, I do not read it off the manufacturer's sheet. A probe measures it: a sequential read of memory, emitted by the JIT as well, on all physical cores, the best of several passes. It tells what the machine delivers to a program that reads the way an engine reads.
On an EPYC 4344P, eight Zen 4 cores and DDR5-5200, the probe gives 54.1 GB/s. On a Ryzen 9 9900X, twelve Zen 5 cores with DDR5 capped at 3,600 MT/s, it gives 44 to 46 GB/s from two threads on and 50 with a single one. On the 9900X, a single core reads faster than twelve, because memory is the bottleneck and the threads step on each other. So the share of bandwidth an engine pulls depends on the probe you take as reference: the probe at the engine's thread count, or the machine's best probe. With the first, Herbert is at 98% on the 9900X for the 4B. With the second, at 86%. I give both. For the headline, I keep the probe at the same thread count, because it says what this engine, with these threads, could hope for. The other says what is left to gain by changing strategy. That will be another article.
The predictions, written beforehand
On 25 September 2026, before measuring the 2B, the 4B and the bandwidth, I wrote five predictions in a dated file, published as is with the data. Each prediction says what would refute it.
The first is the physical bound: no engine decodes faster than memory delivers the weights. The second is the thesis: Herbert pulls at least three quarters of the measured bandwidth on the 2B and the 4B. The third says that decode follows the size of the model. The fourth is about llama.cpp's prefill. The fifth says that decode follows the machine: the ratio between the 9900X and the EPYC follows the ratio of their bandwidths.
None was refuted. For this article, I redid the campaign on 8 October 2026 with the published binary, five repetitions instead of three and two more machines. I added predictions before launching, one of them saying that the published binary reproduces the earlier campaign within 2% at decode, then one saying that at default settings Herbert decodes ahead of the three others everywhere. The first holds at decode, within 0.6%; at prefill it is refuted, the published binary does better; I explain why below. The second is refuted on one machine, I come back to it too. The evaluation file, prediction by prediction, is published with the data.
The figures
To put Herbert in perspective, I compare it to three engines. llama.cpp, the most widespread, the one everybody runs at home, behind LM Studio or Ollama. Remarkable work. OpenVINO, Intel's engine for its own hardware. And vLLM, the reference on servers, designed to serve many people at once. Here, with one request at a time, I am not using it to the best of its power. It remains excellent, even for a single person.
The comparison is not there to name the best. It serves as a yardstick: it tells me whether what I measure is relevant and whether an idea holds up against what others have already done. What is true on 8 October 2026 will not necessarily be true tomorrow. All these engines evolve, each in its own way and at its own pace. There may be a competition, I find that healthy. Above all there is, for each engine, a different point of view on the same problem.
Two AVX-512 machines, three model sizes, four engines, bf16 everywhere, one request at a time, the same 148-token prompt, 64 generated tokens. The versions: llama.cpp b11177, vLLM 0.30.0 for CPU with torch 2.13.0, OpenVINO 2026.4.1 with the SDPA backend, Herbert at tag v0.1.0. Medians of five interleaved repetitions, the order of the engines rotated at each repetition, one process per measure, ten seconds of rest and an idleness guard before each run. Each engine at its default thread count, one per physical core.
At decode, on the EPYC 4344P, the physical ceiling for the 4B is 6.43 tokens per second. Herbert gets 6.15. llama.cpp 5.95, OpenVINO 5.83, vLLM 5.60. All four are under the ceiling. The one closest to it is mine, on 8 October 2026, with these versions. That may no longer be true when you read this. And the gap between the first three is thin. That is the expected result: when memory decides, good engines look alike.
On the 2B and the 4B, Herbert pulls 95 to 98% of the bandwidth measured at the same thread count, on both machines: 95.5 and 95.8% on the EPYC, 98.2 and 98.0% on the 9900X. On the 0.8B, it is 93 to 95%: the model is small, the fixed costs weigh more.
| decode, tokens per second | Herbert | llama.cpp | vLLM | OpenVINO |
|---|---|---|---|---|
| EPYC 4344P, 0.8B | 33.30 | 31.88 | 26.43 | 28.70 |
| EPYC 4344P, 2B | 13.71 | 13.50 | 11.96 | 12.79 |
| EPYC 4344P, 4B | 6.15 | 5.95 | 5.60 | 5.83 |
| Ryzen 9 9900X, 0.8B | 27.71 | 25.87 | 24.02 | 23.74 |
| Ryzen 9 9900X, 2B | 11.51 | 11.15 | 10.63 | 10.54 |
| Ryzen 9 9900X, 4B | 5.12 | 4.91 | 4.88 | 4.81 |
The published binary reproduces the earlier campaign at decode within less than 1% on both machines. At prefill, it does not: it does 13 to 30% better than the September binary. I know where that comes from. One change between the two aligns the weight packs, the KV cache and the activations on a cache line. I isolated it step by step: the whole gap is that change, the emitted code is identical listing for listing and decode does not move. A reader who reproduces will find the figures of this article, not those of September.
At prefill, it is another story, because prefill is compute. On the 4B, Herbert does 282 tokens per second on the EPYC and 461 on the 9900X; vLLM 262 and 418; llama.cpp 133 and 276. llama.cpp is between half and two thirds of Herbert, vLLM a little behind. I give these figures with their machine and their settings, not otherwise. They hold on these two AVX-512 BF16 processors, with these versions, at one thread per core.
On one prompt of thirty-two greedy tokens, Herbert and vLLM produce exactly the tokens of the Hugging Face float32 reference, on the three models and the three machines. llama.cpp flips at the thirteenth token of the 0.8B on the AVX-512 machines, a near tie in bf16, not on AVX2. OpenVINO is right with the SDPA backend and wrong with its default backend, paged attention, from the second token of the 0.8B and from the first of the 4B, in 2026.4.1 as in 2026.4.0. I am preparing a report.
What I do not explain yet
Where the prefill gap comes from, I do not say here. I have a hypothesis, it is not measured, so it is worth nothing. It will be tested by replacing Herbert's kernels one by one, in a separate article.
And three results do not go my way. On the 9900X, llama.cpp set to a single thread decodes faster than Herbert at twelve: 28.60 against 27.71 on the 0.8B, 12.46 against 11.51 on the 2B, 5.37 against 5.12 on the 4B. A single core is enough to saturate this capped memory. A single core pays no barrier. Herbert, for its part, collapses when given more threads than physical cores. I know why, that will be an article too. On a Core Ultra 7 258V, a Lunar Lake laptop with AVX2 only, nobody comes close to the bandwidth, 45 to 69% of the probe. vLLM decodes ahead of Herbert there on the 2B and the 4B, by 0.5 and 2.9%. I had predicted the opposite, it is on record. On AVX2, at prefill, llama.cpp is ahead of Herbert everywhere: by a factor of 1.7 to 1.9 on the Lunar Lake, by 37 to 45% on a Zen 3 Ryzen 5700G. Herbert's AVX2 path did not get the work the AVX-512 path got. Emitted code is not magic. It is good where it was worked on.
That Zen 3 says something else too. With DDR4 at 34 GB/s, Herbert's decode pulls 91, 97 and 95% of the probe there, ahead of llama.cpp by 3 to 6%. The thesis holds on slow memory with an AVX2 path that was not worked on. The same binary, built on Windows 11 with MSVC on a Ryzen 5 5500GT, gives the same picture: 92, 97 and 97% of the probe at decode, the same tokens as on Linux and llama.cpp ahead at prefill. Emitted code does not know which system it runs on.
The three AVX2 machines, same settings, medians of five repetitions; vLLM only ran on the Lunar Lake.
| AVX2, tokens per second | decode Herbert | llama.cpp | vLLM | prefill Herbert | llama.cpp | vLLM |
|---|---|---|---|---|---|---|
| Core Ultra 7 258V, 0.8B | 38.87 | 30.95 | 36.90 | 141 | 260 | 133 |
| Core Ultra 7 258V, 2B | 17.31 | 14.48 | 17.40 | 64 | 107 | 83 |
| Core Ultra 7 258V, 4B | 7.59 | 6.50 | 7.81 | 25 | 42 | 32 |
| Ryzen 7 5700G, 0.8B | 20.78 | 19.61 | · | 287 | 417 | · |
| Ryzen 7 5700G, 2B | 8.63 | 8.38 | · | 134 | 184 | · |
| Ryzen 7 5700G, 4B | 3.87 | 3.74 | · | 52 | 74 | · |
| Ryzen 5 5500GT, Windows, 0.8B | 18.50 | 17.58 | · | 154 | 300 | · |
| Ryzen 5 5500GT, Windows, 2B | 7.83 | 7.64 | · | 88 | 136 | · |
| Ryzen 5 5500GT, Windows, 4B | 3.50 | 3.40 | · | 34 | 55 | · |
The question it raises
If decode is bound by memory and if prefill is compute, then a GPU should be able to do as much, or more. Its memory is faster and it computes faster. That is not the question.
The question is: with the same principle? Code emitted at startup, without the vendor's library, with the same numerical contract, bf16 everywhere and f32 accumulators, outputs identical token for token. Does that hold on a GPU, against the reference GPU engines? I do not answer here. That is what the series is going to look at, backend by backend, starting with Vulkan.
Reproducing at home
The repository is online: https://github.com/xigh/herbert-jit-bf16/. Build, run the probe, run the decode:
git clone --branch v0.1.0 https://github.com/xigh/herbert-jit-bf16.git cd herbert-jit-bf16 && cargo build --release --locked target/release/herbert-bf16-chat ~/models/Qwen3.5-4B --no-think --ask 'What is the capital of France?'
The probe is a small separate program, probe-bw, in the measurements repository: a loop of sequential reads emitted by iced-x86, one thread per physical core, the best of several passes.
The campaign scripts, the predictions and the raw data, machine by machine, are published with this article: github.com/xigh/herbert-cpu-bench, tag c1-2026-10-08, folder campaign-c1/. If a figure looks wrong to you, you have what you need to check it.
One level down, or several...
- One level down: #000 an inference engine that writes its own code
- One level down: #001 a compiler for an LLM engine