← back
Never miss a thing.
One email per new article, plus upcoming course announcements, nothing else. The complete articles, figures and gifts are reserved for Subscriber Area members.

One level down

Herbert-JIT: the code is written at startup

For several months now, I have been working mostly on two projects.

Herbert, my inference engine, runs language models locally.

The latest addition to the mAIstrow family, mAIrness, is the harness that puts those models to work on real projects without handing them the keys to the machine. Think of it as the Claude Code or the Codex that comes with Herbert.

What makes it possible to develop both at once is their pace. A measurement campaign on Herbert takes hours, sometimes several days. A study on mAIrness is the same: it chains dozens of agent runs.

While one campaign is running, I work on the other project.

After a relentless summer of work in 2026, I have finally reached a point where I can talk about these two projects.

In the coming weeks, I will publish results and comments on my work. Think of each article as the log of what moved forward that week.

The harness will come as soon as possible. A few trusted people use it every day, but it is not sturdy enough yet for a wide release.

So to begin with, let me tell you what has changed with Herbert.

The inference engine that writes its own code

On March 15, 2026, I introduced herbert-rs here, an engine in which almost every routine was written by hand, in assembly.

I have changed strategy since then, for two reasons that took me a while to accept.

The first one comes down to the Rust compilers. When you write a compute kernel with intrinsics, that is, functions close to the instructions available in assembly, the compiler keeps the right to reorder and to merge. It even happens to rewrite one variant into another without saying so. I had a bench that concluded that most of the cost came from one precise spot.

The same bench, with code I controlled instruction by instruction, said the opposite. Two identical measurements, no warning: I had to disassemble to understand. And it took time, far too much time.

The second one is about assembly itself. Written by hand, it solves that problem. That is why I had made that choice. But that code is frozen. The shape of a layer, the number of heads, the width of the registers and the instructions available: I only really know them at startup, in front of the model and in front of the machine. All I needed was a kind of template that adapts to the conditions of execution.

That is why, with this new version of Herbert, the code is now written at startup. My program decides on every instruction and it is my program that lays it down. I no longer measure an instruction I am not allowed to emit. Picture something like a compiler for an inference engine.

The version I am going to publish is a little over 40k lines of Rust. It contains no assembly file, no intrinsic function and not a single line of C.

What the engine does

For this engine, a language model works in two regimes. Prefill, when it reads the question. Decode, when it writes the answer, one token after another.

For each of the two, Herbert assembles a single program. All the kernels it needs are in there: the projections, the norms, the attention, the Gated DeltaNet mixer, the output head.

Then there is thread management, to make use of the cores of the processor the engine runs on. Every thread runs that program from end to end. Barriers inside the program keep them in step. That is a choice I will come back to. In short, there is no scheduler on one side and executors on the other.

Rust does only three things: load the weights, emit the code, hold the conversation.

What the repository will hold

At first, a single model, Qwen3.5 in text. A single numerical contract: the weights, the activations and the cache are in bf16, the accumulators in f32. Nothing is quantized.

Three targets:

I will not say more for now, but we are in for a treat.

Reading what the engine wrote

One option writes to disk the code the engine has just emitted. It will be used like this:

cargo build --release
target/release/herbert-bf16-chat ~/models/Qwen3.5-4B --dump-asm listings

Here is the same operation, a projection of the first layer at decode, as the engine wrote it on two machines. First on an AMD processor with AVX-512:

mov           rcx, 0x500
prefetcht0    [rsi+0x40]
prefetcht0    [rsi+0x14040]
vmovdqu16     zmm2,[rsi]
vmovdqu16     zmm3,[rsi+0x14000]
vpbroadcastd  zmm4,[rdi]
vdpbf16ps     zmm0,zmm4,zmm2
vdpbf16ps     zmm1,zmm4,zmm3
add           rdi,4
add           rsi,0x40
dec           rcx
jne           short 0x135

Then on a Mac, with NEON:

movz          x15, #0x140
ldr           q0, [x0], #16
ldp           q4, q5, [x1], #32
ldp           q6, q7, [x1], #32
bfmlalb       v16.4s, v4.8h, v0.h[0]
bfmlalt       v16.4s, v4.8h, v0.h[1]
bfmlalb       v17.4s, v5.8h, v0.h[0]
bfmlalt       v17.4s, v5.8h, v0.h[1]

Look at the first line of each excerpt. The number of loop iterations is hard-coded: 0x500 on one side, 0x140 on the other. So are the offsets to the weights. The engine knew the shape of the layer when it wrote the code, so it has nothing to recompute while it runs.

Once the code is online, you will be able to produce these listings at home and compare them with mine.

What will not be there right away

The 4-bit quantized version of Herbert, Q4, runs in production at customers. I do not consider it publishable yet: I want the same quality of measurement as for bf16 before releasing it.

The Metal and Vulkan versions exist too and they work. They will come later, in their own repository. I do not know whether there will be a CUDA version. I doubt it. We will see.

I am not touching herbert-rs for now. I will archive it once the other versions are available. herbert-jit-bf16 will become the reference. The other models will land there.

Why start with bf16

Q4 is the one in production. Yet bf16 is the one I will publish first. Without quantization, there is no hidden variable left: when a number moves, it is the computation or it is the memory. Quantization will come afterwards. These measurements will tell where it is justified.

The rules of the game

I already wrote it in March: every technical claim rests on reproducible measurements. For this series, that means four things.

If a number looks wrong to you, you will have what you need to check it.

What comes next

The next article opens the measurements. Nothing in it will be written that you cannot replay.