One level down

For several months now, I have been working mostly on two projects.
Herbert, my inference engine, runs language models locally.
The latest addition to the mAIstrow family, mAIrness, is the harness that puts those models to work on real projects without handing them the keys to the machine. Think of it as the Claude Code or the Codex that comes with Herbert.
What makes it possible to develop both at once is their pace. A measurement campaign on Herbert takes hours, sometimes several days. A study on mAIrness is the same: it chains dozens of agent runs.
While one campaign is running, I work on the other project.
After a relentless summer of work in 2026, I have finally reached a point where I can talk about these two projects.
In the coming weeks, I will publish results and comments on my work. Think of each article as the log of what moved forward that week.
The harness will come as soon as possible. A few trusted people use it every day, but it is not sturdy enough yet for a wide release.
So to begin with, let me tell you what has changed with Herbert.
The inference engine that writes its own code
On March 15, 2026, I introduced herbert-rs here, an engine in which almost every routine was written by hand, in assembly.
I have changed strategy since then, for two reasons that took me a while to accept.
The first one comes down to the Rust compilers. When you write a compute kernel with intrinsics, that is, functions close to the instructions available in assembly, the compiler keeps the right to reorder and to merge. It even happens to rewrite one variant into another without saying so. I had a bench that concluded that most of the cost came from one precise spot.
The same bench, with code I controlled instruction by instruction, said the opposite. Two identical measurements, no warning: I had to disassemble to understand. And it took time, far too much time.
The second one is about assembly itself. Written by hand, it solves that problem. That is why I had made that choice. But that code is frozen. The shape of a layer, the number of heads, the width of the registers and the instructions available: I only really know them at startup, in front of the model and in front of the machine. All I needed was a kind of template that adapts to the conditions of execution.
That is why, with this new version of Herbert, the code is now written at startup. My program decides on every instruction and it is my program that lays it down. I no longer measure an instruction I am not allowed to emit. Picture something like a compiler for an inference engine.
The version I am going to publish is a little over 40k lines of Rust. It contains no assembly file, no intrinsic function and not a single line of C.
What the engine does
For this engine, a language model works in two regimes. Prefill, when it reads the question. Decode, when it writes the answer, one token after another.
For each of the two, Herbert assembles a single program. All the kernels it needs are in there: the projections, the norms, the attention, the Gated DeltaNet mixer, the output head.
Then there is thread management, to make use of the cores of the processor the engine runs on. Every thread runs that program from end to end. Barriers inside the program keep them in step. That is a choice I will come back to. In short, there is no scheduler on one side and executors on the other.
Rust does only three things: load the weights, emit the code, hold the conversation.
What the repository will hold
At first, a single model, Qwen3.5 in text. A single numerical contract: the weights, the activations and the cache are in bf16, the accumulators in f32. Nothing is quantized.
Three targets:
- x86-64, Linux and Windows: AVX2 and FMA; the wide kernels with AVX-512 BF16
- AArch64, Apple Silicon processors under macOS: NEON and BF16; SME2 when it is there and even Apple's AMX as an option
- RISC-V, Linux: RVV 1.0 at 256 bits at least, with the bf16 extensions
I will not say more for now, but we are in for a treat.
Reading what the engine wrote
One option writes to disk the code the engine has just emitted. It will be used like this:
cargo build --release target/release/herbert-bf16-chat ~/models/Qwen3.5-4B --dump-asm listings
Here is the same operation, a projection of the first layer at decode, as the engine wrote it on two machines. First on an AMD processor with AVX-512:
mov rcx, 0x500 prefetcht0 [rsi+0x40] prefetcht0 [rsi+0x14040] vmovdqu16 zmm2,[rsi] vmovdqu16 zmm3,[rsi+0x14000] vpbroadcastd zmm4,[rdi] vdpbf16ps zmm0,zmm4,zmm2 vdpbf16ps zmm1,zmm4,zmm3 add rdi,4 add rsi,0x40 dec rcx jne short 0x135
Then on a Mac, with NEON:
movz x15, #0x140 ldr q0, [x0], #16 ldp q4, q5, [x1], #32 ldp q6, q7, [x1], #32 bfmlalb v16.4s, v4.8h, v0.h[0] bfmlalt v16.4s, v4.8h, v0.h[1] bfmlalb v17.4s, v5.8h, v0.h[0] bfmlalt v17.4s, v5.8h, v0.h[1]
Look at the first line of each excerpt. The number of loop iterations is hard-coded: 0x500 on one side, 0x140 on the other. So are the offsets to the weights. The engine knew the shape of the layer when it wrote the code, so it has nothing to recompute while it runs.
Once the code is online, you will be able to produce these listings at home and compare them with mine.
What will not be there right away
The 4-bit quantized version of Herbert, Q4, runs in production at customers. I do not consider it publishable yet: I want the same quality of measurement as for bf16 before releasing it.
The Metal and Vulkan versions exist too and they work. They will come later, in their own repository. I do not know whether there will be a CUDA version. I doubt it. We will see.
I am not touching herbert-rs for now. I will archive it once the other versions are available. herbert-jit-bf16 will become the reference. The other models will land there.
Why start with bf16
Q4 is the one in production. Yet bf16 is the one I will publish first. Without quantization, there is no hidden variable left: when a number moves, it is the computation or it is the memory. Quantization will come afterwards. These measurements will tell where it is justified.
The rules of the game
I already wrote it in March: every technical claim rests on reproducible measurements. For this series, that means four things.
- The code is cited by its tag, never by its main branch, which will move.
- The other engines are cited by their exact version.
- The measurement scripts and the raw data are published with the article.
- My predictions are written before the measurements, dated by a commit.
If a number looks wrong to you, you will have what you need to check it.
What comes next
The next article opens the measurements. Nothing in it will be written that you cannot replay.