Serving Kolibri-1 with a 4-bit KV cache
We introduce a 4-bit KV cache for Kolibri-1's long-context attention layers. On two H100s it fits almost twice as many conversations in memory, for around 70% higher decode throughput.
Kolibri-1 is Aleph Alpha's new open-weight model, trained in the EU and released on 3 October 2026 under Apache 2.0: 78B parameters, 3.5B active, built for German and English. Aleph Alpha serves it on two H100s with FP8 weights and an FP8 KV cache through their vLLM plugin, and that is our stock setup. We store the KV cache of its 10 full-attention layers in 4 bits instead, after rotating the keys by a randomized Hadamard matrix. On the same two GPUs that fits 82% more sessions at 64k context and 95% more at 256k, for about 70% more decode throughput at full memory, with RULER and GPQA Diamond unchanged within noise.
Of Kolibri-1's 50 attention layers, 40 use a 513-token sliding window, so their cache stays small at any context length. The other 10 attend to the full context, and at long context their cache takes most of the GPU memory and most of the bytes read per decode step. We change only those 10. The sliding-window layers stay in FP8, as in stock.
Every session decodes at least as fast as in stock at the same batch size, because each step reads fewer bytes, so the extra capacity turns directly into throughput.
What we do
For each full-attention layer we fix a randomized Hadamard matrix R = H·diag(s)/√128, where H is the 128×128 Hadamard matrix and s a random sign per column, as in QuaRot. Every key is rotated by R before it is cached, and every query by the same R, so q·k is unchanged. Keys and values are then stored as 4-bit codes with one fp16 scale and offset per token and head, 4.25 bits per number. Values are not rotated.
- One key (32 of its 128 dimensions) rounded to 4 bits before and after the rotation. Red is the rounding error.
The rotation is what makes 4 bits work for keys. Kolibri-1's keys put most of their energy in a few fixed channels: across its 10 full-attention layers at 128k context, the largest channel is up to 7.5 times the average, so a handful of outliers set each token's min/max range. After the rotation the largest channel is at most 2.2 times the average, in every layer.
We also cache the keys without Kolibri-1's QK-norm gains. The model computes k = gk ⊙ n(Wkx), with n an RMS normalization and gk learned per-channel gains. Scores only depend on q·k, so we fold gk into the query and cache R·n(Wkx). This lowers the largest channel in 7 of the 10 layers, a small addition to what the rotation does.
At 128k context, 4-bit keys move the attention output by 9.8% without the rotation and by 5.7% with it, averaged over the full-attention layers. Stock's FP8 keys move it by 2.3%.
Normalization, both gains and R fold into one fixed 128×128 matrix per layer for the queries and one for the keys, applied by a single kernel that replaces the stock QK-norm at the same cost. The cache only ever holds rotated keys; decode reads the 4-bit codes directly and nothing is rotated back.
Benchmarks
Everything here runs on one node with two H100 GPUs, the model split across them with tensor parallelism. Stock is the serving command from Kolibri-1's model card, which keeps the KV cache of every layer in FP8. Rotated INT4 adds one option to that command: the 40 sliding-window layers stay in FP8, and the 10 full-attention layers store their keys and values in 4 bits, with one 16-bit scale and offset per token and head, 4.25 bits per number in all.
- Hardware
- 2× H100 SXM 80 GB, TP2
- Software
- vLLM 0.29
aleph-alpha-inference plugin
Rotated INT4 cache: our pull request to the plugin - Stock
vllm serve Aleph-Alpha/Kolibri-1 \ --kv-cache-dtype fp8
- Rotated INT4
vllm serve Aleph-Alpha/Kolibri-1 \ --kv-cache-dtype fp8 \ --additional-config '{"kolibri1": {"kv_cache": "hadamard-int4"}}'- Throughput
- decode only, 512 timed steps per point, every step decoding every session
- Accuracy
- RULER and GPQA Diamond through lm-evaluation-harness, the same prompts for both setups
Throughput
For the throughput chart we fill the cache with sessions until vLLM starts to queue them, then time 512 decode steps at each number of sessions, from the most down to one. A point only counts if every one of its 512 steps produced exactly one token for every session, so no prefill, queueing or eviction happens inside a measurement. All 45 points passed. Within a point the steps are steady: the slowest tenth of them is at most 9% slower than the median. The steps in the curves themselves are not noise. The time per step climbs in plateaus as the batch grows, in both setups and in a repeated run.
Accuracy
A cache half the size is only worth having if the model still finds what it has read. We check that against two benchmarks from Kolibri-1's model card: RULER, for retrieval and reasoning over long contexts, and GPQA Diamond, for reasoning that runs to tens of thousands of generated tokens, every one of them cached in 4 bits in the full-attention layers.
| context | Stock | Rotated INT4 | Δ | 95% interval |
|---|---|---|---|---|
| 4k | 81.4 | 79.4 | −2.0 | −4.9 to +0.8 |
| 8k | 76.2 | 76.5 | +0.3 | −2.4 to +3.2 |
| 16k | 73.6 | 73.2 | −0.4 | −3.0 to +2.1 |
| 32k | 72.1 | 71.7 | −0.5 | −3.1 to +2.4 |
| 64k | 65.4 | 65.5 | +0.0 | −2.5 to +2.6 |
| 128k | 55.2 | 56.1 | +1.0 | −2.4 to +4.3 |
| 256k | 44.8 | 44.5 | −0.3 | −3.1 to +2.7 |
At every length the difference stays between −2.0 and +1.0 points, and every interval includes zero. The model card's RULER rows are for Kolibri Base, which is not released, so we run the released Kolibri-1 on the same tasks, and its scores sit a few points below the card's (81.4 against 86.9 at 4k for stock).
GPQA Diamond asks 198 graduate-level questions in biology, physics and chemistry, and Kolibri-1 reasons for thousands of tokens before each answer, so most of what it reads back from the cache is its own reasoning.
| setup | run 1 | run 2 | run 3 | run 4 | mean |
|---|---|---|---|---|---|
| Stock | 83.3 | 82.3 | 83.3 | 82.8 | 83.0 |
| Rotated INT4 | 82.3 | 85.4 | 81.8 | 84.3 | 83.5 |
Over the four runs, stock scores 83.0 and rotated INT4 83.5, against 84.3 in the model card. When the questions are resampled, the difference of +0.5 points has a 95% interval of −1.8 to +2.7.