Heimr 570M: a local model that stays fast as the context grows
Heimr 570M decodes 2 to 8× faster than comparable open models on long contexts and scores above them on downstream evaluations. It is an experimental model: we are building on top of it and sharing the results early, with more to come as it matures.
Heimr 570M is an experimental model targeting phones and laptops. It is a hybrid mixture-of-experts model built from Mamba-3 [1] layers and Heimr attention layers, a novel attention formulation developed in our lab. It has 2.7B parameters, about 570M of them active per token, and was trained to a 64K context.
Single-stream decode on a device is bound by memory bandwidth, so throughput follows the bytes read per token. In a full-attention transformer those grow with the context through the key-value cache. In Heimr 570M they stay nearly constant: sixteen of its twenty layers carry a fixed-size state, and the four attention layers add 4 KB per token.
five models decode to 64K tokens
at their measured speed
- Single-stream decode speed of five models from 1K to 64K tokens of context, measured on an M5 Pro with 4-bit weights.
- Single-stream decode speed of five models from 1K to 64K tokens of context, measured on an M5 Pro with 4-bit weights.
At 64K, Heimr 570M decodes 14× faster than Qwen3-0.6B, 2.4× faster than Qwen3.5-0.8B and 2.2× faster than Gemma 3 1B, which also keeps full attention on only four layers but reads more weight per token. Decode uses the sparse down-projection kernel from our earlier post, 16% faster than the stock kernel.
Architecture
Twenty layers in an SSSSL pattern: Mamba-3 [1] (18 heads × 64, state dimension 128) on sixteen layers and Heimr attention on layers 4, 9, 14 and 19, each attention layer with a gated value-embedding table [2]. Every feed-forward after layer 0, which is a dense ReLU² MLP, is a mixture of 32 ReLU² experts with top-4 sigmoid routing, aux-loss-free load balancing [3] and one shared expert. The model dimension is 1280 and the vocabulary has 32,768 tokens. Normalisation is a parameter-free RMSNorm. Similar to [2], we update the residual stream as follows: the embedding passes through a smear (x0[t] = e[t] + g · e[t−1]), every layer adds the initial x0 back in (xℓ = λr · xℓ−1 + λ0 · x0), and a backout subtracts a scaled copy of the middle layer's activation from the last layer's output before the head (x = xlast − λb · xmid). The logits are softcapped.
the forward pass, in order
- The forward pass: embedding and smear, four repeats of four Mamba-3 layers and one Heimr attention layer (amber), backout from the middle layer, then the head.
Initialisation and training
Heimr 570M was trained on 200B tokens in two phases: 160B at 4K context with a teacher, then 40B at 64K context without one. It uses the tokenizer of Mistral-7B-v0.3 [4], the model we initialise from and distill from.
Embeddings from a larger model's principal components
Similar to weight selection [5] and GUIDE [6], we initialise the small model from a larger one, in two places. The token embeddings are the centered embedding table of the larger model, projected onto its top 1280 principal components and rescaled to the initialiser's standard deviation. The value embeddings of the four attention layers take the top 256 components of the same basis, which also stabilised the mixture-of-experts training.
a few hundred tokens from random positions
to their PCA positions, in 3D
- A random initialisation of 476 tokens, 316 of them from 14 categories: no structure, and each token's nearest neighbours are unrelated.
- The same tokens after PCA initialisation, viewed along three principal directions of these tokens: months, US states, first names, digits and the other categories each form a cluster.
Distillation and schedule
Phase 1, 160B tokens at 4K context, trains on an equal mix of cross-entropy and KL divergence to the top-256 logits of Mistral 7B [4]. The KL term is annealed out at the end of the phase. Phase 2 continues on cross-entropy alone for 40B tokens at 64K context on a long-document mix. Muon trains the weight matrices at lr 0.02 and AdamW the embeddings, router and scalars at 1.2e-3, with 2M-token batches, a warmup-stable-decay schedule and 16 H100s throughout.
CORE
Base-model quality is measured on DCLM's CORE suite [7], with every model scored by the same harness on the full 91,037 examples.
| task | Heimr 570M | Qwen3-0.6B | Qwen3.5-0.8B | Llama 3.2 1B | Gemma 3 1B |
|---|---|---|---|---|---|
| agi eval lsat ar | 0.274 | 0.261 | 0.252 | 0.248 | 0.235 |
| arc challenge | 0.476 | 0.434 | 0.442 | 0.375 | 0.399 |
| arc easy | 0.755 | 0.729 | 0.716 | 0.683 | 0.702 |
| bb cs algorithms | 0.533 | 0.483 | 0.488 | 0.464 | 0.452 |
| bb dyck languages | 0.175 | 0.123 | 0.249 | 0.133 | 0.226 |
| bb language identification | 0.306 | 0.366 | 0.321 | 0.250 | 0.251 |
| bb operators | 0.543 | 0.600 | 0.524 | 0.409 | 0.271 |
| bb qa wikidata | 0.632 | 0.600 | 0.593 | 0.694 | 0.673 |
| bb repeat copy logic | 0.031 | 0.156 | 0.063 | 0.125 | 0.063 |
| boolq | 0.679 | 0.741 | 0.753 | 0.650 | 0.653 |
| commonsense qa | 0.577 | 0.660 | 0.656 | 0.369 | 0.295 |
| copa | 0.770 | 0.700 | 0.690 | 0.740 | 0.750 |
| coqa | 0.379 | 0.394 | 0.392 | 0.360 | 0.354 |
| hellaswag | 0.687 | 0.525 | 0.540 | 0.651 | 0.613 |
| hellaswag zeroshot | 0.680 | 0.521 | 0.532 | 0.635 | 0.606 |
| jeopardy | 0.343 | 0.251 | 0.198 | 0.348 | 0.360 |
| lambada openai | 0.578 | 0.541 | 0.534 | 0.622 | 0.557 |
| openbook qa | 0.420 | 0.350 | 0.376 | 0.386 | 0.376 |
| piqa | 0.783 | 0.711 | 0.711 | 0.756 | 0.756 |
| squad | 0.532 | 0.578 | 0.598 | 0.520 | 0.476 |
| winograd | 0.806 | 0.758 | 0.769 | 0.809 | 0.802 |
| winogrande | 0.627 | 0.586 | 0.583 | 0.610 | 0.588 |
| CORE | 0.411 | 0.375 | 0.372 | 0.364 | 0.345 |
ClimbMix [8], part of the pretraining mix, was built by its authors against PIQA, ARC-Easy and HellaSwag. Without those three tasks Heimr 570M scores 0.369 against 0.359 and 0.355 for the two Qwens, a narrower lead.
References
- Aakash Lahoti et al., "Mamba-3: Improved Sequence Modeling using State Space Principles", arXiv:2603.15569, 2026.
- Keller Jordan et al., "modded-nanogpt: Speedrunning the NanoGPT baseline", GitHub repository, 2024.
- DeepSeek-AI, "DeepSeek-V3 Technical Report", arXiv:2412.19437, 2024.
- Albert Q. Jiang et al., "Mistral 7B", arXiv:2310.06825, 2023.
- Zhiqiu Xu et al., "Initializing Models with Larger Ones", ICLR, 2024.
- Khoa Trinh et al., "GUIDE: Guided Initialization and Distillation of Embeddings", arXiv:2510.06502, 2025.
- Jeffrey Li et al., "DataComp-LM: In search of the next generation of training sets for language models", NeurIPS Datasets and Benchmarks, 2024.
- Shizhe Diao et al., "Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training", arXiv:2504.13161, 2025.