# Aleph Alpha Releases Kolibri: 78B MoE Architecture and Specs

URL: https://3ammarketer.com/news/aleph-alpha-kolibri-78b-moe-specs-architecture
Published: 2026-10-08

> Aleph Alpha released Kolibri on October 3, 2026, pairing a 78.1B MoE architecture with 3.46B active parameters and context scaling up to 1,048,576 tokens.

Aleph Alpha released Kolibri on October 3, 2026, introducing a 78.1-billion-parameter Mixture-of-Experts transformer that activates 3.46 billion parameters per token. The release marks a major architectural shift toward sparse execution, routing just 4.4% of total capacity during inference across German and English workloads. [1][2]

## How Kolibri routes 3.46B active parameters across 384 experts

Kolibri structures its architecture across 50 layers containing exactly 78,103,074,560 total parameters and 3,457,573,120 active parameters per token. Each layer holds 384 routed experts with a hidden size of 512, accompanied by 1 shared expert through which every token passes. Token assignment relies on a sigmoid router that selects the top 6 routed experts for each step. [2][3]

![Specification table of Kolibri listing total parameters, active parameters per token, context length, and architecture details.](https://dygolixznokmekyvajbb.supabase.co/storage/v1/object/public/images/content/f2618a4fffc772bf81971087.png)

Attention calculations use 48 query heads and 4 key-value heads. The model supports four distinct reasoning levels: none, low, medium, and high. This design succeeds Kolibri Origin, a predecessor model containing 30B total parameters and 3B active parameters that completed pre-training on June 11, 2026, compared to Kolibri's pre-training completion on September 11, 2026. [1][2][3]

## Hardware footprints for running Kolibri in BF16, FP8, and 2-bit

Kolibri's full precision weights require 156.23 GB in BF16 and 78.86 GB in FP8 formats. The smallest community 2-bit quantization build occupies 25.47 GB of storage, placing local execution out of reach for single 24 GB consumer setups. [4]

Kolibri model weight storage sizes across numerical precisions in October 2026. [4]

- BF16: 156.23 GB
- FP8: 78.86 GB
- 2-bit: 25.47 GB

Production deployments running FP8 weights require at least two 80 GB NVIDIA A100 or H100 GPUs, or a single H200, B200, or B300 accelerator. When serving 256,000-token requests across two NVIDIA H100 GPUs, Kolibri handled 18 concurrent users while decoding 28% faster than an evaluated 123B model variant, which supported 3 users. For infrastructure planning alongside other engines, see our [vLLM vs Ollama concurrency benchmarks](/blog/vllm-vs-ollama-benchmark-concurrency). [3][4]

## Extending context to 1M tokens

Kolibri configures a native max_position_embeddings parameter of 262,144 tokens and has been tested up to 1,048,576 tokens. To balance computational overhead, 40 of its 50 layers operate using sliding-window attention restricted to 512 tokens, with rotary position embeddings applied solely to these sliding layers. Every fifth layer applies full-sequence attention without any positional encodings. [2][3]

Aleph Alpha recommends operating Kolibri at or below 262,144 tokens for standard workloads. On the 1-million-token RULER benchmark, the base model scored 63.2, outpacing Qwen3.5 35B-A3B's base score of 57.5. On the American Invitational Mathematics Examination (AIME) 2025 in German, it posted an 87.5 score against 84.4 for NVIDIA's Nemotron 3 Nano. [2][5]

AIME 2025 mathematics scores evaluated in German in October 2026. [2]

- Kolibri: 87.5
- Nemotron 3 Nano: 84.4

Base model scores on the 1-million-token RULER benchmark reported in October 2026. [2][3]

- Kolibri base: 63.2
- Qwen3.5 35B-A3B base: 57.5

## European sovereign training pipeline and licensing terms

Pre-training consumed approximately 24 trillion tokens on 768 NVIDIA B200 GPUs situated in Germany and Finland under German and European law. The primary 20-trillion-token phase ran for 21 days at 16,384 sequence lengths, suffering 38 unplanned interruptions that recovered automatically. Subsequent training included a 3.44-trillion-token phase at 64K context and 200 billion tokens at 256K context. [1][2][3][5]

The corpus incorporated 21.3% organic German text and limited translated text to 6% overall. The 128,000-token UniBPE tokenizer processes German text at 4.7 bytes per token and English at 4.2 bytes per token, requiring 11.2% fewer tokens on German inputs than GPT-5's tokenizer. Kolibri also incorporates the Merlin-Arthur protocol and abstention data to state when answers are missing from context, maintaining a June 18, 2026 knowledge cutoff. [1][2][6]

Open distribution uses an Apache 2.0 license covering model weights and configuration files while excluding training code, model architecture, parameter settings, or training methods. Serving utilizes the separate Apache 2.0 plugin aleph-alpha-inference, which pinned vLLM 0.29 at release with default settings of temperature 1.0, top_p 0.97, and top_k 128. Teams evaluating local execution runtimes can reference our [local AI model sizing guide](/blog/local-ai-models-guide) for adjacent memory constraints. [1][6][7]

Limit: Official hardware requirements mandate at least two 80 GB enterprise GPUs or an H200/B200/B300 class accelerator for FP8 weights, leaving single-GPU 80 GB instances incapable of hosting the model once KV cache allocations are factored in.

## Sources

1. [Kolibri Has Landed: A Sovereign Open-Weight Model: Aleph Alpha](https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/)
2. [Aleph Alpha Kolibri: How the Sovereign German LLM Works](https://tej.as/blog/aleph-alpha-kolibri)
3. [Aleph Alpha Kolibri: The Surprising Truth About 1M Context](https://www.progressiverobot.com/2026/10/04/aleph-alpha-kolibri-78b-moe-1m-context/)
4. [Kolibri Model (Aleph Alpha): Ollama, Size and Hardware](https://www.systemdesign.academy/explained/kolibri-aleph-alpha)
5. [Aleph Alpha launches Kolibri open model with 78B parameters, 3B active and 1M-token context](https://www.datastudios.org/post/aleph-alpha-kolibri-78b-3b-active-1m-token-context)
6. [Aleph Alpha Kolibri-1: EU Open-Weight Model, Benchmarks & Setup | Connic](https://connic.co/blog/aleph-alpha-kolibri-1)
7. [Kolibri Open-Weight LLM: Hands-On Guide [2026] | QWE AI Academy](https://www.qwe.edu.pl/tutorial/kolibri-aleph-alpha-open-weight-llm-tutorial/)
