fbpx

Kolibri AI: Germany’s Free 78B Open-Weight Model, Explained

WANT TO BOOST YOUR SEO TRAFFIC, RANK #1 & Get More CUSTOMERS?

Get free, instant access to our SEO video course, 120 SEO Tips, ChatGPT SEO Course, 999+ make money online ideas and get a 30 minute SEO consultation!

Just Click The Button Below To Get FREE, Instant Access!

Kolibri AI — Kolibri-1 from Germany’s Aleph Alpha — is a brand-new open-weight model you can download free from Hugging Face, with no account. It’s a 78B mixture-of-experts model with a 1M-token context window, built for German and English, and its benchmarks put it up against Qwen, Nemotron and Mistral. I break it down in the video below; here are the specs, the scores and the honest hardware reality.

Key takeaways

  • Kolibri-1 is a 78B mixture-of-experts model with 3.46B active parameters, released 3 October 2026.
  • It’s open weights under Apache 2.0, on Hugging Face at Aleph-Alpha/Kolibri-1.
  • It handles German and English, with up to 1M tokens of context.
  • The full model needs about 78 GB of GPU memory — it’s not a laptop model.

What Kolibri AI is

Aleph Alpha is a German AI company, and Kolibri (“hummingbird” in German) is its new sovereign open-weight model, trained on infrastructure in Germany and Finland. Apart from Mistral, there haven’t been many serious models out of Europe, so this is a big moment.

Kolibri-1 spec Detail
Developer Aleph Alpha (Heidelberg, Germany)
Released 3 October 2026
Architecture Mixture of experts — 78B total parameters, 3.46B active per token
Context window 262K native, extendable to about 1M tokens
Languages German and English
Training data 20T pre-training tokens (about 24% German), plus 3.44T mid-training
Knowledge cutoff June 2026
Licence Apache 2.0 — open weights on Hugging Face
Size About 78 GB in FP8

Kolibri AI benchmarks

Aleph Alpha’s charts show Kolibri scoring highest on its German benchmark average against Nemotron, Gemma 4, Qwen 3.6 and GPT-OSS, and beating Qwen 3.6, Nemotron 3 Super and Mistral Small 4 on a multi-turn tool-calling benchmark. The model card’s headline scores:

Benchmark (Aleph Alpha model card) Kolibri-1
Overall average (English) 75.5%
Overall average (German) 70.8%
MMLU-Pro (English) 80.0%
AIME 2026 96.0%
HumanEval+ 92.7%
SWE-Bench Verified 66.4%

On some benchmarks the leading competitor still scores higher — Kolibri isn’t top everywhere. As always, these are the maker’s own numbers, so test it on your own work.

Want this working in your business, not just bookmarked? Running local and open-weight AI models in your business is exactly what we build together inside the AI Profit Boardroom — 3,400+ members, four live calls a week, and a full one-hour DeepSeek Harness course in the classroom.

Prefer it mapped 1-on-1 first? Book a free strategy session and we’ll plan it for your exact situation.

Join the AI Profit Boardroom →Book a Free Strategy Session →

Built for German, English and compliance

German makes up roughly a quarter of the pre-training data, so Kolibri is genuinely strong in German — useful if you serve German-speaking customers. It’s designed with EU compliance in mind, and I wouldn’t expect much from it in other languages.

It’s text-only (no image generation), supports tool calling and adjustable reasoning effort, and has a June 2026 knowledge cutoff, so it knows about recent events even when it runs offline.

Can you run Kolibri AI locally?

Yes — but be realistic about the hardware. Only 3.46B parameters are active per token, but all 78B have to sit in memory. Several viewers rightly pointed this out.

Hardware Can it run Kolibri-1?
1× H200, B200 or B300 Yes — minimum spec on the model card
2× A100 80GB or 2× H100 Yes — 2× H100 is a recommended setup
Typical laptop or 8–24GB gaming GPU Not the full FP8 model — it needs about 78 GB just for weights
Mac with lots of unified memory Only via community quantised versions, and test it yourself first

On proper GPUs, Aleph Alpha’s recommended route is vLLM:

pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 --tool-call-parser kolibri1

That gives you an OpenAI-compatible endpoint you can point an agent like Hermes at. On a normal computer, look at community quantised versions in LM Studio — or use a smaller local model. On my Mac Studio, the fastest local model I’ve tested with Hermes is LFM 2.5.

Kolibri AI with Hermes Agent

Because it’s built for agentic tool calling, Kolibri is an interesting brain for a free coding or research agent. Serve it with vLLM, then add the endpoint as a custom OpenAI-compatible model in Hermes. My best Hermes setup guide covers profiles and model switching.

Kolibri AI vs other open models

Model Total / active params Context Strength
Kolibri-1 (Aleph Alpha) 78B / 3.46B Up to 1M German + English, tool calling, Apache 2.0
Qwen 3.6 Various sizes Long Strong all-rounder, many sizes
Mistral Small 4 Smaller dense Long European, efficient
Nemotron 3 Super (Nvidia) Large MoE Long Reasoning and agents

Kolibri’s real differentiator is German. If your work is in English only, a smaller Qwen model is often easier to run and very competitive — one viewer argued exactly that, and it’s a fair point. If you serve German-speaking markets or need an EU-built model under Apache 2.0, Kolibri is the one to test.

Who Kolibri AI is for

  • German businesses that want a strong German-language model they can host themselves.
  • Companies with data-sovereignty needs — open weights mean the data never has to leave your servers.
  • Teams with GPU servers running agents that need long context and tool calling.
  • Researchers comparing European open models.

If you’re a solo creator on a laptop, Kolibri isn’t the easy option today — but it’s a strong signal that capable open models are coming from Europe, and smaller or quantised versions may follow.

How to try Kolibri AI without your own GPUs

You don’t need to buy H100s to test it. Options:

  • Rent a GPU server by the hour — a single H200 or B200 instance meets the minimum spec, and you can run the vLLM command above.
  • Wait for hosted endpoints — open-weight models usually appear on inference providers soon after release.
  • Try a community quantised version — check Hugging Face for smaller quantisations, and expect some quality loss.
  • Use a smaller free model in the meantime — for everyday agent work, a free hosted model in Hermes may be all you need.

Test it on your own real tasks — ideally some in German — before deciding.

My verdict on Kolibri AI

Kolibri AI is a genuinely strong open-weight model from Europe — free, Apache 2.0, great at German and built for tool use. If you have the GPUs, it’s well worth testing. If you don’t, watch for good quantised versions, or try a free hosted model like Solar Mini 4 in the meantime.

I’ll update this page as hosted and quantised versions appear.

FAQ: kolibri ai

What is Kolibri AI?

Kolibri-1, a free open-weight model from Germany’s Aleph Alpha: 78B total parameters, 3.46B active, German and English, up to 1M tokens of context.

Is Kolibri free?

Yes — the weights are free on Hugging Face under the Apache 2.0 licence.

Can I run Kolibri on my laptop?

Not the full model — it needs about 78 GB of GPU memory. Community quantised versions may run on high-memory machines.

Is Kolibri better than Qwen?

On Aleph Alpha’s German and tool-calling benchmarks it beats Qwen 3.6, but competitors still lead on some English benchmarks.

Who made Kolibri?

Aleph Alpha, an AI company based in Heidelberg, Germany.

Two ways I can help from here. Join the AI Profit Boardroom — it’s where your local AI setup gets built with 3,400+ members doing the same.

Or grab a free strategy session and bring your questions.

Join the AI Profit Boardroom →Book a Free Strategy Session →

About Julian Goldie

I’m Julian Goldie — founder of Goldie Agency, best-selling author, and an AI educator with 428K+ YouTube subscribers. I’ve spent 10+ years in SEO and online business and I test every AI agent on my own work before I recommend it. Join the AI Profit Boardroom for the daily builds, or book a free strategy session to talk through yours.

Related reading

Last updated October 2026. This page is a living guide to kolibri ai — the facts here move fast and it is updated as they do.

Picture of Julian Goldie

Julian Goldie

Hey, I'm Julian Goldie! I'm an SEO link builder and founder of Goldie Agency. My mission is to help website owners like you grow your business with SEO!

Leave a Comment

WANT TO BOOST YOUR SEO TRAFFIC, RANK #1 & GET MORE CUSTOMERS?

Get free, instant access to our SEO video course, 120 SEO Tips, ChatGPT SEO Course, 999+ make money online ideas and get a 30 minute SEO consultation!

Just Click The Button Below To Get FREE, Instant Access!