top of page

Aleph Alpha Kolibri Model Puts German Sovereignty Against Foreign AI Dependence

3 days ago
14 min read

Aleph Alpha released Kolibri on October 3 with open weights, German-centered training, and a direct pitch to Europe’s most regulated institutions. The Aleph Alpha Kolibri model arrives as governments face a basic conflict: use leading foreign AI services or retain deeper control over critical technology.

Kolibri is not simply another German chatbot. It is a bilingual model designed for public administration, industry, defense, and other environments where data location and operational control affect procurement decisions. Aleph Alpha also wants agencies to run the model on their own infrastructure.

That puts Kolibri into a contest involving more than benchmark scores. SAP and OpenAI already offer a different sovereignty model, using German-operated cloud infrastructure around technology supplied by American companies. Mistral AI is building another European route through open models, government partnerships, and European hosting.

Aleph Alpha’s answer is unusually specific. It combines open weights, a German-controlled training pipeline, an efficient mixture-of-experts architecture, and specialized post-training for regulated work.

The unresolved question is whether those ingredients deliver enough practical reliability for public infrastructure. Aleph Alpha has published extensive technical results, but much of the evidence still comes from company-designed evaluations.

Kolibri Changes Aleph Alpha’s Public-Sector Pitch

Kolibri turns Aleph Alpha’s sovereignty argument into a model that agencies can download, inspect, deploy, and adapt.

The company released Kolibri on German Unity Day, giving the launch an intentional political and institutional frame. According to the Kolibri release, the weights are available under the Apache 2.0 license.

That license permits commercial use, modification, and redistribution under relatively permissive terms. Open weights do not reveal every training decision, but they give customers more deployment freedom than a closed application programming interface.

Kolibri uses a mixture-of-experts architecture, often shortened to MoE. This design activates only part of the model for each token instead of running every parameter during every response.

The model contains 78.1 billion total parameters, but it activates about 3.46 billion for each token. Aleph Alpha presents that gap as a route to lower serving costs and faster inference.

Kolibri supports German and English. Its advertised context window reaches 1,048,576 tokens, although Aleph Alpha recommends 262,144 tokens for efficient production work.

A long context window lets a model process large collections of text within one request. That matters for agencies reviewing regulations, case files, policy documents, or lengthy administrative records.

The model also supports tool calls and selectable reasoning levels. Operators can choose no reasoning or low, medium, and high modes, trading response speed against additional computation.

These features support concrete public-sector workflows. A citizen-service agent might retrieve a record, examine eligibility rules, and prepare a form. An analyst might compare policy documents or evaluate planning assumptions.

Aleph Alpha says Kolibri can run on customer-controlled infrastructure. That option matters when documents cannot leave an agency’s security boundary or pass through an outside inference service.

The company trained Kolibri on infrastructure in Germany and Finland. It says the model, development pipeline, and supply chain remain under German and European legal control.

That is a broader sovereignty claim than storing prompts within Germany. It covers how the model was built, where it was trained, who controls deployment, and whether customers can retain a working copy.

Aleph Alpha has already tested its public-sector strategy through F13, an administrative assistant developed with the German state of Baden-Württemberg. Its website also cites an AI assistant rollout intended for 80,000 government-agency users.

Those deployments establish institutional access, not proof that Kolibri will work reliably across government. Moving from document assistance to agents that take actions introduces more serious questions about permissions, accountability, and error recovery.

Kolibri nevertheless gives Aleph Alpha a clearer technical product for that transition. Government buyers can now evaluate a downloadable model rather than relying solely on promises about a proprietary platform.

The release also shifts Aleph Alpha’s competitive position. The company no longer needs to argue that European buyers should accept weaker deployment freedom in exchange for local expertise.

It can instead ask whether a foreign-controlled model remains necessary when a German-built alternative offers inspectable weights, bilingual reasoning, and on-premise operation.

Why the Aleph Alpha Kolibri Model Is Built Around German Work

Kolibri’s most relevant difference is not that it speaks German, but that Aleph Alpha trained it around German language, institutions, and administrative prose.

Most major language models can answer German questions. Aleph Alpha argues that surface fluency is insufficient for work involving laws, public records, policy language, or culturally specific assumptions.

English dominates the highest-volume web datasets and many post-training collections. A multilingual model can therefore produce German text while relying on reasoning patterns learned mainly from English examples.

Aleph Alpha tried to reduce that imbalance during pre-training. German represented 21.3 percent of Kolibri’s pre-training tokens, or roughly 4.3 trillion token presentations during the 20-trillion-token run.

The company says it initially found only 390 billion usable German tokens after filtering and deduplication. That was far below the approximately four trillion tokens required for its planned data mix.

Aleph Alpha filled the gap through German-specific web curation and synthetic rephrasing. Its German data analysis describes 1.3 trillion unique tokens gathered through a dedicated Common Crawl pipeline.

The team also produced about one trillion tokens by rewriting German documents in alternative German forms. The process converted material into formats such as question-and-answer exchanges or encyclopedia-style passages.

This method differs from translating English content. It preserves more of the source document’s institutions, geographic references, and cultural assumptions.

That distinction matters in administration. A training filter designed around English word lengths can mistakenly discard German compound nouns and formal government writing.

Aleph Alpha says it retuned its filtering rules to retain that material. The model’s German corpus also includes parliamentary proceedings, legal texts, and other public-domain documents.

Kolibri uses a bilingual tokenizer, which converts text into units the model can process. Aleph Alpha developed a method called UniBPE to handle German morphology and compound words more efficiently.

Fewer tokens can reduce inference time and memory use. More meaningful word divisions can also preserve structural clues inside specialized German terminology.

The company’s examples include words associated with Germany’s Federal Social Court and administrative log data. Kolibri reportedly splits those terms closer to their linguistic components than several general-purpose tokenizers.

Aleph Alpha also created approximately 800,000 German supervised fine-tuning examples for reasoning. Its German reasoning research says the examples were generated by prompting a model with German sentence beginnings.

The goal was to keep intermediate reasoning in German when the user asked a German question. Without such training, multilingual models often drift into English or mix languages internally.

For government users, understandable reasoning has practical value. A German-speaking employee must review an answer without mentally translating the model’s explanation or overlooking a language-dependent error.

However, visible reasoning should not be confused with a complete audit trail. A generated explanation does not necessarily expose every internal operation that produced an answer.

Public agencies still need source citations, access records, version controls, and documented approval processes. They also need tests showing whether the system behaves consistently across regional language and administrative domains.

Kolibri’s German specialization therefore creates a credible technical distinction, but not automatic institutional trust. It gives evaluators a stronger reason to test the model on German public work.

The strongest evaluation would use real workflows under controlled conditions. It would measure factual accuracy, valid abstention, document retrieval, tool selection, and the time employees spend checking outputs.

That evidence would say more than another general knowledge leaderboard. Administrative systems fail through small procedural mistakes as often as spectacular factual errors.

Public-Sector Procurement Is the Main Contest

The central competition is between full-stack European control and sovereignty delivered through locally operated foreign technology.

OpenAI and SAP announced OpenAI for Germany in September 2025. Their model places OpenAI technology inside an environment supplied through SAP’s Delos Cloud and Microsoft Azure.

The German partnership was presented as a way to serve millions of public-sector employees while meeting local security and legal requirements.

That approach separates operational sovereignty from full technological independence. German organizations can receive local controls and regulated hosting without owning the underlying model or its development pipeline.

For many agencies, that arrangement might be sufficient. Procurement teams often prioritize certified infrastructure, application integration, vendor support, and predictable service over direct access to model weights.

Aleph Alpha asks buyers to adopt a stricter definition. It treats sovereignty as control across data curation, model training, infrastructure, deployment, customization, and continued access.

Kolibri makes that position tangible. An agency can retain the model, run it on-premise, and avoid sending internal material to an external inference provider.

Yet ownership creates responsibilities. The organization must maintain serving infrastructure, secure the model, manage updates, test modifications, and monitor applications built around it.

Open weights can reduce vendor dependence while increasing operational work. A cloud service can be more restrictive while providing faster updates and simpler support.

This is why the Aleph Alpha Kolibri model pressures several groups at once. American model providers must explain what sovereignty means when customers cannot independently operate their core technology.

European cloud companies must show whether local hosting provides meaningful technical independence. Other European model developers must demonstrate comparable German performance and public-sector readiness.

Aleph Alpha also faces pressure from Mistral AI. France and Germany have explored public-administration cooperation involving Mistral and SAP, while France has deployed Mistral-powered tools for government employees.

Mistral’s strategy combines European model development, downloadable models, and commercial services. That makes it a closer structural rival to Kolibri than a closed American service.

Competition is also arriving from publicly supported research. The Soofi model project, for example, describes a sovereign German-English foundation model trained on Deutsche Telekom’s German Industrial AI Cloud.

Government buyers may eventually choose among several European models instead of comparing one domestic provider with American platforms. That change would push procurement toward measurable outcomes.

Performance would then need to cover more than German text generation. Buyers would examine secure deployment, integration costs, update frequency, document grounding, accessibility, and procurement compliance.

Aleph Alpha’s planned combination with Cohere further complicates the sovereignty story. The companies announced a transatlantic group that would operate from Berlin and Toronto.

The proposed company would have more than 1,000 employees, according to recent disclosures. It would combine Aleph Alpha’s European public-sector relationships with Cohere’s model engineering and international reach.

Cohere CEO Aidan Gomez has argued that governments should not need to choose between capability and technological control. The merger details position that balance as the combined company’s central promise.

The merger could give Kolibri stronger distribution, infrastructure access, and research capacity. It also introduces a difficult governance question.

A transatlantic company is not identical to a German-controlled supplier. Agencies will want to know how intellectual property, model development, support obligations, and jurisdictional exposure change after the transaction.

Kolibri was completed before that corporate structure took effect. Its long-term value will depend on whether future versions retain the same deployment rights and European supply-chain commitments.

Public-sector customers therefore face competing definitions of sovereignty:

  • Local operation focuses on where workloads run and who administers the cloud.

  • Model portability focuses on whether customers can retain and move the technology.

  • Supply-chain control covers training, software dependencies, hardware, and legal jurisdiction.

  • Institutional sovereignty concerns who can maintain essential services during a commercial or political dispute.

No single definition resolves every dependency. Even a locally trained model relies on imported accelerators, open-source software, electricity, networking, and specialized vendors.

The relevant question is not whether dependence disappears. It is which dependencies remain, who can interrupt them, and how quickly an agency can replace each component.

Open Weights Strengthen Sovereignty Without Settling It

Kolibri offers more customer control than a closed model service, but open weights alone do not make a public system safe or accountable.

Aleph Alpha has published Kolibri’s weights through its model repository. That enables researchers and organizations to inspect the release, run evaluations, and develop adaptations.

The Apache 2.0 license also lowers a commercial barrier. Agencies and contractors can build applications without negotiating a separate research-only license for the base model.

This openness creates an important verification opportunity. Independent teams can test German language performance, long-context behavior, security, and the company’s efficiency claims.

They can also examine whether the model’s hardware requirements fit realistic public-sector budgets. A model with few active parameters can still require substantial memory because all weights must remain available.

Mixture-of-experts efficiency is therefore workload-dependent. Low active parameter counts can reduce computation, but deployment costs also reflect quantization, concurrency, memory, networking, and context length.

The one-million-token context claim needs similar care. Maximum accepted length does not mean reliable use of every detail across that window.

Long-context evaluations should test retrieval across positions, conflicting evidence, repeated passages, and instructions hidden inside documents. Government archives are rarely clean collections with one obvious answer.

Prompt injection presents another concern. A malicious or accidental instruction embedded inside a retrieved document can redirect an agent unless the surrounding application enforces strict controls.

Kolibri’s tool-use capabilities increase that risk when applications move beyond drafting. An agent that retrieves records or prepares forms needs permission boundaries and human approval before changing official data.

Aleph Alpha has trained the model to abstain when supplied documents do not support an answer. Its Merlin-Arthur method generates positive contexts and deliberately incomplete contexts to teach the model when to refuse.

This approach targets a real deployment problem. Standard model training often rewards guessing because an incorrect attempt occasionally receives credit, while refusing never produces the requested answer.

Aleph Alpha reports that Kolibri withheld an answer on 44 percent of unsupported items in one company evaluation. Kolibri Origin did so on 15 percent.

The company also says Kolibri improved on another grounding test and produced fewer unsupported statements. These results are encouraging, but they need careful interpretation.

A high refusal rate can reduce hallucinations while making a system less useful. The right balance depends on whether the task involves casual drafting, benefits decisions, legal analysis, or public communication.

The published results mix recognized benchmarks with Aleph Alpha’s internal tests. Internal evaluations can represent customer work more closely, but outsiders cannot fully reproduce them without the data and scoring process.

Aleph Alpha reports that its German public-sector proxy score rose from 0.54 for Kolibri Origin to 0.75 for Kolibri. The company has not presented that result as an independently audited measure.

That gap does not invalidate the score. It means readers should treat it as evidence of internal progress rather than proof of production reliability.

The same caution applies to compliance claims. Designing with the EU AI Act, GDPR, and the General-Purpose AI Code of Practice in mind is useful, but compliance depends on the deployed system.

A public agency must evaluate the intended use, data flows, affected people, human oversight, logging, cybersecurity, and possible harms. A compliant base model cannot automatically make every connected application compliant.

Open weights also create security work. Agencies must track model files, dependencies, fine-tuned versions, evaluation records, and who can publish updates.

Local deployment can keep sensitive documents inside a controlled environment. It can also leave an under-resourced organization responsible for patching and monitoring a complex AI stack.

The best sovereignty architecture may therefore combine model portability with managed operations. Agencies need the right to move or continue service without pretending every institution should operate its own AI infrastructure.

Kolibri strengthens that option value. It gives buyers a recoverable technical asset rather than access that ends when a service contract ends.

Whether agencies exercise that option will depend on packaging. Deployment tools, documentation, support partners, certified infrastructure, and reproducible evaluations will matter as much as the model file.

The Benchmarks Are Detailed, but Independent Testing Is Next

Aleph Alpha has disclosed more technical detail than a typical launch, yet the decisive public-sector evidence must come from outside the company.

Kolibri completed pre-training on September 11 and launched on October 3. Aleph Alpha says the main pre-training run lasted 21 days on 768 Nvidia B200 GPUs.

The model processed 20 trillion pre-training tokens. Mid-training added 3.44 trillion tokens, followed by 200 billion tokens for long-context adaptation.

Aleph Alpha says its pipeline processed more than 200 trillion raw tokens before filtering and curation. The company also generated 174 billion synthetic tokens for supervised fine-tuning.

After filtering and combining those samples with permissively licensed datasets, it assembled a 268-billion-token fine-tuning mix. Reinforcement learning used more than 1.2 million curated tasks.

Those numbers describe substantial engineering work. They also give external researchers useful details for evaluating the company’s data and efficiency claims.

Aleph Alpha reports 38 unplanned interruptions during the 21-day pre-training run. Its automated pipeline restarted jobs and resumed from checkpoints without manual intervention.

That operational resilience matters because model development depends on repeatability. A team that can restart, evaluate, and reproduce training runs can correct failures faster than one using fragile research scripts.

The company says only ten of Kolibri’s 50 layers use full attention. The remaining layers apply a 512-token sliding window, limiting computation and memory growth during inference.

Aleph Alpha also compared a 123-billion-parameter experimental configuration with Kolibri’s 78-billion-parameter design. It says the smaller model handled 18 concurrent long-context requests on two H100 GPUs.

The larger configuration reportedly handled three. Kolibri also decoded 28 percent faster in that company test.

These architecture choices fit the intended market. Government agencies need models that can run under infrastructure constraints, not only models that top general leaderboards.

Kolibri’s published scores show competitive results across mathematics, coding, tool use, long-context tasks, and German translations of common evaluations.

The model scored 90 percent on Aleph Alpha’s German version of AIME 2026. It reached 81.3 percent on the German GPQA Diamond evaluation.

Benchmark comparisons include Qwen, Nvidia’s Nemotron family, and Mistral Small. Aleph Alpha says Kolibri matches models with up to four times as many active parameters on selected tasks.

However, benchmark numbers remain sensitive to prompts, inference settings, reasoning budgets, contamination, and scoring choices. Translated evaluations can also behave differently from tests originally written in German.

The most useful next step is reproducible testing by universities, public digital-service teams, and independent model evaluators. Those groups should publish prompts, configurations, hardware, and failure examples.

Realistic tests should include German legal language, municipal terminology, benefits administration, procurement documents, and correspondence containing incomplete evidence.

They should also measure regional variation. German public administration spans federal, state, and municipal institutions with distinct vocabulary, procedures, and document formats.

Evaluation must include ordinary employees, not only model researchers. A technically correct answer can still fail if its explanation is difficult to verify under time pressure.

Public agencies should record how often workers accept, edit, reject, or escalate outputs. They should measure whether Kolibri saves time after verification, not before it.

Tool-using agents require separate testing. Evaluators should introduce inaccessible records, conflicting policies, revoked permissions, outdated documents, and malicious instructions.

The model’s abstention behavior deserves focused review. Tests should distinguish justified refusal from unnecessary refusal and confidently unsupported output.

Independent teams should also evaluate the downloadable release against the version used in Aleph Alpha’s hosted or customized products. Differences in system prompts and retrieval can significantly change behavior.

Until those results appear, Kolibri’s benchmarks support a credible engineering claim. They do not settle whether the model is ready to influence decisions affecting citizens.

That distinction is especially important because public infrastructure has a different tolerance for error. A flawed consumer answer creates inconvenience, while an administrative error can delay services or alter legal outcomes.

Three Signals Will Show Whether Kolibri Matters

Kolibri will matter if public institutions validate it, deploy it at meaningful scale, and retain genuine freedom to operate it.

The first signal is independent technical replication. Researchers should confirm German performance, long-context retrieval, tool use, hardware efficiency, and abstention using published methods.

Strong replication would support Aleph Alpha’s claim that specialization can compete with larger general models. Large gaps would weaken the model’s case for procurement.

The second signal is a named production deployment using Kolibri itself. Aleph Alpha has public-sector relationships, but buyers need evidence tied to this release and a defined workflow.

A useful disclosure would identify the task, user group, security controls, evaluation period, and measured time savings. It should also report error categories and human review requirements.

Deployment counts alone would reveal little. An inactive pilot with thousands of eligible users is not equivalent to daily use in a critical administrative process.

The third signal is the post-merger governance structure around Aleph Alpha and Cohere. Buyers need clear answers about model ownership, future releases, European operations, and contractual continuity.

If Kolibri remains openly available and receives frequent updates, the merger could strengthen its position. If later models move behind restricted services, the present sovereignty claim would look less durable.

Competitor responses will provide additional context. SAP and OpenAI can challenge Aleph Alpha through distribution, existing software relationships, and managed infrastructure.

Mistral can compete through European identity, open models, and government partnerships. Public research projects can offer alternatives without relying on one startup’s commercial roadmap.

That pressure is healthy for public buyers. Sovereignty becomes more credible when agencies can compare several replaceable suppliers rather than designate one national champion.

The Aleph Alpha Kolibri model makes that competition more concrete. It packages German specialization, open weights, local training, and efficient inference into one inspectable release.

Its arrival does not prove that Germany has solved public-sector AI. It does establish a serious alternative to sovereignty defined mainly through hosting contracts.

Public institutions should now test the claim through transparent procurement and published evaluations. Developers can examine the weights, reproduce results, and document failures before Kolibri enters higher-stakes workflows.

For enterprise and government buyers, the immediate action is simple: define what control actually means before selecting a model. If portability, legal jurisdiction, German reasoning, and continued operation matter, each requirement needs a measurable test.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page