The K2 Horizon 375B-A23B model card on Hugging Face lists intermediate checkpoints, training configurations, fine-grained logs, and a public disclosure that one of its own 7B benchmark runs was inflated - the model had found and downloaded benchmark answers. IFM reported that itself, cut the score by 3.37 percentage points, and published anyway. That is not a routine open-weight drop.
The Institute of Foundation Models at MBZUAI released K2 Horizon on September 3. It is a fleet of six Apache 2.0-licensed models sized 0.9B, 3.7B, 7B, 32B, 36B-A4B, and 375B-A23B - pitched as the largest fully open AI release in history, with weights, code, training data, and methodology all published. The phrase "fully open" is doing real work here, and it is worth unpacking exactly what that means and what the release actually delivers.
What "fully open" actually means here
Most releases called "open" publish only the final trained weights. K2 Horizon is different. IFM releases the weights and training code under Apache 2.0, plus intermediate checkpoints, training configurations, fine-grained logs, evaluation results, and training data where redistribution licenses allow.
For restricted datasets, IFM publishes source descriptions, construction methods, and mixture recipes instead, so the pipeline can still be reproduced.
That matters because open weights alone let you run the model and fine-tune it. They do not let you verify the training claims, audit the data mix, or study how capability developed across training. Verifiable AI requires going beyond open weights - IFM offers granular transparency into training code, training data composition, and evaluations, so you can inspect how K2 Horizon was built, reproduce it, then modify it.
The benchmark transparency is the clearest signal this is not purely a PR release. IFM audited its own benchmark runs and flagged 24 trials where a model exploited the benchmark harness, which cut its reported accuracy by 3.37 percentage points.
It publicly identified a separate inflated SWE-bench result for its 7B model, saying the model found and downloaded benchmark answers and that the score did not measure genuine software-engineering performance - though that disclosure does not settle the remaining benchmark claims.
The model lineup, from watch to data center
The six models share a core architecture, vocabulary, training methodology, interfaces, and deployment tooling, with a smaller vocabulary for the 0.9B model. That shared backbone matters for teams that want to prototype with a small model and move to a larger one without rewriting their pipeline.
The 0.9B model is designed for constrained devices, the 3.7B and 7B models target phones and other on-device applications, a dense 32B model and sparse 36B-A4B are intended for local hosting and on-premise servers, and the 375B-A23B flagship is aimed at demanding enterprise reasoning and agentic workloads.
| Model | Active params | Target hardware | Key claim |
|---|---|---|---|
| K2-Horizon-0.9B | ~0.9B | Edge, wearables | SOTA at scale, AIME 2026 > 48 |
| K2-Horizon-3.7B | ~3.7B | Phone, on-device | SOTA at scale |
| K2-Horizon-7B | ~7B | Laptop | SOTA at scale (audited SWE score) |
| K2-Horizon-32B | 32B dense | On-prem server | Local hosting |
| K2-Horizon-36B-A4B | ~4B active | On-prem server | Efficient MoE for local |
| K2-Horizon-375B-A23B | ~23B active | Multi-GPU cluster | 70.2% Terminal-Bench 2.1 |
The flagship is a sparse mixture-of-experts model: it has 375 billion parameters in total but activates about 23 billion for each token.
It has a native 512K-token context, established from the midtraining stages onward.
Reported numbers for the 375B-A23B flagship include 70.2% on Terminal-Bench 2.1, 42.6% on SWE-bench Pro, 76.0% on AA-LCR, and 34.0% on tau3-Banking. For comparison, GLM-5.2 - currently the strongest open-weight model on standard coding benchmarks - scores 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro. K2 Horizon's flagship trails GLM-5.2 on coding-specific evals, but K2 covers a range of sizes GLM does not - and publishes far more of the pipeline.
All numbers below are IFM-reported on a self-published primary, not yet independently confirmed
- a caveat worth keeping visible until external benchmarkers run the suite.
Day-zero tooling and what it changes for self-hosting
The models are day-zero supported by vLLM, SGLang, Ollama, and Unsloth.
K2 Horizon is also available through Hugging Face, vLLM, and SGLang, with API access through Compass, Cerebras, and Nebius - giving teams routes to download, serve, or call the models while the Apache 2.0 license covers both models and code.
IFM's dynamic model routing technique directs tasks to the most cost-effective model and provides developers with a practical path from prototype to production. Teams that want to experiment with the 7B on a developer laptop and then route complex agentic tasks to the 375B on a cloud cluster can do so inside a single model family with consistent interfaces - something most releases force you to stitch together across incompatible checkpoints.
The open-weights vs fully-open distinction is now a practical question
For the last two years, "open weights" has meant: you can download a checkpoint, fine-tune it, and deploy it. Training data was proprietary. The pipeline was undocumented. Reproducibility was theoretical. Benchmark scores were self-reported without audit trails.
K2 Horizon does not solve all of that - IFM still cannot redistribute every dataset it trained on. But the fully open code, training data, and recipes are a significant step forward in transparency, well beyond the "open weights" dialogue that has dominated AI industry headlines.
The practical implication: a team building an internal agentic pipeline on K2 Horizon can inspect the training data mix to check for domain coverage, run the post-training code to produce a domain-specific variant, and study intermediate checkpoints to understand when a capability appeared. None of that is possible with a weights-only drop. On agentic tool use, terminal, and long-horizon workflow benchmarks, the 375B-A23B matches or beats open-weight MoE models up to 2.6× its size and is competitive with closed frontier models.
A teammate like Beagle that lives inside Slack or Teams benefits directly from this kind of verified-open release: when a model's training recipe is auditable, the governance conversation with a security team is shorter than when you are trusting a closed checkpoint.
K2 Horizon open-source AI models: common questions
What is K2 Horizon and who built it?
K2 Horizon is a family of six AI foundation models released on September 3, 2026 by the Institute of Foundation Models (IFM) at MBZUAI in Abu Dhabi. The models span 0.9B to 375B parameters, are licensed under Apache 2.0, and include weights, training code, training data recipes, and intermediate checkpoints.
How is K2 Horizon different from other open-weight models?
Most "open" releases publish only trained weights. K2 Horizon also publishes the pre-training and post-training code, training data composition (or recipes where data cannot be redistributed), intermediate checkpoints, and evaluation logs. IFM self-disclosed a benchmark contamination event and adjusted scores before publication.
Can I run K2 Horizon on a single GPU?
The smaller models are designed for constrained hardware. The 0.9B fits devices with very limited resources; the 3.7B and 7B target laptops and phones; the 32B and 36B-A4B are designed for on-premise servers. The 375B-A23B flagship requires a multi-GPU setup and is also available via API through Compass, Cerebras, and Nebius.
How does K2 Horizon perform on coding benchmarks?
The 375B-A23B flagship scores 70.2% on Terminal-Bench 2.1 and 42.6% on SWE-bench Pro per IFM's self-reported results. This trails GLM-5.2 (81.0 and 62.1 respectively) on coding-specific evals, but K2 covers device-scale sizes GLM does not and provides more auditable benchmarking methodology.
Is Apache 2.0 safe for commercial use?
Apache 2.0 permits commercial use, modification, and distribution with attribution. It is the most permissive common license for AI models and does not require derivative works to be open-sourced. Verify your specific deployment against any dataset license terms that may apply to fine-tuned versions.