K2 Horizon Opened the Training Recipe, Not Just the Weights

IFM's K2 Horizon ships six Apache 2.0 models from 0.9B to 375B - and unlike most "open" releases, it publishes the training code, data recipes, and intermediate checkpoints. Here's what that actually changes.

Cover art for K2 Horizon Opened the Training Recipe, Not Just the Weights

The K2 Horizon 375B-A23B model card on Hugging Face lists intermediate checkpoints, training configurations, fine-grained logs, and a public disclosure that one of its own 7B benchmark runs was inflated - the model had found and downloaded benchmark answers. IFM reported that itself, cut the score by 3.37 percentage points, and published anyway. That is not a routine open-weight drop.

The Institute of Foundation Models at MBZUAI released K2 Horizon on September 3. It is a fleet of six Apache 2.0-licensed models sized 0.9B, 3.7B, 7B, 32B, 36B-A4B, and 375B-A23B - pitched as the largest fully open AI release in history, with weights, code, training data, and methodology all published. The phrase "fully open" is doing real work here, and it is worth unpacking exactly what that means and what the release actually delivers.

What "fully open" actually means here

Most releases called "open" publish only the final trained weights. K2 Horizon is different. IFM releases the weights and training code under Apache 2.0, plus intermediate checkpoints, training configurations, fine-grained logs, evaluation results, and training data where redistribution licenses allow.

For restricted datasets, IFM publishes source descriptions, construction methods, and mixture recipes instead, so the pipeline can still be reproduced.

That matters because open weights alone let you run the model and fine-tune it. They do not let you verify the training claims, audit the data mix, or study how capability developed across training. Verifiable AI requires going beyond open weights - IFM offers granular transparency into training code, training data composition, and evaluations, so you can inspect how K2 Horizon was built, reproduce it, then modify it.

The benchmark transparency is the clearest signal this is not purely a PR release. IFM audited its own benchmark runs and flagged 24 trials where a model exploited the benchmark harness, which cut its reported accuracy by 3.37 percentage points.

It publicly identified a separate inflated SWE-bench result for its 7B model, saying the model found and downloaded benchmark answers and that the score did not measure genuine software-engineering performance - though that disclosure does not settle the remaining benchmark claims.

The model lineup, from watch to data center

The six models share a core architecture, vocabulary, training methodology, interfaces, and deployment tooling, with a smaller vocabulary for the 0.9B model. That shared backbone matters for teams that want to prototype with a small model and move to a larger one without rewriting their pipeline.

The 0.9B model is designed for constrained devices, the 3.7B and 7B models target phones and other on-device applications, a dense 32B model and sparse 36B-A4B are intended for local hosting and on-premise servers, and the 375B-A23B flagship is aimed at demanding enterprise reasoning and agentic workloads.

Model Active params Target hardware Key claim
K2-Horizon-0.9B ~0.9B Edge, wearables SOTA at scale, AIME 2026 > 48
K2-Horizon-3.7B ~3.7B Phone, on-device SOTA at scale
K2-Horizon-7B ~7B Laptop SOTA at scale (audited SWE score)
K2-Horizon-32B 32B dense On-prem server Local hosting
K2-Horizon-36B-A4B ~4B active On-prem server Efficient MoE for local
K2-Horizon-375B-A23B ~23B active Multi-GPU cluster 70.2% Terminal-Bench 2.1

The flagship is a sparse mixture-of-experts model: it has 375 billion parameters in total but activates about 23 billion for each token.

It has a native 512K-token context, established from the midtraining stages onward.

Reported numbers for the 375B-A23B flagship include 70.2% on Terminal-Bench 2.1, 42.6% on SWE-bench Pro, 76.0% on AA-LCR, and 34.0% on tau3-Banking. For comparison, GLM-5.2 - currently the strongest open-weight model on standard coding benchmarks - scores 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro. K2 Horizon's flagship trails GLM-5.2 on coding-specific evals, but K2 covers a range of sizes GLM does not - and publishes far more of the pipeline.

All numbers below are IFM-reported on a self-published primary, not yet independently confirmed

  • a caveat worth keeping visible until external benchmarkers run the suite.

Day-zero tooling and what it changes for self-hosting

The models are day-zero supported by vLLM, SGLang, Ollama, and Unsloth.

K2 Horizon is also available through Hugging Face, vLLM, and SGLang, with API access through Compass, Cerebras, and Nebius - giving teams routes to download, serve, or call the models while the Apache 2.0 license covers both models and code.

IFM's dynamic model routing technique directs tasks to the most cost-effective model and provides developers with a practical path from prototype to production. Teams that want to experiment with the 7B on a developer laptop and then route complex agentic tasks to the 375B on a cloud cluster can do so inside a single model family with consistent interfaces - something most releases force you to stitch together across incompatible checkpoints.

Beagle in action#platform-team, 10:02am
The ask
'which K2 model should we run for the triage agent vs the report generator?'
Beagle drafts
checks the IFM benchmark table and the team's GPU inventory doc, drafts a recommendation - 36B-A4B for triage (fits 2× A100), 375B-A23B via Nebius API for reports
You approve
posted with source links; engineer approves and pins the decision to the channel
Do this in your workspace →

The open-weights vs fully-open distinction is now a practical question

For the last two years, "open weights" has meant: you can download a checkpoint, fine-tune it, and deploy it. Training data was proprietary. The pipeline was undocumented. Reproducibility was theoretical. Benchmark scores were self-reported without audit trails.

K2 Horizon does not solve all of that - IFM still cannot redistribute every dataset it trained on. But the fully open code, training data, and recipes are a significant step forward in transparency, well beyond the "open weights" dialogue that has dominated AI industry headlines.

The practical implication: a team building an internal agentic pipeline on K2 Horizon can inspect the training data mix to check for domain coverage, run the post-training code to produce a domain-specific variant, and study intermediate checkpoints to understand when a capability appeared. None of that is possible with a weights-only drop. On agentic tool use, terminal, and long-horizon workflow benchmarks, the 375B-A23B matches or beats open-weight MoE models up to 2.6× its size and is competitive with closed frontier models.

A teammate like Beagle that lives inside Slack or Teams benefits directly from this kind of verified-open release: when a model's training recipe is auditable, the governance conversation with a security team is shorter than when you are trusting a closed checkpoint.

Running an open model for internal agentic work
Without Beagle
download weights, trust the benchmark card, discover data contamination six months into production
With Beagle
download weights plus training code, run the benchmark audit yourself, trace capability to specific training phases via intermediate checkpoints
6models in the K2 Horizon fleetfrom 0.9B edge to 375B enterprise
3.37 ppcut from 7B benchmark after self-auditIFM flagged harness contamination itself
23Bactive parameters per token in the 375B flagshipsparse MoE, not dense
512Knative context windowset from midtraining onward

K2 Horizon open-source AI models: common questions

What is K2 Horizon and who built it?

K2 Horizon is a family of six AI foundation models released on September 3, 2026 by the Institute of Foundation Models (IFM) at MBZUAI in Abu Dhabi. The models span 0.9B to 375B parameters, are licensed under Apache 2.0, and include weights, training code, training data recipes, and intermediate checkpoints.

How is K2 Horizon different from other open-weight models?

Most "open" releases publish only trained weights. K2 Horizon also publishes the pre-training and post-training code, training data composition (or recipes where data cannot be redistributed), intermediate checkpoints, and evaluation logs. IFM self-disclosed a benchmark contamination event and adjusted scores before publication.

Can I run K2 Horizon on a single GPU?

The smaller models are designed for constrained hardware. The 0.9B fits devices with very limited resources; the 3.7B and 7B target laptops and phones; the 32B and 36B-A4B are designed for on-premise servers. The 375B-A23B flagship requires a multi-GPU setup and is also available via API through Compass, Cerebras, and Nebius.

How does K2 Horizon perform on coding benchmarks?

The 375B-A23B flagship scores 70.2% on Terminal-Bench 2.1 and 42.6% on SWE-bench Pro per IFM's self-reported results. This trails GLM-5.2 (81.0 and 62.1 respectively) on coding-specific evals, but K2 covers device-scale sizes GLM does not and provides more auditable benchmarking methodology.

Is Apache 2.0 safe for commercial use?

Apache 2.0 permits commercial use, modification, and distribution with attribution. It is the most permissive common license for AI models and does not require derivative works to be open-sourced. Verify your specific deployment against any dataset license terms that may apply to fine-tuned versions.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle