On July 15, 2026, Thinking Machines Lab - the roughly 200-person startup that ex-OpenAI CTO Mira Murati kept in near-total stealth for over a year - dropped Inkling, a 975-billion-parameter open-weight model, with the full weights on Hugging Face the same afternoon. The timing was deliberate; so was what the announcement led with: not a leaderboard position, but a fine-tuning platform.
Thinking Machines is upfront about the ceiling: Inkling is, in its own words, "not the strongest overall model available today, open or closed." So why care? Because the lab didn't build it to top a leaderboard. That honesty makes the model more interesting to evaluate, not less. Here is what actually matters if your team is considering building on it.
What is the Inkling open-weight model, exactly?
Thinking Machines Lab released Inkling as a 975B-parameter open-weight Mixture-of-Experts model. It activates 41B parameters per token and supports a context window of up to 1,048,576 tokens.
The weights ship under the Apache 2.0 license. That combination - frontier scale, permissive license, full weights downloadable - is still rare.
The architecture has several details seen less often in open-weight releases, including short convolutions inside every decoder block and a learned relative-position bias in place of RoPE.
Thinking Machines reports 45 trillion pretraining tokens spanning text, images, audio, and video, with Muon used for the model's large matrix parameters and Adam for the remaining ones. Whether those architectural choices pay off in practice depends entirely on your workload - the short-convolution approach has proponents in long-context retrieval, but it is not yet proven at this scale against independent test sets.
On benchmarks, the comparison with GLM-5.2 illustrates Inkling's overall profile: it scores 79.8 on IFBench versus 73.3 for GLM-5.2, and 43.9 versus 38.1 on SimpleQA Verified.
On Terminal Bench 2.1, Inkling sits mid-pack among open models, trailing GLM-5.2 and Kimi K2.6 by a wide margin at max effort. The instruction-following numbers are genuinely strong; the agentic coding numbers are not.
The effort-control feature and why it changes the cost math
This is the part most coverage glosses over.
Inkling lets you dial reasoning up or down - trading accuracy for speed and cost - and the setting can be adjusted from inside a coding or agent harness. In Hugging Face transformers, it is exposed as a reasoning_effort argument with named levels.
The key detail: this is not a system-prompt toggle. The team collected more than 30 million RL rollouts, and one practical outcome is effort control. A system message tells Inkling how much effort to spend, while a per-token cost during training encourages shorter or longer answers. The behavior is trained in, which means it is more stable than post-hoc prompting and can be used reliably across agent steps.
At lower effort settings, Inkling matches Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. For a team running a high-volume agentic loop - think a triage agent processing hundreds of tickets a shift - that ratio matters more than peak benchmark position. Most steps in an agent loop do not need max reasoning. Routing a simple classification to effort 0.3, then escalating ambiguous cases to 0.9, gives you a cost profile a fixed-effort model cannot match.
Inkling-Small and the case for the smaller sibling
Alongside the flagship, Thinking Machines previewed Inkling-Small - 276B total / 12B active - which reportedly matches the full model on some reasoning benchmarks (Humanity's Last Exam text: 29.6% vs 29.7%) while being cheaper and faster to run.
Inkling-Small matches or beats the full model on several benchmarks, such as 88.3% vs 87.2% on GPQA Diamond, but trails on factual recall (20.9% vs 43.9% on SimpleQA Verified) and hard agentic coding.
That gap tells you where to use which variant:
| Task class | Full Inkling | Inkling-Small |
|---|---|---|
| Hard agentic coding | ✓ | Falls short |
| Reasoning / HLE-style | Comparable | Nearly matches |
| GPQA Diamond | 87.2% | 88.3% |
| Factual recall (SimpleQA) | 43.9% | 20.9% |
| VRAM to self-host | ≥2TB (BF16) | Dramatically less |
If your use case lives in reasoning or multimodal and not deep factual retrieval, Inkling-Small may be the right call - at a fraction of the infrastructure cost. Most teams evaluating "Inkling" should run both against their own task sample before committing.
The three things the benchmarks do not show
1. Self-hosting Inkling is a real infrastructure project.
The model card states the BF16 checkpoint requires at least 2TB of aggregate VRAM, while the NVFP4 checkpoint requires at least 600GB. That is eight H100 80GB cards at minimum for the compressed variant - before you account for KV-cache headroom, tensor parallelism overhead, and spare capacity for rolling restarts. Teams considering self-hosting should model total deployment cost, networking, tensor parallelism, KV-cache headroom, monitoring, and spare capacity. Apache 2.0 does not make the GPUs cheaper.
2. The "Western alternative to Chinese open-weights" framing deserves scrutiny.
One honest caveat from the team: some early post-training data was bootstrapped from other open models - including Kimi K2.5 - with a promise that the next model will use fully self-contained post-training. Kimi K2.5 is Moonshot AI's model, a Beijing-based lab. Teams choosing Inkling for supply-chain or geopolitical reasons should weigh this - the provenance of the post-training data is not the same as the provenance of the weights.
3. Tinker's context window is not the model's context window.
Thinking Machines says the full model supports up to a 1 million-token context window and distributes downloadable checkpoints through Hugging Face. That does not mean every hosted route exposes 1 million tokens: the managed Tinker options are capped at 64K or 256K. If your pipeline depends on long-context retrieval, self-hosting (or a third-party inference provider) is the only way to access it. A tool like Beagle routing long Slack thread histories to a summarization job would hit Tinker's 64K ceiling quickly on a busy channel.
On pricing: independent aggregators list API pricing at roughly $1.87 per 1M input tokens and $4.68 per 1M output tokens, higher than some rivals such as GLM-5.1 at $1.40 input and $4.40 output.
The launch discount has no published end date, so it should not be treated as a permanent project-cost baseline. Build your cost model on the undiscounted rate.
The business model underneath all of this is worth understanding. Because open weights cannot be metered, the company does not sell model access; it sells the training and customization layer (Tinker) plus a cut of the hosting ecosystem. It is a different revenue theory than OpenAI's or Anthropic's, and Inkling is the proof-of-concept. That alignment of incentives - Thinking Machines profits when you fine-tune, not when you query - is structurally different from a closed API vendor, and it shows in how the model is documented.
Inkling open-weight model: common questions
What license does the Inkling model use?
The weights are available under the Apache 2.0 license. This is one of the most permissive options in the open-weight space - it allows commercial use, modification, and redistribution without the user-count restrictions found in Meta's Llama license or Google's Gemma license. Verify the specific model variant on Hugging Face before you build.
How does controllable effort work in Inkling?
Effort control is a trained behavior, not a prompt trick. A system message tells Inkling how much effort to spend, while a per-token cost during RL training encourages shorter or longer answers depending on the setting.
In Hugging Face transformers, it is exposed as a reasoning_effort argument with named levels.
Setting it dynamically per agent step - low for routing, high for generation - is where the cost savings actually appear.
Can a small team self-host Inkling?
Realistically, not the full model. Inkling is a frontier-scale Mixture-of-Experts model, so running it yourself is a real infrastructure decision - not something a single laptop or gaming GPU can do. The NVFP4 checkpoint needs at least 600GB of aggregate VRAM. Most small teams will get further faster by fine-tuning on Tinker or using a third-party inference provider and reserving self-hosting for a later stage when the use case is proven.
How does Inkling-Small compare to the full model?
Inkling-Small nearly matches Inkling across most multimodal evaluations at a lower cost. The main gaps are factual recall - SimpleQA Verified drops from 43.9% to 20.9% - and hard agentic coding tasks. For reasoning, instruction-following, and multimodal work, the delta is small enough that most teams should evaluate Small first and upgrade only if benchmarks on their own data justify it.
What is the Tinker fine-tuning platform?
Inkling is available for fine-tuning on Tinker with 64K and 256K context options, and fine-tuned checkpoints can be deployed via TogetherAI, Fireworks, Modal, Databricks, and Baseten. Tinker is Thinking Machines' primary revenue surface - the model weights are free, but managed fine-tuning and sampling are usage-priced. Checkpoint storage is listed separately at $0.10 per GB per month.