GPT-4-0613 fails to match a given JSON schema on more than 60% of attempts in OpenAI's own evals. Its replacement, gpt-4o-2024-08-06, hits 100%. The model didn't become dramatically smarter at formatting between those two releases - the mechanism changed. Structured Outputs replaced "ask nicely and hope" with something that operates at the math layer of token generation. Understanding that layer tells you exactly where the guarantee holds, where it bends, and where it silently breaks.
What constrained decoding actually does to a model's output
An LLM produces output one token at a time, calculating a probability distribution over its entire token vocabulary at each step. Without any constraint, it samples freely from that distribution. Constrained decoding intercepts exactly that moment.
Constrained decoding is the inference-time technique that forces an LLM's output to conform to a grammar, schema, or regex by masking the next-token distribution at every step. Tokens that would make the partial output invalid are set to logit negative infinity before sampling, leaving only legal continuations.
Concretely: the decoder maintains a state machine - a regex compiled to a deterministic finite automaton, a grammar compiled to a pushdown automaton, or a JSON-schema walker. At each step, the state machine returns the set of token IDs that can legally follow the current state, and the decoder masks all other tokens to logit -∞ before sampling.
Because the constraint is enforced during generation rather than after, the output is guaranteed valid - you never get a parse error, never have to retry, never need a fallback parser.
This is categorically different from JSON mode. JSON mode guarantees syntactically valid JSON but places no constraints on the schema. Structured Outputs enforces a specific JSON Schema via constrained decoding, guaranteeing every field, type, and constraint is satisfied. The difference matters in production: JSON mode has a schema-adherence failure rate of 5-10%, while Structured Outputs has a failure rate of less than 0.1% - two orders of magnitude difference.
| Mechanism | Valid JSON? | Matches your schema? | Overhead |
|---|---|---|---|
| Prompt engineering | Usually | Sometimes | 0ms |
JSON mode (json_object) |
Yes | No guarantee | ~50ms |
Structured Outputs (json_schema, strict) |
Yes | Yes | ~100ms |
| Local constrained decoding (XGrammar) | Yes | Yes | Near-zero after compile |
The token-character mismatch: the part nobody tells you
Here is the non-obvious problem at the center of all constrained decoding implementations.
LLMs do not generate characters - they generate tokens, variable-length byte sequences produced by BPE tokenization. A single token might be "json" (4 characters) or " the" (4 characters including the space). Grammar constraints operate at the character level, but the model generates multi-character tokens, creating a fundamental mismatch.
Say your schema requires a field named "severity". The state machine tracking your partial JSON knows the next legal characters are "s". But the model's vocabulary contains a token for the full string "severity" - it might also contain "sev", "sev"rity, or dozens of other partial overlaps. The mask has to handle all of them correctly, at every step, across a vocabulary of 32K to 128K tokens.
Implementing constrained decoding from scratch requires checking every single token in the LLM's vocabulary against your schema for every output token. This runs on CPUs, while LLMs run on GPUs. If the CPU lags, the GPU sits idle waiting for the token mask.
When constrained decoding forces the model down an unusual token path, it may produce a non-canonical tokenization the model rarely saw during training, subtly degrading output quality. Token healing, pioneered by Microsoft's Guidance library, addresses this by backing up one token at the prompt boundary and constraining the first generated token to begin with the removed token's prefix.
How the engines evolved, and why it matters for your stack
The overhead story for constrained decoding has changed dramatically - and most teams are still making decisions based on the old numbers.
Early Outlines implementations added 50-200% latency overhead, and the compilation cost alone was a show-stopper for schemas that varied per request. That history is why many teams wrote off the approach.
XGrammar is currently the state-of-the-art implementation of constrained decoding, utilizing system optimizations to reduce runtime check via context-independent caching, enabling co-optimizations for end-to-end LLM inference speedup in structured generation settings.
With older backends like Outlines, constrained decoding adds 5-60% latency overhead depending on schema complexity (simple flat schemas ~5-10%, deeply nested schemas 40-60%). Modern backends like XGrammar shift most of that cost to a one-time grammar compilation step of 20-50ms and reduce per-token overhead to near-zero. Grammar caching in SGLang eliminates the compilation cost on repeated calls with the same schema.
JSON Schema features like minItems, maxItems, enum, and Array, while supported in Outlines, often take 40 seconds to 10 minutes to process.
That's not a theoretical concern - it's a production timeout waiting to happen on complex recursive schemas.
Frameworks demonstrate significant differences in their actual support for real-world JSON schemas, with the best framework supporting twice as many schemas as the worst. If you have deeply nested or recursive schemas - think Kubernetes configs, OpenAPI specs, Snowplow event schemas - test your engine against them before shipping.
Structured outputs guarantee syntax - not that the answer is right
This is the part that surprises most teams after they ship.
Across 2,400 calls to four open models in prompt-only and JSON-schema modes, schema-valid output can still have large semantic error rates. In the strongest model tested, both modes achieve 100% schema validity, yet semantic success remains near 80%. JSON Schema and provider-level structured-output modes remove a large class of parse failures, but they do not determine whether the object is a safe, faithful response.
The model will generate a perfectly valid {"severity": "low", "owner": "alice@example.com"} even when the right answer is "critical" and the right owner is someone else. The constraint enforces shape. It says nothing about truth.
Value errors - correct structure, wrong content - are the dominant gap: even models with more than 97% schema compliance show Value Accuracy scores of 0.693 to 0.830, meaning 17-31% of leaf values are incorrect despite valid structure.
This is the argument for keeping a human approval step on any structured output that feeds a downstream action. The guarantee you get from constrained decoding is real and useful - malformed JSON is a solved problem. But the guarantee stops at the closing brace. A teammate like Beagle that drafts and waits for a nod before posting exists precisely in this gap: the structure is machine-enforced, the meaning is human-checked.
How LLM structured outputs work: common questions
What is the difference between JSON mode and Structured Outputs?
JSON mode (response_format: {"type": "json_object"}) guarantees the model returns parseable JSON but imposes no constraints on its shape. Structured Outputs (json_schema with strict: true) enforces your exact schema via constrained decoding - required fields, types, enums - with a failure rate under 0.1% versus 5-10% for JSON mode. Use Structured Outputs for anything a downstream system parses.
Does constrained decoding slow down the model?
It used to. Early Outlines implementations added 50-200% latency overhead. Modern engines like XGrammar shift most cost to a one-time 20-50ms schema compilation step, with near-zero per-token overhead after that. On SGLang with grammar caching, repeated calls with the same schema pay the compilation cost once. The overhead objection is largely obsolete for common schemas.
Can constrained decoding guarantee the values are correct, not just the format?
No. Constrained decoding guarantees syntactic validity - the output matches your schema. It cannot guarantee the values are accurate. Benchmarks show models with 97%+ schema compliance still produce incorrect leaf values 17-31% of the time. Validating values requires either deterministic checks in your application layer or a human in the loop.
When does constrained decoding break down?
Complex schemas are the breaking point. JSONSchemaBench found coverage collapsing from roughly 86% on simple schemas to around 3% on complex ones when tested against nearly 10,000 real-world schemas. FSM-based engines struggle most with recursive structures; specific Outlines features like minItems and enum can take minutes to compile. Use a CFG-based engine and test your actual schemas before shipping.
Does Anthropic support structured outputs natively?
Anthropic's native structured outputs require the beta API with an explicit header and the schema passed as a structured_output parameter - not in the system prompt. Prompt-engineering a schema into the system message lacks the constrained decoding guarantee and is closer to JSON mode reliability.