Learn / Prompting that works / Structured output

Lesson 3 of 5 7 min

Structured output

Getting JSON you can actually parse: schema-constrained decoding vs prompting alone, and what to do when a provider doesn't support the former.

Two ways to get structured output

Prompting alone: instruct the model to return JSON in a specific shape, maybe show an example. This is a request, followed probabilistically like everything else the model generates - it can still wrap the JSON in explanation text, use inconsistent field names across runs, or emit something that’s almost-but-not-quite valid JSON (a trailing comma, an unescaped quote).

Schema-constrained decoding (“structured outputs” / “JSON mode” with a schema, offered by OpenAI, Groq, and others in an OpenAI-compatible way): you pass an actual JSON Schema alongside the request, and the provider constrains the token generation process itself so the output cannot deviate from that shape. This is a guarantee at the mechanism level, not an instruction the model is merely trying to follow.

# Illustrative - the exact parameter name varies by provider/SDK version.
response = client.chat.completions.create(
    model="...",
    messages=[...],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "extraction",
            "schema": {
                "type": "object",
                "properties": {
                    "sentiment": {"type": "string", "enum": ["positive", "negative", "mixed"]},
                    "confidence": {"type": "number"},
                },
                "required": ["sentiment", "confidence"],
            },
        },
    },
)

Prefer schema-constrained decoding whenever your provider supports it and downstream code parses the result - it removes a whole class of parsing bugs for free.

What the guarantee does NOT cover

Schema constraints guarantee shape - valid JSON, right field names, right types. They do not guarantee the values are factually correct, or that confidence: 0.97 means what you think it means. A perfectly schema-valid response can still contain a wrong sentiment label. Schema constraints are a parsing-reliability tool, not an accuracy tool - evaluation (see the RAG track’s lesson 6 for the general pattern) still matters.

Without schema support: parse defensively

Some models/providers don’t offer schema-constrained decoding, or you’re using a mode that doesn’t support it (e.g. streaming free text). Defend the parsing boundary instead of trusting raw output:

  1. Instruct clearly and show a one-shot example of the exact JSON shape (lesson 2).
  2. Strip common wrapping artifacts before parsing - a leading/trailing markdown code fence (```json ... ```) is the most common one.
  3. Parse with a tolerant approach where possible, but ultimately validate against your schema (e.g. with Pydantic in Python, or Zod in TypeScript - the same library this site’s own content collections use for their frontmatter schemas) before trusting any field.
  4. Have an explicit failure path: retry once with an error message fed back to the model (“your last response wasn’t valid JSON: - try again”), or surface a clear error rather than silently passing through malformed data.

Keep schemas simple

A flat schema with a handful of well-typed fields is far more reliable than a deeply nested one with many optional branches - both for constrained decoding (smaller search space per token) and for a model working from instructions alone (less to get subtly wrong). Split a complex extraction into multiple smaller calls before reaching for one giant schema.

Key takeaways

  • Asking a model to 'return JSON' in plain instructions is not reliable on its own - it can add prose before/after the JSON, use the wrong field names, or produce invalid JSON entirely.
  • Schema-constrained decoding (structured outputs / JSON mode with a schema) makes the provider guarantee valid JSON matching your schema at the token-generation level, not just via instruction-following.
  • When a provider doesn't support schema constraints, defensively parse: strip markdown code fences, use a tolerant JSON parser, and validate against your schema before trusting the result.
  • Keep the schema itself simple and flat where possible - deeply nested or very large schemas increase the chance of a subtly malformed field.

Quick check

3 questions - see how much stuck.

1. Why is 'please return valid JSON' in the instructions alone not fully reliable?
2. What does schema-constrained decoding (e.g. OpenAI/Groq's 'structured outputs', or a JSON schema passed to the API) actually guarantee?
3. If a provider doesn't support schema-constrained output, what's the safest fallback?