{"article":{"slug":"introducing-typear-type-safe-decoding-for-autoregressive-llms","title":"Introducing TypeAR: Type-Safe Decoding for Autoregressive LLMs","subtitle":null,"summary":"Type-Safe Decoding for Autoregressive LLMs. Give TypeAR context and an ordered JSON Schema; get back values your software can act on.","content_type":"announcement","language":"en","canonical_url":"https://typear.ai/blog/introducing-typear","author":{"name":"Mingtian Zhang","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"TypeAR","url":"https://typear.ai/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[{"slug":"typesafe-ai","name":"TypeSafe AI","listing_type":"company","url":"https://listedstartups.com/companies/typesafe-ai"},{"slug":"jev","name":"Jev","listing_type":"product","url":"https://listedstartups.com/products/jev"}],"cover_image_url":null,"license":"all-rights-reserved","word_count":1226,"reading_minutes":5,"published_at":"2026-09-17T00:00:00.000Z","added_at":"2026-09-18T06:16:39.078Z","updated_at":"2026-09-18T06:16:39.078Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/introducing-typear-type-safe-decoding-for-autoregressive-llms","markdown_url":"https://listedarticles.com/articles/introducing-typear-type-safe-decoding-for-autoregressive-llms.md","example":false,"citation":"Mingtian Zhang, TypeAR. \"Introducing TypeAR: Type-Safe Decoding for Autoregressive LLMs.\" 17 Sept 2026. https://typear.ai/blog/introducing-typear (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://typear.ai/blog/introducing-typear"},"body_markdown":"[TypeSafe AI's Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)\nhighlighted a useful idea: software needs decisions, not more strings to parse.\nJev answers it with a purpose-built model. TypeAR asks the complementary\nquestion — whether the same typed-decision interface can come from any open\nautoregressive model you already run.\n\nSo that is what it is: you hand TypeAR context and an ordered JSON Schema, and it hands back values your software can act on — enums, booleans, scores — with no proprietary model API, no retraining, and no parser between the model and your code.\n\nThe guarantee it makes is narrow and absolute: **a returned value is always\none of the values your schema declared.** Not usually. Not after a retry. There\nis no validator in the hot path, because there is nothing for a validator to\nreject.\n\n## How a type-safe answer is generated\n\nStart with one field of the schema:\n\nTypeAR renders that field into the prompt and stops mid-literal, exactly where the answer belongs:\n\nThe model completes one token there, and two things make that enough.\n\n**Map every value to one token.** `meal`, `travel` and `equipment` tokenize to\ndifferent lengths, so they cannot be compared at one position. Each gets a\nsingle-token **control label** — `A`, `B`, `C` — written on the `choices` line.\nLabels are checked against the server's tokenizer: one token in, the same text\nback out, or the label is skipped.\n\n**Only those labels are selected.** TypeAR reads the log probability of each\ncandidate token and nothing else, so an out-of-schema token can never be chosen:\nit was never a candidate. Renormalizing over the candidates gives a\ndistribution on your declared values — `meal` 0.04, `travel` 0.93, `equipment`\n0.03 — which sampling draws from and you can threshold on. Argmax does not need\nit: restricting the candidates already fixes the ranking.\n\nAll of this needs the model's per-token probabilities, which closed APIs do not\nexpose. TypeAR gets them from an open-source model served with\n[SGLang](https://github.com/sgl-project/sglang).\n\nOne forward pass covers a decision however many values it allows, so a boolean\nand an eleven-level score cost the same. Temperature acts on the candidate\nscores, never on free generation: turning it up makes `meal` likelier than\n`travel`, never something the schema never declared.\n\n## Two execution modes\n\nEverything above describes one decision. A schema is several of them, and there are two ways to run the set. They differ in exactly one thing: what each decision is conditioned on.\n\n## Sequential: each decision sees the last\n\nThe selected label goes back into the same string along with the value it stands\nfor, so the next decision reads `value=\"travel\"` instead of having to remember\nwhat `B` meant. Then the next question is appended: field order is decision\norder. The second request is therefore the first request, character for\ncharacter, plus a tail:\n\nEverything above `value=\"travel\"` was already prefilled for the first decision\nand is byte-identical now, so the server's prefix cache matches it and reuses\nthat KV state. The client holds one growing string and never touches a KV\ntensor; only the new tail is computed.\n\nThe server reports how many tokens it reused. With `D` decisions, a context of\n`C` tokens, and roughly `S` new tokens per decision:\n\n## Batch: independent fields in one request\n\nNot every schema needs that. Sixteen independent yes/no flags about one document are answerable from the context alone — nothing in flag 12 depends on flag 11 — and running them in sequence makes each wait for a round trip it never needed.\n\nSo TypeAR forks instead of extending. The shared context is prefilled once and\neach field becomes a branch: the same prefix plus its own record, ending at the\nsame `answer=\"` cut. All branches go out in one call and come back as a single\nbatched decode step, one token per field.\n\nBecause every branch starts with identical text, the prefix cache serves all of them. In one local run — a ~1,100-token shared context, sixteen boolean fields, one GPU, sixteen concurrent requests allowed — every branch reused 1,088 cached tokens:\n\n| Execution | End-to-end | Per decision | Relative | \n|---|---|---|---|\n| Sequential | 9.35 s | 0.584 s | 1.0x | \n| Batch | 1.61 s | 0.101 s | 5.8x | \n\nOne measurement on one machine, not a portable benchmark; it moves with `K`, the\nmodel, the context length and the server. The shape is the durable part: prefill\nof about `C + sum(Q_k)` for `K` fields, then one batched decode.\n\nThe cost is the dependency — each branch sees only the context and its own question — so use batch when the fields genuinely are independent. Everything else holds: each branch still stops mid-literal, still scores only single-token labels, still renormalizes over its own declared domain.\n\n## Sequential or batch?\n\nThe two modes are not a speed knob. They compute different conditionals, so the\nquestion is never \"which is faster\" but \"does this field's answer depend on\nanother field's answer\". A severity that should follow from the system already\nidentified, a rollback that should follow from both — those belong in sequence,\nand the cost is `D` round trips. Sixteen flags that each read the same document\nindependently belong in one batch, and the cost is that none of them can see the\nothers.\n\nGetting it wrong in one direction is slow; in the other it is wrong. A field batched away from a dependency it actually has will answer as though that dependency did not exist, and nothing in the output will say so. Nothing stops you from splitting a workflow either — run the dependent fields in sequence, then batch the independent ones over a context that already carries the first results.\n\n## Beyond finite domains\n\nEvery decision so far has come from a finite set, because that is what single- token labels can cover. The schema layer enforces it: an enum, a boolean, or a score field that expands to eleven levels between 0.0 and 1.0. A typed field without an enum is rejected outright rather than guessed at.\n\nNothing in the mechanism requires that, though, and this is where a decoding method has room that a purpose-built model does not. Declare one branch of the field open:\n\nThe typed decision is unchanged — four single-token candidates now instead of\nthree. Pick `A`, `B` or `C` and the value is written back as before. Pick `D`\nand the record opens a string rather than closing one, and generation continues\nunconstrained until the closing quote:\n\nThe guarantee survives because the escape is itself a declared value: free text\nis reachable only through a token you put in the domain, and the closing quote\nbounds what comes out of it. Nor does what comes out have to stay a one-off —\nthe domain is just a list written into the prompt, so `conference registration`\ncan be bound to its own label and offered as a declared choice in the next\ndecision. The vocabulary grows from what the model produced, and every later\ndecision is still one token over a finite set.\n\n## Conclusion\n\nTypeAR shows that a typed-decision interface does not require a purpose-built model. Restricting the candidates at one position, renormalizing over them, and appending the result to the prompt is enough to get values that are typed by construction — one output token per decision, one read of the context per workflow, from an open model you already serve. No proprietary API and no retraining: it reads a distribution the model already produces.\n\nMore benchmarks and worked examples are coming.","body_html":"<p><a href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev\" rel=\"nofollow ugc noopener\">TypeSafe AI&#39;s Jev</a>\nhighlighted a useful idea: software needs decisions, not more strings to parse.\nJev answers it with a purpose-built model. TypeAR asks the complementary\nquestion — whether the same typed-decision interface can come from any open\nautoregressive model you already run.</p>\n<p>So that is what it is: you hand TypeAR context and an ordered JSON Schema, and it hands back values your software can act on — enums, booleans, scores — with no proprietary model API, no retraining, and no parser between the model and your code.</p>\n<p>The guarantee it makes is narrow and absolute: <strong>a returned value is always\none of the values your schema declared.</strong> Not usually. Not after a retry. There\nis no validator in the hot path, because there is nothing for a validator to\nreject.</p>\n<h2 id=\"how-a-type-safe-answer-is-generated\">How a type-safe answer is generated</h2>\n<p>Start with one field of the schema:</p>\n<p>TypeAR renders that field into the prompt and stops mid-literal, exactly where the answer belongs:</p>\n<p>The model completes one token there, and two things make that enough.</p>\n<p><strong>Map every value to one token.</strong> <code>meal</code>, <code>travel</code> and <code>equipment</code> tokenize to\ndifferent lengths, so they cannot be compared at one position. Each gets a\nsingle-token <strong>control label</strong> — <code>A</code>, <code>B</code>, <code>C</code> — written on the <code>choices</code> line.\nLabels are checked against the server&#39;s tokenizer: one token in, the same text\nback out, or the label is skipped.</p>\n<p><strong>Only those labels are selected.</strong> TypeAR reads the log probability of each\ncandidate token and nothing else, so an out-of-schema token can never be chosen:\nit was never a candidate. Renormalizing over the candidates gives a\ndistribution on your declared values — <code>meal</code> 0.04, <code>travel</code> 0.93, <code>equipment</code>\n0.03 — which sampling draws from and you can threshold on. Argmax does not need\nit: restricting the candidates already fixes the ranking.</p>\n<p>All of this needs the model&#39;s per-token probabilities, which closed APIs do not\nexpose. TypeAR gets them from an open-source model served with\n<a href=\"https://github.com/sgl-project/sglang\" rel=\"nofollow ugc noopener\">SGLang</a>.</p>\n<p>One forward pass covers a decision however many values it allows, so a boolean\nand an eleven-level score cost the same. Temperature acts on the candidate\nscores, never on free generation: turning it up makes <code>meal</code> likelier than\n<code>travel</code>, never something the schema never declared.</p>\n<h2 id=\"two-execution-modes\">Two execution modes</h2>\n<p>Everything above describes one decision. A schema is several of them, and there are two ways to run the set. They differ in exactly one thing: what each decision is conditioned on.</p>\n<h2 id=\"sequential-each-decision-sees-the-last\">Sequential: each decision sees the last</h2>\n<p>The selected label goes back into the same string along with the value it stands\nfor, so the next decision reads <code>value=&quot;travel&quot;</code> instead of having to remember\nwhat <code>B</code> meant. Then the next question is appended: field order is decision\norder. The second request is therefore the first request, character for\ncharacter, plus a tail:</p>\n<p>Everything above <code>value=&quot;travel&quot;</code> was already prefilled for the first decision\nand is byte-identical now, so the server&#39;s prefix cache matches it and reuses\nthat KV state. The client holds one growing string and never touches a KV\ntensor; only the new tail is computed.</p>\n<p>The server reports how many tokens it reused. With <code>D</code> decisions, a context of\n<code>C</code> tokens, and roughly <code>S</code> new tokens per decision:</p>\n<h2 id=\"batch-independent-fields-in-one-request\">Batch: independent fields in one request</h2>\n<p>Not every schema needs that. Sixteen independent yes/no flags about one document are answerable from the context alone — nothing in flag 12 depends on flag 11 — and running them in sequence makes each wait for a round trip it never needed.</p>\n<p>So TypeAR forks instead of extending. The shared context is prefilled once and\neach field becomes a branch: the same prefix plus its own record, ending at the\nsame <code>answer=&quot;</code> cut. All branches go out in one call and come back as a single\nbatched decode step, one token per field.</p>\n<p>Because every branch starts with identical text, the prefix cache serves all of them. In one local run — a ~1,100-token shared context, sixteen boolean fields, one GPU, sixteen concurrent requests allowed — every branch reused 1,088 cached tokens:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Execution</th><th>End-to-end</th><th>Per decision</th><th>Relative</th></tr></thead><tbody><tr><td>Sequential</td><td>9.35 s</td><td>0.584 s</td><td>1.0x</td></tr><tr><td>Batch</td><td>1.61 s</td><td>0.101 s</td><td>5.8x</td></tr></tbody></table></div>\n<p>One measurement on one machine, not a portable benchmark; it moves with <code>K</code>, the\nmodel, the context length and the server. The shape is the durable part: prefill\nof about <code>C + sum(Q_k)</code> for <code>K</code> fields, then one batched decode.</p>\n<p>The cost is the dependency — each branch sees only the context and its own question — so use batch when the fields genuinely are independent. Everything else holds: each branch still stops mid-literal, still scores only single-token labels, still renormalizes over its own declared domain.</p>\n<h2 id=\"sequential-or-batch\">Sequential or batch?</h2>\n<p>The two modes are not a speed knob. They compute different conditionals, so the\nquestion is never &quot;which is faster&quot; but &quot;does this field&#39;s answer depend on\nanother field&#39;s answer&quot;. A severity that should follow from the system already\nidentified, a rollback that should follow from both — those belong in sequence,\nand the cost is <code>D</code> round trips. Sixteen flags that each read the same document\nindependently belong in one batch, and the cost is that none of them can see the\nothers.</p>\n<p>Getting it wrong in one direction is slow; in the other it is wrong. A field batched away from a dependency it actually has will answer as though that dependency did not exist, and nothing in the output will say so. Nothing stops you from splitting a workflow either — run the dependent fields in sequence, then batch the independent ones over a context that already carries the first results.</p>\n<h2 id=\"beyond-finite-domains\">Beyond finite domains</h2>\n<p>Every decision so far has come from a finite set, because that is what single- token labels can cover. The schema layer enforces it: an enum, a boolean, or a score field that expands to eleven levels between 0.0 and 1.0. A typed field without an enum is rejected outright rather than guessed at.</p>\n<p>Nothing in the mechanism requires that, though, and this is where a decoding method has room that a purpose-built model does not. Declare one branch of the field open:</p>\n<p>The typed decision is unchanged — four single-token candidates now instead of\nthree. Pick <code>A</code>, <code>B</code> or <code>C</code> and the value is written back as before. Pick <code>D</code>\nand the record opens a string rather than closing one, and generation continues\nunconstrained until the closing quote:</p>\n<p>The guarantee survives because the escape is itself a declared value: free text\nis reachable only through a token you put in the domain, and the closing quote\nbounds what comes out of it. Nor does what comes out have to stay a one-off —\nthe domain is just a list written into the prompt, so <code>conference registration</code>\ncan be bound to its own label and offered as a declared choice in the next\ndecision. The vocabulary grows from what the model produced, and every later\ndecision is still one token over a finite set.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>TypeAR shows that a typed-decision interface does not require a purpose-built model. Restricting the candidates at one position, renormalizing over them, and appending the result to the prompt is enough to get values that are typed by construction — one output token per decision, one read of the context per workflow, from an open model you already serve. No proprietary API and no retraining: it reads a distribution the model already produces.</p>\n<p>More benchmarks and worked examples are coming.</p>","headings":[{"level":2,"text":"How a type-safe answer is generated","id":"how-a-type-safe-answer-is-generated"},{"level":2,"text":"Two execution modes","id":"two-execution-modes"},{"level":2,"text":"Sequential: each decision sees the last","id":"sequential-each-decision-sees-the-last"},{"level":2,"text":"Batch: independent fields in one request","id":"batch-independent-fields-in-one-request"},{"level":2,"text":"Sequential or batch?","id":"sequential-or-batch"},{"level":2,"text":"Beyond finite domains","id":"beyond-finite-domains"},{"level":2,"text":"Conclusion","id":"conclusion"}]}}