TypeSafe AI's Jev highlighted a useful idea: software needs decisions, not more strings to parse. Jev answers it with a purpose-built model. TypeAR asks the complementary question — whether the same typed-decision interface can come from any open autoregressive model you already run.
So that is what it is: you hand TypeAR context and an ordered JSON Schema, and it hands back values your software can act on — enums, booleans, scores — with no proprietary model API, no retraining, and no parser between the model and your code.
The guarantee it makes is narrow and absolute: a returned value is always one of the values your schema declared. Not usually. Not after a retry. There is no validator in the hot path, because there is nothing for a validator to reject.
How a type-safe answer is generated
Start with one field of the schema:
TypeAR renders that field into the prompt and stops mid-literal, exactly where the answer belongs:
The model completes one token there, and two things make that enough.
Map every value to one token. meal, travel and equipment tokenize to
different lengths, so they cannot be compared at one position. Each gets a
single-token control label — A, B, C — written on the choices line.
Labels are checked against the server's tokenizer: one token in, the same text
back out, or the label is skipped.
Only those labels are selected. TypeAR reads the log probability of each
candidate token and nothing else, so an out-of-schema token can never be chosen:
it was never a candidate. Renormalizing over the candidates gives a
distribution on your declared values — meal 0.04, travel 0.93, equipment
0.03 — which sampling draws from and you can threshold on. Argmax does not need
it: restricting the candidates already fixes the ranking.
All of this needs the model's per-token probabilities, which closed APIs do not expose. TypeAR gets them from an open-source model served with SGLang.
One forward pass covers a decision however many values it allows, so a boolean
and an eleven-level score cost the same. Temperature acts on the candidate
scores, never on free generation: turning it up makes meal likelier than
travel, never something the schema never declared.
Two execution modes
Everything above describes one decision. A schema is several of them, and there are two ways to run the set. They differ in exactly one thing: what each decision is conditioned on.
Sequential: each decision sees the last
The selected label goes back into the same string along with the value it stands
for, so the next decision reads value="travel" instead of having to remember
what B meant. Then the next question is appended: field order is decision
order. The second request is therefore the first request, character for
character, plus a tail:
Everything above value="travel" was already prefilled for the first decision
and is byte-identical now, so the server's prefix cache matches it and reuses
that KV state. The client holds one growing string and never touches a KV
tensor; only the new tail is computed.
The server reports how many tokens it reused. With D decisions, a context of
C tokens, and roughly S new tokens per decision:
Batch: independent fields in one request
Not every schema needs that. Sixteen independent yes/no flags about one document are answerable from the context alone — nothing in flag 12 depends on flag 11 — and running them in sequence makes each wait for a round trip it never needed.
So TypeAR forks instead of extending. The shared context is prefilled once and
each field becomes a branch: the same prefix plus its own record, ending at the
same answer=" cut. All branches go out in one call and come back as a single
batched decode step, one token per field.
Because every branch starts with identical text, the prefix cache serves all of them. In one local run — a ~1,100-token shared context, sixteen boolean fields, one GPU, sixteen concurrent requests allowed — every branch reused 1,088 cached tokens:
| Execution | End-to-end | Per decision | Relative |
|---|---|---|---|
| Sequential | 9.35 s | 0.584 s | 1.0x |
| Batch | 1.61 s | 0.101 s | 5.8x |
One measurement on one machine, not a portable benchmark; it moves with K, the
model, the context length and the server. The shape is the durable part: prefill
of about C + sum(Q_k) for K fields, then one batched decode.
The cost is the dependency — each branch sees only the context and its own question — so use batch when the fields genuinely are independent. Everything else holds: each branch still stops mid-literal, still scores only single-token labels, still renormalizes over its own declared domain.
Sequential or batch?
The two modes are not a speed knob. They compute different conditionals, so the
question is never "which is faster" but "does this field's answer depend on
another field's answer". A severity that should follow from the system already
identified, a rollback that should follow from both — those belong in sequence,
and the cost is D round trips. Sixteen flags that each read the same document
independently belong in one batch, and the cost is that none of them can see the
others.
Getting it wrong in one direction is slow; in the other it is wrong. A field batched away from a dependency it actually has will answer as though that dependency did not exist, and nothing in the output will say so. Nothing stops you from splitting a workflow either — run the dependent fields in sequence, then batch the independent ones over a context that already carries the first results.
Beyond finite domains
Every decision so far has come from a finite set, because that is what single- token labels can cover. The schema layer enforces it: an enum, a boolean, or a score field that expands to eleven levels between 0.0 and 1.0. A typed field without an enum is rejected outright rather than guessed at.
Nothing in the mechanism requires that, though, and this is where a decoding method has room that a purpose-built model does not. Declare one branch of the field open:
The typed decision is unchanged — four single-token candidates now instead of
three. Pick A, B or C and the value is written back as before. Pick D
and the record opens a string rather than closing one, and generation continues
unconstrained until the closing quote:
The guarantee survives because the escape is itself a declared value: free text
is reachable only through a token you put in the domain, and the closing quote
bounds what comes out of it. Nor does what comes out have to stay a one-off —
the domain is just a list written into the prompt, so conference registration
can be bound to its own label and offered as a declared choice in the next
decision. The vocabulary grows from what the model produced, and every later
decision is still one token over a finite set.
Conclusion
TypeAR shows that a typed-decision interface does not require a purpose-built model. Restricting the candidates at one position, renormalizing over them, and appending the result to the prompt is enough to get values that are typed by construction — one output token per decision, one read of the context per workflow, from an open model you already serve. No proprietary API and no retraining: it reads a distribution the model already produces.
More benchmarks and worked examples are coming.