---
title: "What Is Jev and How Does It Work?"
slug: what-is-jev-and-how-does-it-work
url: https://listedarticles.com/articles/what-is-jev-and-how-does-it-work
canonical_url: https://shreyshahh.substack.com/p/what-is-jev-and-how-does-it-work
content_type: blog_post
language: en
published_at: 2026-09-17T12:00:00.000Z
updated_at: 2026-09-27T06:22:21.518Z
author: "Shrey Shah"
author_url: https://shreyshahh.substack.com
authored_by: human
publisher: "Shrey Shah"
publisher_url: https://shreyshahh.substack.com
topics: ["AI", "LLMs", "AI Agents", "Developer Tools"]
about: ["https://listedstartups.com/products/jev", "https://listedstartups.com/companies/typesafe-ai"]
license: all-rights-reserved
word_count: 2548
reading_minutes: 11
citation: "Shrey Shah, Shrey Shah. \"What Is Jev and How Does It Work?.\" 17 Sept 2026. https://shreyshahh.substack.com/p/what-is-jev-and-how-does-it-work (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# What Is Jev and How Does It Work?

> Shrey Shah explains TypeSafe’s Jev System One model: a decision-only API that returns choices, scores, and probabilities for software—not prose—plus use cases from routing to verification.

Most AI models are built to answer people. Jev is built to answer software.

Give Jev a customer message and a list of teams, and it chooses a team with probabilities for every option. Give it a code change and a set of risk levels, and it scores the change. Ask whether an agent has completed its task, and it returns a probability between zero and one.

It will not write the customer reply, fix the code, or explain its reasoning. TypeSafe AI built Jev for the smaller judgments that software makes around those actions. The company reports response times between 70 and 500 milliseconds and prices input at $0.042 per million tokens, with no charge for output. Early tests support the basic speed and cost story, while the broader accuracy and calibration claims still need more evidence.

*Source: [TypeSafe AI's launch announcement](https://typesafe.ai/blog/introducing-system-one-models-and-jev).*

## What is Jev?

Jev is the first model in a category TypeSafe calls System One models. The name borrows the idea of fast, intuitive decisions in contrast with slow, deliberate reasoning.

A normal large language model accepts text and generates more text, one token at a time. Jev accepts two things:

1. **State:** The information needed to make a decision. This can be text, a JSON object, or a list.
2. **Questions:** The judgments you want the model to make about that state, with the allowed answer format defined in advance.

It returns typed answers and probability distributions that application code can use directly. Think of an airport dispatcher. A passenger explains the problem, and the dispatcher chooses the right desk. Jev plays a similar role, except your application supplies the possible destinations.

Use a generative model for an email, summary, explanation, or code. Jev fits when software needs a bounded judgment that controls what happens next.

## The three question types

The API exposes three primitives: Choice, Score, and Noul.

### Choice

`Choice` asks the model to select one option from a known set. It returns the selected option, the probability assigned to each option, and a confidence value.

For example, a support system could define four routes:

- Billing
- Technical support
- Sales
- Other

The model cannot invent a fifth department. It must distribute probability across the supplied choices.

Developers can use Choice for model routing, document classification, selecting the next tool, assigning a ticket, or choosing a workflow.

### Score

`Score` places the state along an ordered scale that you define.

A bug report might be scored against these levels:

- Cosmetic
- Inconvenient
- Disruptive
- Blocking

The response includes a score, a probability distribution across the levels, and confidence. This is useful when the answer is not a category but a position on a spectrum, such as urgency, risk, quality, relevance, or customer frustration.

### Noul

`Noul` answers a yes-or-no question by returning the probability that the statement is true.

Examples include:

- Does this message request a refund?
- Does this answer contain evidence for its main claim?
- Has the agent completed the user's request?
- Does this input look like a prompt-injection attempt?

A result near one means likely yes. A result near zero means likely no. A result near 0.5 means the model is uncertain.

Whatever you make of the name, the behavior is straightforward. Noul gives code a probability that can feed an `if` statement or a review threshold.

## A small example

Suppose an analytics product offers CSV exports only on paid plans. A paying customer reports that exports are failing, and the latest log contains a server error.

The application can send the customer message, account state, and export status to Jev. It can then ask which team should own the issue:

```
from typesafe_sdk import Choice, TypeSafeClient
with TypeSafeClient() as client:
    response = client.system_one(
        model="jev-latest",
        state={
            "message": "My CSV export keeps failing.",
            "account": {
                "plan": "paid",
                "payment_status": "current",
            },
            "last_export": {
                "status": "failed",
                "error": "server_error",
            },
        },
        questions={
            "team": Choice(
                instructions="Choose the team that owns this issue.",
                criteria={
                    "billing": "Payments, invoices, and charges",
                    "technical": "Broken features and product errors",
                    "other": "Requests that need another team",
                },
            )
        },
    )
```
The response contains a valid team choice and probabilities for all three teams. Code can assign a clear case automatically and send an uncertain case to a person. Jev makes the semantic judgment, while the application controls the consequences.

## Why Jev can be fast

Generative language models produce output sequentially. Even a small JSON answer requires the model to generate punctuation, property names, values, and closing brackets token by token.

Jev gives up open-ended text generation. TypeSafe says its architecture uses a parallel sampler that produces all declared answers in one query. Questions that share the same state can also be evaluated in parallel.

For a support ticket, one request could ask for the owning team, urgency, refund intent, and whether the customer has already attempted the proposed fix. Those questions see the same state but do not depend on one another.

Parallel questions cannot depend on one another. If a later judgment needs an earlier answer, the application must make a second request. A system might first choose the three most relevant documents, fetch their full text, and then ask which document answers the user's question.

TypeSafe reports that most Jev calls complete within 70 to 500 milliseconds. Its documentation says state and questions share a budget of about 32,000 tokens, and a direct Choice supports up to 255 options. Larger choice sets use an additional stage.

## How Jev differs from structured output in an LLM

The developer experience looks familiar because OpenAI, Anthropic, and other providers already let developers ask a language model to follow a JSON schema. A model can return an enum instead of a paragraph. Underneath that interface, it is still a text generator constrained to a particular shape. Jev is designed around typed decisions as its native output. TypeSafe also says it trains Jev with Reinforcement Learning for Calibrated Decisions, or RLCD, so the returned probabilities better reflect how often the model is correct.

The public material does not yet disclose enough detail to reproduce RLCD or inspect the full architecture. There is no public paper describing the reward function, model size, training data, or calibration procedure. That makes the interface easier to evaluate than the underlying research claim.

Sean Goedecke has argued that existing language models could approximate part of Jev's speed by prefilling most of a structured response and generating only a constrained choice token. His own small test with Qwen2.5-1.5B-Instruct produced a two-to-three-times speedup over ordinary structured output. That does not reproduce Jev's claimed training or parallel question evaluation, but it shows that specialized inference strategies may compete with the category.

It is too early to know whether this gives Jev a lasting technical advantage. The demand for fast structured judgment is easier to see.

## Use case 1: Routing and triage

Routing is the most direct fit because the output space is already bounded. A business can use Jev to route support tickets, classify incoming documents, choose an internal workflow, or decide whether a request needs a specialist. Each option should have a clear description, and the schema should include `other`, `unknown`, or `insufficient_evidence` when the listed choices may not cover every case.

Confidence can control automation. A clear billing question might go straight to the billing queue. An ambiguous request can go to general support or a human triage queue.

At Jev's reported price, the business can consider checking every request rather than sampling a small percentage.

## Use case 2: Model, tool, and agent selection

An AI application often has several capable models and tools. One model may be fast and cheap. Another may reason better. A specialist may be better at SQL, browser work, or code review.

Jev can inspect the task and choose among those options. Vercel lists tool selection, subagent selection, and continue-or-retry decisions among the intended uses for Jev on AI Gateway.

Builders are already experimenting with this pattern. One public project routes work between Claude Code, Codex, and OpenCode. Another production-oriented router reported lower dispatch cost and wall time across 25 test tasks, although matching another model's decisions is a consistency check rather than proof that the choices were correct.

The strongest design keeps the routes explicit. The system should describe what each model or tool is good at, log the decision, and verify whether the chosen route improved the final result.

## Use case 3: Verification and guardrails

Many AI systems need a second judgment after a model produces an answer:

- Does the answer cite evidence for its claims?
- Did the agent stay within the requested scope?
- Does a tool call appear risky?
- Is the response complete?
- Does the input look adversarial?

These checks are often skipped because another large-model call adds too much time and cost. A faster decision model could run during the workflow instead of only at the end.

Mike Taylor at Every tested this idea on writing. He sent 37 documents through 21 checks, producing 777 judgments in under 0.7 seconds for an estimated quarter of a cent. In a smaller comparison, Jev found six of seven planted writing defects. Anthropic's Fable 5.1 found all seven, but took about 25 times longer and was estimated to cost about 580 times more.

Jev missed one defect, so the result does not support using it as the final authority. It does support testing Jev as an inexpensive first reviewer while a stronger model or person handles flagged and high-risk cases.

## Use case 4: Large-scale classification

Low token prices change which datasets can be processed in full. Public experiments have used Jev to classify research papers, scan email, and judge large groups of documents. One builder reported classifying 1,018 AI papers across 24 topics for $0.08 in Jev costs, with a median latency of 256 milliseconds per paper. The author was still running evaluations before replacing the existing labels.

The same approach can classify product reviews, support conversations, invoices, transaction descriptions, content libraries, and application logs. The resulting categories and scores can feed conventional analytics. Before using them, the team still needs a labeled sample. Cheap classification helps only when the labels are good enough for the downstream decision.

## Use case 5: Real-time interfaces and control loops

A decision that takes ten seconds cannot sit comfortably behind every click, keystroke, or game action. A decision that takes a few hundred milliseconds can.

TypeSafe demonstrates Jev playing Doom from structured textual game state at about ten queries per second. The model is not looking at game images, and a purpose-built non-AI bot could play better. The demo shows that a general decision model can operate inside a fast loop.

Browser Use built a flight-search prototype in which Jev chooses browser actions over a changing page state while a small language model handles text entry. One recorded run took about 7.1 seconds. The repository reports only six alternating runs of one task, so it is evidence of feasibility rather than general reliability.

*Source: [Browser Use's Jev ultrafast prototype](https://github.com/browser-use/jev-ultrafast). The repository describes a small task-specific comparison.*

Similar patterns could support live content analysis, adaptive interfaces, robotics, games, and interactive recommendations. The surrounding software must still verify actions, recover from stale state, and handle failures.

## Use case 6: Ranking and next-best actions

Many products need to rank a known set of candidates rather than produce one universal answer. A research tool can score which document is most likely to answer a question. An onboarding product can choose the next instruction based on what the user has completed. A sales system can rank accounts for review, while a moderation system can sort cases by urgency.

Jev's probabilities make these cases more useful than a bare category. An application can act on the top result, compare the gap between the first and second choices, or request more evidence when the distribution is flat.

Question design matters here. A label such as "best document" is vague. Topical relevance, evidence quality, and recency can be scored separately and combined in code.

## What Jev cannot do

Jev does not generate arbitrary text. It cannot write an email, summarize a report, explain a decision, or produce code.

It also cannot select an answer that the developer did not include. If the correct option is missing, it may confidently choose the least-wrong available option. Good schemas need escape routes.

A valid output can still be wrong. TypeSafe's statement that Jev cannot hallucinate is best understood as a guarantee about output shape. The model cannot invent an undeclared field or malformed type. It can misclassify the input.

Probabilities do not remove the need for testing. Calibration is measured across groups of predictions, and it can weaken when the production data changes. A threshold that works for support tickets may be unsafe for financial or medical actions.

The lack of explanations also limits use in audits and appeals. A high-stakes system may need to preserve the evidence, request a written rationale from another model, or require human review.

## What the published evaluation shows

TypeSafe evaluates Jev across four workflows: security incidents, agent-trace observability, invoice processing, and customer service.

Its published aggregate reports 67.8 percent agreement, about $0.0004 per case, and 0.4 seconds per case for Jev. GPT-5.6 Terra reaches 67.9 percent at $0.0304 and 10.1 seconds. GPT-5.6 Sol and Claude Opus 5 score higher at 74.1 and 73.1 percent, with higher cost and latency.

*Source: [TypeSafe workflow evals](https://evals.typesafe.ai/). These are vendor-run results. The reference labels are averages from GPT-6 Astra and Claude Fable 5.1, not independent ground truth.*

The numbers are promising, but this is not a neutral benchmark. TypeSafe created the workflows, and the reference answers come from two other models rather than verified operational outcomes. The company also says its largest reported speed and cost advantages are likely near the high end of real-world gains.

A separate production replay from YTAL shows why full-system testing matters. Jev-only calls for an entity-resolution workflow were estimated to be about 96 percent cheaper. The proposed fallback design still sent uncertain batches through the old process, which raised the estimated total cost by 4.1 percent. The team did not adopt that design because a cheap model call had not produced a cheaper system.

## How I would evaluate Jev

Start with one repeated decision whose possible answers are known. Prefer a recoverable mistake, such as routing or ranking, over an irreversible action.

Build a labeled set from real production examples. Compare Jev with the current rule, model, or human decision. Measure decision quality, probability calibration, end-to-end latency, and the total cost after fallbacks and review.

Then choose thresholds based on consequences. A low-risk recommendation can tolerate more uncertainty than a refund, security action, or account change.

Finally, monitor the distribution over time. New products, policies, customer language, and adversarial behavior can change what the model sees. The system needs a way to detect drift and send uncertain cases elsewhere.

## My read

Jev turns unstructured information into bounded decisions. Its strongest candidates are routing, scoring, verification, large-scale classification, ranking, and real-time loops where a normal LLM call is too slow or expensive.

The architecture and calibration claims still need independent validation. The product idea is easier to judge: a large amount of software does not need another paragraph from AI. It needs a quick answer to a specific question in a form that code can use. Jev is built for that job.
