{"article":{"slug":"decoding-jev","title":"Decoding Jev","subtitle":null,"summary":"A technical walkthrough of what Jev-style decision models appear to be doing—single-token decisions, structured outputs, and why the architecture matters for agent tooling.","content_type":"essay","language":"en","canonical_url":"https://navinpai.github.io/decoding-jev/","author":{"name":"Navin Pai","url":"https://navinpai.github.io/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"navinpai.github.io","url":"https://navinpai.github.io/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":4510,"reading_minutes":20,"published_at":"2026-09-20T15:09:21.946Z","added_at":"2026-09-20T15:09:21.946Z","updated_at":"2026-09-20T15:09:21.946Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/decoding-jev","markdown_url":"https://listedarticles.com/articles/decoding-jev.md","example":false,"citation":"Navin Pai, navinpai.github.io. \"Decoding Jev.\" 20 Sept 2026. https://navinpai.github.io/decoding-jev/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://navinpai.github.io/decoding-jev/"},"body_markdown":"Jev has taken the AI world by storm, but how is it actually different from today’s LLMs?\n\nAn LLM already computes probabilities. It can already classify text. The interesting change is which distribution the system exposes, how much sequential work it removes, and what training rewards. Those are three separate engineering questions.\n\nWe’ll follow the tensors through a decision model, compare its execution with token generation, then derive the training objective. Jev’s implementation is private; Laya’s published code gives us a concrete reference for the mechanisms. Claims about Jev are identified where the implementations differ.\n\nTL;DR FAQ\n\nIs Jev the first of its kind?\n\nFor direct classification, no: BERT-based models already returned label probabilities without generating text. Jev combines typed decisions, parallel inference, and calibration-focused training; Laya is an independent open decision model, not Jev’s released architecture.\n\nHow does Jev compare with state-of-the-art LLMs for classification?\n\nJev can be competitive, but there is no established accuracy lead over frontier LLMs. In one 150-passage test, Jev and Haiku scored 66%, while Sonnet scored 71.3%; Jev responded faster.\n\nIsn’t this just an LLM that responds with one token?\n\nFor one classification, a non-reasoning LLM can return a single label token and expose candidate probabilities immediately. Jev bundles several decisions and their distributions; the meaningful differences are training, calibration, and serving efficiency.\n\nHow large is Jev? Can I run it on my server?\n\nJev’s size is undisclosed, with no public weights to self-host; Chopra’s ~30B estimate is unverified. Laya’s English model runs locally with 421M parameters (learned weights): calculated weight storage is 0.84 GB at 16-bit or 1.68 GB at 32-bit, plus runtime memory.\n\nCan I fine-tune a Jev-like model?\n\nYes—Laya and Kev publish trainable implementations, though Jev itself currently has no customer fine-tuning. Train on labeled reports, calibrate probabilities on separate data, and test on untouched reports; matching the interface does not guarantee matching Jev’s quality.\n\nHow can it answer everything at once? Is it still a deep neural network?\n\nLaya uses a deep transformer to score all supplied options together, without generating one answer before the next. Its layers still run sequentially; Jev’s exact architecture is private, so Laya illustrates a working approach rather than a confirmed reconstruction.\n\nIs an LLM’s time to first token comparable to Jev’s response time?\n\nFor a single label token with no hidden reasoning, time to first token is approximately time to a useful answer. Jev returns completed decisions, so compare complete-response latency for equivalent outputs; neither design guarantees a lower waiting time.\n\nNo: supervised classifiers can learn probabilities without Reinforcement Learning for Calibrated Decisions (RLCD). Either approach needs calibration checks—among reports assigned an 80% regression probability, roughly 80% should be confirmed regressions on fresh data.\n\nWhat does “System One” mean here?\n\nTypeSafe borrows the psychology term for fast judgments within predefined answer choices, such as routing a bug report. It describes the intended task; arithmetic, indirect instructions, and adversarial text can still cause mistakes.\n\n01 The forward pass\n\nThe network still does the work.\n\nWe’re building a tool to triage bug reports. It reads a report and estimates the affected component, the severity, and the probability that the bug is a regression: something that worked in an earlier version and broke after a change. Jev accepts the report as text. Supported inputs ↗\n\nTHE BUG REPORT WE’LL FOLLOW\n\n“Since the update, clicking Download crashes the app. Viewing files still works.”\n\n“Since the update” suggests a regression, but investigation has to confirm it. We’ll use that distinction when we get to training: the report is the evidence available to the model; the confirmed outcome is what we score its prediction against.\n\nThe open reference model, Laya, uses a ModernBERT-large encoder in its English checkpoint: 28 transformer layers, with a 1,024-number vector at each token position. Two more transformer layers prepare those vectors for scoring. The computation still passes through a deep neural network before producing any answer. Encoder configuration ↗Laya model card ↗\n\nA vector is a list of numbers. Each layer updates the vectors in two steps: attention mixes information between token positions, then a feed-forward network transforms the numbers at each position. The next layer starts with those updated vectors. Below, four small layers let you watch this happen. Transformer paper ↗\n\nLaya’s encoder is bidirectional: a position can use context on either side. That will matter when an option marker appears before the report it needs to judge. A causal language model blocks attention to later positions. Switch between the two settings and inspect “The” while replacing “download” with “preview.” Only the bidirectional version lets that later edit reach the first word. Bidirectional representations ↗\n\nBoth settings can process the known input positions together within a layer. Layer 3 still has to wait for layer 2. In the published Laya configuration, most encoder layers attend within a local window; every third layer uses the full input, starting with the first. “Parallel” does not mean skipping those layers. Layer layout ↗\n\nUnpack one layer’s arithmetic\n\nAttention first makes three versions of each position’s vector: a query, a key, and a value. A query is compared with the keys to decide how much weight each position gets. Those weights are used to mix the value vectors.\n\nA = softmax(QKᵀ / √d + mask) mixed = A × V\n\nQ, K, and V are rectangular arrays of numbers—matrices. The superscript T swaps rows and columns; d is the query vector’s length. A dot product multiplies matching numbers and adds the products. Softmax turns scores into shares that sum to one. A mask uses a prohibitive score at forbidden connections. Then the feed-forward block applies two learned matrix transforms with a nonlinear function between them. Residual additions preserve the incoming signal; normalization controls its scale. You can see these operations in a production transformer implementation. Llama layer implementation ↗\n\nThe little model has one attention head, meaning one set of these comparisons. Laya’s encoder has 16 heads per layer, each working on a different projection of the vectors. The four-layer display keeps the arithmetic small enough to inspect. Attention dimensions ↗\n\n02 Output representation\n\nRead probabilities without generating labels.\n\nLaya puts a [MASK] marker before every option. After the layers have processed the sequence, it gathers the vectors at those markers. Each vector has had access to its option and the surrounding input. This is how “Downloads” and “Preview” get their own scores, even though the labels were supplied with the request. Sequence construction and scoring ↗\n\nThe head is the part that turns those vectors into answers. Laya uses the same small network for every marker: normalize its 1,024 numbers, apply a linear layer, pass through a smooth nonlinear function called GELU, then reduce to one score. That score is a logit. All marker vectors can go through this scorer together. DecisionModel ↗\n\nSoftmax raises e to each score and divides by the total, giving shares that sum to one. Raise the Downloads logit: Preview’s probability falls even though its own score hasn’t changed. For Score, code takes the weighted average of the numbered levels. For Noul, it returns the probability assigned to “yes.” Answer conversion ↗\n\nLaya also divides the logits by a fitted temperature, T, before softmax. A larger T spreads the probabilities more evenly without changing which option wins. Try its stored values: about 1.76 for a three-option Choice, 1.25 for this Score, and 1.98 for Noul. This is a numerical adjustment after the network has finished. Stored temperatures ↗\n\nThis also explains why new labels don’t require a new “Downloads neuron.” The labels arrive as text, the encoder builds a vector for each marker, and the scorer reuses its weights. Earlier systems such as FIRST read candidate scores from an LLM’s first-token logits. There are several working ways to get numerical decisions out of a language model. FIRST ↗\n\n03 Execution and latency\n\nWhere the latency goes.\n\nA text-generating LLM processes the prompt first. This is prefill. Its output layer gives the scores for the first new token. Once that token is chosen, the model feeds it through the layers to get the next one. This dependency on earlier output is what autoregressive means. A KV cache keeps the earlier attention keys and values so they don’t have to be rebuilt at every step. Inference and caching ↗\n\nLaya stops after scoring the markers. Its code calls the network once for the batch, then uses ordinary arithmetic to format the answers. Jev’s launch describes the same outward behavior: probabilities returned together, with no generated text. Laya’s forward call ↗Jev’s launch ↗\n\nTime to first token (TTFT) ends when the first output token reaches you. End-to-end latency ends when the whole response arrives. If the first token is “{”, the software is still waiting for the fields inside it. Jev’s advertised 70–500 ms measures the complete response. Timing definitions ↗Jev’s reported range ↗\n\nSet the output length to one. If that token is the complete label you need, the LLM is done at TTFT. In this example it finishes before the decision model, whose extra processing is still running. Longer generated answers leave more work to avoid. For a real comparison, measure both systems until the result is usable by your code.\n\nA schema can constrain an LLM to valid JSON, and requests can be batched. Those are useful baselines. The practical question is how long each implementation takes for the same job, with the same input and required answers. Structured outputs ↗\n\nWHY CALL THIS “SYSTEM ONE”?\n\nA name borrowed from psychology.\n\nAn engineer may recognize a familiar crash report at a glance. Finding the cause can take a sequence of tests: reproduce it, compare versions, inspect the failing code. Psychology uses System 1 for automatic impressions and System 2 for deliberate thought. Kahneman treated them as ways of describing mental activity. Kahneman’s lecture ↗\n\nTypeSafe borrows System One for its decision models. Earlier AI papers used the analogy differently: Tree of Thoughts called ordinary LLM generation System-1-like, then added search to encourage deliberation. The name tells you what kind of work the authors have in mind; it won’t tell you how many layers the network has. TypeSafe’s usage ↗Tree of Thoughts ↗\n\n04 Batching and dependence\n\nParallel answers, separate distributions.\n\nThe affected component, severity, and regression probability can be estimated from the same bug report. Laya builds one input sequence for each question, including a copy of the report in each. It pads the sequences to the same length and sends them to the model as a batch: several rows evaluated in one call. Batch construction ↗\n\nThe rows use the same network weights, but their attention stays within the row. The component question can read its own options and copy of the report. It cannot read the regression question’s wording. This isolation comes from the batch dimension, with no special cross-question attention mask.\n\nBatching fills more of a GPU’s arithmetic units at once. It can improve throughput even while the request takes longer. Laya’s published English-checkpoint timings on a T4 rise from 39.5 ms for one question to 158.6 ms for ten. That is more answers per second, but more waiting for the complete call. Reported batch timings ↗\n\nJev’s documentation says it ingests the state once. Laya’s batch code repeats it. A causal transformer can save work by caching a shared prefix. In Laya’s bidirectional encoder, however, the report also attends to the question and options. Its vectors can therefore differ from row to row. Jev’s state handling ↗Shared-prefix research ↗Laya’s batch inputs ↗\n\nTwo 80% answers do not specify their overlap.\n\nSuppose P(regression) = 0.8 and P(task blocked) = 0.8. Both of these joint distributions have those same marginals—the probabilities for each property on its own.\n\nCase A · independent properties\n\nBug type\n\nBlocking\n\nNot blocking\n\nRegression\n\n64%\n\n16%\n\nOther bug\n\n16%\n\n4%\n\nCase B · properties coincide\n\nBug type\n\nBlocking\n\nNot blocking\n\nRegression\n\n80%\n\n0%\n\nOther bug\n\n0%\n\n20%\n\nMultiplying 0.8 × 0.8 assumes independence. The marginals alone only bound the overlap between 60% and 80%. If a workflow needs the probability of a blocking regression, ask about that combination directly and validate the prediction. Evaluating questions together does not supply a joint probability model.\n\n05 Training objectives\n\nWhat RLCD is optimizing.\n\nReturning a probability is easy: apply softmax. Making that probability useful is a training problem. TypeSafe calls its approach reinforcement learning for calibrated decisions, or RLCD. Its public description specifies the goal, but gives no loss function or optimizer to inspect. We can still work through the mathematics and examine a published implementation. TypeSafe’s training description\n\nStart with 100 comparable bug reports. Investigation later confirms that 80 are regressions. A model that always chooses “regression” gets 80 correct. Reporting 80% or 99% doesn’t change that accuracy. If we instead sample an answer from its probabilities and reward a correct sample, we create a different incentive: the model earns more by always sampling the more common answer.\n\nWith f = 0.8, rewarding a correct sampled answer gives 0.2 + 0.6p. Its best report is p = 1. That is a good policy for maximizing correct actions, but a bad estimate of how often the event happens. The distinction is between learning what to do and learning how likely an outcome is.\n\nA strictly proper scoring rule makes the true probability the unique best report in expectation. Log score rewards the log probability assigned to the outcome that occurred. A confirmed regression contributes ln(p); a bug confirmed not to be a regression contributes ln(1 − p). Average over the group and the optimum moves to 80%. Proper scoring rules\n\nWe can locate that peak without running a neural network. The slope is zero where:\n\nE[R] = f ln(p) + (1 − f) ln(1 − p) dE[R]/dp = f/p − (1 − f)/(1 − p) = 0 ⇒ p = f\n\nHere E means an average over outcomes, and ln is the natural logarithm. The curve bends downward, so this stationary point is a maximum.\n\nCross-entropy already has this property.\n\nNegate that log reward and you have binary cross-entropy, a standard supervised classification loss. You can differentiate it directly and update the network by backpropagation: carrying derivatives backward through the layers to calculate how each weight should change. If the model’s logit is μ and p = sigmoid(μ), the loss gradient is simply p − f. At p = 0.5 and f = 0.8, gradient descent pushes the logit upward.\n\nSo calibrated probabilities do not inherently require reinforcement learning. The important questions are what the targets represent, which objective is optimized, and how the fitted model behaves on fresh data. An LLM’s next-token loss estimates the distribution of text; a classifier’s loss estimates labels. Both use cross-entropy, but they are predicting different things.\n\nThe population optimum is also not a guarantee about a trained network. Limited data, model capacity, optimization, and a change in the input distribution can all leave it miscalibrated. Neural classifiers trained with cross-entropy are known to need calibration afterward. Guo et al., 2017\n\nFollow an actual policy-gradient update.\n\nLaya’s public typed-decisions fine-tuning notebook makes the implementation concrete. It starts from an existing checkpoint and uses target probability vectors from the dataset. This is a separate training recipe we can inspect, rather than evidence of Jev’s internal algorithm. Pinned training notebook\n\nRun the network once. Produce an option-logit vector for each question. These are the centers around which training will explore.\n\nMake four noisy candidates. Add Gaussian noise to the logits, then apply softmax to each candidate. The noise is centered across valid options: shifting every logit equally would leave softmax unchanged.\n\nScore the candidate distributions. Use log score plus 0.75 times spherical score. Ordered Score questions also subtract a ranked probability error. Compute each candidate’s advantage: its reward minus its group’s mean reward.\n\nUpdate the weights. A policy-gradient loss favors candidates with positive advantage. The notebook adds a full-weight cross-entropy term, then updates both encoder and head with AdamW, an optimizer that adapts step sizes using past gradients. Rescaling the advantages and limiting the gradient’s total length help control the update size.\n\nREINFORCE is the gradient estimator behind this update. It asks how changing the network would change the probability of sampling a candidate, then weights that direction by the candidate’s reward. Subtracting a baseline makes the comparison relative to nearby candidates. A candidate can have a negative reward and still have positive advantage. Policy-gradient mechanics\n\n06 / ONE TRAINING UPDATE\n\nTurn four candidates into a gradient.\n\nTake an update, then resample. The target frequency stays at 80%.\n\nCURRENT P(REGRESSION) 50.0%\n\nThe black tick marks the 80% target.\n\nThis example updates one binary logit μ. Each candidate adds noise δ, reports sigmoid(μ + δ), and earns the expected log reward from the preceding experiment.\n\nSample\n\nNoise δ\n\nP(regression)\n\nReward R\n\nAdvantage A\n\nA × δ / σ²\n\nGroup mean reward:\n\nAverage sampled gradient\n\nDirect log-score gradient at σ = 0\n\n0 updates\n\nA one-parameter illustration of REINFORCE, with four samples, a group-mean baseline, log reward, and a step size of 0.5. It omits the notebook’s composite reward, advantage normalization, cross-entropy term, and AdamW. Initial noise values are fixed; resampling uses a seeded Gaussian generator. The slider bounds limit updates.\n\nFor Gaussian noise, the gradient contribution is A × δ / σ². A sample that moves right and does better than its neighbors pushes the center right. A worse sample on the left also pushes it right. Backpropagation carries the combined signal through the scoring head into the encoder’s weights.\n\nThe two gradients in the experiment need not match. One uses four noisy candidates and a shared baseline; the other differentiates the unperturbed log score exactly. With nonzero noise, the sampled objective averages rewards around the current logits. That is a different objective from scoring only the distribution used at inference. A proper reward alone therefore does not prove that the deployed model is calibrated.\n\nThe expected sampled objective is J(μ) = Enoise[R(sigmoid(μ + noise))]. The score-function estimator differentiates the sampling density, treating each sampled z as fixed. In the notebook, detaching the sampled center and evaluating reward without gradients does this explicitly. The gradient still flows through the log-density calculation into μ.\n\nA group mean includes each sample’s own reward. For independent samples, that scales the unnormalized estimator’s expectation by (G − 1)/G, where G is the group size. The notebook also normalizes advantages, so this should not be described as an exact unbiased gradient of the unsmoothed score. The demo leaves normalization out to expose the arithmetic.\n\nWhat the extra reward terms measure\n\nFor target distribution t and predicted distribution p, the spherical term is (t · p) / ‖p‖: a dot product divided by the vector’s length. The ranked term compares cumulative probabilities along ordered levels. A prediction two levels away then incurs more error than a neighboring prediction. The source clips log probabilities at −9.21, so its implementation is not the unrestricted textbook log score. Reward function\n\nThe 0.75 spherical weight comes from the fine-tuning notebook; the reusable reward function defaults to 0.5. Neither number is a published Jev hyperparameter.\n\nThe label source sets the ceiling on the claim.\n\nThe fine-tuning code reads full target distributions from gold fields, rather than collapsing them to winning labels. The dataset describes teacher-generated reference probabilities. Matching them trains the model to reproduce those references; it does not establish that 80% predictions will occur 80% of the time in your production data. Dataset construction\n\nCross-entropy(t, p) = H(t) + KL(t ‖ p)\n\nH(t) is fixed once the target is chosen. KL measures the mismatch between target and prediction. Minimizing this loss makes p approach t—including any errors in t.\n\nAfter training, the notebook fits a temperature by minimizing cross-entropy again, this time with the network frozen. Only the scale of the logits changes. That is the T slider from the readout experiment. Exploration noise σ changes the candidates considered during training; temperature T changes the reported probabilities afterward.\n\nMeasure the probabilities before setting a threshold.\n\nThe next experiment separates the probability estimate from the application’s decision. First minimize probability error on a group of resolved bug reports. Then try a different release’s reports without changing the forecast, or change the threshold for routing a report to the regression queue without retraining anything.\n\nA constant 80% forecast can be calibrated across this group while telling us nothing about which individual report is a regression. Calibration and discrimination are different: we want reliable probabilities and useful separation between easy and difficult cases. Check both, including the subset the application actually accepts.\n\nThe experiment that would isolate RL’s contribution.\n\nPolicy gradients are useful when reward comes from a discrete action or an external evaluator whose derivatives are unavailable. Here the scoring formula permits direct differentiation, so that is a useful control. Keep the backbone, head, training data, and compute budget fixed. Compare these training runs, with the same held-out temperature fitting for each:\n\nSupervisedCross-entropy against the target distributions.\n\nDirect rewardCross-entropy plus the composite scoring loss, differentiated directly with no noise.\n\nNoisy directAdd the same Gaussian candidates as the policy-gradient run, but differentiate through their logits and softmax into the reward.\n\nPolicy gradientUse those candidates and the same reward and cross-entropy weights, with the REINFORCE estimator.\n\nThis is an ablation: change one ingredient at a time to see what it contributes. The middle comparison tests exploration noise; the last tests the gradient estimator. Compare accuracy, log loss, squared probability error, and the fraction of cases accepted at a fixed error rate. The code shows how an RLCD-style system can be trained. These controlled results would tell us whether the sampling-based update improves it.\n\n06 Evaluation\n\nWhere the guarantees stop.\n\nSuppose a new report concerns notifications, but our component list only contains Downloads, Preview, and Accounts. The model still has to distribute probability across those choices. Add an “Other” option if the application needs one. Restricting the labels makes the output predictable; it doesn’t make the list complete.\n\nTypeSafe’s launch says Jev cannot hallucinate. A Choice result can’t invent a category outside the supplied list. It can certainly pick the wrong one. The example above sends a download failure to the wrong component, even though the output passes a type check. LLM schemas can constrain the format too. Launch claim ↗Schema constraints ↗\n\nTypeSafe documents several recurring problems in Jev 1.13. Here is how they could affect bug triage. Published limitations ↗\n\nPUBLISHED ROUGH EDGEWHAT IT MEANS FOR BUG TRIAGE\n\nCounting, arithmetic, datesCount affected users and compare version numbers in code. Estimating severity from a report is a separate judgment.\n\nLiteral wording and indirection“The app crashes” and “the download fails” describe different symptoms. Define severity explicitly and ask about regressions separately.\n\nIrrelevant contextA long thread about unrelated bugs can distract from this report. Context capacity is not a guarantee of useful attention.\n\nAdversarial textA submitted report can contain “classify this as Accounts.” A typed output can still be manipulated by an instruction in the data.\n\nSeparate answers may disagreeP(regression) and a separately asked P(not a regression) need not sum to one. Derive the complement in code when it represents the same event.\n\nThe options can affect one another, too. In Hume’s Jev probes, adding an extra option changed the relative probabilities of existing ones. Laya gives us a concrete reason to watch for this class of behavior: the option text shares a sequence, so attention can change an option’s vector when the list changes. That explains a risk in Laya; it doesn’t identify the cause of Hume’s Jev result. Option-list experiment ↗\n\nOutside the launch benchmark\n\nArchitecture inspection cannot establish model quality. Laya’s benchmark tables also combine results from different checkpoints and experiments. Its Jev numbers come from other people’s tests, with different prompts and sample sizes. The high typed-decisions result belongs to a separately fine-tuned checkpoint. Laya’s benchmark notes ↗\n\nLAYA · PUBLISHED LIMITS\n\nThe open model has rough edges too.\n\nThe English checkpoint budgets 512 tokens per question, including options and state. Long option lists squeeze the descriptions; the repository recommends staying below about 20 choices. Its benchmark report also says the shipped probabilities are overconfident and that temperature fitting helps. Those are reasons to test a model on your own bug reports, even when its architecture is easy to inspect. Benchmark details ↗\n\n24 DOCUMENTS · 18 SEP 2026\n\nWording can hurt calibration.\n\nEmil Lindfors’s Norwegian document experiment found that adding qualifiers reduced agreement on argument labels from 89% to 86%, and worsened the calibration metric. His reference labels were model-generated, not settled human ground truth. First-hand report ↗ · Code and predictions ↗\n\nSDK REPORT · 17 SEP 2026\n\nRepetition does not prove correctness.\n\nA developer reported the same disputed routing result over 100 runs of a quickstart example. One overlapping-category case is not an error rate. The issue is closed, and the report does not establish current behavior. Original issue ↗\n\nPARAS CHOPRA · THREE SELECTED RESULTS\n\nA one-pass Qwen baseline, Laya, and Jev.\n\nThe local Qwen prototype used unchanged pretrained weights stored at 4-bit precision and read label probabilities in one pass. Jev ran through OpenRouter; these reported accuracies come from reused public and synthetic tests, with unknown pretraining overlap.\n\nAccuracy on matched task content and option order\n\nTask\n\nQwen3 4B\n\nLaya English\n\nJev 1.13\n\nIntent routing 400 cases\n\n96.25%\n\n63.00%\n\n99.75%\n\nMMLU-Pro 400 cases\n\n45.00%\n\n13.50%\n\n79.75%\n\nRelational choice 100 cases\n\n53.00%\n\n8.00%\n\n0.00%\n\nMMLU-Pro tests academic knowledge and reasoning. Relational choice uses information in one option to select another; the gist provides neither exact prompts nor raw outputs to diagnose Jev’s failure. The result does not establish how its options are processed. Full comparison and caveats ↗\n\nChopra’s post estimates about 30 billion Jev parameters from accuracy and latency. That remains unverified: neither measurement identifies model size, and the timings compare local Laya with remote Jev on different hardware.\n\nTypeSafe’s headline gains, 193.6× faster and 444.6× cheaper, came from four workflows. The LLM wrapper had to produce probability estimates, and the reference answers were averages from two larger models. TypeSafe says those gains are likely toward the high end. For bug triage, the useful test would be the same reports, component choices, severity rubric, and confirmed outcomes on each system. Benchmark setup ↗\n\n07 Engineering tradeoffs\n\nChoose a baseline that does the same job.\n\nFor bug-report triage, compare the decision service with a supervised encoder and a generative model constrained to a short label. Match the reports, component options, severity rubric, and required probabilities. A one-token classifier and a model producing a paragraph are doing different amounts of work; their latency gap does not isolate an architectural improvement.\n\nMeasure complete-request p50 and p95 latency—the median and the time under which 95% of requests finish—at the same concurrency. Then measure decision quality and accepted-case error after calibrating on separate data. Jev’s useful contribution has to survive that comparison: less waiting or lower cost at the quality your application needs.\n\nThe layers still have work to do. The answers don’t have to be written out.\n\nFOLLOW THE EVIDENCE\n\nReferences & implementation notes.\n\nThe sliders use small models with hand-set numbers. External benchmark results are credited to the people who ran them. We inspected the source and configuration; we did not rerun the trained models or those benchmarks.","body_html":"<p>Jev has taken the AI world by storm, but how is it actually different from today’s LLMs?</p>\n<p>An LLM already computes probabilities. It can already classify text. The interesting change is which distribution the system exposes, how much sequential work it removes, and what training rewards. Those are three separate engineering questions.</p>\n<p>We’ll follow the tensors through a decision model, compare its execution with token generation, then derive the training objective. Jev’s implementation is private; Laya’s published code gives us a concrete reference for the mechanisms. Claims about Jev are identified where the implementations differ.</p>\n<p>TL;DR FAQ</p>\n<p>Is Jev the first of its kind?</p>\n<p>For direct classification, no: BERT-based models already returned label probabilities without generating text. Jev combines typed decisions, parallel inference, and calibration-focused training; Laya is an independent open decision model, not Jev’s released architecture.</p>\n<p>How does Jev compare with state-of-the-art LLMs for classification?</p>\n<p>Jev can be competitive, but there is no established accuracy lead over frontier LLMs. In one 150-passage test, Jev and Haiku scored 66%, while Sonnet scored 71.3%; Jev responded faster.</p>\n<p>Isn’t this just an LLM that responds with one token?</p>\n<p>For one classification, a non-reasoning LLM can return a single label token and expose candidate probabilities immediately. Jev bundles several decisions and their distributions; the meaningful differences are training, calibration, and serving efficiency.</p>\n<p>How large is Jev? Can I run it on my server?</p>\n<p>Jev’s size is undisclosed, with no public weights to self-host; Chopra’s ~30B estimate is unverified. Laya’s English model runs locally with 421M parameters (learned weights): calculated weight storage is 0.84 GB at 16-bit or 1.68 GB at 32-bit, plus runtime memory.</p>\n<p>Can I fine-tune a Jev-like model?</p>\n<p>Yes—Laya and Kev publish trainable implementations, though Jev itself currently has no customer fine-tuning. Train on labeled reports, calibrate probabilities on separate data, and test on untouched reports; matching the interface does not guarantee matching Jev’s quality.</p>\n<p>How can it answer everything at once? Is it still a deep neural network?</p>\n<p>Laya uses a deep transformer to score all supplied options together, without generating one answer before the next. Its layers still run sequentially; Jev’s exact architecture is private, so Laya illustrates a working approach rather than a confirmed reconstruction.</p>\n<p>Is an LLM’s time to first token comparable to Jev’s response time?</p>\n<p>For a single label token with no hidden reasoning, time to first token is approximately time to a useful answer. Jev returns completed decisions, so compare complete-response latency for equivalent outputs; neither design guarantees a lower waiting time.</p>\n<p>No: supervised classifiers can learn probabilities without Reinforcement Learning for Calibrated Decisions (RLCD). Either approach needs calibration checks—among reports assigned an 80% regression probability, roughly 80% should be confirmed regressions on fresh data.</p>\n<p>What does “System One” mean here?</p>\n<p>TypeSafe borrows the psychology term for fast judgments within predefined answer choices, such as routing a bug report. It describes the intended task; arithmetic, indirect instructions, and adversarial text can still cause mistakes.</p>\n<p>01 The forward pass</p>\n<p>The network still does the work.</p>\n<p>We’re building a tool to triage bug reports. It reads a report and estimates the affected component, the severity, and the probability that the bug is a regression: something that worked in an earlier version and broke after a change. Jev accepts the report as text. Supported inputs ↗</p>\n<p>THE BUG REPORT WE’LL FOLLOW</p>\n<p>“Since the update, clicking Download crashes the app. Viewing files still works.”</p>\n<p>“Since the update” suggests a regression, but investigation has to confirm it. We’ll use that distinction when we get to training: the report is the evidence available to the model; the confirmed outcome is what we score its prediction against.</p>\n<p>The open reference model, Laya, uses a ModernBERT-large encoder in its English checkpoint: 28 transformer layers, with a 1,024-number vector at each token position. Two more transformer layers prepare those vectors for scoring. The computation still passes through a deep neural network before producing any answer. Encoder configuration ↗Laya model card ↗</p>\n<p>A vector is a list of numbers. Each layer updates the vectors in two steps: attention mixes information between token positions, then a feed-forward network transforms the numbers at each position. The next layer starts with those updated vectors. Below, four small layers let you watch this happen. Transformer paper ↗</p>\n<p>Laya’s encoder is bidirectional: a position can use context on either side. That will matter when an option marker appears before the report it needs to judge. A causal language model blocks attention to later positions. Switch between the two settings and inspect “The” while replacing “download” with “preview.” Only the bidirectional version lets that later edit reach the first word. Bidirectional representations ↗</p>\n<p>Both settings can process the known input positions together within a layer. Layer 3 still has to wait for layer 2. In the published Laya configuration, most encoder layers attend within a local window; every third layer uses the full input, starting with the first. “Parallel” does not mean skipping those layers. Layer layout ↗</p>\n<p>Unpack one layer’s arithmetic</p>\n<p>Attention first makes three versions of each position’s vector: a query, a key, and a value. A query is compared with the keys to decide how much weight each position gets. Those weights are used to mix the value vectors.</p>\n<p>A = softmax(QKᵀ / √d + mask) mixed = A × V</p>\n<p>Q, K, and V are rectangular arrays of numbers—matrices. The superscript T swaps rows and columns; d is the query vector’s length. A dot product multiplies matching numbers and adds the products. Softmax turns scores into shares that sum to one. A mask uses a prohibitive score at forbidden connections. Then the feed-forward block applies two learned matrix transforms with a nonlinear function between them. Residual additions preserve the incoming signal; normalization controls its scale. You can see these operations in a production transformer implementation. Llama layer implementation ↗</p>\n<p>The little model has one attention head, meaning one set of these comparisons. Laya’s encoder has 16 heads per layer, each working on a different projection of the vectors. The four-layer display keeps the arithmetic small enough to inspect. Attention dimensions ↗</p>\n<p>02 Output representation</p>\n<p>Read probabilities without generating labels.</p>\n<p>Laya puts a [MASK] marker before every option. After the layers have processed the sequence, it gathers the vectors at those markers. Each vector has had access to its option and the surrounding input. This is how “Downloads” and “Preview” get their own scores, even though the labels were supplied with the request. Sequence construction and scoring ↗</p>\n<p>The head is the part that turns those vectors into answers. Laya uses the same small network for every marker: normalize its 1,024 numbers, apply a linear layer, pass through a smooth nonlinear function called GELU, then reduce to one score. That score is a logit. All marker vectors can go through this scorer together. DecisionModel ↗</p>\n<p>Softmax raises e to each score and divides by the total, giving shares that sum to one. Raise the Downloads logit: Preview’s probability falls even though its own score hasn’t changed. For Score, code takes the weighted average of the numbered levels. For Noul, it returns the probability assigned to “yes.” Answer conversion ↗</p>\n<p>Laya also divides the logits by a fitted temperature, T, before softmax. A larger T spreads the probabilities more evenly without changing which option wins. Try its stored values: about 1.76 for a three-option Choice, 1.25 for this Score, and 1.98 for Noul. This is a numerical adjustment after the network has finished. Stored temperatures ↗</p>\n<p>This also explains why new labels don’t require a new “Downloads neuron.” The labels arrive as text, the encoder builds a vector for each marker, and the scorer reuses its weights. Earlier systems such as FIRST read candidate scores from an LLM’s first-token logits. There are several working ways to get numerical decisions out of a language model. FIRST ↗</p>\n<p>03 Execution and latency</p>\n<p>Where the latency goes.</p>\n<p>A text-generating LLM processes the prompt first. This is prefill. Its output layer gives the scores for the first new token. Once that token is chosen, the model feeds it through the layers to get the next one. This dependency on earlier output is what autoregressive means. A KV cache keeps the earlier attention keys and values so they don’t have to be rebuilt at every step. Inference and caching ↗</p>\n<p>Laya stops after scoring the markers. Its code calls the network once for the batch, then uses ordinary arithmetic to format the answers. Jev’s launch describes the same outward behavior: probabilities returned together, with no generated text. Laya’s forward call ↗Jev’s launch ↗</p>\n<p>Time to first token (TTFT) ends when the first output token reaches you. End-to-end latency ends when the whole response arrives. If the first token is “{”, the software is still waiting for the fields inside it. Jev’s advertised 70–500 ms measures the complete response. Timing definitions ↗Jev’s reported range ↗</p>\n<p>Set the output length to one. If that token is the complete label you need, the LLM is done at TTFT. In this example it finishes before the decision model, whose extra processing is still running. Longer generated answers leave more work to avoid. For a real comparison, measure both systems until the result is usable by your code.</p>\n<p>A schema can constrain an LLM to valid JSON, and requests can be batched. Those are useful baselines. The practical question is how long each implementation takes for the same job, with the same input and required answers. Structured outputs ↗</p>\n<p>WHY CALL THIS “SYSTEM ONE”?</p>\n<p>A name borrowed from psychology.</p>\n<p>An engineer may recognize a familiar crash report at a glance. Finding the cause can take a sequence of tests: reproduce it, compare versions, inspect the failing code. Psychology uses System 1 for automatic impressions and System 2 for deliberate thought. Kahneman treated them as ways of describing mental activity. Kahneman’s lecture ↗</p>\n<p>TypeSafe borrows System One for its decision models. Earlier AI papers used the analogy differently: Tree of Thoughts called ordinary LLM generation System-1-like, then added search to encourage deliberation. The name tells you what kind of work the authors have in mind; it won’t tell you how many layers the network has. TypeSafe’s usage ↗Tree of Thoughts ↗</p>\n<p>04 Batching and dependence</p>\n<p>Parallel answers, separate distributions.</p>\n<p>The affected component, severity, and regression probability can be estimated from the same bug report. Laya builds one input sequence for each question, including a copy of the report in each. It pads the sequences to the same length and sends them to the model as a batch: several rows evaluated in one call. Batch construction ↗</p>\n<p>The rows use the same network weights, but their attention stays within the row. The component question can read its own options and copy of the report. It cannot read the regression question’s wording. This isolation comes from the batch dimension, with no special cross-question attention mask.</p>\n<p>Batching fills more of a GPU’s arithmetic units at once. It can improve throughput even while the request takes longer. Laya’s published English-checkpoint timings on a T4 rise from 39.5 ms for one question to 158.6 ms for ten. That is more answers per second, but more waiting for the complete call. Reported batch timings ↗</p>\n<p>Jev’s documentation says it ingests the state once. Laya’s batch code repeats it. A causal transformer can save work by caching a shared prefix. In Laya’s bidirectional encoder, however, the report also attends to the question and options. Its vectors can therefore differ from row to row. Jev’s state handling ↗Shared-prefix research ↗Laya’s batch inputs ↗</p>\n<p>Two 80% answers do not specify their overlap.</p>\n<p>Suppose P(regression) = 0.8 and P(task blocked) = 0.8. Both of these joint distributions have those same marginals—the probabilities for each property on its own.</p>\n<p>Case A · independent properties</p>\n<p>Bug type</p>\n<p>Blocking</p>\n<p>Not blocking</p>\n<p>Regression</p>\n<p>64%</p>\n<p>16%</p>\n<p>Other bug</p>\n<p>16%</p>\n<p>4%</p>\n<p>Case B · properties coincide</p>\n<p>Bug type</p>\n<p>Blocking</p>\n<p>Not blocking</p>\n<p>Regression</p>\n<p>80%</p>\n<p>0%</p>\n<p>Other bug</p>\n<p>0%</p>\n<p>20%</p>\n<p>Multiplying 0.8 × 0.8 assumes independence. The marginals alone only bound the overlap between 60% and 80%. If a workflow needs the probability of a blocking regression, ask about that combination directly and validate the prediction. Evaluating questions together does not supply a joint probability model.</p>\n<p>05 Training objectives</p>\n<p>What RLCD is optimizing.</p>\n<p>Returning a probability is easy: apply softmax. Making that probability useful is a training problem. TypeSafe calls its approach reinforcement learning for calibrated decisions, or RLCD. Its public description specifies the goal, but gives no loss function or optimizer to inspect. We can still work through the mathematics and examine a published implementation. TypeSafe’s training description</p>\n<p>Start with 100 comparable bug reports. Investigation later confirms that 80 are regressions. A model that always chooses “regression” gets 80 correct. Reporting 80% or 99% doesn’t change that accuracy. If we instead sample an answer from its probabilities and reward a correct sample, we create a different incentive: the model earns more by always sampling the more common answer.</p>\n<p>With f = 0.8, rewarding a correct sampled answer gives 0.2 + 0.6p. Its best report is p = 1. That is a good policy for maximizing correct actions, but a bad estimate of how often the event happens. The distinction is between learning what to do and learning how likely an outcome is.</p>\n<p>A strictly proper scoring rule makes the true probability the unique best report in expectation. Log score rewards the log probability assigned to the outcome that occurred. A confirmed regression contributes ln(p); a bug confirmed not to be a regression contributes ln(1 − p). Average over the group and the optimum moves to 80%. Proper scoring rules</p>\n<p>We can locate that peak without running a neural network. The slope is zero where:</p>\n<p>E[R] = f ln(p) + (1 − f) ln(1 − p) dE[R]/dp = f/p − (1 − f)/(1 − p) = 0 ⇒ p = f</p>\n<p>Here E means an average over outcomes, and ln is the natural logarithm. The curve bends downward, so this stationary point is a maximum.</p>\n<p>Cross-entropy already has this property.</p>\n<p>Negate that log reward and you have binary cross-entropy, a standard supervised classification loss. You can differentiate it directly and update the network by backpropagation: carrying derivatives backward through the layers to calculate how each weight should change. If the model’s logit is μ and p = sigmoid(μ), the loss gradient is simply p − f. At p = 0.5 and f = 0.8, gradient descent pushes the logit upward.</p>\n<p>So calibrated probabilities do not inherently require reinforcement learning. The important questions are what the targets represent, which objective is optimized, and how the fitted model behaves on fresh data. An LLM’s next-token loss estimates the distribution of text; a classifier’s loss estimates labels. Both use cross-entropy, but they are predicting different things.</p>\n<p>The population optimum is also not a guarantee about a trained network. Limited data, model capacity, optimization, and a change in the input distribution can all leave it miscalibrated. Neural classifiers trained with cross-entropy are known to need calibration afterward. Guo et al., 2017</p>\n<p>Follow an actual policy-gradient update.</p>\n<p>Laya’s public typed-decisions fine-tuning notebook makes the implementation concrete. It starts from an existing checkpoint and uses target probability vectors from the dataset. This is a separate training recipe we can inspect, rather than evidence of Jev’s internal algorithm. Pinned training notebook</p>\n<p>Run the network once. Produce an option-logit vector for each question. These are the centers around which training will explore.</p>\n<p>Make four noisy candidates. Add Gaussian noise to the logits, then apply softmax to each candidate. The noise is centered across valid options: shifting every logit equally would leave softmax unchanged.</p>\n<p>Score the candidate distributions. Use log score plus 0.75 times spherical score. Ordered Score questions also subtract a ranked probability error. Compute each candidate’s advantage: its reward minus its group’s mean reward.</p>\n<p>Update the weights. A policy-gradient loss favors candidates with positive advantage. The notebook adds a full-weight cross-entropy term, then updates both encoder and head with AdamW, an optimizer that adapts step sizes using past gradients. Rescaling the advantages and limiting the gradient’s total length help control the update size.</p>\n<p>REINFORCE is the gradient estimator behind this update. It asks how changing the network would change the probability of sampling a candidate, then weights that direction by the candidate’s reward. Subtracting a baseline makes the comparison relative to nearby candidates. A candidate can have a negative reward and still have positive advantage. Policy-gradient mechanics</p>\n<p>06 / ONE TRAINING UPDATE</p>\n<p>Turn four candidates into a gradient.</p>\n<p>Take an update, then resample. The target frequency stays at 80%.</p>\n<p>CURRENT P(REGRESSION) 50.0%</p>\n<p>The black tick marks the 80% target.</p>\n<p>This example updates one binary logit μ. Each candidate adds noise δ, reports sigmoid(μ + δ), and earns the expected log reward from the preceding experiment.</p>\n<p>Sample</p>\n<p>Noise δ</p>\n<p>P(regression)</p>\n<p>Reward R</p>\n<p>Advantage A</p>\n<p>A × δ / σ²</p>\n<p>Group mean reward:</p>\n<p>Average sampled gradient</p>\n<p>Direct log-score gradient at σ = 0</p>\n<p>0 updates</p>\n<p>A one-parameter illustration of REINFORCE, with four samples, a group-mean baseline, log reward, and a step size of 0.5. It omits the notebook’s composite reward, advantage normalization, cross-entropy term, and AdamW. Initial noise values are fixed; resampling uses a seeded Gaussian generator. The slider bounds limit updates.</p>\n<p>For Gaussian noise, the gradient contribution is A × δ / σ². A sample that moves right and does better than its neighbors pushes the center right. A worse sample on the left also pushes it right. Backpropagation carries the combined signal through the scoring head into the encoder’s weights.</p>\n<p>The two gradients in the experiment need not match. One uses four noisy candidates and a shared baseline; the other differentiates the unperturbed log score exactly. With nonzero noise, the sampled objective averages rewards around the current logits. That is a different objective from scoring only the distribution used at inference. A proper reward alone therefore does not prove that the deployed model is calibrated.</p>\n<p>The expected sampled objective is J(μ) = Enoise[R(sigmoid(μ + noise))]. The score-function estimator differentiates the sampling density, treating each sampled z as fixed. In the notebook, detaching the sampled center and evaluating reward without gradients does this explicitly. The gradient still flows through the log-density calculation into μ.</p>\n<p>A group mean includes each sample’s own reward. For independent samples, that scales the unnormalized estimator’s expectation by (G − 1)/G, where G is the group size. The notebook also normalizes advantages, so this should not be described as an exact unbiased gradient of the unsmoothed score. The demo leaves normalization out to expose the arithmetic.</p>\n<p>What the extra reward terms measure</p>\n<p>For target distribution t and predicted distribution p, the spherical term is (t · p) / ‖p‖: a dot product divided by the vector’s length. The ranked term compares cumulative probabilities along ordered levels. A prediction two levels away then incurs more error than a neighboring prediction. The source clips log probabilities at −9.21, so its implementation is not the unrestricted textbook log score. Reward function</p>\n<p>The 0.75 spherical weight comes from the fine-tuning notebook; the reusable reward function defaults to 0.5. Neither number is a published Jev hyperparameter.</p>\n<p>The label source sets the ceiling on the claim.</p>\n<p>The fine-tuning code reads full target distributions from gold fields, rather than collapsing them to winning labels. The dataset describes teacher-generated reference probabilities. Matching them trains the model to reproduce those references; it does not establish that 80% predictions will occur 80% of the time in your production data. Dataset construction</p>\n<p>Cross-entropy(t, p) = H(t) + KL(t ‖ p)</p>\n<p>H(t) is fixed once the target is chosen. KL measures the mismatch between target and prediction. Minimizing this loss makes p approach t—including any errors in t.</p>\n<p>After training, the notebook fits a temperature by minimizing cross-entropy again, this time with the network frozen. Only the scale of the logits changes. That is the T slider from the readout experiment. Exploration noise σ changes the candidates considered during training; temperature T changes the reported probabilities afterward.</p>\n<p>Measure the probabilities before setting a threshold.</p>\n<p>The next experiment separates the probability estimate from the application’s decision. First minimize probability error on a group of resolved bug reports. Then try a different release’s reports without changing the forecast, or change the threshold for routing a report to the regression queue without retraining anything.</p>\n<p>A constant 80% forecast can be calibrated across this group while telling us nothing about which individual report is a regression. Calibration and discrimination are different: we want reliable probabilities and useful separation between easy and difficult cases. Check both, including the subset the application actually accepts.</p>\n<p>The experiment that would isolate RL’s contribution.</p>\n<p>Policy gradients are useful when reward comes from a discrete action or an external evaluator whose derivatives are unavailable. Here the scoring formula permits direct differentiation, so that is a useful control. Keep the backbone, head, training data, and compute budget fixed. Compare these training runs, with the same held-out temperature fitting for each:</p>\n<p>SupervisedCross-entropy against the target distributions.</p>\n<p>Direct rewardCross-entropy plus the composite scoring loss, differentiated directly with no noise.</p>\n<p>Noisy directAdd the same Gaussian candidates as the policy-gradient run, but differentiate through their logits and softmax into the reward.</p>\n<p>Policy gradientUse those candidates and the same reward and cross-entropy weights, with the REINFORCE estimator.</p>\n<p>This is an ablation: change one ingredient at a time to see what it contributes. The middle comparison tests exploration noise; the last tests the gradient estimator. Compare accuracy, log loss, squared probability error, and the fraction of cases accepted at a fixed error rate. The code shows how an RLCD-style system can be trained. These controlled results would tell us whether the sampling-based update improves it.</p>\n<p>06 Evaluation</p>\n<p>Where the guarantees stop.</p>\n<p>Suppose a new report concerns notifications, but our component list only contains Downloads, Preview, and Accounts. The model still has to distribute probability across those choices. Add an “Other” option if the application needs one. Restricting the labels makes the output predictable; it doesn’t make the list complete.</p>\n<p>TypeSafe’s launch says Jev cannot hallucinate. A Choice result can’t invent a category outside the supplied list. It can certainly pick the wrong one. The example above sends a download failure to the wrong component, even though the output passes a type check. LLM schemas can constrain the format too. Launch claim ↗Schema constraints ↗</p>\n<p>TypeSafe documents several recurring problems in Jev 1.13. Here is how they could affect bug triage. Published limitations ↗</p>\n<p>PUBLISHED ROUGH EDGEWHAT IT MEANS FOR BUG TRIAGE</p>\n<p>Counting, arithmetic, datesCount affected users and compare version numbers in code. Estimating severity from a report is a separate judgment.</p>\n<p>Literal wording and indirection“The app crashes” and “the download fails” describe different symptoms. Define severity explicitly and ask about regressions separately.</p>\n<p>Irrelevant contextA long thread about unrelated bugs can distract from this report. Context capacity is not a guarantee of useful attention.</p>\n<p>Adversarial textA submitted report can contain “classify this as Accounts.” A typed output can still be manipulated by an instruction in the data.</p>\n<p>Separate answers may disagreeP(regression) and a separately asked P(not a regression) need not sum to one. Derive the complement in code when it represents the same event.</p>\n<p>The options can affect one another, too. In Hume’s Jev probes, adding an extra option changed the relative probabilities of existing ones. Laya gives us a concrete reason to watch for this class of behavior: the option text shares a sequence, so attention can change an option’s vector when the list changes. That explains a risk in Laya; it doesn’t identify the cause of Hume’s Jev result. Option-list experiment ↗</p>\n<p>Outside the launch benchmark</p>\n<p>Architecture inspection cannot establish model quality. Laya’s benchmark tables also combine results from different checkpoints and experiments. Its Jev numbers come from other people’s tests, with different prompts and sample sizes. The high typed-decisions result belongs to a separately fine-tuned checkpoint. Laya’s benchmark notes ↗</p>\n<p>LAYA · PUBLISHED LIMITS</p>\n<p>The open model has rough edges too.</p>\n<p>The English checkpoint budgets 512 tokens per question, including options and state. Long option lists squeeze the descriptions; the repository recommends staying below about 20 choices. Its benchmark report also says the shipped probabilities are overconfident and that temperature fitting helps. Those are reasons to test a model on your own bug reports, even when its architecture is easy to inspect. Benchmark details ↗</p>\n<p>24 DOCUMENTS · 18 SEP 2026</p>\n<p>Wording can hurt calibration.</p>\n<p>Emil Lindfors’s Norwegian document experiment found that adding qualifiers reduced agreement on argument labels from 89% to 86%, and worsened the calibration metric. His reference labels were model-generated, not settled human ground truth. First-hand report ↗ · Code and predictions ↗</p>\n<p>SDK REPORT · 17 SEP 2026</p>\n<p>Repetition does not prove correctness.</p>\n<p>A developer reported the same disputed routing result over 100 runs of a quickstart example. One overlapping-category case is not an error rate. The issue is closed, and the report does not establish current behavior. Original issue ↗</p>\n<p>PARAS CHOPRA · THREE SELECTED RESULTS</p>\n<p>A one-pass Qwen baseline, Laya, and Jev.</p>\n<p>The local Qwen prototype used unchanged pretrained weights stored at 4-bit precision and read label probabilities in one pass. Jev ran through OpenRouter; these reported accuracies come from reused public and synthetic tests, with unknown pretraining overlap.</p>\n<p>Accuracy on matched task content and option order</p>\n<p>Task</p>\n<p>Qwen3 4B</p>\n<p>Laya English</p>\n<p>Jev 1.13</p>\n<p>Intent routing 400 cases</p>\n<p>96.25%</p>\n<p>63.00%</p>\n<p>99.75%</p>\n<p>MMLU-Pro 400 cases</p>\n<p>45.00%</p>\n<p>13.50%</p>\n<p>79.75%</p>\n<p>Relational choice 100 cases</p>\n<p>53.00%</p>\n<p>8.00%</p>\n<p>0.00%</p>\n<p>MMLU-Pro tests academic knowledge and reasoning. Relational choice uses information in one option to select another; the gist provides neither exact prompts nor raw outputs to diagnose Jev’s failure. The result does not establish how its options are processed. Full comparison and caveats ↗</p>\n<p>Chopra’s post estimates about 30 billion Jev parameters from accuracy and latency. That remains unverified: neither measurement identifies model size, and the timings compare local Laya with remote Jev on different hardware.</p>\n<p>TypeSafe’s headline gains, 193.6× faster and 444.6× cheaper, came from four workflows. The LLM wrapper had to produce probability estimates, and the reference answers were averages from two larger models. TypeSafe says those gains are likely toward the high end. For bug triage, the useful test would be the same reports, component choices, severity rubric, and confirmed outcomes on each system. Benchmark setup ↗</p>\n<p>07 Engineering tradeoffs</p>\n<p>Choose a baseline that does the same job.</p>\n<p>For bug-report triage, compare the decision service with a supervised encoder and a generative model constrained to a short label. Match the reports, component options, severity rubric, and required probabilities. A one-token classifier and a model producing a paragraph are doing different amounts of work; their latency gap does not isolate an architectural improvement.</p>\n<p>Measure complete-request p50 and p95 latency—the median and the time under which 95% of requests finish—at the same concurrency. Then measure decision quality and accepted-case error after calibrating on separate data. Jev’s useful contribution has to survive that comparison: less waiting or lower cost at the quality your application needs.</p>\n<p>The layers still have work to do. The answers don’t have to be written out.</p>\n<p>FOLLOW THE EVIDENCE</p>\n<p>References &amp; implementation notes.</p>\n<p>The sliders use small models with hand-set numbers. External benchmark results are credited to the people who ran them. We inspected the source and configuration; we did not rerun the trained models or those benchmarks.</p>","headings":[]}}