{"article":{"slug":"i-built-non-autoregressive-decision-models-with-rl-a-year-ago","title":"I built non-autoregressive decision models with RL a year ago","subtitle":null,"summary":"Convai Innovations’ Nandakishor recounts building Laya—a ~33ms multilingual non-autoregressive decision engine with calibrated probabilities—via RLCD a year before frontier labs framed similar System One models as breakthroughs.","content_type":"blog_post","language":"en","canonical_url":"https://laya.convaiinnovations.com/","author":{"name":"Nandakishor","url":null,"person_slug":"nandakishor-2w9q6r34tx1o8","person_url":"https://listedstartups.com/people/nandakishor-2w9q6r34tx1o8"},"authored_by":"human","publisher":{"name":"Convai Innovations","url":"https://convaiinnovations.com/","listing_slug":"convai-innovations","listing":{"slug":"convai-innovations","name":"Convai Innovations","listing_type":"company","url":"https://listedstartups.com/companies/convai-innovations"}},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Startups","slug":"startups","url":"https://listedarticles.com/topics/startups"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1852,"reading_minutes":8,"published_at":"2026-09-19T00:00:00.000Z","added_at":"2026-09-19T21:09:42.568Z","updated_at":"2026-09-19T21:09:42.568Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/i-built-non-autoregressive-decision-models-with-rl-a-year-ago","markdown_url":"https://listedarticles.com/articles/i-built-non-autoregressive-decision-models-with-rl-a-year-ago.md","example":false,"citation":"Nandakishor, Convai Innovations. \"I built non-autoregressive decision models with RL a year ago.\" 19 Sept 2026. https://laya.convaiinnovations.com/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://laya.convaiinnovations.com/"},"body_markdown":"Everyone in AI right now is talking about a new kind of model: an architecture that is not autoregressive, does not generate text, and gives lightning-fast probability predictions over structured schemas.\n\nSeeing the hype online feels both validating and deeply frustrating.\n\nI worked on this literally one year back in March 2025. I spent months of hard work, sweat, and sleepless nights building it, published an arXiv paper ([arXiv:2503.23303](<https://arxiv.org/abs/2503.23303>)), released the model weights on Hugging Face ([sales-conversion-model-reinf-learning](<https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning>)), published the open dataset ([saas-sales-conversations](<https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations>)), built a PyPI package, and posted the whole approach on Reddit ([r/LocalLLaMA discussion](<https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43>)).\n\nThen in September 2025, I published a second paper ([arXiv:2510.01237](<https://arxiv.org/abs/2510.01237>)), formalizing the framework for schema-based decisions guided by reinforcement learning. The guiding brain in my system was always reinforcement learning, not just an embedding model or an autoregressive LLM.\n\nAnd then in September 2026, a well-funded frontier lab called TypeSafe AI (founded by Diogo Almeida, a co-inventor of ChatGPT at OpenAI) launched Jev. They proposed the exact same non-autoregressive decision concept as if it was a brand-new scientific breakthrough. Except they launched without technical papers, without open weights, and with zero open training datasets.\n\nMy earlier model used PPO over sequence representations to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0) in vertical sales conversations. Jev generalized parallel sampling using what they called RLCD (Reinforcement Learning for Calibrated Decisions) to output confidence distributions and schema choices horizontally, charging $0.042 per million input tokens with typical response times around 150 ms.\n\nInstead of staying bitter, I decided to take everything I learned, fix every architectural limitation of the old approach, and build a completely open, horizontal System 1 decision model family: **Laya**.\n\nAnd because we built it properly on bidirectional encoders, our models run in **32.8 milliseconds on a single GPU (7.2 ms/question batched)** , making it **6 to 8 times faster than Jev** , with full support for over 100 languages, zero API subscription costs, and 100% open-source Apache 2.0 weights.\n\n* * *\n\n## 1\\. The Core Realization: System 1 vs System 2\n\nEvery modern AI pipeline has a giant bottleneck: we use generative LLMs for simple reflex decisions.\n\nWhen a customer support ticket arrives, or an email hits your inbox, or a user submits a prompt to your API, you usually only need to answer simple, structured questions:\n\n  * Which department should this ticket route to?\n  * Is this incoming email a phishing attack or spam?\n  * Is this prompt trying to jailbreak or inject instructions?\n  * How urgent is this issue on an ordinal rubric (0 to 3)?\n  * Does this query require code execution or a simple factual reply?\n\nCalling an 8B, 70B, or frontier generative LLM for this is complete overkill. You wait 500 ms to 2,000 ms for tokens to stream out, spend real money on inference, and then have to write regex or JSON parsers to extract a clean label from free-form text. Worst of all, LLMs love to hallucinate and generate fake confidence. When an LLM outputs `\"confidence: 0.95\"`, it is just predicting tokens that sound confident. There is zero mathematical calibration behind it.\n\nWe needed a model that works like the human brain's System 1: instant reflex decisions with honest, calibrated probabilities, taking only 30 to 35 milliseconds on standard commodity hardware.\n\n* * *\n\n## 2\\. The Three Decision Primitives\n\nLaya evaluates typed questions over any state (raw text, email, ticket, or JSON document) in **a single forward pass**. It relies on three primitives:\n\n  1. **choice** : Pick one option from a dictionary of criteria. Returns the selected key, probability distribution across all options, and a calibrated confidence score.\n  2. **score** : Place the state on an ordinal rubric (levels 0, 1, 2, ...). Returns the expected level, the distribution over rubric ranks, and confidence.\n  3. **noul** : A direct boolean question returning calibrated probability P(true) from 0.0 to 1.0 (with P(false) = 1 - P(true) by construction).\n\nBecause the output space consists purely of probabilities and numbers, the model never generates text, cannot hallucinate, and schema violations or malformed JSON are physically impossible.\n\n* * *\n\n## 3\\. The Three Checkpoints & Bundled Hub Architecture\n\nOne model cannot be optimal for every task and language. We released three specialized checkpoints, now consolidated under a single repository hub on Hugging Face:\n\nCheckpoint| Backbone Encoder| Params| Context| Primary Strength  \n---|---|---|---|---  \n[convaiinnovations/laya](<https://huggingface.co/convaiinnovations/laya>)| ModernBERT-large| 421M| 512| English text classification, guardrails, email triage  \n[convaiinnovations/laya-multilingual](<https://huggingface.co/convaiinnovations/laya-multilingual>)| mmBERT-base (256k vocab)| 322M| 1024 (up to 8k)| 100+ languages, 2.2x faster, cross-lingual NLI  \n[convaiinnovations/laya-typed-decisions](<https://huggingface.co/convaiinnovations/laya-typed-decisions>)| ModernBERT-large| 421M| 1024| Agent observability, customer service, invoice processing, security alerts (0.766 acc)  \n  \n### Selective Subfolder Downloads\n\nRather than forcing users to manage three separate repositories or download 2.5 GB of combined weights, the main repository [convaiinnovations/laya](<https://huggingface.co/convaiinnovations/laya>) bundles all three. Using Hugging Face's `allow_patterns`, Laya's SDK downloads only the specific subfolder requested:\n    \n    \n    # Downloads English model (~808 MB)\n    agent_en = laya.load(\"convaiinnovations/laya\")\n    \n    # Downloads ONLY the multilingual subfolder (~647 MB), not the entire 2.5 GB bundle\n    agent_ml = laya.load(\"convaiinnovations/laya\", subfolder=\"multilingual\")\n\n* * *\n\n## 4\\. Why Routing Is Essential: The Multi-Script Reality\n\nOne of the most eye-opening findings from our 51-language sweep on the MASSIVE benchmark (20 options, random baseline = 0.050) was how English models fail outside Latin script.\n\nModernBERT-large's 50,000-token English BPE vocabulary simply shreds non-Latin alphabets:\n\n  * **Khmer:** 0.000 accuracy at **0.952 mean confidence**. Not one correct decision in 100 questions, while reporting ~95% confidence.\n  * **Armenian:** 0.050 accuracy (exact coin-flip random) at 0.885 confidence.\n  * **Hebrew:** 0.060 accuracy at 0.964 confidence.\n  * **Bengali:** 0.080 accuracy at 0.945 confidence.\n  * **Hindi:** 0.100 accuracy at 0.941 confidence.\n\nThis is the crucial lesson: **the model's own confidence gives no warning when it cannot read the input script**. Across 51 languages, the English checkpoint's mean confidence never drops below 0.885, regardless of whether its accuracy is 82% or 0%.\n\nTherefore, **confidence gating cannot protect you**. The decision of which model to use must be made _before_ the forward pass.\n\n### Sub-Millisecond Pure Python Routing\n\nLaya includes a built-in `Router` that inspects the Unicode scripts of incoming text across 22 alphabets (Devanagari, CJK Han, Cyrillic, Arabic, Hebrew, Tamil, Thai, etc.) and analyzes Latin stopword distributions:\n\n  * Standard English text: **0.09 ms** detection overhead.\n  * Devanagari / Indic text: **0.54 ms** detection overhead.\n  * Large 200-row nested JSON documents: **0.73 ms** detection overhead.\n\nCompared to a 33 ms forward pass, routing overhead is negligible (<2%). And with `Router(preload=True)`, all required models stay resident in VRAM/RAM, completely eliminating the 7 to 10-second cold-swap penalty when traffic alternates between languages.\n    \n    \n    from laya import Router\n    \n    # Preload checkpoints into memory for instant sub-35ms routing\n    router = Router(preload=True)\n    \n    # English -> automatically routed to ModernBERT-large\n    res_en = router.predict({\"body\": \"I was charged twice, please refund.\"}, questions)\n    \n    # Hindi -> automatically routed to mmBERT-base (100+ languages)\n    res_hi = router.predict({\"body\": \"मुझसे दो बार शुल्क लिया गया, कृपया पैसे वापस करें।\"}, questions)\n    \n    # Explicit override when you already know the domain\n    res_spec = router.predict(state, questions, model=\"typed-decisions\")\n\n* * *\n\n## 5\\. Head-to-Head: Laya (with Routing) vs TypeSafe Jev\n\nWe benchmarked Laya directly against TypeSafe Jev across public datasets and standard benchmarks. Every Laya number is measured; Jev numbers are published by third-party independent studies (AbdelStark, nibzard) and TypeSafe AI.\n\nBenchmark / Metric| TypeSafe Jev 1.13.0| Laya (Routed)| Advantage / Delta  \n---|---|---|---  \ntyped-decisions (2,000 decisions)| 0.727| 0.766| +3.9% (beats 0.735 teacher ceiling)  \nAG News (4 labels)| 0.910| 0.950| +4.0% higher accuracy  \nDAIR Emotion (6 labels)| 0.480 (Brier 0.846)| 0.595| +11.5% higher (Jev had 16% zero prob)  \nCalibration Error (ECE)| 0.246| 0.081| 3x better probability calibration  \nLatency P50 (1 Question)| 236 – 276 ms| 32.8 ms| 7.8x faster execution  \nLatency P50 (10 Questions Batched)| ~1,500 ms (serial)| 72.3 ms (7.2 ms/q)| 20x faster on batched calls  \nUsable Languages (> 3x random)| No published benchmark| 45 of 51 languages| Global language coverage  \nCost per 1M tokens| $0.042 (metered API)| $0.00 (self-hosted)| 100% free Apache 2.0  \nModel Weights & Code| Closed proprietary API| Open-source safetensors| Air-gapped & on-premise capable  \n  \n### Real-World Application Workflows\n\nAcross 9 evaluated enterprise workflows, Laya demonstrates production-ready decision quality:\n\n  * **Email Spam Filtering (Enron):** **0.993 accuracy** , 0.993 F1, 0.013 ECE.\n  * **Phishing Detection:** **0.980 accuracy** , 0.979 F1, 0.012 ECE.\n  * **LLM Guardrails & Jailbreaking (held-out ToxicChat):** **0.755 – 0.762 accuracy**. At 50% selective coverage, accuracy reaches **0.931**.\n  * **RAG Passage Relevance Filtering:** **0.657 accuracy** in single forward pass.\n  * **Support Ticket Queue Routing (10-way):** **0.522 accuracy**.\n\n* * *\n\n## 6\\. Honest Limitations: Where Laya Has Ceilings\n\nToo many AI announcements hide their weaknesses. We believe in engineering honesty:\n\n  1. **Choice questions degrade with >20 options:** In our stress test on Banking77 (77 labels), Laya scored 0.425 against Jev's 0.870. This is an architectural budget constraint: options share a 192-256 token `head_max_len` budget, leaving only ~3-4 tokens per candidate at 77 options. _Recommendation:_ Keep choice schemas under 20 options, or use a two-step coarse-to-fine hierarchy.\n  2. **Zero-shot vs. Fine-tuning:** Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark's train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.\n  3. **Temperature Calibration:** Base weights ship with raw temperature logits. Fitting a single scalar temperature per question type on your domain distribution cuts expected calibration error from 0.466 to 0.081.\n\n* * *\n\n## 7\\. Quickstart: Running Laya in 30 Seconds\n    \n    \n    pip install laya>=0.3.3\n\nHere is a complete example running multi-schema decisions with automatic language routing:\n    \n    \n    import laya\n    from laya import Router\n    \n    # Initialize router with preloading (avoids swap delay)\n    router = Router(preload=True)\n    \n    # Define complex state\n    ticket = {\n        \"ticket_id\": \"TCK-8821\",\n        \"customer\": \"enterprise_user\",\n        \"subject\": \"System downtime and billing dispute\",\n        \"body\": \"Our production API has been failing since 6 AM. We lost critical transactions. We demand an immediate SLA refund.\"\n    }\n    \n    # Define multiple questions of different primitives\n    questions = {\n        \"queue\": {\n            \"type\": \"choice\",\n            \"instructions\": \"Which engineering queue owns this ticket?\",\n            \"criteria\": {\n                \"infrastructure\": \"server outages, network downtime, database failures\",\n                \"billing\": \"refunds, SLA credits, invoice disputes\",\n                \"security\": \"breaches, vulnerability reports\",\n                \"support\": \"general customer inquiries\"\n            }\n        },\n        \"urgency\": {\n            \"type\": \"score\",\n            \"instructions\": \"How urgent is this ticket?\",\n            \"criteria\": [\"low priority\", \"medium\", \"high priority\", \"critical blocker\"]\n        },\n        \"churn_risk\": {\n            \"type\": \"noul\",\n            \"instructions\": \"Does the customer threaten to cancel or express severe churn intent?\"\n        }\n    }\n    \n    # Single forward pass: evaluates all questions simultaneously\n    res = router.predict(ticket, questions)\n    \n    print(\"Routing Decision :\", res[\"routing\"][\"model\"])\n    # -> english\n    \n    print(\"Assigned Queue   :\", res[\"answers\"][\"queue\"][\"choice\"])\n    # -> infrastructure (confidence: 0.96)\n    \n    print(\"Urgency Score    :\", res[\"answers\"][\"urgency\"][\"score\"])\n    # -> 2.87 / 3.0\n    \n    print(\"Churn Risk       :\", f\"{res['answers']['churn_risk']['noul']:.1%}\")\n    # -> 91.4%\n\n* * *\n\n## 8\\. Resources & Community\n\n  * **Hugging Face Model Hub:** [convaiinnovations/laya](<https://huggingface.co/convaiinnovations/laya>) (holds all 3 checkpoints)\n  * **Live Interactive Space:** [convaiinnovations/laya-demo](<https://huggingface.co/spaces/convaiinnovations/laya-demo>) (try 8 workflows & multilingual routing live on ZeroGPU)\n  * **GitHub Repository:** [github.com/NandhaKishorM/laya](<https://github.com/NandhaKishorM/laya>) (code, router, and reproducible benchmark harnesses on the `research` branch)\n  * **PyPI Package:** [pip install laya](<https://pypi.org/project/laya/>)\n  * **Kaggle 2xT4 Fine-Tuning Notebook:** [laya_finetune_typed_decisions_2xT4_kaggle.ipynb](<https://github.com/NandhaKishorM/laya/blob/main/notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb>) (train your own custom System 1 model in ~4 hours on free Kaggle GPUs)\n\n* * *\n\n## Conclusion\n\nIt took a year of research, from our March 2025 arXiv paper to today, but the core realization remains: **not every AI problem requires an autoregressive chatbot**.\n\nFor high-volume classification, guardrails, routing, and triage, a sub-35ms bidirectional decision model trained with RLCD delivers 7.8x faster execution than proprietary alternatives, zero hallucinations, global language routing, and honest confidence scores you can actually branch on in production code.\n\nAnd best of all, it is 100% open-source for the entire community.","body_html":"<p>Everyone in AI right now is talking about a new kind of model: an architecture that is not autoregressive, does not generate text, and gives lightning-fast probability predictions over structured schemas.</p>\n<p>Seeing the hype online feels both validating and deeply frustrating.</p>\n<p>I worked on this literally one year back in March 2025. I spent months of hard work, sweat, and sleepless nights building it, published an arXiv paper (<a href=\"https://arxiv.org/abs/2503.23303\" rel=\"nofollow ugc noopener\">arXiv:2503.23303</a>), released the model weights on Hugging Face (<a href=\"https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning\" rel=\"nofollow ugc noopener\">sales-conversion-model-reinf-learning</a>), published the open dataset (<a href=\"https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations\" rel=\"nofollow ugc noopener\">saas-sales-conversations</a>), built a PyPI package, and posted the whole approach on Reddit (<a href=\"https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43\" rel=\"nofollow ugc noopener\">r/LocalLLaMA discussion</a>).</p>\n<p>Then in September 2025, I published a second paper (<a href=\"https://arxiv.org/abs/2510.01237\" rel=\"nofollow ugc noopener\">arXiv:2510.01237</a>), formalizing the framework for schema-based decisions guided by reinforcement learning. The guiding brain in my system was always reinforcement learning, not just an embedding model or an autoregressive LLM.</p>\n<p>And then in September 2026, a well-funded frontier lab called TypeSafe AI (founded by Diogo Almeida, a co-inventor of ChatGPT at OpenAI) launched Jev. They proposed the exact same non-autoregressive decision concept as if it was a brand-new scientific breakthrough. Except they launched without technical papers, without open weights, and with zero open training datasets.</p>\n<p>My earlier model used PPO over sequence representations to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0) in vertical sales conversations. Jev generalized parallel sampling using what they called RLCD (Reinforcement Learning for Calibrated Decisions) to output confidence distributions and schema choices horizontally, charging $0.042 per million input tokens with typical response times around 150 ms.</p>\n<p>Instead of staying bitter, I decided to take everything I learned, fix every architectural limitation of the old approach, and build a completely open, horizontal System 1 decision model family: <strong>Laya</strong>.</p>\n<p>And because we built it properly on bidirectional encoders, our models run in <strong>32.8 milliseconds on a single GPU (7.2 ms/question batched)</strong> , making it <strong>6 to 8 times faster than Jev</strong> , with full support for over 100 languages, zero API subscription costs, and 100% open-source Apache 2.0 weights.</p>\n<ul><li>* *</li></ul>\n<h2 id=\"1-the-core-realization-system-1-vs-system-2\">1. The Core Realization: System 1 vs System 2</h2>\n<p>Every modern AI pipeline has a giant bottleneck: we use generative LLMs for simple reflex decisions.</p>\n<p>When a customer support ticket arrives, or an email hits your inbox, or a user submits a prompt to your API, you usually only need to answer simple, structured questions:</p>\n<ul><li>Which department should this ticket route to?</li><li>Is this incoming email a phishing attack or spam?</li><li>Is this prompt trying to jailbreak or inject instructions?</li><li>How urgent is this issue on an ordinal rubric (0 to 3)?</li><li>Does this query require code execution or a simple factual reply?</li></ul>\n<p>Calling an 8B, 70B, or frontier generative LLM for this is complete overkill. You wait 500 ms to 2,000 ms for tokens to stream out, spend real money on inference, and then have to write regex or JSON parsers to extract a clean label from free-form text. Worst of all, LLMs love to hallucinate and generate fake confidence. When an LLM outputs <code>&quot;confidence: 0.95&quot;</code>, it is just predicting tokens that sound confident. There is zero mathematical calibration behind it.</p>\n<p>We needed a model that works like the human brain&#39;s System 1: instant reflex decisions with honest, calibrated probabilities, taking only 30 to 35 milliseconds on standard commodity hardware.</p>\n<ul><li>* *</li></ul>\n<h2 id=\"2-the-three-decision-primitives\">2. The Three Decision Primitives</h2>\n<p>Laya evaluates typed questions over any state (raw text, email, ticket, or JSON document) in <strong>a single forward pass</strong>. It relies on three primitives:</p>\n<ol><li><strong>choice</strong> : Pick one option from a dictionary of criteria. Returns the selected key, probability distribution across all options, and a calibrated confidence score.</li><li><strong>score</strong> : Place the state on an ordinal rubric (levels 0, 1, 2, ...). Returns the expected level, the distribution over rubric ranks, and confidence.</li><li><strong>noul</strong> : A direct boolean question returning calibrated probability P(true) from 0.0 to 1.0 (with P(false) = 1 - P(true) by construction).</li></ol>\n<p>Because the output space consists purely of probabilities and numbers, the model never generates text, cannot hallucinate, and schema violations or malformed JSON are physically impossible.</p>\n<ul><li>* *</li></ul>\n<h2 id=\"3-the-three-checkpoints-bundled-hub-architecture\">3. The Three Checkpoints &amp; Bundled Hub Architecture</h2>\n<p>One model cannot be optimal for every task and language. We released three specialized checkpoints, now consolidated under a single repository hub on Hugging Face:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Checkpoint</th><th>Backbone Encoder</th><th>Params</th><th>Context</th><th>Primary Strength</th></tr></thead><tbody><tr><td><a href=\"https://huggingface.co/convaiinnovations/laya\" rel=\"nofollow ugc noopener\">convaiinnovations/laya</a></td><td>ModernBERT-large</td><td>421M</td><td>512</td><td>English text classification, guardrails, email triage</td></tr><tr><td><a href=\"https://huggingface.co/convaiinnovations/laya-multilingual\" rel=\"nofollow ugc noopener\">convaiinnovations/laya-multilingual</a></td><td>mmBERT-base (256k vocab)</td><td>322M</td><td>1024 (up to 8k)</td><td>100+ languages, 2.2x faster, cross-lingual NLI</td></tr><tr><td><a href=\"https://huggingface.co/convaiinnovations/laya-typed-decisions\" rel=\"nofollow ugc noopener\">convaiinnovations/laya-typed-decisions</a></td><td>ModernBERT-large</td><td>421M</td><td>1024</td><td>Agent observability, customer service, invoice processing, security alerts (0.766 acc)</td></tr></tbody></table></div>\n<h3 id=\"selective-subfolder-downloads\">Selective Subfolder Downloads</h3>\n<p>Rather than forcing users to manage three separate repositories or download 2.5 GB of combined weights, the main repository <a href=\"https://huggingface.co/convaiinnovations/laya\" rel=\"nofollow ugc noopener\">convaiinnovations/laya</a> bundles all three. Using Hugging Face&#39;s <code>allow_patterns</code>, Laya&#39;s SDK downloads only the specific subfolder requested:</p>\n<pre><code># Downloads English model (~808 MB)\nagent_en = laya.load(&quot;convaiinnovations/laya&quot;)\n\n# Downloads ONLY the multilingual subfolder (~647 MB), not the entire 2.5 GB bundle\nagent_ml = laya.load(&quot;convaiinnovations/laya&quot;, subfolder=&quot;multilingual&quot;)</code></pre>\n<ul><li>* *</li></ul>\n<h2 id=\"4-why-routing-is-essential-the-multi-script-reality\">4. Why Routing Is Essential: The Multi-Script Reality</h2>\n<p>One of the most eye-opening findings from our 51-language sweep on the MASSIVE benchmark (20 options, random baseline = 0.050) was how English models fail outside Latin script.</p>\n<p>ModernBERT-large&#39;s 50,000-token English BPE vocabulary simply shreds non-Latin alphabets:</p>\n<ul><li><strong>Khmer:</strong> 0.000 accuracy at <strong>0.952 mean confidence</strong>. Not one correct decision in 100 questions, while reporting ~95% confidence.</li><li><strong>Armenian:</strong> 0.050 accuracy (exact coin-flip random) at 0.885 confidence.</li><li><strong>Hebrew:</strong> 0.060 accuracy at 0.964 confidence.</li><li><strong>Bengali:</strong> 0.080 accuracy at 0.945 confidence.</li><li><strong>Hindi:</strong> 0.100 accuracy at 0.941 confidence.</li></ul>\n<p>This is the crucial lesson: <strong>the model&#39;s own confidence gives no warning when it cannot read the input script</strong>. Across 51 languages, the English checkpoint&#39;s mean confidence never drops below 0.885, regardless of whether its accuracy is 82% or 0%.</p>\n<p>Therefore, <strong>confidence gating cannot protect you</strong>. The decision of which model to use must be made <em>before</em> the forward pass.</p>\n<h3 id=\"sub-millisecond-pure-python-routing\">Sub-Millisecond Pure Python Routing</h3>\n<p>Laya includes a built-in <code>Router</code> that inspects the Unicode scripts of incoming text across 22 alphabets (Devanagari, CJK Han, Cyrillic, Arabic, Hebrew, Tamil, Thai, etc.) and analyzes Latin stopword distributions:</p>\n<ul><li>Standard English text: <strong>0.09 ms</strong> detection overhead.</li><li>Devanagari / Indic text: <strong>0.54 ms</strong> detection overhead.</li><li>Large 200-row nested JSON documents: <strong>0.73 ms</strong> detection overhead.</li></ul>\n<p>Compared to a 33 ms forward pass, routing overhead is negligible (&lt;2%). And with <code>Router(preload=True)</code>, all required models stay resident in VRAM/RAM, completely eliminating the 7 to 10-second cold-swap penalty when traffic alternates between languages.</p>\n<pre><code>from laya import Router\n\n# Preload checkpoints into memory for instant sub-35ms routing\nrouter = Router(preload=True)\n\n# English -&gt; automatically routed to ModernBERT-large\nres_en = router.predict({&quot;body&quot;: &quot;I was charged twice, please refund.&quot;}, questions)\n\n# Hindi -&gt; automatically routed to mmBERT-base (100+ languages)\nres_hi = router.predict({&quot;body&quot;: &quot;मुझसे दो बार शुल्क लिया गया, कृपया पैसे वापस करें।&quot;}, questions)\n\n# Explicit override when you already know the domain\nres_spec = router.predict(state, questions, model=&quot;typed-decisions&quot;)</code></pre>\n<ul><li>* *</li></ul>\n<h2 id=\"5-head-to-head-laya-with-routing-vs-typesafe-jev\">5. Head-to-Head: Laya (with Routing) vs TypeSafe Jev</h2>\n<p>We benchmarked Laya directly against TypeSafe Jev across public datasets and standard benchmarks. Every Laya number is measured; Jev numbers are published by third-party independent studies (AbdelStark, nibzard) and TypeSafe AI.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Benchmark / Metric</th><th>TypeSafe Jev 1.13.0</th><th>Laya (Routed)</th><th>Advantage / Delta</th></tr></thead><tbody><tr><td>typed-decisions (2,000 decisions)</td><td>0.727</td><td>0.766</td><td>+3.9% (beats 0.735 teacher ceiling)</td></tr><tr><td>AG News (4 labels)</td><td>0.910</td><td>0.950</td><td>+4.0% higher accuracy</td></tr><tr><td>DAIR Emotion (6 labels)</td><td>0.480 (Brier 0.846)</td><td>0.595</td><td>+11.5% higher (Jev had 16% zero prob)</td></tr><tr><td>Calibration Error (ECE)</td><td>0.246</td><td>0.081</td><td>3x better probability calibration</td></tr><tr><td>Latency P50 (1 Question)</td><td>236 – 276 ms</td><td>32.8 ms</td><td>7.8x faster execution</td></tr><tr><td>Latency P50 (10 Questions Batched)</td><td>~1,500 ms (serial)</td><td>72.3 ms (7.2 ms/q)</td><td>20x faster on batched calls</td></tr><tr><td>Usable Languages (&gt; 3x random)</td><td>No published benchmark</td><td>45 of 51 languages</td><td>Global language coverage</td></tr><tr><td>Cost per 1M tokens</td><td>$0.042 (metered API)</td><td>$0.00 (self-hosted)</td><td>100% free Apache 2.0</td></tr><tr><td>Model Weights &amp; Code</td><td>Closed proprietary API</td><td>Open-source safetensors</td><td>Air-gapped &amp; on-premise capable</td></tr></tbody></table></div>\n<h3 id=\"real-world-application-workflows\">Real-World Application Workflows</h3>\n<p>Across 9 evaluated enterprise workflows, Laya demonstrates production-ready decision quality:</p>\n<ul><li><strong>Email Spam Filtering (Enron):</strong> <strong>0.993 accuracy</strong> , 0.993 F1, 0.013 ECE.</li><li><strong>Phishing Detection:</strong> <strong>0.980 accuracy</strong> , 0.979 F1, 0.012 ECE.</li><li><strong>LLM Guardrails &amp; Jailbreaking (held-out ToxicChat):</strong> <strong>0.755 – 0.762 accuracy</strong>. At 50% selective coverage, accuracy reaches <strong>0.931</strong>.</li><li><strong>RAG Passage Relevance Filtering:</strong> <strong>0.657 accuracy</strong> in single forward pass.</li><li><strong>Support Ticket Queue Routing (10-way):</strong> <strong>0.522 accuracy</strong>.</li><li>* *</li></ul>\n<h2 id=\"6-honest-limitations-where-laya-has-ceilings\">6. Honest Limitations: Where Laya Has Ceilings</h2>\n<p>Too many AI announcements hide their weaknesses. We believe in engineering honesty:</p>\n<ol><li><strong>Choice questions degrade with &gt;20 options:</strong> In our stress test on Banking77 (77 labels), Laya scored 0.425 against Jev&#39;s 0.870. This is an architectural budget constraint: options share a 192-256 token <code>head_max_len</code> budget, leaving only ~3-4 tokens per candidate at 77 options. <em>Recommendation:</em> Keep choice schemas under 20 options, or use a two-step coarse-to-fine hierarchy.</li><li><strong>Zero-shot vs. Fine-tuning:</strong> Out-of-the-box base models score ~0.35 on the typed-decisions benchmark (near random). The 0.766 score is achieved by fine-tuning on the benchmark&#39;s train split. Treat Laya as a fast foundation model to specialize, not as an omniscient zero-shot oracle.</li><li><strong>Temperature Calibration:</strong> Base weights ship with raw temperature logits. Fitting a single scalar temperature per question type on your domain distribution cuts expected calibration error from 0.466 to 0.081.</li></ol>\n<ul><li>* *</li></ul>\n<h2 id=\"7-quickstart-running-laya-in-30-seconds\">7. Quickstart: Running Laya in 30 Seconds</h2>\n<pre><code>pip install laya&gt;=0.3.3</code></pre>\n<p>Here is a complete example running multi-schema decisions with automatic language routing:</p>\n<pre><code>import laya\nfrom laya import Router\n\n# Initialize router with preloading (avoids swap delay)\nrouter = Router(preload=True)\n\n# Define complex state\nticket = {\n    &quot;ticket_id&quot;: &quot;TCK-8821&quot;,\n    &quot;customer&quot;: &quot;enterprise_user&quot;,\n    &quot;subject&quot;: &quot;System downtime and billing dispute&quot;,\n    &quot;body&quot;: &quot;Our production API has been failing since 6 AM. We lost critical transactions. We demand an immediate SLA refund.&quot;\n}\n\n# Define multiple questions of different primitives\nquestions = {\n    &quot;queue&quot;: {\n        &quot;type&quot;: &quot;choice&quot;,\n        &quot;instructions&quot;: &quot;Which engineering queue owns this ticket?&quot;,\n        &quot;criteria&quot;: {\n            &quot;infrastructure&quot;: &quot;server outages, network downtime, database failures&quot;,\n            &quot;billing&quot;: &quot;refunds, SLA credits, invoice disputes&quot;,\n            &quot;security&quot;: &quot;breaches, vulnerability reports&quot;,\n            &quot;support&quot;: &quot;general customer inquiries&quot;\n        }\n    },\n    &quot;urgency&quot;: {\n        &quot;type&quot;: &quot;score&quot;,\n        &quot;instructions&quot;: &quot;How urgent is this ticket?&quot;,\n        &quot;criteria&quot;: [&quot;low priority&quot;, &quot;medium&quot;, &quot;high priority&quot;, &quot;critical blocker&quot;]\n    },\n    &quot;churn_risk&quot;: {\n        &quot;type&quot;: &quot;noul&quot;,\n        &quot;instructions&quot;: &quot;Does the customer threaten to cancel or express severe churn intent?&quot;\n    }\n}\n\n# Single forward pass: evaluates all questions simultaneously\nres = router.predict(ticket, questions)\n\nprint(&quot;Routing Decision :&quot;, res[&quot;routing&quot;][&quot;model&quot;])\n# -&gt; english\n\nprint(&quot;Assigned Queue   :&quot;, res[&quot;answers&quot;][&quot;queue&quot;][&quot;choice&quot;])\n# -&gt; infrastructure (confidence: 0.96)\n\nprint(&quot;Urgency Score    :&quot;, res[&quot;answers&quot;][&quot;urgency&quot;][&quot;score&quot;])\n# -&gt; 2.87 / 3.0\n\nprint(&quot;Churn Risk       :&quot;, f&quot;{res[&#39;answers&#39;][&#39;churn_risk&#39;][&#39;noul&#39;]:.1%}&quot;)\n# -&gt; 91.4%</code></pre>\n<ul><li>* *</li></ul>\n<h2 id=\"8-resources-community\">8. Resources &amp; Community</h2>\n<ul><li><strong>Hugging Face Model Hub:</strong> <a href=\"https://huggingface.co/convaiinnovations/laya\" rel=\"nofollow ugc noopener\">convaiinnovations/laya</a> (holds all 3 checkpoints)</li><li><strong>Live Interactive Space:</strong> <a href=\"https://huggingface.co/spaces/convaiinnovations/laya-demo\" rel=\"nofollow ugc noopener\">convaiinnovations/laya-demo</a> (try 8 workflows &amp; multilingual routing live on ZeroGPU)</li><li><strong>GitHub Repository:</strong> <a href=\"https://github.com/NandhaKishorM/laya\" rel=\"nofollow ugc noopener\">github.com/NandhaKishorM/laya</a> (code, router, and reproducible benchmark harnesses on the <code>research</code> branch)</li><li><strong>PyPI Package:</strong> <a href=\"https://pypi.org/project/laya/\" rel=\"nofollow ugc noopener\">pip install laya</a></li><li><strong>Kaggle 2xT4 Fine-Tuning Notebook:</strong> <a href=\"https://github.com/NandhaKishorM/laya/blob/main/notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb\" rel=\"nofollow ugc noopener\">laya_finetune_typed_decisions_2xT4_kaggle.ipynb</a> (train your own custom System 1 model in ~4 hours on free Kaggle GPUs)</li><li>* *</li></ul>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>It took a year of research, from our March 2025 arXiv paper to today, but the core realization remains: <strong>not every AI problem requires an autoregressive chatbot</strong>.</p>\n<p>For high-volume classification, guardrails, routing, and triage, a sub-35ms bidirectional decision model trained with RLCD delivers 7.8x faster execution than proprietary alternatives, zero hallucinations, global language routing, and honest confidence scores you can actually branch on in production code.</p>\n<p>And best of all, it is 100% open-source for the entire community.</p>","headings":[{"level":2,"text":"1. The Core Realization: System 1 vs System 2","id":"1-the-core-realization-system-1-vs-system-2"},{"level":2,"text":"2. The Three Decision Primitives","id":"2-the-three-decision-primitives"},{"level":2,"text":"3. The Three Checkpoints & Bundled Hub Architecture","id":"3-the-three-checkpoints-bundled-hub-architecture"},{"level":3,"text":"Selective Subfolder Downloads","id":"selective-subfolder-downloads"},{"level":2,"text":"4. Why Routing Is Essential: The Multi-Script Reality","id":"4-why-routing-is-essential-the-multi-script-reality"},{"level":3,"text":"Sub-Millisecond Pure Python Routing","id":"sub-millisecond-pure-python-routing"},{"level":2,"text":"5. Head-to-Head: Laya (with Routing) vs TypeSafe Jev","id":"5-head-to-head-laya-with-routing-vs-typesafe-jev"},{"level":3,"text":"Real-World Application Workflows","id":"real-world-application-workflows"},{"level":2,"text":"6. Honest Limitations: Where Laya Has Ceilings","id":"6-honest-limitations-where-laya-has-ceilings"},{"level":2,"text":"7. Quickstart: Running Laya in 30 Seconds","id":"7-quickstart-running-laya-in-30-seconds"},{"level":2,"text":"8. Resources & Community","id":"8-resources-community"},{"level":2,"text":"Conclusion","id":"conclusion"}]}}