{"article":{"slug":"new-in-llama-cpp-decision-models","title":"New in llama.cpp: Decision Models","subtitle":null,"summary":"Text Classification • 0.1B • Updated • 297 • 4 Community Article /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed questions.","content_type":"changelog","language":"en","canonical_url":"https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp","author":{"name":"Xuan-Son Nguyen, Victor Mustar","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Hugging Face","url":"https://huggingface.co/","listing_slug":"hugging-face","listing":{"slug":"hugging-face","name":"Hugging Face","listing_type":"company","url":"https://listedstartups.com/companies/hugging-face"}},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":812,"reading_minutes":4,"published_at":"2026-10-02T00:00:00.000Z","added_at":"2026-10-02T15:16:49.040Z","updated_at":"2026-10-02T15:16:49.040Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/new-in-llama-cpp-decision-models","markdown_url":"https://listedarticles.com/articles/new-in-llama-cpp-decision-models.md","example":false,"citation":"Xuan-Son Nguyen, Victor Mustar, Hugging Face. \"New in llama.cpp: Decision Models.\" 2 Oct 2026. https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp"},"body_markdown":"[Text Classification •  0.1B • Updated   •  297  •  4](/ggml-org/Julia-1-GGUF)  \n\n# \n\t\n\t\t\n\t\n\t\n\t\tNew in llama.cpp: Decision Models\n\t\n\n [Community Article](/blog/community)\n\n`/v1/systemone` endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.\nThe API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL. Implementation details are in [PR #29818](https://github.com/ggml-org/llama.cpp/pull/29818).\n\n**What is a decision model?** A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.\n\n\n## \n\t\n\t\t\n\t\n\t\n\t\tSupported models\n\t\n\n| Model | Size | Based on | Languages | Images | License | Speed* | \n|---|---|---|---|---|---|---|\n| [Julia-1](https://huggingface.co/ggml-org/Julia-1-GGUF) | 144M | mmBERT-small | 50+ | no | Apache 2.0 | 3 ms | \n| [Laya](https://huggingface.co/ggml-org/Laya-GGUF) | 421M | ModernBERT-large | English | no | Apache 2.0 | 5 ms | \n| [Kev-4B](https://huggingface.co/ggml-org/Kev-4B-GGUF) | 4B | Qwen3.5-4B-Base | English | no | Apache 2.0 | 12 ms | \n| [lev](https://huggingface.co/ggml-org/lev-GGUF) | 4B | Qwen3.5-4B | English | no | Apache 2.0 | 36 ms | \n| [OpenJev](https://huggingface.co/ggml-org/OpenJev-GGUF) | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | yes | CC BY-NC 4.0 | 43 ms | \n\n<sub>*Median time to answer one question, on one NVIDIA RTX PRO 6000.</sub>\n\nFind these models in the [Decision models collection](https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769), with more coming. The community [Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) shows how they compare.\n\n## \n\t\n\t\t\n\t\n\t\n\t\tQuick start\n\t\n\nGet the latest llama.cpp from [llama.app](https://llama.app) (or run `llama update`), then start a model:\n\n```\nllama serve -hf ggml-org/Kev-4B-GGUF\n```\nA request contains a state and one or more questions. There are three question types:\n\n| Type | You send | You get | \n|---|---|---|\n| `choice` | options, with optional descriptions | the top option, plus a probability per option | \n| `score` | 2 to 10 levels, lowest first | the expected level (can fall between two) | \n| `noul` | a yes/no question | the probability of yes | \n\nSend a request with your state and questions:\n\n```\ncurl http://localhost:8080/v1/systemone \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"state\": \"Customer message: I was charged twice for my order last week and nobody has replied.\",\n    \"questions\": {\n      \"route\": {\n        \"type\": \"choice\",\n        \"instructions\": \"Which team should handle this?\",\n        \"criteria\": {\n          \"billing\": \"payments, charges, refunds, invoices\",\n          \"shipping\": \"delivery, tracking, lost or late parcels\",\n          \"technical\": \"bugs, errors, login problems\"\n        }\n      },\n      \"angry\": {\n        \"type\": \"noul\",\n        \"instructions\": \"Is the customer angry?\"\n      },\n      \"urgency\": {\n        \"type\": \"score\",\n        \"instructions\": \"How urgent is this?\",\n        \"criteria\": [\"can wait\", \"this week\", \"today\", \"right now\"]\n      }\n    }\n  }'\n```\nResponse (values rounded):\n\n```\n{\n  \"model\": \"ggml-org/Kev-4B-GGUF\",\n  \"answers\": {\n    \"route\": {\n      \"type\": \"choice\",\n      \"choice\": \"billing\",\n      \"probabilities\": {\"billing\": 0.9049, \"shipping\": 0.0275, \"technical\": 0.0676},\n      \"confidence\": 0.8574\n    },\n    \"angry\": {\n      \"type\": \"noul\",\n      \"noul\": 0.8208\n    },\n    \"urgency\": {\n      \"type\": \"score\",\n      \"score\": 2.2821,\n      \"legend\": {\"0\": \"can wait\", \"1\": \"this week\", \"2\": \"today\", \"3\": \"right now\"},\n      \"probabilities\": {\"0\": 0.036, \"1\": 0.1937, \"2\": 0.2225, \"3\": 0.5478},\n      \"confidence\": 0.2821\n    }\n  },\n  \"usage\": {\"input_tokens\": 130, \"output_tokens\": 0}\n}\n```\nThe full reference is in the [server docs](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).\n\n## \n\t\n\t\t\n\t\n\t\n\t\tImages\n\t\n\nSome models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:\n\n```\nllama serve -hf ggml-org/OpenJev-GGUF\n```\nFor example, to classify an uploaded document:\n\n```\nimport base64\nimport requests\nwith open(\"document.png\", \"rb\") as f:\n    image = \"data:image/png;base64,\" + base64.b64encode(f.read()).decode()\nresponse = requests.post(\"http://localhost:8080/v1/systemone\", json={\n    \"state\": \"A file uploaded by a customer.\",\n    \"images\": [image],\n    \"questions\": {\n        \"kind\": {\n            \"type\": \"choice\",\n            \"instructions\": \"What kind of document is this?\",\n            \"criteria\": {\"invoice\": None, \"receipt\": None, \"contract\": None, \"other\": None},\n        },\n    },\n})\nprint(response.json()[\"answers\"][\"kind\"][\"choice\"])  # invoice\n```\nThe `state` can also be a list of chat messages. Any `image_url` part (data URL) is read as an image, same as chat completions.\n\n## \n\t\n\t\t\n\t\n\t\n\t\tSeveral models, one server\n\t\n\nIn [router mode](https://huggingface.co/blog/ggml-org/model-management-in-llamacpp), models load on demand and you pick one per request:\n\n```\nllama serve\ncurl http://localhost:8080/v1/systemone \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\": \"ggml-org/Julia-1-GGUF:Q8_0\", \"state\": \"...\", \"questions\": {...}}'\n```\n`/v1/models` lists the ids. With a single model loaded, the `model` field is ignored.\n\n## \n\t\n\t\t\n\t\n\t\n\t\tTips\n\t\n\n- **Try several models, of different sizes.** Small models are faster, large ones know more. The[Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) compares them.\n- **Describe your options.** Julia-1 routed \"I was charged twice\" to`shipping` with bare labels, and to`billing` (0.99) once each option had a description.\n- **Pick your confidence cutoff per model.** A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket (\"Hi, quick question about my account\") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.\n- **Batch your questions.** They are answered independently, and Kev-4B, lev and OpenJev process the state only once.\n- **Try different quantizations.** Like any GGUF, these models come in several precisions, for example`llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0` .\n\n## \n\t\n\t\t\n\t\n\t\n\t\tWhat's next\n\t\n\nNew open decision models come out every week, and we'll keep adding the best ones. Cloudflare's Clef is next. Is there one you particularly want? Tell us in the comments.","body_html":"<p><a href=\"/ggml-org/Julia-1-GGUF\">Text Classification •  0.1B • Updated   •  297  •  4</a>  </p>\n<p># </p>\n<pre><code>    New in llama.cpp: Decision Models</code></pre>\n<p> <a href=\"/blog/community\">Community Article</a></p>\n<p><code>/v1/systemone</code> endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.\nThe API follows the System One format introduced with TypeSafe&#39;s Jev model, so existing clients only need a new base URL. Implementation details are in <a href=\"https://github.com/ggml-org/llama.cpp/pull/29818\" rel=\"nofollow ugc noopener\">PR #29818</a>.</p>\n<p><strong>What is a decision model?</strong> A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent&#39;s step worked, or choosing its next action.</p>\n<p>## </p>\n<pre><code>    Supported models</code></pre>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Size</th><th>Based on</th><th>Languages</th><th>Images</th><th>License</th><th>Speed*</th></tr></thead><tbody><tr><td><a href=\"https://huggingface.co/ggml-org/Julia-1-GGUF\" rel=\"nofollow ugc noopener\">Julia-1</a></td><td>144M</td><td>mmBERT-small</td><td>50+</td><td>no</td><td>Apache 2.0</td><td>3 ms</td></tr><tr><td><a href=\"https://huggingface.co/ggml-org/Laya-GGUF\" rel=\"nofollow ugc noopener\">Laya</a></td><td>421M</td><td>ModernBERT-large</td><td>English</td><td>no</td><td>Apache 2.0</td><td>5 ms</td></tr><tr><td><a href=\"https://huggingface.co/ggml-org/Kev-4B-GGUF\" rel=\"nofollow ugc noopener\">Kev-4B</a></td><td>4B</td><td>Qwen3.5-4B-Base</td><td>English</td><td>no</td><td>Apache 2.0</td><td>12 ms</td></tr><tr><td><a href=\"https://huggingface.co/ggml-org/lev-GGUF\" rel=\"nofollow ugc noopener\">lev</a></td><td>4B</td><td>Qwen3.5-4B</td><td>English</td><td>no</td><td>Apache 2.0</td><td>36 ms</td></tr><tr><td><a href=\"https://huggingface.co/ggml-org/OpenJev-GGUF\" rel=\"nofollow ugc noopener\">OpenJev</a></td><td>27B</td><td>Qwen3.8-27B</td><td>en, de, fr, hi, zh, ja</td><td>yes</td><td>CC BY-NC 4.0</td><td>43 ms</td></tr></tbody></table></div>\n<p>&lt;sub&gt;*Median time to answer one question, on one NVIDIA RTX PRO 6000.&lt;/sub&gt;</p>\n<p>Find these models in the <a href=\"https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769\" rel=\"nofollow ugc noopener\">Decision models collection</a>, with more coming. The community <a href=\"https://multimodalart-jev-decision-index.static.hf.space/index.html\" rel=\"nofollow ugc noopener\">Decision Index</a> shows how they compare.</p>\n<p>## </p>\n<pre><code>    Quick start</code></pre>\n<p>Get the latest llama.cpp from <a href=\"https://llama.app\" rel=\"nofollow ugc noopener\">llama.app</a> (or run <code>llama update</code>), then start a model:</p>\n<pre><code>llama serve -hf ggml-org/Kev-4B-GGUF</code></pre>\n<p>A request contains a state and one or more questions. There are three question types:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Type</th><th>You send</th><th>You get</th></tr></thead><tbody><tr><td><code>choice</code></td><td>options, with optional descriptions</td><td>the top option, plus a probability per option</td></tr><tr><td><code>score</code></td><td>2 to 10 levels, lowest first</td><td>the expected level (can fall between two)</td></tr><tr><td><code>noul</code></td><td>a yes/no question</td><td>the probability of yes</td></tr></tbody></table></div>\n<p>Send a request with your state and questions:</p>\n<pre><code>curl http://localhost:8080/v1/systemone \\\n  -H &quot;Content-Type: application/json&quot; \\\n  -d &#39;{\n    &quot;state&quot;: &quot;Customer message: I was charged twice for my order last week and nobody has replied.&quot;,\n    &quot;questions&quot;: {\n      &quot;route&quot;: {\n        &quot;type&quot;: &quot;choice&quot;,\n        &quot;instructions&quot;: &quot;Which team should handle this?&quot;,\n        &quot;criteria&quot;: {\n          &quot;billing&quot;: &quot;payments, charges, refunds, invoices&quot;,\n          &quot;shipping&quot;: &quot;delivery, tracking, lost or late parcels&quot;,\n          &quot;technical&quot;: &quot;bugs, errors, login problems&quot;\n        }\n      },\n      &quot;angry&quot;: {\n        &quot;type&quot;: &quot;noul&quot;,\n        &quot;instructions&quot;: &quot;Is the customer angry?&quot;\n      },\n      &quot;urgency&quot;: {\n        &quot;type&quot;: &quot;score&quot;,\n        &quot;instructions&quot;: &quot;How urgent is this?&quot;,\n        &quot;criteria&quot;: [&quot;can wait&quot;, &quot;this week&quot;, &quot;today&quot;, &quot;right now&quot;]\n      }\n    }\n  }&#39;</code></pre>\n<p>Response (values rounded):</p>\n<pre><code>{\n  &quot;model&quot;: &quot;ggml-org/Kev-4B-GGUF&quot;,\n  &quot;answers&quot;: {\n    &quot;route&quot;: {\n      &quot;type&quot;: &quot;choice&quot;,\n      &quot;choice&quot;: &quot;billing&quot;,\n      &quot;probabilities&quot;: {&quot;billing&quot;: 0.9049, &quot;shipping&quot;: 0.0275, &quot;technical&quot;: 0.0676},\n      &quot;confidence&quot;: 0.8574\n    },\n    &quot;angry&quot;: {\n      &quot;type&quot;: &quot;noul&quot;,\n      &quot;noul&quot;: 0.8208\n    },\n    &quot;urgency&quot;: {\n      &quot;type&quot;: &quot;score&quot;,\n      &quot;score&quot;: 2.2821,\n      &quot;legend&quot;: {&quot;0&quot;: &quot;can wait&quot;, &quot;1&quot;: &quot;this week&quot;, &quot;2&quot;: &quot;today&quot;, &quot;3&quot;: &quot;right now&quot;},\n      &quot;probabilities&quot;: {&quot;0&quot;: 0.036, &quot;1&quot;: 0.1937, &quot;2&quot;: 0.2225, &quot;3&quot;: 0.5478},\n      &quot;confidence&quot;: 0.2821\n    }\n  },\n  &quot;usage&quot;: {&quot;input_tokens&quot;: 130, &quot;output_tokens&quot;: 0}\n}</code></pre>\n<p>The full reference is in the <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md\" rel=\"nofollow ugc noopener\">server docs</a>.</p>\n<p>## </p>\n<pre><code>    Images</code></pre>\n<p>Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:</p>\n<pre><code>llama serve -hf ggml-org/OpenJev-GGUF</code></pre>\n<p>For example, to classify an uploaded document:</p>\n<pre><code>import base64\nimport requests\nwith open(&quot;document.png&quot;, &quot;rb&quot;) as f:\n    image = &quot;data:image/png;base64,&quot; + base64.b64encode(f.read()).decode()\nresponse = requests.post(&quot;http://localhost:8080/v1/systemone&quot;, json={\n    &quot;state&quot;: &quot;A file uploaded by a customer.&quot;,\n    &quot;images&quot;: [image],\n    &quot;questions&quot;: {\n        &quot;kind&quot;: {\n            &quot;type&quot;: &quot;choice&quot;,\n            &quot;instructions&quot;: &quot;What kind of document is this?&quot;,\n            &quot;criteria&quot;: {&quot;invoice&quot;: None, &quot;receipt&quot;: None, &quot;contract&quot;: None, &quot;other&quot;: None},\n        },\n    },\n})\nprint(response.json()[&quot;answers&quot;][&quot;kind&quot;][&quot;choice&quot;])  # invoice</code></pre>\n<p>The <code>state</code> can also be a list of chat messages. Any <code>image_url</code> part (data URL) is read as an image, same as chat completions.</p>\n<p>## </p>\n<pre><code>    Several models, one server</code></pre>\n<p>In <a href=\"https://huggingface.co/blog/ggml-org/model-management-in-llamacpp\" rel=\"nofollow ugc noopener\">router mode</a>, models load on demand and you pick one per request:</p>\n<pre><code>llama serve\ncurl http://localhost:8080/v1/systemone \\\n  -H &quot;Content-Type: application/json&quot; \\\n  -d &#39;{&quot;model&quot;: &quot;ggml-org/Julia-1-GGUF:Q8_0&quot;, &quot;state&quot;: &quot;...&quot;, &quot;questions&quot;: {...}}&#39;</code></pre>\n<p><code>/v1/models</code> lists the ids. With a single model loaded, the <code>model</code> field is ignored.</p>\n<p>## </p>\n<pre><code>    Tips</code></pre>\n<ul><li><strong>Try several models, of different sizes.</strong> Small models are faster, large ones know more. The<a href=\"https://multimodalart-jev-decision-index.static.hf.space/index.html\" rel=\"nofollow ugc noopener\">Decision Index</a> compares them.</li><li><strong>Describe your options.</strong> Julia-1 routed &quot;I was charged twice&quot; to<code>shipping</code> with bare labels, and to<code>billing</code> (0.99) once each option had a description.</li><li><strong>Pick your confidence cutoff per model.</strong> A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket (&quot;Hi, quick question about my account&quot;) scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.</li><li><strong>Batch your questions.</strong> They are answered independently, and Kev-4B, lev and OpenJev process the state only once.</li><li><strong>Try different quantizations.</strong> Like any GGUF, these models come in several precisions, for example<code>llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0</code> .</li></ul>\n<p>## </p>\n<pre><code>    What&#39;s next</code></pre>\n<p>New open decision models come out every week, and we&#39;ll keep adding the best ones. Cloudflare&#39;s Clef is next. Is there one you particularly want? Tell us in the comments.</p>","headings":[]}}