---
title: "New in llama.cpp: Decision Models"
slug: new-in-llama-cpp-decision-models
url: https://listedarticles.com/articles/new-in-llama-cpp-decision-models
canonical_url: https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp
content_type: changelog
language: en
published_at: 2026-10-02T00:00:00.000Z
updated_at: 2026-10-02T15:16:49.040Z
author: "Xuan-Son Nguyen, Victor Mustar"
authored_by: human
publisher: "Hugging Face"
publisher_url: https://listedstartups.com/companies/hugging-face
topics: ["AI", "LLMs", "Open Source", "Programming"]
license: all-rights-reserved
word_count: 812
reading_minutes: 4
citation: "Xuan-Son Nguyen, Victor Mustar, Hugging Face. \"New in llama.cpp: Decision Models.\" 2 Oct 2026. https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# New in llama.cpp: Decision Models

> Text Classification • 0.1B • Updated • 297 • 4 Community Article /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed questions.

[Text Classification •  0.1B • Updated   •  297  •  4](/ggml-org/Julia-1-GGUF)  

# 
	
		
	
	
		New in llama.cpp: Decision Models
	

 [Community Article](/blog/community)

`/v1/systemone` endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.
The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL. Implementation details are in [PR #29818](https://github.com/ggml-org/llama.cpp/pull/29818).

**What is a decision model?** A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.


## 
	
		
	
	
		Supported models
	

| Model | Size | Based on | Languages | Images | License | Speed* | 
|---|---|---|---|---|---|---|
| [Julia-1](https://huggingface.co/ggml-org/Julia-1-GGUF) | 144M | mmBERT-small | 50+ | no | Apache 2.0 | 3 ms | 
| [Laya](https://huggingface.co/ggml-org/Laya-GGUF) | 421M | ModernBERT-large | English | no | Apache 2.0 | 5 ms | 
| [Kev-4B](https://huggingface.co/ggml-org/Kev-4B-GGUF) | 4B | Qwen3.5-4B-Base | English | no | Apache 2.0 | 12 ms | 
| [lev](https://huggingface.co/ggml-org/lev-GGUF) | 4B | Qwen3.5-4B | English | no | Apache 2.0 | 36 ms | 
| [OpenJev](https://huggingface.co/ggml-org/OpenJev-GGUF) | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | yes | CC BY-NC 4.0 | 43 ms | 

<sub>*Median time to answer one question, on one NVIDIA RTX PRO 6000.</sub>

Find these models in the [Decision models collection](https://huggingface.co/collections/ggml-org/decision-models-6abf80cca3c83f127060a769), with more coming. The community [Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) shows how they compare.

## 
	
		
	
	
		Quick start
	

Get the latest llama.cpp from [llama.app](https://llama.app) (or run `llama update`), then start a model:

```
llama serve -hf ggml-org/Kev-4B-GGUF
```
A request contains a state and one or more questions. There are three question types:

| Type | You send | You get | 
|---|---|---|
| `choice` | options, with optional descriptions | the top option, plus a probability per option | 
| `score` | 2 to 10 levels, lowest first | the expected level (can fall between two) | 
| `noul` | a yes/no question | the probability of yes | 

Send a request with your state and questions:

```
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'
```
Response (values rounded):

```
{
  "model": "ggml-org/Kev-4B-GGUF",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
      "confidence": 0.8574
    },
    "angry": {
      "type": "noul",
      "noul": 0.8208
    },
    "urgency": {
      "type": "score",
      "score": 2.2821,
      "legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
      "probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
      "confidence": 0.2821
    }
  },
  "usage": {"input_tokens": 130, "output_tokens": 0}
}
```
The full reference is in the [server docs](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md).

## 
	
		
	
	
		Images
	

Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:

```
llama serve -hf ggml-org/OpenJev-GGUF
```
For example, to classify an uploaded document:

```
import base64
import requests
with open("document.png", "rb") as f:
    image = "data:image/png;base64," + base64.b64encode(f.read()).decode()
response = requests.post("http://localhost:8080/v1/systemone", json={
    "state": "A file uploaded by a customer.",
    "images": [image],
    "questions": {
        "kind": {
            "type": "choice",
            "instructions": "What kind of document is this?",
            "criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
        },
    },
})
print(response.json()["answers"]["kind"]["choice"])  # invoice
```
The `state` can also be a list of chat messages. Any `image_url` part (data URL) is read as an image, same as chat completions.

## 
	
		
	
	
		Several models, one server
	

In [router mode](https://huggingface.co/blog/ggml-org/model-management-in-llamacpp), models load on demand and you pick one per request:

```
llama serve
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'
```
`/v1/models` lists the ids. With a single model loaded, the `model` field is ignored.

## 
	
		
	
	
		Tips
	

- **Try several models, of different sizes.** Small models are faster, large ones know more. The[Decision Index](https://multimodalart-jev-decision-index.static.hf.space/index.html) compares them.
- **Describe your options.** Julia-1 routed "I was charged twice" to`shipping` with bare labels, and to`billing` (0.99) once each option had a description.
- **Pick your confidence cutoff per model.** A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.
- **Batch your questions.** They are answered independently, and Kev-4B, lev and OpenJev process the state only once.
- **Try different quantizations.** Like any GGUF, these models come in several precisions, for example`llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0` .

## 
	
		
	
	
		What's next
	

New open decision models come out every week, and we'll keep adding the best ones. Cloudflare's Clef is next. Is there one you particularly want? Tell us in the comments.
