##
Supported models
| Model | Size | Based on | Languages | Images | License | Speed* |
|---|
| Julia-1 | 144M | mmBERT-small | 50+ | no | Apache 2.0 | 3 ms |
| Laya | 421M | ModernBERT-large | English | no | Apache 2.0 | 5 ms |
| Kev-4B | 4B | Qwen3.5-4B-Base | English | no | Apache 2.0 | 12 ms |
| lev | 4B | Qwen3.5-4B | English | no | Apache 2.0 | 36 ms |
| OpenJev | 27B | Qwen3.8-27B | en, de, fr, hi, zh, ja | yes | CC BY-NC 4.0 | 43 ms |
<sub>*Median time to answer one question, on one NVIDIA RTX PRO 6000.</sub>
Find these models in the Decision models collection, with more coming. The community Decision Index shows how they compare.
##
Quick start
Get the latest llama.cpp from llama.app (or run llama update), then start a model:
llama serve -hf ggml-org/Kev-4B-GGUF
A request contains a state and one or more questions. There are three question types:
| Type | You send | You get |
|---|
choice | options, with optional descriptions | the top option, plus a probability per option |
score | 2 to 10 levels, lowest first | the expected level (can fall between two) |
noul | a yes/no question | the probability of yes |
Send a request with your state and questions:
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer message: I was charged twice for my order last week and nobody has replied.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "payments, charges, refunds, invoices",
"shipping": "delivery, tracking, lost or late parcels",
"technical": "bugs, errors, login problems"
}
},
"angry": {
"type": "noul",
"instructions": "Is the customer angry?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["can wait", "this week", "today", "right now"]
}
}
}'
Response (values rounded):
{
"model": "ggml-org/Kev-4B-GGUF",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
"confidence": 0.8574
},
"angry": {
"type": "noul",
"noul": 0.8208
},
"urgency": {
"type": "score",
"score": 2.2821,
"legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
"probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
"confidence": 0.2821
}
},
"usage": {"input_tokens": 130, "output_tokens": 0}
}
The full reference is in the server docs.
##
Images
Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:
llama serve -hf ggml-org/OpenJev-GGUF
For example, to classify an uploaded document:
import base64
import requests
with open("document.png", "rb") as f:
image = "data:image/png;base64," + base64.b64encode(f.read()).decode()
response = requests.post("http://localhost:8080/v1/systemone", json={
"state": "A file uploaded by a customer.",
"images": [image],
"questions": {
"kind": {
"type": "choice",
"instructions": "What kind of document is this?",
"criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
},
},
})
print(response.json()["answers"]["kind"]["choice"]) # invoice
The state can also be a list of chat messages. Any image_url part (data URL) is read as an image, same as chat completions.
##
Several models, one server
In router mode, models load on demand and you pick one per request:
llama serve
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'
/v1/models lists the ids. With a single model loaded, the model field is ignored.
##
Tips
- Try several models, of different sizes. Small models are faster, large ones know more. TheDecision Index compares them.
- Describe your options. Julia-1 routed "I was charged twice" to
shipping with bare labels, and tobilling (0.99) once each option had a description. - Pick your confidence cutoff per model. A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.
- Batch your questions. They are answered independently, and Kev-4B, lev and OpenJev process the state only once.
- Try different quantizations. Like any GGUF, these models come in several precisions, for example
llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0 .
##
What's next
New open decision models come out every week, and we'll keep adding the best ones. Cloudflare's Clef is next. Is there one you particularly want? Tell us in the comments.