{"article":{"slug":"build-your-own-decision-model","title":"Build your own decision model","subtitle":null,"summary":"Nish Tahir explains \"system one\" decision models like TypeSafe AI's Jev, which answer in a single forward pass by masking the vocabulary to a fixed set of options, then walks through fine-tuning a small model on multiple-choice data and calibrating its confidence with temperature scaling, with a companion GitHub repo.","content_type":"tutorial","language":"en","canonical_url":"https://nishtahir.com/build-your-own-decision-model/","author":{"name":"Nish Tahir","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Another Dev's Two Cents","url":"https://nishtahir.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1019,"reading_minutes":4,"published_at":"2026-10-10T21:24:15.000Z","added_at":"2026-10-10T23:12:52.478Z","updated_at":"2026-10-10T23:12:52.478Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/build-your-own-decision-model","markdown_url":"https://listedarticles.com/articles/build-your-own-decision-model.md","example":false,"citation":"Nish Tahir, Another Dev's Two Cents. \"Build your own decision model.\" 10 Oct 2026. https://nishtahir.com/build-your-own-decision-model/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://nishtahir.com/build-your-own-decision-model/"},"body_markdown":"[\"System one\"](https://thedecisionlab.com/reference-guide/philosophy/system-1-and-system-2-thinking?ref=nishtahir.com) decision models are models that infer and respond with calibrated probabilities or every allowed answer.\n\nConsider your everyday language model, to get typed output from it (JSON), you may use [Structured Output](https://nishtahir.com/how-llm-structured-decoding-works/) to constrain the output to guaranteed valid JSON. While model prefills the input in one pass, it still has to go perform a pass for every token in order to generate a valid response.\n\nIn this example, 11 passes are required to generate the final output. (We're not accounting for speculative decoding and other inference optimization techniques.)\n\n\nDecision models such as [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev?ref=nishtahir.com), make the assumption that there are fixed options we can select from and we can do so quickly by making a single pass. In this example we constrain the set of possible outputs to the options `A`, `B`, `C`, `D`, `E`. By masking other items in the vocabulary, the model can only emit those tokens. By selecting the highest probability output, we get our answer.\n\n\nSince the outputs are constrained to only a fixed set of options, the model can't select anything outside of those. This however does not guarantee that the output will be correct. It's also common to treat the output token probabilities a confidence scores in this context, but without additional training, those scores likely reflect its confidence in what the next token will be rather than the true probability of the response being the correct answer.\n\n# Build your own\n\nWe can emulate this behavior by constraining output tokens using an LLM. Here I'm using `Qwen/Qwen3-1.7B`\n\n```\nimport argparse\nimport json\n\nimport torch\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\n\nmodel_name = \"Qwen/Qwen3-1.7B\"\noptions = [\"A\", \"B\", \"C\", \"D\", \"E\"]\n\nparser = argparse.ArgumentParser()\nparser.add_argument(\"--input\", default=\"question.json\")\nargs = parser.parse_args()\n\n# load the tokenizer and the model\ntokenizer = AutoTokenizer.from_pretrained(model_name)\nmodel = AutoModelForCausalLM.from_pretrained(\n    model_name,\n    torch_dtype=\"auto\",\n    device_map=\"auto\"\n)\n\n# the token the model would emit for each option as the first assistant token\noption_token_ids = [tokenizer.encode(opt, add_special_tokens=False)[0] for opt in options]\n\ndef format_prompt(item):\n    prompt = item[\"question\"] + \"\\n\"\n    for opt in options:\n        prompt += f\"{opt}. {item[opt]}\\n\"\n    prompt += \"Answer:\"\n    messages = [\n        {\"role\": \"user\", \"content\": prompt}\n    ]\n    return tokenizer.apply_chat_template(\n        messages,\n        tokenize=False,\n        add_generation_prompt=True,\n        enable_thinking=False\n    )\n\nwith open(args.input) as f:\n    item = json.load(f)\n\nmodel_inputs = tokenizer(format_prompt(item), return_tensors=\"pt\").to(model.device)\nwith torch.no_grad():\n    logits = model(**model_inputs).logits[0, -1]\n# constrained decoding: only the option tokens are allowed\nprobs = torch.softmax(logits[option_token_ids].float(), dim=-1)\n\nprint(f\"prediction: {options[probs.argmax().item()]}\")\nfor opt, prob in zip(options, probs.tolist()):\n    print(f\"{opt}: {prob:.4f}  {item[opt]}\")\n```\n\nRunning it with an simple question to test it yeilds the following output\n\n```\n// input\n\n{\n    \"question\": \"What color is the sky?\",\n    \"A\": \"Red\",\n    \"B\": \"Blue\",\n    \"C\": \"Green\",\n    \"D\": \"Purple\",\n    \"E\": \"I don't know\"\n}\n\n// output\n\nprediction: B\nA: 0.0000  Red\nB: 0.9988  Blue\nC: 0.0000  Green\nD: 0.0000  Purple\nE: 0.0012  I don't know\n```\n\nThe model was able to make sense of our input and make a prediction that reasonably corresponds to the correct answer.\n\nWe can test the accuracy of the model by running it against public datasets. I ran this against a random sample holdout of [CommonsenseQA](https://huggingface.co/datasets/tau/commonsense_qa?ref=nishtahir.com)\n\n```\n    precision    recall        f1   support\nA    0.5733    0.7197    0.6382       239\nB    0.5506    0.7686    0.6416       255\nC    0.5372    0.6598    0.5922       241\nD    0.7206    0.3904    0.5065       251\nE    0.7519    0.4255    0.5435       235\naccuracy: 725/1221 = 0.5938\nmacro f1: 0.5844\n```\n\nNot bad for a 1.7B model, Running a quick finetune on the dataset gives us slightly better performance\n\n```\n    precision    recall        f1   support\nA    0.6475    0.6611    0.6542       239\nB    0.6113    0.6784    0.6431       255\nC    0.6234    0.5975    0.6102       241\nD    0.6700    0.5418    0.5991       251\nE    0.5808    0.6426    0.6101       235\naccuracy: 762/1221 = 0.6241\nmacro f1: 0.6234\n```\n\n# Calibrating your model\n\nTesting the model against a very ambiguous problem demonstrates an interesting problem.\n\n```\n// input\n{\n    \"question\": \"Where would you most likely find a bat?\",\n    \"A\": \"Cave\",\n    \"B\": \"Baseball game\",\n    \"C\": \"Attic\",\n    \"D\": \"Zoo\",\n    \"E\": \"Sporting goods store\"\n}\n\n// output\n\nprediction: A\nA: 0.9978  Cave\nB: 0.0004  Baseball game\nC: 0.0017  Attic\nD: 0.0000  Zoo\nE: 0.0001  Sporting goods store\n```\n\nThere should be no clear answer here, but treating the output probabilities as a pseudo \"confidence\" score, shows that the model is extremely overconfident in this answer.\n\nIf we bin the confidence score ranges in the eval I ran earlier, we can see that the model's confidence does not match its accuracy. This means that the model is not [calibrated](https://towardsdatascience.com/a-comprehensive-guide-on-model-calibration-part-1-of-4-73466eb5e09a/?ref=nishtahir.com).\n\n```\n         bin   count  confidence  accuracy\n(0.00, 0.10]       0      0.0000    0.0000\n(0.10, 0.20]       0      0.0000    0.0000\n(0.20, 0.30]       3      0.2834    0.0000\n(0.30, 0.40]      26      0.3761    0.2692\n(0.40, 0.50]      41      0.4538    0.2683\n(0.50, 0.60]      70      0.5490    0.3286\n(0.60, 0.70]      74      0.6476    0.3649\n(0.70, 0.80]      77      0.7495    0.4286\n(0.80, 0.90]     121      0.8555    0.4711\n(0.90, 1.00]     809      0.9855    0.7009\n```\n\nWe can notice that the model tends to be extremely overconfident in the 0.9 - 1.0 bin but it's only correct 70% of the time. When it makes a prediction with 0.8 - 0.9 confidence it's only accurate ~40% of the time. This means that the model is generally overconfident in its predictions.\n\nSince our goal is to have the model output scores that is reflective of its accuracy, one method we can use to callibrate it is through temperature scaling. By modifying the temperature value, we can flatten its output probability distribution curve and scale it to approximate its accuracy.\n\n\nCurve fitting fit the temperature parameter to the model's accuracy, I found `3.797280788421631` as a temp value.\n\n```\n         bin   count  confidence  accuracy\n(0.00, 0.10]       0      0.0000    0.0000\n(0.10, 0.20]       0      0.0000    0.0000\n(0.20, 0.30]      82      0.2712    0.2317\n(0.30, 0.40]     217      0.3507    0.3917\n(0.40, 0.50]     199      0.4472    0.5126\n(0.50, 0.60]     166      0.5475    0.5482\n(0.60, 0.70]     139      0.6562    0.5827\n(0.70, 0.80]     140      0.7492    0.7714\n(0.80, 0.90]     169      0.8507    0.7988\n(0.90, 1.00]     109      0.9333    0.9541\n```\n\nThis gets us a much better calibration. If you want to play around with this, I made a [GitHub repo](https://github.com/nishtahir/build-your-own-jev?ref=nishtahir.com) with scripts that walk you through building a dataset, evaluating, finetuning and calibrating your own model. I encourage pulling it and trying it on other bigger models.\n","body_html":"<p><a href=\"https://thedecisionlab.com/reference-guide/philosophy/system-1-and-system-2-thinking?ref=nishtahir.com\" rel=\"nofollow ugc noopener\">&quot;System one&quot;</a> decision models are models that infer and respond with calibrated probabilities or every allowed answer.</p>\n<p>Consider your everyday language model, to get typed output from it (JSON), you may use <a href=\"https://nishtahir.com/how-llm-structured-decoding-works/\" rel=\"nofollow ugc noopener\">Structured Output</a> to constrain the output to guaranteed valid JSON. While model prefills the input in one pass, it still has to go perform a pass for every token in order to generate a valid response.</p>\n<p>In this example, 11 passes are required to generate the final output. (We&#39;re not accounting for speculative decoding and other inference optimization techniques.)</p>\n<p>Decision models such as <a href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev?ref=nishtahir.com\" rel=\"nofollow ugc noopener\">Jev</a>, make the assumption that there are fixed options we can select from and we can do so quickly by making a single pass. In this example we constrain the set of possible outputs to the options <code>A</code>, <code>B</code>, <code>C</code>, <code>D</code>, <code>E</code>. By masking other items in the vocabulary, the model can only emit those tokens. By selecting the highest probability output, we get our answer.</p>\n<p>Since the outputs are constrained to only a fixed set of options, the model can&#39;t select anything outside of those. This however does not guarantee that the output will be correct. It&#39;s also common to treat the output token probabilities a confidence scores in this context, but without additional training, those scores likely reflect its confidence in what the next token will be rather than the true probability of the response being the correct answer.</p>\n<h1 id=\"build-your-own\">Build your own</h1>\n<p>We can emulate this behavior by constraining output tokens using an LLM. Here I&#39;m using <code>Qwen/Qwen3-1.7B</code></p>\n<pre><code>import argparse\nimport json\n\nimport torch\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\n\nmodel_name = &quot;Qwen/Qwen3-1.7B&quot;\noptions = [&quot;A&quot;, &quot;B&quot;, &quot;C&quot;, &quot;D&quot;, &quot;E&quot;]\n\nparser = argparse.ArgumentParser()\nparser.add_argument(&quot;--input&quot;, default=&quot;question.json&quot;)\nargs = parser.parse_args()\n\n# load the tokenizer and the model\ntokenizer = AutoTokenizer.from_pretrained(model_name)\nmodel = AutoModelForCausalLM.from_pretrained(\n    model_name,\n    torch_dtype=&quot;auto&quot;,\n    device_map=&quot;auto&quot;\n)\n\n# the token the model would emit for each option as the first assistant token\noption_token_ids = [tokenizer.encode(opt, add_special_tokens=False)[0] for opt in options]\n\ndef format_prompt(item):\n    prompt = item[&quot;question&quot;] + &quot;\\n&quot;\n    for opt in options:\n        prompt += f&quot;{opt}. {item[opt]}\\n&quot;\n    prompt += &quot;Answer:&quot;\n    messages = [\n        {&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}\n    ]\n    return tokenizer.apply_chat_template(\n        messages,\n        tokenize=False,\n        add_generation_prompt=True,\n        enable_thinking=False\n    )\n\nwith open(args.input) as f:\n    item = json.load(f)\n\nmodel_inputs = tokenizer(format_prompt(item), return_tensors=&quot;pt&quot;).to(model.device)\nwith torch.no_grad():\n    logits = model(**model_inputs).logits[0, -1]\n# constrained decoding: only the option tokens are allowed\nprobs = torch.softmax(logits[option_token_ids].float(), dim=-1)\n\nprint(f&quot;prediction: {options[probs.argmax().item()]}&quot;)\nfor opt, prob in zip(options, probs.tolist()):\n    print(f&quot;{opt}: {prob:.4f}  {item[opt]}&quot;)</code></pre>\n<p>Running it with an simple question to test it yeilds the following output</p>\n<pre><code>// input\n\n{\n    &quot;question&quot;: &quot;What color is the sky?&quot;,\n    &quot;A&quot;: &quot;Red&quot;,\n    &quot;B&quot;: &quot;Blue&quot;,\n    &quot;C&quot;: &quot;Green&quot;,\n    &quot;D&quot;: &quot;Purple&quot;,\n    &quot;E&quot;: &quot;I don&#39;t know&quot;\n}\n\n// output\n\nprediction: B\nA: 0.0000  Red\nB: 0.9988  Blue\nC: 0.0000  Green\nD: 0.0000  Purple\nE: 0.0012  I don&#39;t know</code></pre>\n<p>The model was able to make sense of our input and make a prediction that reasonably corresponds to the correct answer.</p>\n<p>We can test the accuracy of the model by running it against public datasets. I ran this against a random sample holdout of <a href=\"https://huggingface.co/datasets/tau/commonsense_qa?ref=nishtahir.com\" rel=\"nofollow ugc noopener\">CommonsenseQA</a></p>\n<pre><code>    precision    recall        f1   support\nA    0.5733    0.7197    0.6382       239\nB    0.5506    0.7686    0.6416       255\nC    0.5372    0.6598    0.5922       241\nD    0.7206    0.3904    0.5065       251\nE    0.7519    0.4255    0.5435       235\naccuracy: 725/1221 = 0.5938\nmacro f1: 0.5844</code></pre>\n<p>Not bad for a 1.7B model, Running a quick finetune on the dataset gives us slightly better performance</p>\n<pre><code>    precision    recall        f1   support\nA    0.6475    0.6611    0.6542       239\nB    0.6113    0.6784    0.6431       255\nC    0.6234    0.5975    0.6102       241\nD    0.6700    0.5418    0.5991       251\nE    0.5808    0.6426    0.6101       235\naccuracy: 762/1221 = 0.6241\nmacro f1: 0.6234</code></pre>\n<h1 id=\"calibrating-your-model\">Calibrating your model</h1>\n<p>Testing the model against a very ambiguous problem demonstrates an interesting problem.</p>\n<pre><code>// input\n{\n    &quot;question&quot;: &quot;Where would you most likely find a bat?&quot;,\n    &quot;A&quot;: &quot;Cave&quot;,\n    &quot;B&quot;: &quot;Baseball game&quot;,\n    &quot;C&quot;: &quot;Attic&quot;,\n    &quot;D&quot;: &quot;Zoo&quot;,\n    &quot;E&quot;: &quot;Sporting goods store&quot;\n}\n\n// output\n\nprediction: A\nA: 0.9978  Cave\nB: 0.0004  Baseball game\nC: 0.0017  Attic\nD: 0.0000  Zoo\nE: 0.0001  Sporting goods store</code></pre>\n<p>There should be no clear answer here, but treating the output probabilities as a pseudo &quot;confidence&quot; score, shows that the model is extremely overconfident in this answer.</p>\n<p>If we bin the confidence score ranges in the eval I ran earlier, we can see that the model&#39;s confidence does not match its accuracy. This means that the model is not <a href=\"https://towardsdatascience.com/a-comprehensive-guide-on-model-calibration-part-1-of-4-73466eb5e09a/?ref=nishtahir.com\" rel=\"nofollow ugc noopener\">calibrated</a>.</p>\n<pre><code>         bin   count  confidence  accuracy\n(0.00, 0.10]       0      0.0000    0.0000\n(0.10, 0.20]       0      0.0000    0.0000\n(0.20, 0.30]       3      0.2834    0.0000\n(0.30, 0.40]      26      0.3761    0.2692\n(0.40, 0.50]      41      0.4538    0.2683\n(0.50, 0.60]      70      0.5490    0.3286\n(0.60, 0.70]      74      0.6476    0.3649\n(0.70, 0.80]      77      0.7495    0.4286\n(0.80, 0.90]     121      0.8555    0.4711\n(0.90, 1.00]     809      0.9855    0.7009</code></pre>\n<p>We can notice that the model tends to be extremely overconfident in the 0.9 - 1.0 bin but it&#39;s only correct 70% of the time. When it makes a prediction with 0.8 - 0.9 confidence it&#39;s only accurate ~40% of the time. This means that the model is generally overconfident in its predictions.</p>\n<p>Since our goal is to have the model output scores that is reflective of its accuracy, one method we can use to callibrate it is through temperature scaling. By modifying the temperature value, we can flatten its output probability distribution curve and scale it to approximate its accuracy.</p>\n<p>Curve fitting fit the temperature parameter to the model&#39;s accuracy, I found <code>3.797280788421631</code> as a temp value.</p>\n<pre><code>         bin   count  confidence  accuracy\n(0.00, 0.10]       0      0.0000    0.0000\n(0.10, 0.20]       0      0.0000    0.0000\n(0.20, 0.30]      82      0.2712    0.2317\n(0.30, 0.40]     217      0.3507    0.3917\n(0.40, 0.50]     199      0.4472    0.5126\n(0.50, 0.60]     166      0.5475    0.5482\n(0.60, 0.70]     139      0.6562    0.5827\n(0.70, 0.80]     140      0.7492    0.7714\n(0.80, 0.90]     169      0.8507    0.7988\n(0.90, 1.00]     109      0.9333    0.9541</code></pre>\n<p>This gets us a much better calibration. If you want to play around with this, I made a <a href=\"https://github.com/nishtahir/build-your-own-jev?ref=nishtahir.com\" rel=\"nofollow ugc noopener\">GitHub repo</a> with scripts that walk you through building a dataset, evaluating, finetuning and calibrating your own model. I encourage pulling it and trying it on other bigger models.</p>","headings":[{"level":1,"text":"Build your own","id":"build-your-own"},{"level":1,"text":"Calibrating your model","id":"calibrating-your-model"}]}}