{"article":{"slug":"microsoft-decision-1-our-model-for-fast-decision-making","title":"Microsoft-Decision-1: Our model for fast decision-making","subtitle":null,"summary":"Microsoft introduces Microsoft-Decision-1, a decision-scoring model in Foundry built to return structured outputs software can act on immediately, and reports that it leads on latency and quality across structured decision tasks against LLMs and other decision models.","content_type":"announcement","language":"en","canonical_url":"https://commandline.microsoft.com/microsoft-decision-1-model-foundry/","author":{"name":"Lisa Brown Jaloza","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Microsoft","url":"https://www.microsoft.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Announcements","slug":"announcements","url":"https://listedarticles.com/topics/announcements"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1358,"reading_minutes":6,"published_at":"2026-10-09T00:00:00.000Z","added_at":"2026-10-09T20:24:36.632Z","updated_at":"2026-10-09T20:24:36.632Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/microsoft-decision-1-our-model-for-fast-decision-making","markdown_url":"https://listedarticles.com/articles/microsoft-decision-1-our-model-for-fast-decision-making.md","example":false,"citation":"Lisa Brown Jaloza, Microsoft. \"Microsoft-Decision-1: Our model for fast decision-making.\" 9 Oct 2026. https://commandline.microsoft.com/microsoft-decision-1-model-foundry/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://commandline.microsoft.com/microsoft-decision-1-model-foundry/"},"body_markdown":"Decision models are quickly emerging as an important new category in AI. Unlike LLMs, which are designed to generate text or reason through complex problems, decision models are purpose-built to deliver structured outputs that software can immediately act on. And once you understand that capability—making decisions and classifying things at very low cost with high performance—all kinds of useful tasks get unlocked.\n\nToday we’re introducing **Microsoft-Decision-1**, our new model for fast decision-scoring, available in [Microsoft Foundry](https://ai.azure.com/catalog/models/Microsoft-Decision-1) and coming soon through OpenRouter. This model is designed for routing, classification, prioritization, verification, and workflow control, making it easier to incorporate decision intelligence into existing applications, agents, and workflows in a secure, trusted environment. Microsoft-Decision-1 delivers top performance in latency and quality on structured decision tasks to outperform both LLMs and other decision models.\n\nMicrosoft-Decision-1 achieved the highest accuracy in our 36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training. And in our benchmarking, it was the fastest measured: 4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol.\n\n## How Microsoft-Decision-1 works\n\nTo build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI. When given a fixed set of answer options, Microsoft-Decision-1 provides a calibrated probability score for each option. The model supports yes/no, multiple-choice, and rating options, as well as rubric-based grading of AI responses and agent actions, all through a simple structured API call.\n\nTo build a reliable decision model, we had to address several challenges:\n\n### 1. Speed\n\nEach decision adds delay, especially when one step depends on another. For example, adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow.\n\nMicrosoft-Decision-1 P50 latency is ~35x faster than GPT-6 Sol.\n\n### 2. Quality that generalizes\n\nIt’s easy to overfit a model for one benchmark or one type of decision task. We need to know whether that quality carries over to various tasks the model wasn’t trained on. That’s why we evaluated Microsoft-Decision-1 across dozens of benchmarks kept blinded from training, spanning routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning, and safety. We also took several of the top public models on the [popular open leaderboard JevBench](https://benchmarkheaven.com/jev-models) and tested them across 36 additional public and private benchmarks. Microsoft-Decision-1 performed the best across these broader sets of benchmarks, demonstrating strong generalization.\n\n### 3. Robustness\n\nEquivalent inputs should produce equivalent decisions. In production, states and instructions get paraphrased, option descriptions change, choices are reordered, keys change, and harmless formatting noise appears. None of those changes should materially alter the decision.\n\nWe perturb the same request in eight ways and measure how often the decision flips. Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled.\n\n### 4. Probability and confidence calibration\n\nThe probability itself is part of the API, not just a ranking score. Applications use confidence to decide when to act, defer, or ask for review, so a 90% prediction should be right about nine times out of 10 on representative cases.\n\n### 5. Safety\n\nA decision model should recognize harmful requests without needlessly blocking harmless ones. We tested Microsoft-Decision-1 on 5,250 requests across 11 benchmarks, covering harmful content, jailbreak attempts, and prompt injection, and found that the model successfully refused harmful behavior while retaining a high degree of utility.\n\n## Demo examples\n\nClassification is a key use case for decision models. Check out how accurately and quickly Microsoft-Decision-1 can categorize a variety of queries compared to GPT-6 Sol:\n\n  \nDecision models can also be efficient for computer use scenarios. This demo shows how fast Microsoft-Decision-1 can complete the task of buying a backpack compared to GPT-6 Sol:\n\n  \n## How we’re testing Microsoft-Decision-1 internally\n\nHere are some of the ways we’ve been testing Microsoft-Decision-1 internally, with a lot more to come.\n\n### Labeling data\n\nXBOX Research used Microsoft-Decision-1 to process more than 10,000 open-ended pieces of feedback and reviews from surveys, STEAM, and Twitter/X and sort them into a fixed set of themes established by researchers to understand what people are saying about different games, launches, streams, and more. They found Microsoft-Decision-1 to be competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive.\n\n### Quality control\n\nThe Copilot team measures the quality of chat and agentic responses. Their testing found Microsoft-Decision-1 to be competitive with GPT5.6 Luna and 100 times faster.\n\n### Incident response\n\nOur on-call engineers use AI to retrieve relevant knowledge to respond to live incidents across logs, ticketing systems, calls, messages, and other data sources. Microsoft-Decision-1 performed better and faster than an LLM for knowledge retrieval.\n\n### Scientific discovery\n\nMicrosoft Discovery implements an adaptive replanning feature where an agent evaluates a previous experiment, revises its approach based on rubric grades, and repeats until it has completed its objectives. Microsoft-Decision-1 scored as 46 times more consistent than the LLM-based score at three times the speed and resulted in nearly four times the speed on adaptive replanning\n\nFaster, more reliable planning could significantly impact outcomes for long-running scientific experiments.\n\nThose are just a small handful of examples. There are many more potential use cases for Microsoft-Decision-1. Consider trying out the following:\n\n- **Agent controls:** Evaluate an agent’s proposed next step and decide whether to continue, stop, retry, or hand off to a model, tool, or human.\n- **Model routing:** Evaluate an incoming request and select the best model for the task based on quality, cost, and latency requirements.\n- **Skill-based decisions:** Apply the rules from an agent skill to select the appropriate next action without repeatedly processing a long list of instructions.\n- **Data labeling:** Assign consistent labels to social media posts, customer feedback, and other data for analysis or training.\n- **AI judging:** Evaluate an AI-generated response against defined quality criteria and decide whether to accept, revise, or reject it.\n- **Intent analysis:** Identify what a user is trying to accomplish and match their request to a supported intent.\n- **Incident response routing:** Classify an incident by type and urgency, then route it to the appropriate team or workflow.\n- **Data validation:** Check whether an input meets defined requirements and decide whether to accept it, reject it, or flag it for review.\n- **Recommendations:** Select the most relevant item, offer, or next action from a set of candidates.\n- **Search relevance:** Assess how well a search result matches a query and assign a relevance label or priority.\n- **Content classification and filtering:** Categorize content by topic or policy and decide whether to display it, filter it, or send it for review.\n- **Code scanning:** Evaluate code against defined criteria and flag potential defects or policy violations for review.\n- **Safety and security screening:** Classify requests, outputs, or proposed actions by risk and decide whether to allow, block, or escalate them.\n- **Computer and UI use:** Select the next interface action from a set of options based on the current screen and task.\n- **Robotics:** Select among predefined robot actions based on observations, task goals, and operating constraints.\n- **Scientific discovery:** Screen candidate hypotheses, compounds, or experiments against defined criteria and prioritize them for further evaluation.\n\n## Getting started\n\nDevelopers can get started with Microsoft-Decision-1 today in Microsoft Foundry here: [https://ai.azure.com/catalog/models/Microsoft-Decision-1](https://ai.azure.com/catalog/models/Microsoft-Decision-1).\n\n### Pricing\n\nInput tokens cost $0.042 USD per million tokens. Output tokens are free.\n\n## Looking ahead\n\nNow that agentic AI is a reality, we’ve seen that cost plays a major role in how people decide to use AI. And it’s increasingly important to choose the right model for the right job. With agents taking action and making an impact in the real world, decision models have the potential to help people guide and control those agents through complex environments.\n\nWe look forward to seeing what developers build with Microsoft-Decision-1 and hearing their feedback. We’ll continue to release updates to the model, including by incorporating evaluations and data to further optimize quality, confidence, and cost.\n\n### Appendix: Models benchmarked\n\nQuyet-1.0-Large — [https://huggingface.co/chinhnc/Quyet-1.0-Large](https://huggingface.co/chinhnc/Quyet-1.0-Large) \n\nSurogate Rune 26B-A4B — [https://huggingface.co/surogate/rune-26b-a4b-GGUF](https://huggingface.co/surogate/rune-26b-a4b-GGUF) \n\nGPT-6 Luna Decisions — [https://developers.openai.com/api/docs/guides/decisions](https://developers.openai.com/api/docs/guides/decisions) \n\ndeck-31B — [https://github.com/krishna-gogineni-765/deck31b](https://github.com/krishna-gogineni-765/deck31b) \n\nH2O-Lightning-4B — [https://huggingface.co/h2oai/h2o-lightning-4b](https://huggingface.co/h2oai/h2o-lightning-4b) \n\nStrands-Decider 2B — [https://github.com/strands-labs/strands-decider](https://github.com/strands-labs/strands-decider)\n","body_html":"<p>Decision models are quickly emerging as an important new category in AI. Unlike LLMs, which are designed to generate text or reason through complex problems, decision models are purpose-built to deliver structured outputs that software can immediately act on. And once you understand that capability—making decisions and classifying things at very low cost with high performance—all kinds of useful tasks get unlocked.</p>\n<p>Today we’re introducing <strong>Microsoft-Decision-1</strong>, our new model for fast decision-scoring, available in <a href=\"https://ai.azure.com/catalog/models/Microsoft-Decision-1\" rel=\"nofollow ugc noopener\">Microsoft Foundry</a> and coming soon through OpenRouter. This model is designed for routing, classification, prioritization, verification, and workflow control, making it easier to incorporate decision intelligence into existing applications, agents, and workflows in a secure, trusted environment. Microsoft-Decision-1 delivers top performance in latency and quality on structured decision tasks to outperform both LLMs and other decision models.</p>\n<p>Microsoft-Decision-1 achieved the highest accuracy in our 36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training. And in our benchmarking, it was the fastest measured: 4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol.</p>\n<h2 id=\"how-microsoft-decision-1-works\">How Microsoft-Decision-1 works</h2>\n<p>To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI. When given a fixed set of answer options, Microsoft-Decision-1 provides a calibrated probability score for each option. The model supports yes/no, multiple-choice, and rating options, as well as rubric-based grading of AI responses and agent actions, all through a simple structured API call.</p>\n<p>To build a reliable decision model, we had to address several challenges:</p>\n<h3 id=\"1-speed\">1. Speed</h3>\n<p>Each decision adds delay, especially when one step depends on another. For example, adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow.</p>\n<p>Microsoft-Decision-1 P50 latency is ~35x faster than GPT-6 Sol.</p>\n<h3 id=\"2-quality-that-generalizes\">2. Quality that generalizes</h3>\n<p>It’s easy to overfit a model for one benchmark or one type of decision task. We need to know whether that quality carries over to various tasks the model wasn’t trained on. That’s why we evaluated Microsoft-Decision-1 across dozens of benchmarks kept blinded from training, spanning routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning, and safety. We also took several of the top public models on the <a href=\"https://benchmarkheaven.com/jev-models\" rel=\"nofollow ugc noopener\">popular open leaderboard JevBench</a> and tested them across 36 additional public and private benchmarks. Microsoft-Decision-1 performed the best across these broader sets of benchmarks, demonstrating strong generalization.</p>\n<h3 id=\"3-robustness\">3. Robustness</h3>\n<p>Equivalent inputs should produce equivalent decisions. In production, states and instructions get paraphrased, option descriptions change, choices are reordered, keys change, and harmless formatting noise appears. None of those changes should materially alter the decision.</p>\n<p>We perturb the same request in eight ways and measure how often the decision flips. Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled.</p>\n<h3 id=\"4-probability-and-confidence-calibration\">4. Probability and confidence calibration</h3>\n<p>The probability itself is part of the API, not just a ranking score. Applications use confidence to decide when to act, defer, or ask for review, so a 90% prediction should be right about nine times out of 10 on representative cases.</p>\n<h3 id=\"5-safety\">5. Safety</h3>\n<p>A decision model should recognize harmful requests without needlessly blocking harmless ones. We tested Microsoft-Decision-1 on 5,250 requests across 11 benchmarks, covering harmful content, jailbreak attempts, and prompt injection, and found that the model successfully refused harmful behavior while retaining a high degree of utility.</p>\n<h2 id=\"demo-examples\">Demo examples</h2>\n<p>Classification is a key use case for decision models. Check out how accurately and quickly Microsoft-Decision-1 can categorize a variety of queries compared to GPT-6 Sol:</p>\n<p>Decision models can also be efficient for computer use scenarios. This demo shows how fast Microsoft-Decision-1 can complete the task of buying a backpack compared to GPT-6 Sol:</p>\n<h2 id=\"how-we-re-testing-microsoft-decision-1-internally\">How we’re testing Microsoft-Decision-1 internally</h2>\n<p>Here are some of the ways we’ve been testing Microsoft-Decision-1 internally, with a lot more to come.</p>\n<h3 id=\"labeling-data\">Labeling data</h3>\n<p>XBOX Research used Microsoft-Decision-1 to process more than 10,000 open-ended pieces of feedback and reviews from surveys, STEAM, and Twitter/X and sort them into a fixed set of themes established by researchers to understand what people are saying about different games, launches, streams, and more. They found Microsoft-Decision-1 to be competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive.</p>\n<h3 id=\"quality-control\">Quality control</h3>\n<p>The Copilot team measures the quality of chat and agentic responses. Their testing found Microsoft-Decision-1 to be competitive with GPT5.6 Luna and 100 times faster.</p>\n<h3 id=\"incident-response\">Incident response</h3>\n<p>Our on-call engineers use AI to retrieve relevant knowledge to respond to live incidents across logs, ticketing systems, calls, messages, and other data sources. Microsoft-Decision-1 performed better and faster than an LLM for knowledge retrieval.</p>\n<h3 id=\"scientific-discovery\">Scientific discovery</h3>\n<p>Microsoft Discovery implements an adaptive replanning feature where an agent evaluates a previous experiment, revises its approach based on rubric grades, and repeats until it has completed its objectives. Microsoft-Decision-1 scored as 46 times more consistent than the LLM-based score at three times the speed and resulted in nearly four times the speed on adaptive replanning</p>\n<p>Faster, more reliable planning could significantly impact outcomes for long-running scientific experiments.</p>\n<p>Those are just a small handful of examples. There are many more potential use cases for Microsoft-Decision-1. Consider trying out the following:</p>\n<ul><li><strong>Agent controls:</strong> Evaluate an agent’s proposed next step and decide whether to continue, stop, retry, or hand off to a model, tool, or human.</li><li><strong>Model routing:</strong> Evaluate an incoming request and select the best model for the task based on quality, cost, and latency requirements.</li><li><strong>Skill-based decisions:</strong> Apply the rules from an agent skill to select the appropriate next action without repeatedly processing a long list of instructions.</li><li><strong>Data labeling:</strong> Assign consistent labels to social media posts, customer feedback, and other data for analysis or training.</li><li><strong>AI judging:</strong> Evaluate an AI-generated response against defined quality criteria and decide whether to accept, revise, or reject it.</li><li><strong>Intent analysis:</strong> Identify what a user is trying to accomplish and match their request to a supported intent.</li><li><strong>Incident response routing:</strong> Classify an incident by type and urgency, then route it to the appropriate team or workflow.</li><li><strong>Data validation:</strong> Check whether an input meets defined requirements and decide whether to accept it, reject it, or flag it for review.</li><li><strong>Recommendations:</strong> Select the most relevant item, offer, or next action from a set of candidates.</li><li><strong>Search relevance:</strong> Assess how well a search result matches a query and assign a relevance label or priority.</li><li><strong>Content classification and filtering:</strong> Categorize content by topic or policy and decide whether to display it, filter it, or send it for review.</li><li><strong>Code scanning:</strong> Evaluate code against defined criteria and flag potential defects or policy violations for review.</li><li><strong>Safety and security screening:</strong> Classify requests, outputs, or proposed actions by risk and decide whether to allow, block, or escalate them.</li><li><strong>Computer and UI use:</strong> Select the next interface action from a set of options based on the current screen and task.</li><li><strong>Robotics:</strong> Select among predefined robot actions based on observations, task goals, and operating constraints.</li><li><strong>Scientific discovery:</strong> Screen candidate hypotheses, compounds, or experiments against defined criteria and prioritize them for further evaluation.</li></ul>\n<h2 id=\"getting-started\">Getting started</h2>\n<p>Developers can get started with Microsoft-Decision-1 today in Microsoft Foundry here: <a href=\"https://ai.azure.com/catalog/models/Microsoft-Decision-1\" rel=\"nofollow ugc noopener\"><a href=\"https://ai.azure.com/catalog/models/Microsoft-Decision-1\" rel=\"nofollow ugc noopener\">https://ai.azure.com/catalog/models/Microsoft-Decision-1</a></a>.</p>\n<h3 id=\"pricing\">Pricing</h3>\n<p>Input tokens cost $0.042 USD per million tokens. Output tokens are free.</p>\n<h2 id=\"looking-ahead\">Looking ahead</h2>\n<p>Now that agentic AI is a reality, we’ve seen that cost plays a major role in how people decide to use AI. And it’s increasingly important to choose the right model for the right job. With agents taking action and making an impact in the real world, decision models have the potential to help people guide and control those agents through complex environments.</p>\n<p>We look forward to seeing what developers build with Microsoft-Decision-1 and hearing their feedback. We’ll continue to release updates to the model, including by incorporating evaluations and data to further optimize quality, confidence, and cost.</p>\n<h3 id=\"appendix-models-benchmarked\">Appendix: Models benchmarked</h3>\n<p>Quyet-1.0-Large — <a href=\"https://huggingface.co/chinhnc/Quyet-1.0-Large\" rel=\"nofollow ugc noopener\"><a href=\"https://huggingface.co/chinhnc/Quyet-1.0-Large\" rel=\"nofollow ugc noopener\">https://huggingface.co/chinhnc/Quyet-1.0-Large</a></a> </p>\n<p>Surogate Rune 26B-A4B — <a href=\"https://huggingface.co/surogate/rune-26b-a4b-GGUF\" rel=\"nofollow ugc noopener\"><a href=\"https://huggingface.co/surogate/rune-26b-a4b-GGUF\" rel=\"nofollow ugc noopener\">https://huggingface.co/surogate/rune-26b-a4b-GGUF</a></a> </p>\n<p>GPT-6 Luna Decisions — <a href=\"https://developers.openai.com/api/docs/guides/decisions\" rel=\"nofollow ugc noopener\"><a href=\"https://developers.openai.com/api/docs/guides/decisions\" rel=\"nofollow ugc noopener\">https://developers.openai.com/api/docs/guides/decisions</a></a> </p>\n<p>deck-31B — <a href=\"https://github.com/krishna-gogineni-765/deck31b\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/krishna-gogineni-765/deck31b\" rel=\"nofollow ugc noopener\">https://github.com/krishna-gogineni-765/deck31b</a></a> </p>\n<p>H2O-Lightning-4B — <a href=\"https://huggingface.co/h2oai/h2o-lightning-4b\" rel=\"nofollow ugc noopener\"><a href=\"https://huggingface.co/h2oai/h2o-lightning-4b\" rel=\"nofollow ugc noopener\">https://huggingface.co/h2oai/h2o-lightning-4b</a></a> </p>\n<p>Strands-Decider 2B — <a href=\"https://github.com/strands-labs/strands-decider\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/strands-labs/strands-decider\" rel=\"nofollow ugc noopener\">https://github.com/strands-labs/strands-decider</a></a></p>","headings":[{"level":2,"text":"How Microsoft-Decision-1 works","id":"how-microsoft-decision-1-works"},{"level":3,"text":"1. Speed","id":"1-speed"},{"level":3,"text":"2. Quality that generalizes","id":"2-quality-that-generalizes"},{"level":3,"text":"3. Robustness","id":"3-robustness"},{"level":3,"text":"4. Probability and confidence calibration","id":"4-probability-and-confidence-calibration"},{"level":3,"text":"5. Safety","id":"5-safety"},{"level":2,"text":"Demo examples","id":"demo-examples"},{"level":2,"text":"How we’re testing Microsoft-Decision-1 internally","id":"how-we-re-testing-microsoft-decision-1-internally"},{"level":3,"text":"Labeling data","id":"labeling-data"},{"level":3,"text":"Quality control","id":"quality-control"},{"level":3,"text":"Incident response","id":"incident-response"},{"level":3,"text":"Scientific discovery","id":"scientific-discovery"},{"level":2,"text":"Getting started","id":"getting-started"},{"level":3,"text":"Pricing","id":"pricing"},{"level":2,"text":"Looking ahead","id":"looking-ahead"},{"level":3,"text":"Appendix: Models benchmarked","id":"appendix-models-benchmarked"}]}}