Back

Microsoft's new AI model doesn't write, it picks: a fast, cheap way for apps to make quick yes-or-no and sorting decisions

Microsoft's demo screen comparing Microsoft-Decision-1 with OpenAI's GPT-6 Sol as both sort the same ten work requests into categories, finishing in 182 milliseconds and 6.24 seconds in this demo
Image: Microsoft

Microsoft has released Microsoft-Decision-1, an AI model that makes quick calls inside apps and AI agents, such as sorting a request or checking an agent's next step, by choosing from a fixed set of answers instead of writing text. Developers can use it now, for less than half what OpenAI charges for a similar service. Microsoft says it beat rivals on speed and accuracy, but those are its own tests, and no one outside the company has checked them yet.

Microsoft on Friday, Oct. 9, released Microsoft-Decision-1, an AI model that does not write anything. Give it a question and a fixed set of answers, such as yes or no, a list of categories or a rating scale, and it returns how likely each answer is to be right, in a form other software can act on straight away. Microsoft calls these decision models and says they are quickly becoming an important new kind of AI. They handle the small, quick judgments inside apps and AI agents, programs that carry out a task step by step on their own, such as which team should get an urgent support ticket, or whether an agent should go ahead with its next step, stop or hand the job to a person.

Developers can use it now in Microsoft Foundry, the company's service for building with AI models. Microsoft charges 4.2 cents for every million tokens it reads, and nothing for its answers. Tokens are the small chunks of text, often parts of words, that AI models read and write. OpenAI charges 10 cents per million tokens for its own decision service, still a test version, which runs on its GPT-6 Luna model. OpenAI charges more for some regions and for very long inputs.

Microsoft says its model scored highest on average across 36 tests it ran, nearly 150,000 questions kept out of its training: 83.5 percent, against 81.9 percent for the next best, Quyet-1.0-Large, a rival decision model and the top one on a public ranking of these tools, and 79.4 percent for OpenAI's decision service. Its typical answer takes 85 milliseconds, under a tenth of a second, measured on Microsoft's own service, while rivals' times come from that public ranking. Microsoft says that makes it is 4.5 times quicker than that runner-up and 35 times quicker than OpenAI's GPT-6 Sol, a full AI model that writes and reasons. An agent may make many of these decisions one after another, so small delays add up: Microsoft notes that a tenth of a second on each of 20 steps adds two seconds to the job. Amazon's Strands-Decider 2B could answer only 23 of the 36 tests, averaging 54.8 percent on those.

Microsoft says the odds the model gives can be taken at face value, so an app can act when the model is sure and pass the case to a person when it is not: an answer given 90 percent odds should be right about nine times out of 10. On Microsoft's own measure of this, it came second, just behind Quyet-1.0-Large. Inside the company, the Xbox research team used it to sort more than 10,000 game reviews, survey answers and social media posts into themes, and found it about as good as GPT-6 Sol while running more than 14 times faster at a 200th of the cost, Microsoft says.

These results are Microsoft's own, and no one outside the company has checked them yet. The model is not built from scratch: Microsoft took an existing model, Qwen3.5-9B, and trained it further for this one job, and says it will soon rebuild it on other models, including its own and OpenAI's.

It is not yet clear how well the model holds up on companies' own tasks, whether its results shift as Microsoft keeps updating it, or how OpenAI's decision service compares in tests that neither company runs.

Correction (Oct. 10): an earlier version said both companies charge more in some cases. Microsoft lists one flat price; only OpenAI charges more for some regions and for very long inputs.

More on Microsoft