OpenAI's new Decisions API runs many AI requests in parallel. Ari Weinstein says model gains, the orchestration code around the model, and direct screen and app access are stacking, and computer use is 180 degrees different than three months ago.
Three months ago, Dwarkesh Patel posted a question that ruffled feathers in the Computer Use community: Why has progress on computer use been so slow? Computer use is so clearly verifiable. In a podcast recorded at OpenAI DevDay 2026, Ari Weinstein argued directly against that claim. "They said that a few months ago," Weinstein said of Dwarkesh's claim. "I think Computer Use is 180 degrees different than it was."
Weinstein's core argument is that the past year has brought compounding improvements across multiple axes simultaneously — model capability, harness engineering, and multimodal sensing — rather than any single breakthrough. "The models, just in the last one year, have become extraordinarily capable at Computer Use," he said. "The biggest delta is that before they could reliably start tasks, but then they would run into problems. Now they're really good at debugging, really good at trying again, introspecting what is and isn't working."
A key distinction Weinstein draws is between improvements driven by the model itself and improvements driven by the surrounding deployment harness. On September 4, 2026, he posted publicly: "With Astra, ChatGPT is nearly 2x faster at computer use than before. In addition to the amazing new model, we've optimized the harness to make computer use run much faster. These optimizations work in existing models too, so tasks are sped up by ~60% in GPT-5.6 Sol." The nearly-2x figure represents end-to-end experience with the new Astra model; the ~60% figure isolates the harness optimization contribution on the prior-generation model, holding the model fixed. The two claims measure different things and are not interchangeable. (source)
The practical impact shows up in anecdote: Weinstein described spending two hours manually customizing a meal-prep order with granular macro-nutrient specifications. "I found that I could ask Computer Use to do it for me, and it did it in 15 minutes," he said. "It both did it eight times faster than I could, and it saved me two hours on GPT-6.1 Sol." This is a single self-reported data point, not a benchmark.
The podcast also covers OpenAI's Decisions API, a new fast-inference layer the company announced at DevDay. Weinstein describes it as architecturally distinct from the larger reasoning models used for complex Computer Use tasks: "It does inference in parallel. It doesn't have reasoning. It's a smaller model than the ones we use for Computer Use, and so those capabilities make it really fast." He added that the tradeoff is that it is "a little bit less good at doing long-horizon, sort of sophisticated tasks," and that "it's still an open area of research for how we bring those approaches together." (source)
The second half of the episode, featuring Nikunj Handa from OpenAI's API team, reveals the unusually fast development sprint behind the Decisions API. According to the transcript, OpenAI developed what is described as a Luna-model wrapper implementing a Jev-like pattern in approximately one week and announced it as a limited-preview launch. (source — confirmed via transcript segments 00:23:21–00:32:23)
When asked about the trajectory from human-level to superhuman Computer Use, Weinstein pointed to the shift from single-step screenshot-and-act loops to richer environmental representation. "With accessibility and direct access to the DOM and other things, the language model can actually see an entire page or an entire application," he said. "It can write code that can do multiple steps at once." He cited App Shots — capturing rich context from an application and bringing it into Codex or ChatGPT quickly — as a concrete example of the multimodal approach. "The model may use screenshots, it may use accessibility, it may use Playwright. It can use a lot of different mechanisms based on the task at hand."