TypeSafe's fast action "System One" model, Jev, already returns probabilities over a fixed menu of options, and whether the cheap layer's confidence scores actually hold up on real work is still unproven.
OpenAI showed off a new tool at Dev Day 2026 that does something the company's own researchers have argued against for years: it deliberately routes most of the decisions inside an AI agent through a smaller, cheaper model instead of a frontier LLM. The product is called the Decisions API, and it works by handing the model a predefined menu of options and getting back a probability for each one, instead of asking the model to generate free-form text from scratch.
That pattern has a name. TypeSafe, a small AI safety startup, has been calling it "System One," a reference to the psychologist Daniel Kahneman's split between fast, intuitive thinking and slow, deliberate reasoning. The company shipped a model called Jev that does exactly this. The structural implication is the same one TypeSafe has been publishing about for months: a single big model per call is the wrong shape for most of the small decisions an agent has to make.
The Decisions API is built around a tradeoff. A large language model that can reason through a complex question is expensive and slow, measured in seconds and cents per call. For a software agent that has to make hundreds of small decisions a minute ("should I open this file, run this command, post this result"), that bill and latency add up fast. The cheaper alternative is to ask a smaller model, often fine-tuned for the task, to score a short list of possible actions and return probabilities. The smaller model is wrong more often, but it is also fast enough and cheap enough to run on every action an agent takes.
The question OpenAI and TypeSafe are both trying to answer is whether the cheap layer can be made reliable enough to trust. TypeSafe's Jev has been put through a public demo called Jev Sentinel, which checks proposed actions against the task's stated scope and routes them to allow, block, or human review. The demo's recorded statistics show 53,870 actions classified as "swarm" traffic, with 45,190 routed to block, 7,781 to ask a human, and 899 allowed. The median check took 230 milliseconds. On a separate coding-agent hook, the demo reports 838 coding actions with 748 allowed, 76 escalated, and 14 blocked at a 156 ms median.
Those numbers are not an independent ground truth. The Jev Sentinel demo itself notes that the share of traffic that is actually attack traffic is uncertain, and the early counters in the UI started at zero as placeholders. What the demo does show is that a small decision layer can be wired into a coding agent hook and produce classifications at useful speeds. The 7,781 "ask" cases are also not the same thing as automatic prevention; they are requests for human review, and an agent whose decision layer escalates a large share of its work has not solved the cost problem it set out to solve.
TypeSafe's introduction to Jev describes training the model with reinforcement learning from constrained decoding, which forces its output to be valid structured data every time. The company is careful not to call that a guarantee against all model errors. The structured output means the model cannot return malformed JSON; it does not mean the probabilities it returns are correct judgments about the world. TypeSafe's own workflow evaluation uses averaged probabilities from other models as a stand-in for ground truth and acknowledges that the workload and comparison methods are limited.
The structural read here is that the agent stack is bifurcating. A frontier model still does the slow work: the planning, the hard reasoning, the conversations. A cheap classifier handles the flood of small decisions in between. The competitive question shifts from "how big is your biggest model" to "how well-calibrated is your cheap layer," where calibration means a returned probability of, say, 0.7 actually corresponds to 70% of those actions being correct across thousands of real deployments.
The hard part is proving calibration on someone else's workload. A model that returns "70% chance this action is allowed" needs that number to actually mean something outside the demo. The Decisions API has not yet been tested in public the way Jev Sentinel has, and TechCrunch's reporting notes that similarity to Jev is unconfirmed and no third-party developer testing has been observed.
What to watch next: whether OpenAI publishes real workload numbers, latency floors, and pricing for the Decisions API that anyone outside the company can verify, and whether the calibration question gets a public answer before the next wave of agent deployments goes live.