OpenAI's Decisions API Mirrors TypeSafe's Jev as Low-Cost Agent Monitoring Emerges
OpenAI unveiled a Decisions API at Dev Day that mirrors TypeSafe's Jev, a fast, cheap classifier now being tested as a low-cost way to monitor and secure AI agents.
Like Jev, the Decisions API works as a kind of classifier built on a language model. Developers supply a set of choices and the model returns them as probabilities, cheaply and at high speed. "By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections," Altman said.
TypeSafe did not respond to TechCrunch's questions about the new product. Its chief executive, Diogo Almeida, a former OpenAI engineer who co-invented reinforcement learning, joked on X about the beginning of the clone wars. He added that OpenAI's interest could be "a sign…that building in a System One compatible way is the future." System One is TypeSafe's term for fast, intuitive thinking, as opposed to System Two, which it applies to deliberate reasoning.
The subtext is that large language models as they are currently built are not the right solution for a lot of software, because they are comparatively slow and expensive. Developers have been using Jev to augment LLMs and have found the combination faster and cheaper.
How closely the Decisions API resembles Jev is not yet clear. OpenAI released it as a limited preview, and TechCrunch reported that it has so far not spotted developers running it through its paces, though the conversation on X indicates interest. The Decisions API is not the only Jev-like product available; other startups are rolling out similar models, and OpenAI is unlikely to be the last technology giant to produce one. According to TechCrunch, a key question is how well calibrated each of these decision models' outputs will be to real life.
Almeida said TypeSafe's moat is the synthetic data it creates to generate statistically useful outputs. "Fast and cheap is very easy, you know," he told TechCrunch last week. "If you want it really fast and cheap, use dice, right? Intelligence is the hard part, and my North Star is always pushing the intelligence-per-dollar Pareto curve."
Monitoring and securing AI agents looks like one likely application. Following a series of incidents in which its agents misbehaved on the open internet, OpenAI adopted new security measures that include using a separate model to watch for bad actions at what the company describes as significant compute cost.
Shapor Naghibzadeh, a longtime cybersecurity professional who leads the startup QueryStory, argues that a model like Jev could make such monitoring far cheaper. He built a demo for a hackathon held last weekend that uses Jev to check each agentic action against the task it was given, blocking actions it had high confidence were bad, flagging others for review and permitting the rest. In theory, monitoring of that kind could have stopped the Hugging Face incident, and it costs $2.94 with Jev versus $372 with a frontier LLM.
The observation from that experiment is that Jev is arguably cheap enough to run on every agentic action, which offers a layer of review that could improve the reliability of agents more broadly. That is the outcome TypeSafe was hoping to achieve, and OpenAI now appears to see the value as well.