Models

OpenAI's Decisions API chases TypeSafe's Jev in fast agent control

OpenAI's Decisions API mirrors TypeSafe AI's Jev. Fast decision models could monitor swarming agents for $2.94 versus $372 with a frontier LLM, developers say.

By Sophie Lindqvist5 min read

Updated

Why it matters

  • OpenAI announced a 'Decisions API' at Dev Day on Tuesday, described by Sam Altman as giving the company's Luna model a predefined set of options to choose between.
  • TypeSafe AI CEO Diogo Almeida says his company's moat is synthetic data that produces statistically useful outputs; his North Star is 'pushing the intelligence-per-dollar Pareto curve.'
  • A hackathon demo by QueryStory's Shapor Naghibzadeh used Jev to monitor agentic actions for $2.94, versus $372 with a frontier LLM.

OpenAI CEO Sam Altman used an aside at the company's Dev Day event on Tuesday to announce a "Decisions API" — a product that appears to replicate the core function of Jev, a decision model released by TypeSafe AI earlier this month.

The announcement matters because it marks the first time a frontier lab has moved to productize a new class of model: small, fast classifiers built on LLM technology that developers can use to cheaply route, gate, and police software. That class of model, barely weeks old, is already shaping up as a building block for controlling AI agents at scale.

What the Decisions API does

Altman described the API as a way to give OpenAI's Luna model a predefined set of options to choose between — categories for classifying an image, for example, or different agent behaviors.

"By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections," Altman said at the event.

That description tracks closely with Jev, which TypeSafe AI bills as a model explicitly designed for software automation. Jev works as a kind of super-powered classifier built on an LLM: developers hand it a set of choices, and it outputs those choices as probabilities, cheaply and at high speeds.

OpenAI has released Decisions API only as a limited preview, and TechCrunch reports it has not yet spotted developers running the product through its paces. How closely it matches Jev's capabilities, including its calibration quality, remains unknown. Conversations on X, however, show clear interest in the product.

"The clone wars"

TypeSafe did not respond to TechCrunch's questions about OpenAI's new product. But the startup's CEO, Diogo Almeida — a former OpenAI engineer who, according to the source, co-invented reinforcement learning — joked on X about the beginning of the clone wars.

He added that OpenAI's interest could be "a sign…that building in a System One compatible way is the future." "System One" is TypeSafe's term of art for fast, intuitive thinking, as opposed to "System 2," which the company applies to deliberate reasoning.

The subtext is straightforward. LLMs as we know them are the wrong tool for a lot of software because they are comparatively slow and expensive. Developers have already been using Jev to augment LLMs in their pipelines, and in doing so they have found their systems run faster and cheaper.

Decisions API is not the only Jev-like API on the internet. Other startups are rolling out similar models, and OpenAI will almost certainly not be the last tech giant to build one. The key open question, per the source, is how well calibrated each of these decision models' outputs will be to real life — how closely the probabilities they emit correspond to actual outcomes.

Intelligence per dollar

Almeida argues that calibration, not speed, is the differentiator. His company's moat, he says, is the synthetic data TypeSafe creates to generate statistically useful outputs.

"Fast and cheap is very easy, you know," Almeida told TechCrunch last week. "If you want it really fast and cheap, use dice, right? Intelligence is the hard part, and my North Star is always pushing the intelligence-per-dollar Pareto curve."

The distinction matters for buyers. A decision model that returns probabilities quickly but poorly calibrated is worse than useless in agent pipelines, because it injects confident-looking noise into systems that act autonomously. Whoever wins the calibration race — measured against real-world outcomes — likely wins the category.

The agent security application

The most concrete near-term application for these models may be monitoring and securing AI agents. OpenAI has a direct incentive here: following a series of incidents in which its agents misbehaved on the open internet, the company has introduced new security measures, one of which uses a separate model to watch for bad actions at what the source describes as "significant compute cost."

Shapor Naghibzadeh, a long-time cybersecurity professional who leads the startup QueryStory, believes a model like Jev can deliver that kind of oversight far more cheaply. He built a demo for a hackathon held last weekend that uses Jev to check each agentic action against the task the agent was given. The system blocks actions it has high confidence are bad, flags others for human review, and permits the rest.

The numbers are stark. In theory, such monitoring could have stopped the Hugging Face incident — and monitoring of that kind costs $2.94 with Jev, versus $372 with a frontier LLM.

That cost gap is the whole story. At $2.94, a review layer on every single agentic action becomes economically viable. At $372, it does not. If agent deployments continue to grow, per-action monitoring at the cheaper price point offers a path to more reliable autonomous systems across the board — the kind of outcome TypeSafe was hoping to achieve when it built Jev, and one OpenAI now appears to have seen the value in as well.

Why it matters

The stakes run in two directions. Commercially, a new category is forming around decision models, with a three-week-old startup, OpenAI, and unnamed others competing on calibration and intelligence per dollar. Operationally, the technology offers a credible answer to one of the most urgent problems in deployed AI: agents that act on the open internet faster and more erratically than human oversight can track.

OpenAI's entry, even in limited preview, validates the approach. The open question is whether a frontier lab's distribution can outrun a startup's data advantage — or whether calibration, as Almeida bets, proves to be the moat that matters. With agent incidents already on the record and monitoring costs now quantified at $2.94 versus $372 per check, the market for cheap, reliable decision models looks set to grow rather than shrink.

Original: openai.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

148 articles

Related articles

  1. OpenAI Adds Remote MCP Support and New Built-In Tools to Responses API
  2. OpenAI ships a model-native harness and native sandboxes for its Agents SDK
  3. OpenAI Ships AgentKit, Expanded Evals, and Reinforcement Fine-Tuning for Agents
  4. OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
  5. OpenAI details codex-1: an o3 variant tuned for real coding work

« Previous article