OpenAI's Decisions API Enters Public Beta With 10x Faster Typed Answers
OpenAI pushed its Decisions API into public beta on Oct. 6, 2026, returning typed answers — probabilities, choices, and scores — in roughly 150 ms, about ten times faster than the Responses API, at $0.10 per 1M input tokens.
Updated
Why it matters
- Decisions API entered public beta on Oct. 6, 2026, per OpenAI's developer documentation
- OpenAI claims Decisions runs about 10x faster than the Responses API; DevDay coverage placed one decision near 150 ms versus roughly 1.6 seconds
- Pricing is $0.10 per 1M input tokens; output, cache-read, and cache-write charges are $0
- The only model today is gpt-6-luna, with a 1,050,000-token context window and 2x input pricing above 272,000 tokens
- Closest rival TypeSafe Jev costs $0.042 per 1M input tokens, about 2.4x cheaper than OpenAI Decisions
OpenAI pushed its Decisions API into public beta on Oct. 6, 2026, returning typed answers — probabilities, choices, and scores — in roughly 150 milliseconds per call, about ten times faster than the company's Responses API on the same gpt-6-luna model, per OpenAI's developer documentation.
The endpoint targets a recurring pain point for application developers: prompt a large language model, then parse its free-text output into a label a program can branch on. Decisions replaces the round trip with a single request that returns structured values directly.
Why it matters: OpenAI is recasting Luna, a multimodal model with a 1,050,000-token context window, as a primitive for classification-style work, going up against purpose-built rivals on cost and latency. The product lands three weeks after GPT-6 Luna reached general availability on Sept. 22, 2026.
What does the Decisions API actually return?
The endpoint does not generate prose. It accepts text, images (inlined as base64), or both, and answers a list of questions in one call. A request carries shared evidence plus named questions; the response returns an answers array keyed by those names.
Three question types ship today:
- predicate — checks a condition and returns a
probabilitybetween 0 and 1. Example: does a product photo show a crack, tear, or dent? - choice — picks one value from a developer-supplied list. It also returns per-option
probabilitiesand aconfidencefield. - score — rates input against ordered
levelsindexed from 0. The score is a probability-weighted average of those indices.
OpenAI's severity example makes the math concrete. Level probabilities of 0.1, 0.7, and 0.2 yield a score of 1.1, a value that sits between "Workaround available" and "Fully blocked." That weighting lets callers set thresholds on a continuous scale rather than over discrete labels.
How fast is it, and what is the evidence?
OpenAI's documentation claims "about 10x faster" responses than the Responses API. DevDay coverage placed a decision call near 150 ms versus roughly 1.6 seconds for a regular gpt-6-luna request through Responses — a figure consistent with OpenAI's claim.
The vendor has not published accuracy or calibration data for the endpoint. The developer docs tell callers to set thresholds against labeled examples drawn from their own application traffic.
Calibration is the unresolved question. Probabilities are useful only when they track observed hit rates. Without independent benchmarks, teams shipping Decisions into customer-facing systems will have to run their own.
What does it cost, and where can it run?
Pricing for gpt-6-luna on the Decisions endpoint is $0.10 per 1M input tokens. There are no output-token, cache-read, or cache-write charges. Regional processing premiums and long-context multipliers still apply; gpt-6-luna's model card prices prompts above 272,000 tokens at 2x input rates.
Deployment details, per OpenAI's documentation:
- Endpoint:
POST /v1/decisions, plus a Playground. - SDK minimums: Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, Java 4.78.0.
- Compliance: Zero Data Retention and HIPAA, for eligible customers.
- Residency: United States and Europe (EEA plus Switzerland).
- Voice: Decisions can drive actions via client delegation with the Live API.
The absence of output-token charges flips the cost calculus. A classification pass that bundles prompts and images now bills like text alone, changing break-even math for image-heavy moderation or product-tagging pipelines.
When should developers still use Structured Outputs or function calling?
OpenAI draws explicit lines between its three structured interfaces:
- Use Decisions for probabilities, choices, or scores.
- Use Structured Outputs when callers need to fill a developer-supplied JSON schema or want the model to write explanations.
- Use function calling when the model must select a tool and supply arguments.
That three-tier menu signals OpenAI is partitioning API surface by output type rather than by task. Each tier hands downstream code a single claim — probability, schema, or tool — without a parser.
How does OpenAI Decisions compare with TypeSafe Jev?
The closest rival is TypeSafe Jev, launched Sept. 15, 2026. Jev is what TypeSafe calls a "System One" model, returning typed values with calibrated probabilities. TypeSafe prices input at $0.042 per 1M tokens, with output free, putting OpenAI's base rate about 2.4x higher. TypeSafe reports 70 to 500 ms end-to-end latency, measured from the US West Coast.
Jev supports up to 255 choices per question. OpenAI has not disclosed the per-question option cap on Decisions. Jev remains in early access; Decisions is in public beta.
| Feature | OpenAI Decisions API | TypeSafe Jev 1.13 | GPT-6 Luna via Responses |
|---|---|---|---|
| Release | Beta, Oct 6, 2026 | Early access, Sep 15, 2026 | GA, Sep 22, 2026 |
| Underlying model | gpt-6-luna |
Jev (System One) | gpt-6-luna |
| Output | Probability, choice, or score | Typed values with probabilities | Generated text (JSON via Structured Outputs) |
| Max options per question | Not disclosed | Up to 255 | N/A |
| Speed (vendor-stated) | ~10x faster than Responses | 70–500 ms end-to-end | Baseline |
| Input price per 1M tokens | $0.10 | $0.042 | $0.10 |
| Output price per 1M tokens | $0 | $0 | $0.50 |
| Compliance | ZDR, HIPAA (eligible); US/EU residency | Not disclosed | EU data residency available |
OpenAI's edge: image input, named compliance offerings, and a beta any developer with API access can hit. TypeSafe's edge is price and a disclosed latency range drawn from public tests. Jev has shipped less into production than Luna has, so the head-to-head will wait for independent benchmarks.
What comes next?
OpenAI's documentation frames general availability as "in the coming weeks" from beta. Pricing of $0.10 input and $0 output gives OpenAI room to drop rates if Jev or another System One entrant pulls classification workloads off Luna before GA. The single-model beta also creates concentration risk: until OpenAI ships another backend, every Decisions call funnels through Luna.
The bigger watch item is calibration. A probability that consistently returns 0.7 on a binary predicate with a 50% true rate is worse than useless — it is silently wrong. OpenAI's call for application-specific labeled examples amounts to an admission that the company has not yet published independent eval results on Decisions accuracy. Until those arrive, adoption inside customer-facing flows will hinge on the thresholds each team sets itself.
Original: developers.openai.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
185 articles
Related articles
- OpenAI Launches GPT-6 Sol and Luna at Half the Price
- OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
- OpenAI Adds Remote MCP Support and New Built-In Tools to Responses API
- OpenAI ships GPT-5.4 mini and nano for coding, tool use, and agent workloads
- OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board