OpenAI Ships o1 to Developers With 60% Cheaper Audio
OpenAI's o1 reasoning model hits the API with function calling and vision, as GPT-4o audio prices drop 60% and Preference Fine-Tuning debuts for subjective customization tasks.

Updated
Why it matters
- OpenAI released o1-2024-12-17 in the API for tier 5 developers, with function calling, Structured Outputs, developer messages, vision, and 60% fewer reasoning tokens than o1-preview
- GPT-4o audio prices dropped 60% to $40/1M input and $80/1M output tokens; cached audio input fell 87.5% to $2.50/1M, and GPT-4o mini realtime audio runs at roughly one-tenth of GPT-4o rates
- Preference Fine-Tuning, built on Direct Preference Optimization, is live for gpt-4o-2024-08-06; partner Rogo AI improved accuracy from 75% to over 80% on its Rogo-Golden benchmark
OpenAI has released its o1 reasoning model to developers in the API, pairing the launch with a 60% price cut on GPT-4o audio, native WebRTC support, and a new fine-tuning technique based on Direct Preference Optimization. The December 17 announcement, aimed at usage tier 5 developers, turns o1 from a research preview into a production-ready model with function calling, Structured Outputs, developer messages, and vision capabilities.
The rollout matters for a developer ecosystem increasingly split between fast, cheap general-purpose models and slower, more accurate reasoning models. OpenAI is betting o1 can serve both camps: it claims the new snapshot uses on average 60% fewer reasoning tokens than o1-preview for a given request, directly addressing the latency and cost complaints that made the preview impractical for many production workloads.
o1 goes production-ready
The model shipping in the API, o1-2024-12-17, is a new post-trained version of the model OpenAI released in ChatGPT two weeks earlier. According to the company, it improves model behavior based on feedback while maintaining the frontier capabilities documented in the o1 System Card. OpenAI says it will update o1 in ChatGPT to this version soon.
The API release adds the plumbing developers need to ship real applications. Function calling lets o1 connect to external data and APIs. Structured Outputs makes the model adhere to custom JSON schemas. Developer messages allow engineers to specify instructions, tone, and behavioral context for the model. Vision capabilities let o1 reason over images, which OpenAI frames as opening applications in science, manufacturing, and coding where visual inputs matter.
A new reasoning_effort API parameter gives developers direct control over how long the model thinks before answering — a lever for balancing accuracy against cost and latency that was missing from the preview.
The benchmark numbers
OpenAI reports that o1-2024-12-17 sets new state-of-the-art results on several benchmarks. The gains over o1-preview are largest on the hardest evaluations.
On AIME 2024, a competition mathematics benchmark, pass@1 accuracy jumped from 42.0 to 79.2. On the MATH benchmark, it rose from 85.5 to 96.4. Coding performance improved sharply: SWE-bench Verified climbed from 41.3 to 48.9, and LiveBench Coding jumped from 52.3 to 76.6. On GPQA diamond, a graduate-level science question set, the new snapshot scored 75.7 against the preview's 73.3.
The vision-capable model posts 77.3 on MMMU and 71.0 on MathVista — evaluations o1-preview did not report. On agentic benchmarks, o1 scores 73.5 on TAU-bench (retail) and 54.2 on TAU-bench (airline). One figure barely moved: SimpleQA factuality inched from 42.4 to 42.6, a reminder that reasoning strength does not automatically translate into factual grounding. OpenAI also says the new snapshot "significantly outperforms gpt-4o" in its internal function calling and Structured Outputs testing, though it did not publish those numbers.
OpenAI notes that developers have already used o1-preview to build agentic applications for customer support, supply chain optimization, and financial forecasting — the use cases the production release is designed to serve.
Realtime API gets WebRTC and steep price cuts
OpenAI's Realtime API, used for low-latency voice applications such as assistants, live translation, and virtual tutors, received three upgrades: WebRTC integration, lower prices, and more control over responses.
WebRTC support is the headline feature. The open standard handles audio encoding, streaming, noise suppression, and congestion control, and OpenAI says its integration keeps interactions smooth "even with variable network quality." The integration works across browser apps, mobile clients, IoT devices, and server-to-server setups. According to the announcement, developers can add Realtime capabilities "with just a handful of lines of Javascript."
The pricing changes are substantial. GPT-4o audio token prices drop 60%, to $40 per million input tokens and $80 per million output tokens. Cached audio input falls 87.5%, to $2.50 per million tokens. The new snapshot, gpt-4o-realtime-preview-2024-12-17, also brings improved voice quality and more reliable input — OpenAI specifically calls out better handling of dictated numbers.
GPT-4o mini joins the Realtime API beta as gpt-4o-mini-realtime-preview-2024-12-17, at roughly one-tenth of GPT-4o's audio rates: $10 per million audio input tokens and $20 per million audio output tokens. Text tokens cost $0.60 per million input and $2.40 per million output, with cached audio and text at $0.30 per million tokens. Both snapshots are also available through the Chat Completions API as gpt-4o-audio-preview-2024-12-17 and gpt-4o-mini-audio-preview-2024-12-17.
For builders of voice products, the economics just shifted meaningfully — a mini-based realtime voice agent is now priced at a level that makes consumer and high-volume enterprise applications far more viable.
Preference Fine-Tuning arrives
The fine-tuning API now supports Preference Fine-Tuning, a customization method built on Direct Preference Optimization. Instead of training on fixed input-output pairs, the technique compares pairs of model responses — one preferred, one not — and teaches the model to favor the preferred output. OpenAI positions it as especially effective for subjective tasks where tone, style, and creativity matter, such as creative writing or summarization, and notes that training pairs can come from human annotation, A/B testing, or synthetic data generation.
Early partner results are promising. Rogo AI, which builds an AI assistant for financial analysts, tested the method on its expert-built benchmark Rogo-Golden. The company found that Supervised Fine-Tuning struggled with out-of-distribution query expansion — for example, missing metrics like ARR for queries such as "how fast is company X growing" — while Preference Fine-Tuning resolved these issues, improving performance from 75% accuracy in the base model to over 80%.
Preference Fine-Tuning is available today for gpt-4o-2024-08-06, with support for gpt-4o-mini-2024-07-18 coming soon. It costs the same per trained token as Supervised Fine-Tuning. Support for OpenAI's newest models, presumably including o1, arrives early next year.
Go and Java SDKs enter beta
OpenAI also released Go and Java SDKs in beta, extending its official library coverage beyond Python and JavaScript. The Java SDK provides typed request and response objects plus utilities for managing API requests — a direct pitch at enterprise development teams, where Java remains what OpenAI calls "a staple of enterprise software development, favored for its type system and massive ecosystem of open-source libraries." The Go SDK is documented in a README on GitHub.
What it means
Taken together, the releases sketch OpenAI's developer platform strategy: reasoning models made affordable, voice interfaces made cheap, and customization made subjective-taste-friendly. The 60% reduction in reasoning tokens and the 60-87.5% audio price cuts attack the two biggest cost objections to production deployment, while Preference Fine-Tuning and new SDKs widen the aperture of who can build on the platform and how tightly they can tailor model behavior.
The competitive stakes are real. Google, Anthropic, and others are pushing their own reasoning and voice models to developers, and price-per-token arithmetic increasingly decides procurement. With support for o1 fine-tuning flagged for early next year, OpenAI's next move is letting developers customize its reasoning models — a step that could determine whether o1 becomes infrastructure or stays a showcase.
Original: platform.openai.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
- OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
- OpenAI swaps GPT-4o for o3 inside Operator, leaves API unchanged
- OpenAI publishes o3-mini system card detailing safety work
- OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board