Products & Tools

OpenAI Launches Model Distillation Suite in Its API

OpenAI's new Model Distillation suite combines Stored Completions, Evals, and fine-tuning so developers can build cheaper GPT-4o mini models that rival frontier performance on specific tasks.

Model Distillation in the API
Model Distillation in the APIAutomotive Rhythms / Openverse
By Elena Vasquez3 min read

Updated

Why it matters

  • Model Distillation is available today to all developers and works with any OpenAI model, including GPT-4o and o1-preview.
  • OpenAI offers 2M free training tokens per day on GPT-4o mini and 1M per day on GPT-4o until October 31; after that, standard fine-tuning prices apply.
  • Stored Completions is free; Evals (beta) is token-billed but free for up to 7 runs per week through year-end for developers who opt in to share their Evals with OpenAI.

OpenAI has launched Model Distillation, an integrated API offering that lets developers fine-tune cost-efficient models like GPT-4o mini using outputs from frontier models like GPT-4o and o1-preview — all within the OpenAI platform. The suite is available today to all developers and works with any of OpenAI's models.

Model distillation fine-tunes smaller, cheaper models on the outputs of more capable ones, allowing them to match advanced model performance on specific tasks at a much lower cost. That economics matters for the broader AI market: if a $0.33-per-million-token model can approach frontier quality on a narrow workload, companies can shift production traffic away from expensive flagship APIs. But until now, OpenAI says, distillation has been a multi-step, error-prone process requiring developers to manually orchestrate operations across disconnected tools — generating datasets, fine-tuning models, and measuring performance improvements. Because distillation is inherently iterative, developers had to repeatedly run each step, adding significant effort and complexity.

The new suite bundles three components.

Stored Completions. Developers can automatically capture and store input-output pairs generated by models like GPT-4o or o1-preview through the API. Setting a store:true flag in the Chat Completions API stores these pairs with no latency impact, and developers can build datasets from production data to evaluate and fine-tune models. Stored completions can be reviewed, filtered, and tagged to create high-quality datasets. The feature is free.

Evals (beta). Developers can create and run custom evaluations on the platform to measure model performance on specific tasks. Instead of manually writing evaluation scripts and integrating disparate logging tools, users can set up evaluations from Stored Completions data or uploaded existing datasets. Evals also work independently of fine-tuning for quantitative benchmarking. Evals are charged at standard model prices based on tokens used, but through the end of the year developers can run up to 7 evaluations per week for free if they opt in to share their Evals with OpenAI. Shared Evals will be used to improve and evaluate future models — a data flywheel that gives OpenAI visibility into real-world task performance.

Fine-tuning. Stored Completions and Evals integrate with OpenAI's existing fine-tuning offering. Datasets created with Stored Completions can feed fine-tuning jobs directly, and evaluations can run against fine-tuned models, all within one platform.

The intended workflow runs in three steps. First, create an evaluation to measure baseline performance of the model to be distilled into — GPT-4o mini in OpenAI's example — which becomes the ongoing test for deciding whether to deploy the distilled model. Next, use Stored Completions to build a distillation dataset of real-world examples from GPT-4o outputs on the target tasks. Finally, fine-tune GPT-4o mini on that dataset and return to Evals to check whether the fine-tuned model meets performance criteria compared to GPT-4o.

OpenAI is candid that one pass rarely suffices. If initial results fall short, developers may need to refine the dataset, adjust training parameters, or capture more specific examples where the model underperforms. The goal is incremental improvement until the distilled model performs well enough for production use.

To drive adoption, OpenAI is offering 2 million free training tokens per day on GPT-4o mini and 1 million free training tokens per day on GPT-4o until October 31. Beyond those limits, training and running a distilled model costs the same as OpenAI's standard fine-tuning prices listed on its API pricing page.

The move consolidates what was previously a fragmented toolchain into OpenAI's own platform, raising switching costs for developers who adopt it while lowering the barrier to cheaper custom models. The free-token promotion and the year-end Evals offer signal that OpenAI wants distillation volume now — and the shared evaluation data it collects along the way will feed directly into its next generation of models.

Source: OpenAI News

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
  2. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  3. OpenAI Adds Remote MCP Support and New Built-In Tools to Responses API
  4. OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
  5. OpenAI Ships o1 to Developers With 60% Cheaper Audio

Next article »