OpenAI Ships AgentKit, Expanded Evals, and Reinforcement Fine-Tuning for Agents
OpenAI releases AgentKit, expanded evals capabilities, and reinforcement fine-tuning for agents, a three-part toolkit aimed at moving AI agents from prototype to production faster.

Updated
Why it matters
- OpenAI today released AgentKit, expanded evals capabilities, and reinforcement fine-tuning (RFT) for agents.
- The company frames the three releases as tools to help developers go from prototype to production faster.
- RFT for agents extends fine-tuning to agentic tasks, complementing new agent-building and evaluation tooling.
OpenAI has released three new developer tools designed to help teams move AI agents from prototype to production faster: AgentKit, expanded evals capabilities, and reinforcement fine-tuning (RFT) for agents.
The company announced the releases today, framing them as a single package aimed at the stage of agent development that has proven hardest: the transition from a working demo to a system reliable enough to deploy.
What was announced
The release covers three components:
AgentKit. A new toolkit for building agents. The name positions it as a bundled set of building blocks rather than a single API, and it arrives as developers increasingly assemble agents from multiple interacting parts — model calls, tool use, orchestration logic, and memory.
Expanded evals. OpenAI is widening its evals capabilities, the measurement tools developers use to score how well a model or agent performs on a task. Evals have become a central pain point in agent work: an agent that succeeds on a demo can fail unpredictably across the longer, multi-step trajectories that real deployments demand. Measuring that gap systematically, before shipping, is what expanded evals are for.
Reinforcement fine-tuning for agents. RFT for agents extends OpenAI's fine-tuning program to agentic tasks. Instead of optimizing a model purely on static examples, reinforcement fine-tuning rewards the model for outcomes — a training approach that fits agents, where success is defined by completing a task, not by producing a single correct token sequence.
Why the prototype-to-production gap matters
The announcement's framing — "from prototype to production faster" — points at the current bottleneck in the agent market. Building a demonstration agent has become relatively cheap. Operating one in production, where failures compound across steps and costs, has not.
Agents differ from single-turn model calls in a structural way: they act over time, use tools, and make decisions with downstream consequences. That makes evaluation and tuning fundamentally harder than in chat-style applications. A model that answers 95 percent of questions correctly can still fail a multi-step task most of the time, because errors accumulate across steps.
The three releases map directly onto that problem. AgentKit addresses construction. Expanded evals address measurement. RFT addresses improvement — giving developers a supervised path to make an agent better at its specific job rather than prompting around its weaknesses.
Packaging these as a coordinated release also signals where OpenAI sees developer attention going. The company is not just shipping individual capabilities; it is building out the production toolchain around agentic AI, the layer where enterprises decide whether to commit.
The stakes for developers
For development teams, the practical question is whether these tools reduce the engineering burden enough to change build-versus-buy decisions. Agent frameworks have proliferated across the industry, and evaluation tooling in particular has remained fragmented, often built in-house at companies that can afford it.
An official evals expansion from a major model provider moves some of that work into the platform layer. RFT for agents goes further: it offers a way to specialize models for agentic tasks through reinforcement, which historically required capabilities and infrastructure most teams lack.
What to watch
OpenAI says the tools are available starting today. The open questions are adoption and depth: whether AgentKit covers enough of the agent stack to become a default, how far the expanded evals reach into long-horizon, tool-using behavior, and how accessible RFT proves for teams without ML engineering resources.
The release lands at a moment when agent reliability is the defining constraint on the market's growth. Tools that genuinely shorten the prototype-to-production cycle would remove one of the widest bottlenecks between today's demonstrations and deployed, revenue-generating agent products.
Source: OpenAI News
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles
Related articles
- OpenAI ships a model-native harness and native sandboxes for its Agents SDK
- AgentKit, Expanded Evals, and RFT for Agents Ship Together
- OpenAI Upgrades Codex: Faster, More Reliable, More Autonomous
- OpenAI details codex-1: an o3 variant tuned for real coding work
- OpenAI Models Are Coming to Amazon Bedrock via a Stateful Agent Runtime