Models

OpenAI launches ChatGPT agent, folding Operator into ChatGPT

OpenAI's new ChatGPT agent unifies Operator's browsing and deep research in one model with a 41.6 HLE SOTA, rolling out to Pro, Plus and Team with new prompt-injection safeguards.

Introducing ChatGPT agent
Introducing ChatGPT agentAI-generated
By Elena Vasquez7 min read

Updated

Why it matters

  • ChatGPT agent rolls out to Pro, Plus and Team users, merging Operator's web interaction with deep research; Operator's site will be sunset within weeks.
  • The model posts 41.6 pass@1 SOTA on Humanity's Last Exam (44.4 with parallel rollouts), 27.4% on FrontierMath, 68.9% on BrowseComp, and 45.5% on SpreadsheetBench versus Copilot in Excel's 20.0%.
  • OpenAI classifies the model as High Biological and Chemical capability under its Preparedness Framework and flags a higher overall risk profile due to prompt injection and real-world actions; Pro users get 400 agent messages per month, other paid tiers 40.

OpenAI has launched ChatGPT agent, a unified agentic system that lets ChatGPT complete complex tasks end-to-end using its own virtual computer — and the model behind it posts a new state-of-the-art 41.6 pass@1 score on Humanity's Last Exam, plus 27.4% on FrontierMath with terminal access.

The rollout started July 17 for Pro, Plus and Team users via the "agent mode" option in the composer's tools dropdown. Pro users get access first, Plus and Team follow over the next few days, and Enterprise and Education customers gain access in the coming weeks. Usage is capped at 400 messages per month for Pro and 40 for other paid tiers, with additional consumption available through flexible credit-based options. Users in the European Economic Area and Switzerland are still waiting — OpenAI says it is "still working on enabling access" there.

One agent, three capabilities

ChatGPT agent merges two earlier OpenAI products that each solved half of a problem. Operator, launched as a research preview, could scroll, click and type on the web but couldn't produce deep analysis or detailed reports. Deep research could synthesize information but couldn't interact with websites or access content behind user authentication. OpenAI says it observed that "many queries users attempted with Operator were actually better suited for deep research," which motivated combining the two.

The result is a system that shifts between reasoning and action inside a single chat. A user can ask it to "look at my calendar and brief me on upcoming client meetings based on recent news," "plan and buy ingredients to make Japanese breakfast for four," or "analyze three competitors and create a slide deck," and the agent navigates websites, filters results, prompts for secure logins, runs code and returns editable slideshows and spreadsheets.

The toolkit includes a visual browser for graphical web interaction, a text-based browser for cheaper reasoning-heavy queries, a terminal, and direct API access. The agent also plugs into ChatGPT connectors for apps like Gmail and GitHub, letting it pull relevant information into responses. Users can take over the browser to log into any site themselves, giving the agent deeper access without handing over credentials. The virtual computer preserves task context across tools: the model can open a page in the text browser, download a file, manipulate it in the terminal, then view the output in the visual browser.

The Operator research preview site will keep running for a few more weeks before it is sunset. Deep research remains available as a standalone option for users who want longer, more detailed responses by default.

Benchmark claims

The launch post reports a battery of results for the model powering ChatGPT agent:

  • Humanity's Last Exam: 41.6 pass@1, a new SOTA on the expert-level cross-subject evaluation. Because the agent plans dynamically and picks its own tools, it can solve the same task differently across runs; with a parallel rollout strategy of up to eight attempts and self-reported confidence selection, the score rises to 44.4. OpenAI notes it mitigated browsing-based cheating by blocking known cheat domains and using a monitor model to flag attempts that found exact answers online.
  • FrontierMath: 27.4% accuracy with tool use, which OpenAI describes as outperforming previous models by a wide margin on what it calls the hardest known math benchmark. The footnote discloses that OpenAI has exclusive access to 237 of 290 private Tier 1-3 questions, and that agent results were elicited by OpenAI and graded by Epoch AI.
  • BrowseComp: 68.9%, a new SOTA and 17.4 percentage points above deep research.
  • SpreadsheetBench: 35.27% overall, rising to 45.54% when the agent can edit .xlsx files directly, versus 20.0% for Copilot in Excel and 23.25% for o3. Humans still lead at 71.33%.
  • WebArena: improvement over o3-powered CUA, the model behind Operator.

On an internal benchmark of "complex, economically valuable knowledge-work tasks" — graded by experts against human baselines built by top performers in each field — OpenAI says ChatGPT agent's output is comparable to or better than human work in roughly half the cases across a range of completion times, while significantly outperforming o3 and o4-mini. Sample tasks include competitive analyses of on-demand urgent care providers, amortization schedules, and identifying viable water wells for a green hydrogen facility. On DSBench, which covers realistic data science work, the company says the agent surpasses human performance by a significant margin. On an internal investment-banking benchmark covering first to third-year analyst tasks, such as three-statement models and leveraged buyout models graded on hundreds of correctness and formula-use criteria, it significantly outperforms deep research and o3.

Safety: a higher risk profile, acknowledged

The release is the first time ChatGPT can take actions on the web, and OpenAI is explicit that this raises the stakes. "While these mitigations significantly reduce risk, ChatGPT agent's expanded tools and broader user reach mean its overall risk profile is higher," the company writes.

The central concern is prompt injection — malicious instructions hidden in webpages, in invisible elements or metadata, that could trick the agent into sharing private connector data or acting harmfully on a site the user has logged into. "Because ChatGPT agent can take direct actions, successful attacks can have greater impact and pose higher risks," OpenAI warns. Mitigations include training and testing for injection resistance, monitoring to detect and respond to attacks, mandatory user confirmation before consequential actions, and user-side advice: disable connectors when they aren't needed for a task.

Other guardrails include explicit confirmation before real-world actions like purchases, a "Watch Mode" requiring active supervision for critical tasks such as sending emails, and training the model to refuse high-risk tasks like bank transfers. Privacy controls let users delete all browsing data and log out of all website sessions with one click; in browser takeover mode, OpenAI says it does not collect or store anything the user types, including passwords.

OpenAI has also classified ChatGPT agent as having High Biological and Chemical capabilities under its Preparedness Framework, activating the associated safeguards. The company states it lacks "definitive evidence that the model could meaningfully help a novice create severe biological harm" — its threshold for High — but is implementing safeguards now. That stack includes comprehensive threat modeling, dual-use refusal training, always-on classifiers and reasoning monitors, and enforcement pipelines. External involvement includes biology-trained reviewers, domain-expert red teamers, and a Biodefense workshop convened this month with experts from government, academia, national labs and NGOs. A bug bounty program is launching alongside the agent.

Rough edges remain

OpenAI is candid about limitations. Slideshow generation is in beta, and "outputs can sometimes feel rudimentary in its formatting and polish, particularly when starting without an existing document." There are occasional discrepancies between slides shown in the viewer and the exported PowerPoint file, and users can't yet upload an existing deck as a template the way they can with spreadsheets. OpenAI says it is already training the next iteration of slideshow creation for more polished output.

The workflow is built to be interruptible: users can pause a task, request a progress summary, stop it entirely and keep partial results, or redirect mid-run without the agent losing prior progress. The ChatGPT mobile app sends a notification when a task finishes. Completed tasks can be scheduled to recur, such as a weekly metrics report every Monday morning.

Why it matters

The launch collapses OpenAI's separate agent experiments into its flagship product, putting autonomous browsing, code execution and multi-step task completion in front of millions of paying users rather than a research-preview audience. It also lands in a market where agentic AI is the main competitive frontier, and where the spreadsheet and knowledge-work benchmark comparisons — particularly the direct 45.5% versus 20.0% spreadsheet win over Copilot in Excel — put OpenAI in direct conflict with Microsoft's enterprise productivity positioning.

The safety disclosures cut both ways. OpenAI is publishing more risk detail than most competitors, including an admission of a higher overall risk profile and a High biosecurity classification applied out of caution. Whether confirmation prompts, Watch Mode and injection-resistance training are enough to make autonomous web action safe at this scale is now the industry's open question.

For now, OpenAI frames this as the start of an iterative process, promising continued improvements to the agent's efficiency, depth and versatility — including progressively reducing the amount of user oversight required, a direction that will test whether convenience and safety can keep advancing together.

Source: OpenAI News

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. OpenAI Launches Workspace Agents in ChatGPT for Teams
  2. OpenAI Says a Quarter of U.S. Workers Now Use ChatGPT on the Job
  3. OpenAI Will Show Ads in ChatGPT and Bring $8 ChatGPT Go to the US
  4. OpenAI Says Over a Quarter of U.S. Workers Now Use ChatGPT on the Job
  5. OpenAI Signals Data Shows ChatGPT Use Deepening Worldwide

« Previous articleNext article »