Enterprise & Work

OpenAI Pitches GPT-5 To The Enterprise With ChatGPT Work Agent

OpenAI is shipping GPT-5 to corporate teams through ChatGPT Work, citing 74.9% on SWE-Bench, 99.6% on AIME 2025, 45% fewer hallucinations than GPT-4o, and signed pilots at BBVA, Lowe's, and Bain.

By Marcus Bennett5 min read

Updated

Why it matters

  • GPT-5 scores 99.6% on AIME 2025 with tools, 74.9% on SWE-Bench, and 84.2% on MMMU according to OpenAI's launch materials.
  • OpenAI claims GPT-5 produces 45% fewer hallucinations than GPT-4o and uses 50-80% fewer output tokens than o3 across capabilities.
  • BBVA's Elena Alfaro said GPT-5 inside ChatGPT helped her team complete a 'very strategic task' in 'a couple of hours' instead of '2-3 weeks.'
  • ChatGPT Work ships with AES-256 encryption, SAML SSO, SCIM provisioning, and SOC 2 Type 2 certification across seven data residency regions.
  • Signed enterprise pilots at launch include BBVA (banking), Lowe's (retail), and Bain (consulting).

OpenAI is selling GPT-5 directly to corporate teams through a new product called ChatGPT Work, bundling the model with single-sign-on governance, a 45% hallucination reduction versus GPT-4o, and signed testimonials from BBVA, Lowe's, and Bain. The pitch reframes ChatGPT from a chat surface into an agent that handles multi-step workflows across marketing, engineering, finance, strategy, legal, and IT.

What exactly is in the launch?

ChatGPT Work is the enterprise SKU OpenAI built around GPT-5. The release eliminates model selection from the user flow. "One powerful model for every task—no model selection, no setup," the company states in launch materials. Employees interact with a single assistant that handles writing, research, analysis, coding, and problem-solving against internal files and connected apps.

OpenAI positions the experience as "a trusted PhD-level expert in every employee's pocket." That framing matters because it tells procurement teams the same vendor will absorb roles previously split across copywriters, analysts, and junior engineers.

How does GPT-5 perform on the benchmarks OpenAI is citing?

Three numbers anchor the launch:

  • 99.6% on AIME 2025 (with tools) for mathematics
  • 74.9% on SWE-Bench for real-world coding
  • 84.2% on MMMU for multimodal understanding

OpenAI also claims 45% fewer hallucinated responses than GPT-4o and 50–80% fewer output tokens across capabilities versus o3.

Those benchmarks matter more to CIOs than to casual users. SWE-Bench evaluates whether a model resolves actual GitHub issues rather than toy prompts, and it has become a procurement gate for engineering-led buyers. A 74.9% pass rate places GPT-5 in the same neighborhood as Anthropic's Claude 4 family and Google's Gemini 2.5 Pro — the two rivals OpenAI is contesting inside IT departments today.

AIME 2025 at 99.6% removes routine mathematics as a failure mode. MMMU at 84.2% keeps chart and document reasoning in the conversation. The 50–80% token reduction against o3 directly addresses the line-item cost concern that stalled some o-series enterprise pilots earlier in 2025.

Who is already running GPT-5 inside ChatGPT Work?

OpenAI published three named customer quotes. Elena Alfaro, Head of Global AI Adoption at BBVA, reported a measurable productivity swing. "GPT-5 is showing real promise, especially when it comes to writing code and handling technical tasks," she said, adding that "in one case, the model in ChatGPT even helped us accomplish a very strategic task that would have taken 2-3 weeks to just a couple of hours, with amazing proactivity." Alfaro also flagged Spanish-language performance, saying GPT-5 "beat older models on accuracy and working twice as fast."

Seemantini Godbole, CIO at Lowe's, framed the deployment as an extension of existing retail AI investment. "With GPT-5, corporate teams now have access to an ideal balance of reasoning and responsiveness for tasks like planning, analysis, research, and multi-step workflows," Godbole said. Lowe's plans to use the assistant for "planning, analysis, research, and multi-step workflows" across corporate teams.

Gene Rapoport, Partner and Co-Head of the Private Equity AI Practice at Bain, focused on client outcomes. "ChatGPT enables our teams to enhance their analysis and research, leading us to sharper insights faster with greater confidence," Rapoport said.

Three verticals show up in the rollout: a multinational bank, a Fortune 100 retailer, and a top-tier consultancy. Each operates in regulated workflows where audit trails, residency controls, and vendor lock-in drive decisions.

What governance and security ships in the box?

ChatGPT Work arrives with the controls enterprise security teams have demanded since the Samsung leak in 2023:

  • AES-256 encryption at rest, TLS 1.2+ in transit
  • SAML single sign-on and SCIM provisioning
  • Role-based access with real-time usage analytics
  • GDPR, CCPA, CSA STAR, and SOC 2 Type 2 certifications
  • Data residency in seven global regions

OpenAI also states the platform does not train on business data by default — a guarantee that has become table stakes for regulated buyers but still tilts deals at pharmaceutical and financial firms.

What does the assistant actually do?

OpenAI published a use-case grid mapping GPT-5 to specific C-suite pain points:

  • Marketing: A CMO needing a board-ready launch plan receives market analysis, draft plan, messaging, and sales content — a GTM brief with talking points.
  • Engineering: An SVP requesting a live incident dashboard gets a built-from-scratch application sourced from a plain-English prompt.
  • Finance: A CFO modeling a rate-change impact receives simulations, recommended levers, and a one-slide summary plus model.
  • Strategy: Responding to a new market entrant yields research, benchmarks, a leadership deck, and sales assets.
  • Legal: New regulations trigger policy updates that compare state laws and surface commonalities for compliance controls.
  • IT: An IT incident is diagnosed from logs, with a generated troubleshooting plan.

The grid is marketing, but it also signals how OpenAI plans to compete with workflow vendors like Salesforce Agentforce, Microsoft Copilot, and Palantir Foundry — by absorbing the prompt-and-template layer that consultants usually build for clients.

Why this matters now

Three forces collide in this release. The first is model consolidation: OpenAI is committing customers to GPT-5 instead of letting them toggle between GPT-4o, o3, and o4-mini. The second is cost: 50–80% fewer tokens per task rewrites the unit economics against pure o-series deployments. The third is governance: SAML SSO, SOC 2 Type 2, and seven-region residency answer the procurement objections that pushed 2024 pilots into legal review.

What's next for enterprise AI buyers

ChatGPT Work turns the GPT-5 launch into a B2B sales motion rather than a consumer subscription pivot. Buyers comparing OpenAI to Anthropic, Google, or Microsoft will weigh the AIME and SWE-Bench deltas against per-seat pricing, residency coverage, and the depth of app connectors (Google Drive, SharePoint, GitHub) the assistant can read against. The BBVA report of a multi-week task collapsing into hours sets a high expectation bar that any competing deployment will now have to clear in proof-of-concept benchmarks — and the first quarter of that contest starts the moment ChatGPT Work reaches general availability.

Source: OpenAI News

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

182 articles

Related articles

  1. OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work
  2. OpenAI Puts ChatGPT Inside Excel With GPT-5.4 Built for Finance
  3. First Look at GPT-5: Leading Developers Test OpenAI's New Model
  4. OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
  5. OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board

« Previous article