Safety & Security

OpenAI and Paradigm Launch EVMbench for Smart Contract Security

OpenAI and Paradigm launch EVMbench: 117 real audit vulnerabilities testing whether AI agents can detect, patch, or exploit smart contracts. GPT-5.3-Codex drains funds in 71% of scenarios.

Introducing EVMbench
Introducing EVMbenchseanrnicholson / Openverse
By Rebecca Stone5 min read

Updated

Why it matters

  • GPT-5.3-Codex scores 71.0% on EVMbench's exploit mode, up from GPT-5's 33.3% released just over six months earlier.
  • EVMbench contains 117 curated high-severity vulnerabilities from 40 audits, mostly from Code4rena competitions plus scenarios from the Tempo blockchain audit.
  • OpenAI commits $10M in API credits through its Cybersecurity Grant Program to accelerate cyber defense for open-source software and critical infrastructure.

OpenAI's latest frontier model drains funds from vulnerable smart contracts in 71.0% of test scenarios, more than doubling the 33.3% score of GPT-5, released just over six months ago. The result comes from EVMbench, a new benchmark OpenAI built together with crypto investment firm Paradigm to measure how well AI agents detect, patch, and exploit high-severity smart contract vulnerabilities.

The stakes are substantial. Smart contracts routinely secure more than $100 billion in open-source crypto assets, according to OpenAI. As AI agents improve at reading, writing, and executing code, measuring their capabilities in economically meaningful environments becomes a way to track emerging cyber risk — and, OpenAI argues, to push the ecosystem toward using AI defensively to audit and strengthen deployed contracts.

What EVMbench contains

EVMbench draws on 117 curated vulnerabilities from 40 audits, with most sourced from open code audit competitions such as Code4rena. The benchmark additionally includes several vulnerability scenarios drawn from the security auditing process for the Tempo blockchain, a purpose-built layer-1 designed for high-throughput, low-cost stablecoin payments. These scenarios extend the benchmark into payment-oriented smart contract code, a domain where OpenAI expects agentic stablecoin payments to grow.

To build the task environments, the team adapted existing proof-of-concept exploit tests and deployment scripts where they existed and manually wrote the rest. For the patch mode, the researchers verified that each vulnerability is exploitable and can be mitigated without introducing compilation-breaking changes. For the exploit mode, they wrote custom graders and red-teamed the environments to close off methods by which an agent might cheat the grader. Paradigm provided domain expertise for task quality control, and automated task auditing agents helped increase the soundness of the environments.

Three capability modes

EVMbench evaluates agents across three modes:

  • Detect: Agents audit a smart contract repository and are scored on recall of ground-truth vulnerabilities and associated audit rewards.
  • Patch: Agents modify vulnerable contracts and must preserve intended functionality while eliminating exploitability, verified through automated tests and exploit checks.
  • Exploit: Agents execute end-to-end fund-draining attacks against deployed contracts in a sandboxed blockchain environment, graded programmatically via transaction replay and on-chain verification.

To keep evaluation objective and reproducible, the team built a Rust-based harness that deploys contracts, replays agent transactions deterministically, and restricts unsafe RPC methods. Exploit tasks run in an isolated local Anvil environment rather than on live networks. All vulnerabilities in the benchmark are historical and publicly documented.

Results and model behavior

GPT-5.3-Codex running via Codex CLI tops the exploit leaderboard at 71.0%, a significant gain over GPT-5's 33.3%. But the picture is uneven. Detect recall and patch success rates remain below full coverage, with a large fraction of vulnerabilities still difficult for agents to find and fix.

The benchmark also exposes behavioral differences across task types. Agents perform best in the exploit setting, where the objective is explicit and the agent can iterate until funds are drained. Performance weakens on detect and patch. In detect mode, agents sometimes stop after identifying a single issue rather than exhaustively auditing the codebase. In patch mode, maintaining full functionality while removing subtle vulnerabilities remains challenging.

Stated limitations

OpenAI is candid about what EVMbench does not measure. The vulnerabilities come from Code4rena auditing competitions; while realistic and high-severity, many heavily deployed and widely used contracts undergo significantly more scrutiny and may be harder to exploit.

The grading system has known gaps. In detect mode, the benchmark checks whether an agent finds the same vulnerabilities human auditors identified. If the agent flags additional issues, OpenAI has no reliable way to determine whether they are true vulnerabilities humans missed or false positives.

The exploit setting carries structural limits too. Transactions are replayed sequentially in the grading container, so behaviors that depend on precise timing mechanics are out of scope. The chain state is a clean local Anvil instance rather than a mainnet fork, only single-chain environments are supported, and some scenarios require mock contracts instead of mainnet deployments.

Defensive agenda

Alongside the benchmark, OpenAI frames EVMbench as "both a measurement tool and as a call to action." The company says agents are likely to be transformative for both attackers and defenders, and urges developers and security researchers to incorporate AI-assisted auditing into their workflows as agents improve.

Because cybersecurity is inherently dual-use, OpenAI says it is taking an evidence-based, iterative approach that accelerates defenders' ability to find and fix vulnerabilities while slowing misuse. Its mitigations include safety training, automated monitoring, trusted access for advanced capabilities, and enforcement pipelines including threat intelligence. The company has also been preparing strengthened cyber safeguards to support defensive use and broader ecosystem resilience.

On the ecosystem side, OpenAI is expanding the private beta of Aardvark, its security research agent, and partnering with open-source maintainers to provide free codebase scanning for widely used projects. Building on its Cybersecurity Grant Program launched in 2023, the company is committing $10 million in API credits to accelerate cyber defense with its most capable models, especially for open-source software and critical infrastructure systems. Organizations engaged in good-faith security research can apply for API credits and support through the program.

OpenAI has released EVMbench's tasks, tooling, and evaluation framework to support continued research on measuring and managing emerging AI cyber capabilities. With exploit-mode performance doubling in roughly six months, the benchmark is positioned to become a standing gauge of how quickly agentic capability in high-value code environments — and the defensive response to it — continues to advance.

Original: paradigm.xyz

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Publishes Policy for Disclosing Bugs It Finds in Others' Software
  2. OpenAI's Daybreak Launches Patch the Planet for Open-Source Security
  3. OpenAI Launches Safety Bug Bounty to Pay for AI Abuse Findings
  4. OpenAI's Codex Security Enters Research Preview
  5. OpenAI to Acquire Promptfoo and Build Red-Teaming Into Frontier

« Previous articleNext article »