OpenAI's o3 and o4-mini Now Reason With Images, Not Just See Them
OpenAI's o3 and o4-mini manipulate images — cropping, zooming, rotating — inside their chain of thought, hitting 95.7% on V* visual search. The catch: long reasoning chains and inconsistent results.

Updated
Why it matters
- OpenAI's o3 and o4-mini, released April 16, 2025, are the first OpenAI models to think with images in their chain of thought — cropping, zooming, and rotating user uploads natively, without separate specialized models.
- The models set new state-of-the-art results on MMMU, MathVista, CharXiv, VLMs are Blind, and V*, achieving 95.7% accuracy on V* visual search, which OpenAI describes as largely solving the benchmark.
- OpenAI discloses three current limitations: excessively long reasoning chains with redundant tool calls, basic perception errors, and unreliable results across multiple attempts at the same problem.
OpenAI's o3 and o4-mini models, announced April 16, 2025, are the company's latest visual reasoning models in its o-series — and for the first time, the models can think with images in their chain of thought, not just see them.
The distinction matters. Similar to OpenAI's earlier o1 model, o3 and o4-mini are trained to think for longer before answering, using a long internal chain of thought before responding to the user. What is new is that o3 and o4-mini extend this capability into the visual domain. According to OpenAI, this is "achieved by transforming user uploaded images with tools, allowing them to crop, zoom in, and rotate, in addition to other simple image processing techniques."
"More importantly, these capabilities come natively, without relying on separate specialized models," OpenAI states.
How the models manipulate images mid-reasoning
In practice, ChatGPT's enhanced visual intelligence combines advanced reasoning with tools like web search and image manipulation. OpenAI says the models automatically zoom, crop, flip, or enhance user images to extract insights even from imperfect photos. The company offers two concrete use cases: uploading a photo of an economics problem set to receive step-by-step explanations, or sharing a screenshot of a build error to get a quick root-cause analysis.
The capability also changes how users interact with ChatGPT. OpenAI says users can ask questions by taking a photo without worrying about the positioning of objects — whether text is upside down or multiple physics problems share a single frame. When objects are not obvious at first glance, visual reasoning lets the model zoom in to see more clearly.
OpenAI's published examples, all completed with o3, show the mechanism at work. Asked to read a notebook photographed upside down, the model reasoned for 20 seconds, recognized the orientation problem, rotated the image, cropped the relevant section, and returned the answer: "4th February – finish roadmap."
A second example is more revealing about what happens under the hood. Asked to solve a maze and plot a path with a red line, o3 reasoned for 1 minute 44 seconds. The published chain of thought shows the model writing Python code, inspecting pixel arrays, discovering that the maze walls were black lines on a transparent background that PIL rendered on white, and ultimately running a breadth-first search across the pixel grid before drawing the red solution path — 1,144 pixels long — onto the image.
That maze trace matters because it shows the model debugging its own perception. When grayscale values came back all zeros, contradicting what the model had seen in an earlier Matplotlib display, it investigated the alpha channel, worked out the display discrepancy, and adjusted its approach. This is reasoning about images, not just pattern-matching on them.
A new axis for test-time compute
The technical significance extends beyond demos. OpenAI says the approach "enables a new axis for test-time compute scaling that seamlessly blends visual and textual reasoning." Test-time compute — letting a model think longer at inference in exchange for better answers — has been the central bet of the o-series line. Extending that scaling axis to images ties the visual modality into the same performance curve that has driven recent reasoning-model gains.
The stakes are competitive. Multimodal reasoning is the current frontier for frontier labs, and benchmarks that test whether models genuinely perceive — rather than guess from context — have exposed weaknesses in existing systems. OpenAI's claim that it can now scale visual reasoning in the same way it scales textual reasoning positions image understanding as a compute problem rather than a fixed architectural limitation.
Benchmark results
To measure visual reasoning improvements against its previous multimodal models, OpenAI tested o3 and o4-mini on a diverse set of human exams and ML benchmarks. The company reports the new models "significantly outperform their predecessors on all multimodal tasks we tested." All models were evaluated at high reasoning effort settings — similar to variants like "o4-mini-high" in ChatGPT.
OpenAI attributes significant gains across all perception benchmarks it evaluated to thinking with images, without relying on browsing. The company says its models set new state-of-the-art performance in:
- STEM question-answering (MMMU, MathVista)
- Chart reading and reasoning (CharXiv)
- Perception primitives (VLMs are Blind)
- Visual search (V*)
The standout number: on V*, OpenAI's visual reasoning approach achieves 95.7% accuracy, which the company describes as "largely solving the benchmark." V* is designed to test fine-grained visual search within complex images — a task where models have historically failed in ways that suggested fundamental perceptual blind spots.
Notably, OpenAI updated its results on April 16 for o3 on Charxiv-r, Mathvista, and vlmsareblind to reflect a system prompt change that was not present in the original evaluation — an acknowledgment that even benchmark runs of these models are sensitive to prompting conditions.
An agentic step
The visual reasoning models also work in tandem with other tools: Python data analysis, web search, and image generation. OpenAI says this combination lets them "creatively and effectively solve more complex problems, delivering our first multimodal agentic experience to users."
That framing signals where OpenAI sees this going. A model that can look at an input, decide it needs a closer look, manipulate the image, run code on the pixels, and search the web — all within a single reasoning process — behaves less like a question-answering system and more like an agent that acts on visual evidence.
OpenAI's own list of limitations
OpenAI is candid about where thinking with images currently falls short. The company lists three limitations:
- Excessively long reasoning chains: Models may perform redundant or unnecessary tool calls and image manipulation steps, resulting in overly long chains of thought. The maze example, with its 1-minute-44-second solve and multiple abandoned approaches, illustrates the cost.
- Perception errors: Models can still make basic perception mistakes. Even when tool calls correctly advance the reasoning process, visual misinterpretations may lead to incorrect final answers.
- Reliability: Models may attempt different visual reasoning processes across multiple tries of the same problem, some of which lead to incorrect results.
The reliability point deserves attention. Non-deterministic reasoning paths mean the same image and question can yield different answers on different attempts — a property that complicates deployment in any setting where consistency matters.
Why it matters
OpenAI positions o3 and o4-mini as significantly advancing state-of-the-art visual reasoning and "representing an important step toward broader multimodal reasoning." The company claims the models deliver best-in-class accuracy on visual perception tasks, enabling them to solve questions that were previously out of reach.
For the broader market, the release raises the bar for every competitor building multimodal systems. If perception benchmarks like V* can be largely solved through test-time compute and image manipulation during reasoning, the differentiator shifts from raw perception accuracy to how effectively models orchestrate tools inside their thinking.
OpenAI says it is "continually refining the models' reasoning capabilities with images to be more concise, less redundant, and more reliable," and frames the release as an invitation: "We're excited to continue our research in multimodal reasoning, and for people to explore how these improvements can enhance their everyday work." The next question is whether the reliability and efficiency gaps close fast enough for visual reasoning to move from impressive demos to dependable daily use.
Original: openaifoundation.org
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
144 articles
Related articles
- OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
- OpenAI Ships ChatGPT Images 2.5 With Sketch Tool and 50% Faster Generation
- OpenAI Ships gpt-image-1 to API After 700M-Image Debut
- OpenAI Adds Remote MCP Support and New Built-In Tools to Responses API
- OpenAI Ships o1 to Developers With 60% Cheaper Audio