Cognition Turns Devin Loose on Its Own Code with GPT‑6 Astra
Cognition is using GPT‑6 Astra to have Devin test its own code, returning recordings, test reports, and screenshot-verified fixes as evidence — and shrinking manual code review.

Updated
Why it matters
- Cognition is applying GPT‑6 Astra across its full product lineup: the Devin cloud agent, its CLI, and its desktop products.
- Devin used Astra to test the iPhone game Otter Run, returning a simulator recording plus a report of checks that passed and areas left untested.
- Co-founder Walden Yan says Astra helps Cognition respond to customers 'much quicker': a bug screenshot goes to Devin, which returns a screenshot of the fix.
Cognition, the company behind the autonomous software engineer Devin, is using GPT‑6 Astra to have its AI test its own work — and to hand engineers proof that the code actually runs.
The stakes are straightforward. Devin is deployed at businesses around the world, ranging from big banks to tech-native startups, and as Cognition's engineering teams write more code, reviewing that output has become a bottleneck. Cognition sees GPT‑6 Astra's ability to test its work and show the results as a way to make code review more efficient.
"One of the big pieces that Astra improves on is its ability to test and prove that its work actually functions the way you expect," said Walden Yan, co-founder of Cognition.
Testing software and showing the results
Cognition is not running a narrow pilot. The company is applying GPT‑6 Astra across its entire product lineup. "We've been using it to make improvements across our product, including the core cloud agent that is Devin, but also our CLI and desktop products," Yan said.
The most concrete example involves an iPhone game called Otter Run. Devin uses Astra to test the game and returns two artifacts: a recording of the game running in a simulator, and a report identifying which checks passed and which areas were left untested.
Those two outputs do different jobs. The recording shows the application's behavior — what the software actually does when it runs. The report documents the scope of the testing, making explicit what was covered and what was not. Engineers can use the recording to inspect how the software works and use the report to understand what still needs attention.
That pairing matters. A video of a running app is evidence; a scope report is a boundary. Together they let a human reviewer verify behavior without re-running every test themselves, while still seeing exactly where confidence ends. For an autonomous agent that writes code at machine speed, the gap between "it claims to work" and "here is it working" is the entire problem, and this workflow is built to close it.
Astra is also compressing Cognition's customer support loop. Yan says the model is helping Cognition get back to customers "much quicker." When a customer sends a screenshot of a bug, the team can pass it to Devin running on Astra, which fixes the issue and returns a screenshot showing the result — the same format the customer used to report the problem in the first place.
The symmetry is worth noting. A bug arrives as a screenshot; a fix returns as a screenshot. The customer never has to read a diff or a changelog to confirm their issue was addressed. For enterprise customers, who range from large financial institutions to startups, that kind of closed verification loop can decide whether an AI coding tool is trusted with real work or confined to toy projects.
Working toward less manual code review
The near-term efficiency gains point at a larger structural shift: reducing how much code human engineers must read at all.
Cognition frames stronger testing with GPT‑6 Astra as a path toward a more efficient review process. If Devin tests its own work and provides evidence of the results, engineers can evaluate changes with less manual examination of the code itself. The reviewer's job moves from line-by-line reading toward auditing artifacts — recordings, reports, before-and-after screenshots — that demonstrate what the change does.
"We expect over time that we have to manually look at less code and end up shipping more at the end of the day," Yan said. "This is one of the things we're really excited about when it comes to GPT‑6."
That expectation cuts at one of the central tensions in AI-assisted software development. Autonomous agents can already generate large volumes of code, but generation is cheap only if verification doesn't scale linearly with it. If every AI-written change still requires a human to read it in full, the agent simply moves the cost from writing to reviewing. Cognition's bet is that self-testing with verifiable evidence — recordings of behavior, documented test scope, screenshot-confirmed fixes — breaks that link, letting output grow faster than review effort.
Why it matters
Cognition occupies a specific position in this story: it is both the maker of a widely deployed autonomous engineer and an early, intensive user of GPT‑6 Astra inside that product. Devin's customers include large banks, which face some of the strictest software verification and audit requirements of any industry. A workflow where an agent produces not just a change but a recorded demonstration of that change working aligns with how regulated engineering teams already think about evidence and validation.
The Otter Run example also shows the current boundary of the approach. The report explicitly identifies areas left untested, meaning the system flags its own coverage gaps rather than hiding them. That transparency is what makes the evidence usable: an engineer reviewing Devin's output knows both what was verified and what remains a judgment call.
Yan's framing suggests the direction of travel. Manual code review shrinks; shipping volume grows; the human role concentrates on evaluating demonstrated behavior rather than reading syntax. If GPT‑6 Astra's testing capabilities hold up across Cognition's cloud, CLI, and desktop products at production scale, the company will have a working template for the industry's next argument: that the answer to AI-written code isn't more human reading, but better machine-generated proof.
Source: OpenAI News
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
144 articles
Related articles
- OpenAI Says Two API Settings Tripled GPT-5.6's ARC-AGI-3 Score
- Every Frontier AI Agent Cheats, New CAIS Benchmark Shows
- OpenAI Cancels GPT-6.1 Astra Release Over Deceptive Behavior in Testing
- OpenAI and Anthropic Ship Cheaper, Faster Models Same Day
- OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context