Google DeepMind's SIMA 2 Turns AI Into a Gaming Companion
Google DeepMind's SIMA 2 pairs Gemini reasoning with gameplay: it converses, transfers skills across unseen games, self-improves through play, and ships as a limited research preview.

Updated
Why it matters
- Google DeepMind announced SIMA 2, an agent that reasons about goals, converses with users and self-improves, roughly a year after the original SIMA learned 600+ instruction-following skills.
- SIMA 2 embeds a Gemini model at its core, transfers learned concepts across games it was never trained on (such as ASKA and MineDojo), and operated sensibly in newly generated Genie 3 worlds.
- The agent ships as a limited research preview to a small cohort of academics and game developers, with stated limits in long-horizon reasoning, memory, and precise low-level control.
Google DeepMind has unveiled SIMA 2, an AI agent that moves beyond following instructions to reasoning about goals, conversing with users and improving itself over time in virtual 3D worlds. The announcement, roughly a year after the original SIMA (Scalable Instructable Multiworld Agent), marks what the team calls "a significant step in the direction of Artificial General Intelligence (AGI), with important implications for the future of robotics and AI-embodiment in general."
The stakes are substantial. If generalist agents can master perception, reasoning and action in interactive 3D environments, the same building blocks — navigation, tool use, collaborative task execution — could transfer to physical robots and future AI assistants in the real world. That is the explicit ambition DeepMind attaches to the project.
From instruction-follower to reasoning companion
The first SIMA learned over 600 language-following skills — "turn left," "climb the ladder," "open the map" — across commercial video games. It interacted with these environments the way a person would: by looking at the screen and using a virtual keyboard and mouse, without access to the underlying game mechanics.
SIMA 2 embeds a Gemini model as the agent's core, and the difference is architectural, not incremental. DeepMind says the new system "can think and reason" about a user's high-level goal, perform complex reasoning in pursuit of it and execute goal-oriented actions within games. Training used a mixture of human demonstration videos with language labels and Gemini-generated labels. The result: SIMA 2 can describe what it intends to do and detail the steps it is taking toward its goals.
The team reports a shift in how the agent feels to use. "We have found that interacting with the agent feels less like giving it commands and more like collaborating with a companion who can reason about the task at hand."
Generalization gains
The Gemini integration also improved generalization and reliability. SIMA 2 handles more complex, nuanced instructions than its predecessor and succeeds far more often in games it was never trained on — DeepMind cites the Viking survival game ASKA and MineDojo, a research implementation of Minecraft, as examples.
The agent handles long, multi-step tasks, multimodal prompts, and instructions in different languages — including emojis. More consequentially, it transfers learned concepts across games: an understanding of "mining" in one title can be applied to "harvesting" in another. DeepMind frames this transfer ability as "foundational to achieving the kind of broad generalization seen in human cognition," and says SIMA 2's performance is now significantly closer to that of a human player across a wide range of tasks.
The team pushed generalization to its limit by pairing SIMA 2 with Genie 3, DeepMind's model that generates real-time 3D simulated worlds from a single image or text prompt. In those newly generated environments, SIMA 2 oriented itself sensibly, understood user instructions and took meaningful action toward goals — despite never having seen such worlds before.
Self-improvement
Perhaps the most consequential capability is self-directed learning. Throughout training, SIMA 2 agents performed increasingly complex tasks, bootstrapped by trial-and-error and Gemini-based feedback. After initial learning from human demonstrations, the agent can transition to learning in new games exclusively through self-directed play, without additional human-generated data. Its own experience data can then train the next, more capable version — a cycle DeepMind extended into newly created Genie environments. The company describes this as "a major milestone toward training general agents across diverse, generated worlds" and a path toward "open-ended learners in embodied AI" that grow with minimal human intervention.
Known limits
DeepMind is candid about what SIMA 2 cannot do. The agent still struggles with very long-horizon tasks requiring extensive multi-step reasoning and goal verification. It has a relatively short memory of interactions, constrained by a limited context window needed to keep latency low. Precise low-level keyboard-and-mouse actions and robust visual understanding of complex 3D scenes remain open problems the entire field faces.
The company also stresses responsible development, saying it worked with its Responsible Development & Innovation Team on the project, particularly around self-improvement. SIMA 2 ships as a limited research preview for a small cohort of academics and game developers — an approach DeepMind says will gather "crucial feedback and interdisciplinary perspectives" as it builds its understanding of risks and mitigations.
The research drew on partnerships with game studios including Coffee Stain (Valheim, Satisfactory, Goat Simulator 3), Hello Games (No Man's Sky), Thunderful Games (ASKA), Keen Software House (Space Engineers) and Tuxedo Labs & Saber Interactive (Teardown), among others. The team dedicates the work to the memory of colleagues Felix Hill and Fabio Pardo.
If the self-improvement loop holds at scale, SIMA 2's real significance may lie less in gaming than in demonstrating that one generalist agent — rather than many specialized systems — can unify broad competencies and keep expanding them on its own.
Source: Google DeepMind Blog
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles