Imagine robots that don’t merely follow pre-programmed commands but actively consider their environment and decide their next steps. This significant leap forward arrived on September 25, 2025, when Google DeepMind unveiled its first “thinking” robotics AI.
Table of Contents
- Key Takeaways
- The Dawn of Agentic Robots
- Introducing Gemini Robotics: Thinking and Doing Models
- Simulated Reasoning: How DeepMind’s Robots “Think”
- From Thought to Action: Gemini Robotics in Practice
- Unlocking General Functionality in Robotics
- The Future of DeepMind’s Robotics AI
Importantly, deepMind researchers believe this marks the dawn of “agentic robots,” fundamentally transforming how AI systems interact with the physical world and promising a future where robots possess general functionality rather than being limited to specific, intensively trained tasks.
Google DeepMind has introduced a pioneering project known as Gemini Robotics, leveraging generative AI to empower robots with simulated reasoning before action.
This innovation addresses a long-standing challenge in the field: traditional robots require extensive training for singular tasks and struggle with adaptability.
By enabling robots to “think” and process new situations, DeepMind aims to unlock versatile capabilities, moving beyond the bespoke and difficult-to-deploy nature of current robotic systems.
Key Takeaways
- Google DeepMind unveiled its inaugural “thinking” robotics AI, featuring two new models: Gemini Robotics 1.5 and Gemini Robotics-ER 1.5.
- Gemini Robotics-ER 1.5 is a vision-language model (VLM) capable of simulated reasoning, enabling robots to generate complex task steps before acting.
- The new technology aims to provide general functionality for robots, overcoming the limitations of intensive, task-specific training required by traditional systems.
- The two models work in tandem: Gemini Robotics-ER 1.5 generates natural language instructions, and Gemini Robotics 1.5 then translates these into robot actions.
The Dawn of Agentic Robots
Generative AI systems are increasingly common for creating diverse data types like text and images. Google DeepMind recognized that this powerful capability could extend to outputting robot actions, forming the core of its Gemini Robotics project.
DeepMind researchers believe this development represents the advent of agentic robots, ushering in a new era for robotics applications as reported by the original article.
Traditional robots face significant challenges, primarily their need for intensive training on specific tasks, which makes them inefficient at anything beyond their programmed functions.
Specifically, Carolina Parada, head of robotics at Google DeepMind, observed that “Robots today are highly bespoke and difficult to deploy, often taking many months in order to install a single cell that can do a single task.”” This limitation highlights the critical need for a more adaptable approach in robotics.
Introducing Gemini Robotics: Thinking and Doing Models
The Gemini Robotics project features two innovative models designed to work together, creating the first robots that can “think” before executing actions.
These models are known as Gemini Robotics 1.5 and Gemini Robotics-ER 1.5 according to newsbytesapp.com.
Each model plays a distinct yet complementary role in enabling intelligent robotic behavior.
Gemini Robotics 1.5 operates as a vision-language-action (VLA) model. It processes both visual and text data to generate concrete robot actions. In contrast, Gemini Robotics-ER 1.5 specializes in embodied reasoning (ER) and functions as a vision-language model (VLM).
This ER model takes visual and text input, then generates the necessary steps required to complete a complex task, laying the groundwork for the robot’s subsequent physical movements.
Simulated Reasoning: How DeepMind’s Robots “Think”
Gemini Robotics-ER 1.5 distinguishes itself as the first robotics AI capable of simulated reasoning, a feature akin to modern text-based chatbots. Google describes this capability as “thinking,” even if it’s a specific form of generative AI processing.
This model excels at making accurate decisions about how a robot should interact with its physical surroundings, consistently achieving top scores in both academic and internal benchmarks.
Crucially, the ER model, while capable of sophisticated reasoning, does not perform any physical actions itself. Its role is purely cognitive: to interpret requests, understand the environment, and plan.
This division of labor allows for specialized optimization, where one AI focuses on strategic planning and the other on precise execution, enhancing overall system reliability and intelligence.
From Thought to Action: Gemini Robotics in Practice
The practical application of DeepMind’s dual-model system is straightforward yet powerful. Consider a scenario where a user wants a robot to sort laundry into whites and colors.
Gemini Robotics-ER 1.5 initiates this process by taking the request and analyzing images of the physical environment, such as a pile of clothing. This AI can even leverage external tools like Google Search to gather additional data pertinent to the task.
Following its assessment, the ER model generates natural language instructions—a detailed sequence of specific steps the robot should follow to successfully complete the given task. Subsequently, Gemini Robotics 1.5, the action model, receives these instructions.
It then translates these human-readable directives into precise robot actions, enabling the robot to carry out the complex task effectively and autonomously as highlighted by blog.google.
Unlocking General Functionality in Robotics
DeepMind asserts that generative AI is an exceptionally vital technology for the field of robotics because it directly addresses the critical need for general functionality.
Current robotic systems are notoriously limited; they undergo intensive training for specific tasks and typically lack the ability to perform other functions effectively.
This makes them expensive and time-consuming to deploy, as each new task often necessitates months of reprogramming and installation.
The inherent design of generative systems makes AI-powered robots far more versatile. They can encounter and adapt to entirely new situations and workspaces without requiring extensive reprogramming or specific task-based training.
This paradigm shift, from highly specialized machines to generally capable agents, promises to make robotics more accessible, adaptable, and integrated into diverse environments, moving beyond the current “bespoke” limitations.
The Future of DeepMind’s Robotics AI
The unveiling of Google DeepMind’s “thinking” robotics AI, spearheaded by Gemini Robotics 1.5 and Gemini Robotics-ER 1.5, represents a significant milestone in artificial intelligence.
This pioneering system’s ability to combine simulated reasoning with actionable commands offers a compelling vision for the future. By moving beyond traditional, task-specific robot programming, DeepMind champions a new era of adaptable, generally functional robotic agents.
DeepMind researchers believe this technology heralds the “dawn of agentic robots,” emphasizing a future where machines autonomously understand and interact with their environment with unprecedented flexibility.
Specifically, the collaborative framework of the Gemini Robotics models, where one thinks and the other acts, holds the potential to redefine robotic deployment and utility, making sophisticated AI-powered robots a more versatile and integral part of various industries.
| Latest From Us
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure


