Video Friday: Meet Google DeepMind’s Gemini Robotics 2
Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few month...
WhatIsFuture AI Editor
Contributor
For years, the artificial intelligence revolution has unfolded primarily behind glass screens. We have marvelled at large language models capable of drafting essays, generating photorealistic artwork, and writing functional code in seconds. Yet, the physical world has remained a stubborn frontier for autonomous systems. Now, Google DeepMind’s latest showcase, Gemini Robotics 2, marks a decisive turning point in bridging the gap between digital cognition and physical manifestation, bringing multimodal intelligence directly into mechanical limbs and embodied hardware.
By combining state-of-the-art vision-language-action (VLA) architectures with real-time robotic motor control, DeepMind is demonstrating that embodied AI is no longer a distant theoretical goal. Gemini Robotics 2 isn't merely executing rigid, pre-programmed industrial routines; it is perceiving dynamic human environments, reasoning through complex spatial tasks, and adjusting its physical grip and force in real time. This breakthrough signals the arrival of a new era where artificial general intelligence begins to operate hands-on in factories, laboratories, and eventually, our daily lives.
Join 15,000+ tech leaders
Get instant alerts on the most critical AI breakthroughs on our WhatsApp channel. No spam, just pure alpha.
From Chatbots to Hardware: The Evolution of Embodied AI
The shift from pure text-and-image models to physical embodiment represents one of the most complex scaling challenges in modern computer science. Standard language models operate in predictable, discrete token spaces where a computational mistake results in nothing more than confusing text on a monitor. Physical environments, by contrast, are continuous, chaotic, and governed by strict Newtonian physics. Gemini Robotics 2 solves this fundamental disconnect by treating physical actions—such as joint rotations, torque applied to a gripper, and spatial trajectories—as native tokens within a single, unified multimodal neural network.
This holistic integration allows the system to synthesize visual inputs from camera feeds, natural language instructions from human supervisors, and tactile feedback from motor sensors into a single continuous decision-making loop. As the broader technology industry remains locked in intense debates over rapid acceleration versus safety restraint, reminiscent of Sam Altman and AI’s decel debate, DeepMind is demonstrating that physical integration is accelerating regardless of philosophical friction. The ability for an autonomous robot to adapt to an unscripted environment
Supercharge Your Workflow with Claude AI
The AI assistant used by 100K+ professionals. Write, code, analyse — all in one place.



