Article may be outdated

This article is 62 days old. Some details may have changed since publication.

Wired·3 min read·medium

Gemini Robotics 2 Brings Google's AI Into the Physical World

W
Will Knight
Gemini Robotics 2 Brings Google's AI Into the Physical World
✦AI Summary

Google DeepMind has released Gemini Robotics 2, a system that integrates vision language models with physical action models to enable autonomous robot tasks. The technology allows robots to reason through complex environments, though experts warn of potential safety risks.

Why it matters

Advancements in physical AI are necessary for robots to move beyond controlled environments and perform useful tasks in human-centric spaces.

✦Dive DeeperCreate a free account to unlock

Gemini Robotics 2 combines several different AI models into a single system. Taken together, they allow a robot to make sense of its surroundings and how to act in it. A vision language model (VLM), which understands images and video, can communicate with humans and reason how to perform different tasks. Two vision language action (VLA) models, trained to understand how to move in physical space, control the robot’s full-body movement as well as the movements of grippers or hands.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in