Flux 3 X Mimic: The Next Generation of Video-Action Models
Researchers have developed FLUX-mimic, a multimodal foundation model that integrates video, audio, and action prediction to control robots. By training on video prediction, the model learns physical world behaviors, enabling robots to perform tasks with greater awareness of cause and effect.
Why it matters
This represents a significant leap in robotics, moving from pre-programmed tasks to models that understand physical world dynamics through multimodal learning.
Back to blog Research Models FLUX 3 x mimic: The Next Generation of Video-Action Models July 23, 2026 9 min read An early version of FLUX 3, our new multimodal foundation model , is now running on robots. We gave mimic robotics early access to FLUX.3. Their strength in robot learning and deployment, combined with the model's world knowledge and BFL's foundation model expertise, produced FLUX-mimic: the next generation of video-action models.
FLUX 1 and FLUX 2 generate images. FLUX 3 expands into multimodality and generates audio-visual content jointly - and, at the same time, provides the foundation of FLUX-mimic: A video-action model, developed in collaboration with mimic, running robots that have been tested and deployed at Audi.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in