Black Forest Labs is betting that the future of AI won't stop at generating images and videos.
The German startup has unveiled FLUX 3, a multimodal AI model designed to generate images, video, and audio while also serving as the foundation for robotic systems. Unlike many multimodal platforms that combine separate models, FLUX 3 is trained within a single architecture intended to understand and predict changes across digital and physical environments.
If the approach works, enterprises could eventually rely on one AI backbone for creative content, simulation, and robotics rather than deploying separate models for each task.
“FLUX 3 is our first model built entirely on that principle,” Black Forest Labs said, describing the model as a step toward AI systems that can “perceive, predict, and act.”
Video generation with native audio
FLUX 3 expands Black Forest Labs’ technology beyond still images by generating video clips of up to 20 seconds with synchronized audio.
The model supports text-to-video generation, image-to-video animation, video editing using reference clips, keyframe-controlled transitions, and multilingual dialogue. Black Forest Labs said it can also chain shorter clips into longer sequences while maintaining character consistency.
In early internal comparisons using 10-second, 720p text-to-video clips, Black Forest Labs said FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%. It also reported wins against Grok Imagine Video, Kling v3 Pro and other models.
However, those results remain preliminary. Black Forest Labs has not yet published full testing methods, including details such as evaluator numbers, prompt sets, or independent benchmark results.
Moving from content creation to physical AI
The bigger bet behind FLUX 3 is that understanding how objects move in video could also help robots predict and execute physical actions.
Black Forest Labs developed FLUX-mimic with robotics company mimic, using the FLUX 3 backbone to create a video-action model that predicts robot movements. The system is being tested in manufacturing environments, including with Audi.
The companies said FLUX-mimic is being used for tasks such as assembling components, handling flexible materials and placing parts into production fixtures. Audi’s Production Lab said the technology has helped robots perform complex soft-material manipulation tasks.
“We have seen these robots solve complex soft body manipulation work that would have been simply impossible with conventional robotics,” said Christoph Schneider, Audi Production Lab.
Why the architecture matters
Black Forest Labs’ strategy is based on the idea that video contains information about how the real world works. A model that learns movement, timing, and cause-and-effect relationships could potentially become useful for more than generating media.
The company said video prediction accounts for more than 95% of FLUX 3’s training compute, while audio represents a much smaller portion. It argues that once AI understands visual changes over time, learning relationships between movement and sound becomes easier.
For businesses, this could mean fewer AI systems dedicated to separate tasks. A single model could eventually support marketing content creation, product visualization, simulation and robotics. But the approach is still unproven at scale. Companies will need to evaluate reliability, cost, safety and performance before using such systems in production environments, especially when robots are involved.
Open models could shape adoption
The model is now entering early access, with video and action capabilities rolling out first. Image generation features are expected to follow, while an open-weight version called FLUX 3 Dev is planned for later.
Open access helped earlier FLUX models gain popularity among developers and researchers. Black Forest Labs said FLUX 3 Dev will eventually bring a multimodal backbone to developers, including support for image, video, audio and action prediction.
FLUX 3 is still in its early rollout, and enterprises will likely wait for independent benchmarks, pricing details and broader availability before committing to production deployments. But by extending its technology beyond image generation into robotics, Black Forest Labs is positioning itself to compete in the emerging race to build foundation models for both digital content and physical AI.
Other News: A Canadian startup is preparing home trials of its AI-guided Arlo robotic arm, highlighting how assistive robotics could support independent living and help address workforce shortages in disability care and aged-care services.


