Black Forest Labs announced FLUX-mimic on July 23, 2026, a video-action model built on an early version of its FLUX 3 multimodal foundation model and developed with Swiss robotics company mimic robotics, which was given early access to FLUX 3. FLUX 3 is jointly trained on images, video and audio, with video prediction accounting for over 95% of training compute, using tens of millions of hours of general video plus hundreds of thousands of hours of human and robot manipulation footage. A lightweight action decoder reads intermediate features from the video prediction path; adding action prediction cut human-rated video quality by up to 10% before recovering fully after 3,500 steps. BFL reports the backbone runs input-to-world-representation in under 80ms on a single NVIDIA RTX 5090, with a full robot system reaction time of 101ms.
Sources