Black Forest Labs launched FLUX 3, a single model that generates images, video and audio, with a robot action-prediction capability described as coming soon. Video clips can run up to 20 seconds with speech, sound effects and ambience produced alongside the frames, and can start from text, an image or key frames. A cheaper draft mode returns a quick low-quality preview that can then be re-rendered at full quality.
What changed
Black Forest Labs' earlier FLUX models generated still images, with video, audio and robot action prediction handled separately or not at all.
What it unlocks
Generating a video with matching speech, sound effects and ambience, several scenes and camera angles, from one prompt, image or set of key frames.
- up to 20 second clips in a single generation
What you need to act on it
- access via the BFL playground, API or a partner platform
- action-prediction listed as coming soon
- bfl.ai2026-08-04