Release2026-08-04

Black Forest Labs launched FLUX 3, a single model that generates images, video and audio, with a robot action-prediction capability described as coming soon. Video clips can run up to 20 seconds with speech, sound effects and ambience produced alongside the frames, and can start from text, an image or key frames. A cheaper draft mode returns a quick low-quality preview that can then be re-rendered at full quality.

What changed

Black Forest Labs' earlier FLUX models generated still images, with video, audio and robot action prediction handled separately or not at all.

What it unlocks

Generating a video with matching speech, sound effects and ambience, several scenes and camera angles, from one prompt, image or set of key frames.

  • up to 20 second clips in a single generation

What you need to act on it

  • access via the BFL playground, API or a partner platform
  • action-prediction listed as coming soon

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.