Black Forest Labs introduced FLUX 3, a single model trained together on images, video and sound that can also predict physical movement for robots. The video part, which produces clips with matching audio, opened in early access, while the image, action and open-weight versions are promised over the coming weeks and months. The company describes its comparison results as preliminary and still improving.
What changed
Earlier FLUX models generated still images only, without video, sound or movement prediction.
What it unlocks
Generating video clips with matching sound from a text prompt, a starting image or a reference clip, and chaining clips into longer multi-shot sequences.
- video with audio up to 20 seconds in a single generation
- preferred over Runway Gen-4.5 in 77% of comparisons
- preferred over Luma Ray 3.2 in 93% of comparisons
- preferred over Kling v3 Pro in 60% of comparisons
What you need to act on it
- approved request for early access
- video capability only at launch; image and open-weight versions come later
Sources