Flux 3 AI Video and Image Model

Flux 3 AI Video and Image Model

Flux 3 is Black Forest Labs' multimodal model for video, images, native audio, and action prediction. Explore confirmed Flux 3 features and early-access plans.

  • Native Video with Synchronized Audio: Flux 3 can generate diverse videos with native audio in a single pass, with clips up to 20 seconds during early access. Black Forest Labs highlights synchronized speech, sound effects, and physical events, but the published evaluations remain preliminary.
  • Multimodal Inputs and References: Flux 3 supports text-to-video, image-to-video, video-to-video, and video-audio continuation. Images can act as starting frames or visual references, while source clips can carry a character or central element into a new scene.
  • Keyframes, Longer Sequences, and Typography: Flux 3 includes keyframe-to-video transitions, multilingual dialogue, animated typography, varied aspect ratios, and agentic chaining of clips. Black Forest Labs says visual references can help preserve characters across multi-shot sequences, though independent testing is still limited.
  • One Backbone Beyond Video: Flux 3 also targets image synthesis and editing plus action prediction from the same multimodal backbone. Image early access is planned after the video preview, while action access begins with selected research and commercial partners such as mimic robotics.
  • A Planned Open-Weight Multimodal Model: Flux 3 Dev is planned as an open-weight backbone for video, audio, images, and action prediction. Black Forest Labs has not yet published the release date, license, model size, hardware requirements, API pricing, or final production benchmarks.
  • Define the Scene and References: Start with the subject, environment, action, camera movement, duration, aspect ratio, dialogue, and sound. When Flux 3 access supports references, add a starting image, keyframe, or source clip and identify the details that must remain consistent.
Black Forest Labs Multimodal Model

Flux 3 AI Video and Image Model

Flux 3 is Black Forest Labs' multimodal model for video, images, native audio, and action prediction. Explore confirmed Flux 3 features and early-access plans.

Explore AI Video

Official Flux 3 Video Showcase

Watch four launch samples published by Black Forest Labs on July 23, 2026. These official Flux 3 clips show coastal weather, first-person racing, a mirrored dance sequence, and warm clay animation.

Flux 3 Ocean Dynamics

An official Flux 3 launch sample showing storm-driven waves, spray, and seabirds along a coastal scene.

Flux 3 First-Person Motion

An official Flux 3 launch sample following a motorcycle around a rain-soaked circuit with fast camera movement, wet asphalt, and trackside barriers.

Mirrors, Water, and Dance

An official Flux 3 launch sample showing a ballroom dance among mirrored architecture, water, candlelight, and layered reflections.

Flux 3 Stylized Character Animation

An official Flux 3 launch sample using warm clay animation, tactile materials, and character-focused physical motion.

What Makes Flux 3 Different

01 / Video • Audio

Native Video with Synchronized Audio

Flux 3 can generate diverse videos with native audio in a single pass, with clips up to 20 seconds during early access. Black Forest Labs highlights synchronized speech, sound effects, and physical events, but the published evaluations remain preliminary.

02 / Text • Image • Video

Multimodal Inputs and References

Flux 3 supports text-to-video, image-to-video, video-to-video, and video-audio continuation. Images can act as starting frames or visual references, while source clips can carry a character or central element into a new scene.

03 / Control • Continuity

Keyframes, Longer Sequences, and Typography

Flux 3 includes keyframe-to-video transitions, multilingual dialogue, animated typography, varied aspect ratios, and agentic chaining of clips. Black Forest Labs says visual references can help preserve characters across multi-shot sequences, though independent testing is still limited.

04 / Image • Action

One Backbone Beyond Video

Flux 3 also targets image synthesis and editing plus action prediction from the same multimodal backbone. Image early access is planned after the video preview, while action access begins with selected research and commercial partners such as mimic robotics.

05 / Dev • Open Weights

A Planned Open-Weight Multimodal Model

Flux 3 Dev is planned as an open-weight backbone for video, audio, images, and action prediction. Black Forest Labs has not yet published the release date, license, model size, hardware requirements, API pricing, or final production benchmarks.

How Flux 3 Fits a Video Production Workflow

This workflow is based on capabilities announced by Black Forest Labs and is intended for approved early-access users. Features, interfaces, and usage terms may change before general availability.

Define the Scene and References

Start with the subject, environment, action, camera movement, duration, aspect ratio, dialogue, and sound. When Flux 3 access supports references, add a starting image, keyframe, or source clip and identify the details that must remain consistent.

Generate and Evaluate the Core Shot

Use Flux 3 to create the main shot, then review anatomy, object persistence, lip sync, audio timing, physical motion, typography, and camera continuity. Change one variable at a time instead of rewriting the entire prompt after every preview.

Extend, Chain, and Finish

Black Forest Labs says visual references and keyframes can help chain clips into longer, multi-shot sequences. Production use and public sharing remain subject to the applicable Flux 3 access terms and permissions.

Flux 3 Frequently Asked Questions

Explore AI Video While Flux 3 Rolls Out

Explore AI Video