Generate video with synchronized audio from text, images, or video. FLUX 3 is Black Forest Labs' multimodal model (early access preview).
Frequently asked questions
What is FLUX 3 and who made it?
FLUX 3 is a multimodal AI video generation model created by Black Forest Labs. It uses the Self-Flow architecture to generate synchronized video and audio from text prompts, images, or existing video clips in a single inference pass. It is currently available as an early access preview on the CRAISEE platform via Replicate.
What can FLUX 3 generate?
FLUX 3 can generate text-to-video, image-to-video (single or two images), storyboard-to-video using 3–10 images, and video continuations. It also produces synchronized audio alongside video by default, outputting a single .mp4 file with matching sound — no separate audio pipeline or post-processing is required.
What resolution and aspect ratios does FLUX 3 support?
FLUX 3 supports output resolutions up to 1080p and video lengths up to 20 seconds. It supports multiple aspect ratios including 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16, covering cinematic widescreen, standard, social media vertical, and square formats for diverse creative use cases.
What is Draft Mode in FLUX 3?
Draft Mode is a fast, lower-cost 720p preview option in FLUX 3 designed for rapid prompt iteration. It lets creators quickly test and refine their prompts before committing to a full-quality 1080p render, saving both time and compute costs during the creative development process.
How does FLUX 3 generate audio with video?
FLUX 3 uses the Self-Flow architecture, a unified flow-matching model trained simultaneously across image, video, and audio modalities. This allows it to generate synchronized audio alongside video in a single inference pass, producing a complete .mp4 file with matching sound without requiring separate audio generation tools or post-processing steps.
What is the Self-Flow architecture used by FLUX 3?
Self-Flow is a unified flow-matching architecture developed by Black Forest Labs that trains across image, video, and audio modalities simultaneously. This enables FLUX 3 to produce cohesive audiovisual content in one inference pass, making it one of the first models capable of generating synchronized video and audio from a single model run.
Who is FLUX 3 designed for?
FLUX 3 is designed for creators, filmmakers, marketers, and developers who need high-quality video output without assembling separate audio and video pipelines. Its flexible input modes — text, single image, dual images, storyboards, and video continuation — make it suitable for a wide range of professional and creative production workflows.
How does FLUX 3 handle storyboard-to-video generation?
In storyboard-to-video mode, FLUX 3 accepts 3 to 10 images. The first and last images anchor the start and end of the clip, while intermediate images are distributed evenly across the timeline. The model interpolates motion between all frames, allowing creators to guide the visual narrative with a sequence of reference images.



