FLUX 3 Dev is the open-weights multimodal AI model from Black Forest Labs for image generation, video creation with native audio, and action prediction. Built on the Self-Flow architecture, FLUX 3 Dev unifies images, video, audio, and robotics in one model.
0/2000
Preview
Result will appear here
Multimodal open-weights AI
Why create with FLUX.2 Dev?
FLUX 3 Dev is the open-weights variant of FLUX 3, Black Forest Labs' multimodal foundation model announced in July 2026. Unlike earlier FLUX models that focused only on image generation, FLUX 3 Dev unifies image synthesis, video generation with native audio, and action prediction in a single architecture. Built on the Self-Flow framework, FLUX 3 Dev jointly learns from images, video, and audio to build a unified representation of the physical world. FLUX 3 Dev gives developers and creators open-weight access to this multimodal backbone for content creation and robotics applications.
Text to ImageImage Edit · up to 3 references
Multimodal content creation
Video generation with audio
Open-weights fine-tuning
Standard
Up to 1024px on the longest edge
Text 2 credits · Edit 2 credits
High
Up to 1280px on the longest edge
Text 3 credits · Edit 3 credits
Ultra
Up to 1536px on the longest edge
Text 4 credits · Edit 4 credits
Real model outputs
See FLUX.2 Dev generate and edit
Both examples were generated through the exact WaveSpeed model used by the tool above and are served from this site.
Text to Image
FLUX 3 Dev generated this studio-grade product visual with enhanced material realism, demonstrating the image generation quality of the multimodal backbone.
Prompt
Premium studio product photograph of a sculptural amber perfume bottle on dark travertine, controlled rim lighting, realistic glass refraction, restrained luxury campaign styling, crisp commercial detail
Image Edit
The FLUX 3 Dev edit preserves the product while transforming the entire environment, showcasing precise image editing within the multimodal architecture.
Prompt
Keep the perfume bottle unchanged and replace the dark studio with a warm ivory stone set, soft morning sunlight, and delicate botanical shadows; premium commercial photography
01
Image, video, and audio in one FLUX 3 Dev model
FLUX 3 Dev generates photorealistic images, creates videos up to 20 seconds with synchronized audio, and handles image editing — all within a single unified model. Use FLUX 3 Dev for end-to-end multimedia production without switching between separate tools.
02
Native video generation with FLUX 3 Dev
FLUX 3 Dev produces videos from text prompts, animates still images, generates video from keyframes, and creates multilingual dialogue with lip-synced audio. FLUX 3 Dev video generation includes styles from cinematic footage to animation and camcorder aesthetics.
03
FLUX 3 Dev open weights for full control
FLUX 3 Dev provides open-weight access to the multimodal backbone. Fine-tune FLUX 3 Dev with custom datasets, train LoRA adapters, build domain-specific pipelines for video production, or develop robotics applications using the FLUX 3 Dev action prediction capabilities.
How it works
Create with FLUX.2 Dev in three steps
1
Choose your FLUX 3 Dev modality
Start with FLUX 3 Dev Text to Image for still visuals, or explore FLUX 3 Dev video generation for motion content with native audio. Image Edit lets you transform existing assets with up to three reference images.
2
Write a detailed FLUX 3 Dev prompt
FLUX 3 Dev rewards specific prompts. For images, name the subject, materials, lighting, and lens style. For video, describe the scene, camera movement, audio elements, and dialogue. FLUX 3 Dev follows complex multimodal instructions.
3
Generate and refine with FLUX 3 Dev
Use Standard quality for quick drafts, then increase to High or Ultra for production candidates. FLUX 3 Dev supports multiple aspect ratios and output formats for both image and video workflows.
FLUX 3 Dev is the open-weights variant of FLUX 3, Black Forest Labs' multimodal foundation model. FLUX 3 Dev unifies image generation, video creation with native audio, and action prediction in a single architecture built on the Self-Flow framework. FLUX 3 Dev provides developers open-weight access to this multimodal backbone.
FLUX 3 Dev goes far beyond image generation. FLUX 3 Dev creates videos up to 20 seconds with synchronized audio, generates multilingual dialogue with lip-synced speech, animates still images into video, produces video from keyframes, and supports action prediction for robotics applications. FLUX 3 Dev is a true multimodal model.
FLUX.2 Dev was an image-only model. FLUX 3 Dev adds video generation, native audio synthesis, and action prediction built on the Self-Flow architecture. FLUX 3 Dev jointly learns from images, video, and audio to build a unified world representation, while FLUX.2 Dev only handled still images.
Yes. FLUX 3 Dev generates video with native synchronized audio, including speech synced to lip movement and sound effects synced to physical events. FLUX 3 Dev video generation supports multiple styles such as cinematic, animation, and camcorder footage, with videos up to 20 seconds per generation.
Self-Flow is the modality-agnostic, self-supervised flow matching framework behind FLUX 3 Dev. It unifies generation and representation learning across images, video, audio, and actions. Video prediction accounts for over 95% of FLUX 3 Dev compute costs, while audio adds less than 0.5% of tokens — making multimodal learning efficient.
Yes. FLUX 3 Dev includes native action prediction capabilities. The FLUX-mimic model, built on the FLUX 3 Dev backbone and developed with mimic robotics, is deployed at Audi for real production tasks including kitting, assembly, and handling flexible materials. FLUX 3 Dev achieves up to 10x sample efficiency over vision-language-action models.
Yes. FLUX 3 Dev provides open-weight access to the FLUX 3 multimodal backbone for content creation (video, audio, image) and action prediction. Developers can fine-tune FLUX 3 Dev with custom datasets, train LoRA adapters, and build specialized pipelines for any visual domain or robotics application.
Choose FLUX 3 Dev when you need open-weights access for fine-tuning, custom pipelines, or robotics development. Choose FLUX 3 Pro for the highest-quality commercial output with API access. FLUX 3 Dev gives you full control over the multimodal backbone, while FLUX 3 Pro offers polished production-ready results.
Create your next image with FLUX.2 Dev
Start from a prompt or upload references, choose the output quality, and generate directly on this page.