Black Forest Labs Release Notes
22 release notes curated from 30 sources by the Releasebot Team. Last updated: Aug 20, 2026
- Aug 20, 2026
- Date parsed from source:Aug 20, 2026
- First seen by Releasebot:Aug 20, 2026
FLUX Upscale: 2K and 4K for Video
Black Forest Labs adds FLUX Upscale, a standalone tool and API endpoint that regenerates video up to native 4K. It targets broadcast, campaign, and big screen delivery with sharper detail, faster processing options, and better repair of common artifacts.
A shot generated in hd still needs another step before it’s ready for delivery. Broadcast, campaign, and big screen formats all require more pixels, and local upscalers tend to struggle at those resolutions.
You can now use FLUX Upscale, available as a standalone FLUX Tool and endpoint, to regenerate any video at a higher resolution, up to native 4K
It understands FLUX 3 output natively, and fixes things a general-purpose upscaler can miss such as smudged faces or gridded artifacts on textures such as water and grass.
The problem we’re solving
- Local upscalers can lose quality as you move toward broadcast and campaign resolutions
- High-quality upscalers available today can be slow when processing video at scale
What FLUX Upscale does
- Takes a video from 480p up and regenerates it at up to 4K
- Fixes common imperfections in generated video
- Runs in two modes, so you can choose between speed and cost or more repair and detail
Two modes
Precise: 4 steps, $0.07/mp/s. This is faster and cheaper, and the better choice when you need to keep identity or reference details consistent.
Creative: 8 steps, $0.10/mp/s. Uses more steps to push repair and detail generation. It can change or replace identity, so reference consistency is lower than with Precise.
How it works
upscale_factor supports 1.5x, 2x, and 3x. From an hd input, that lands at roughly 1080p, 2K, and 4K.
FLUX Upscale is available via the BFL API and playground
- Try now →
- View docs →
- View Pricing →
- Aug 4, 2026
- Date parsed from source:Aug 4, 2026
- First seen by Releasebot:Aug 5, 2026
FLUX 3 Video, Part 1: Generation
Black Forest Labs releases FLUX 3 Video, making its first video generation capabilities generally available through the BFL API and select partners. It creates up to 20-second clips from text or images in HD and Full HD, with native audio and a fast Draft Mode for easier iteration.
Starting today, an initial version of FLUX 3 Video for generation from text and images is generally available via the BFL API and select partners. The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video.
FLUX 3 is our frontier multimodal model for generating and predicting video, audio, images, and actions. With this release, its video generation capabilities are being made generally available for the first time. Further details about FLUX 3 here.
Video Models Must be Reality Models.
Reality is inherently multimodal; but every snapshot representation (e.g. image, video, sound) captures only a fragment of it. No single fragment is complete on its own. This is why FLUX 3 - a model designed to model reality as accurately as possible - is natively multimodal, and designed to model reality without collapsing into a particular uniform subset or style. The results are video outputs that are not limited to a cinematic aesthetic, but can appear raw, natural, playful, nostalgic, or strange.
This makes FLUX 3 a highly flexible and controllable model for content creation. FLUX 3 is able to process both simple and complex prompts with a deep understanding of the world. It can switch between scenes and camera angles within a single generation, render typography as a natural part of the scene, and generate dialogue in multiple languages with natural accents and lip-syncing. The same model can produce something personal and natural, highly stylized and cinematic, or simply entertaining and fun.
FLUX 3 Video - Generation Capabilities
We make FLUX 3 Video available to a general audience today. In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. We are releasing our model at HD (720p) and Full HD (1080p) resolutions, and we’re providing the following set of capabilities today:
Text-to-Video: Describe a scene in simple language or using a detailed prompt. FLUX 3 follows complex instructions while generating natural movements, scene logic, and audio.
Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. FLUX 3 Video connects these in sequence while following the intended visual language.
Video Continuation: Provide FLUX 3 Video with up to four seconds of existing video and audio and tell it what should happen next. The model takes both components into account to continue movement, camera behavior, dialogue, and audio across the video seam.
Creating Multiple Shots: Create multiple scenes and camera angles within a single video, while keeping the sequence coherent.
Audio and Dialogue: Generate dialogue, sound effects, and ambient sounds along with the individual frames.
Multi-Linguality: FLUX 3 Video is built to be a powerful tool for people of many different ethnicities and languages. Supported languages include English (various dialects), Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more, with precise lip-syncing.
World Knowledge and Grounding: FLUX 3 Video combines world knowledge acquired in pretraining with real-time grounding, making it a powerful tool for documentaries, and short-form educational content, just from as little as a handful of words in a short prompt.
Draft Mode: Enables the exploration of creative directions easily. A draft generation returns a fast preview of your prompt at a fraction of the cost, so you can iterate on ideas instead of waiting for a full high-quality generation every time. When a draft is satisfactory, FLUX 3 renders the video at full quality. It includes the same subjects, same composition and same motion so the final output matches the version you approved.
FLUX 3 Video provides a SOTA Video Experience
FLUX 3 Video provides SOTA capabilities in both text-to-video and image-to-video generation. We extensively evaluated our released model against existing state-of-the-art models. Since our initial announcement the video capabilities of FLUX 3 have progressed further. Human raters found it to be the preferred model for both text-to-video and image-to-video generation.
In our internal evaluation, FLUX 3 outperforms existing SOTA models in text to video generation by a solid margin. It ties Seedance 2.0 and beats all other existing SOTA models in image-to-video.
Responsible development and deployment Assurance
We are committed to responsible development and deployment of AI models, and apply multiple layers of mitigation before, during, and after release to combat the risk of misuse. Working with a trusted third-party partners, Cinder, we evaluated FLUX 3 Video for a range of risks prior to release to validate our mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM).
What comes next
Our next releases will expand FLUX 3 Video for enhanced controllability and ship capabilities for new modalities. We will enable video generation from combinations of image, video and audio references. Our roadmap further includes FLUX 3 Image for image generation and editing, and FLUX 3 Dev as an open-weight variant.
Make things with FLUX 3.
FLUX 3 Video is available now through the BFL API and selected partners. We are excited to see what you make.
Learn more about FLUX 3 → https://bfl.ai/models/flux-3
Original source
Docs here → https://docs.bfl.ai/flux_3 All of your release notes in one feed
Join Releasebot and get updates from Black Forest Labs and hundreds of other software products.
- Jul 23, 2026
- Date parsed from source:Jul 23, 2026
- First seen by Releasebot:Jul 24, 2026
FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.
Black Forest Labs releases FLUX 3 in Early Access, a new multimodal foundation model for images, video and audio that can generate and edit across modalities with native audio, multilingual dialogue and action prediction capabilities.
FLUX 3 is now available in Early Access
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.
Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.
FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.
FLUX 3: One model, multiple capabilities.
FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.
Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).
Capabilities & Early Evaluations
As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below.
Video
FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.
Its core capabilities include the following (all outputs come with native audio generation):
- Text-to-video generation.
- Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references.
- Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context.
- Generative video-audio continuation from input video and audio.
- Keyframe-to-video generation for controlled transitions between defined moments.
- Multilingual dialogue.
- A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.
- Agentic chaining of individual clips into longer, multi-shot sequences.
- High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
- Strong typography generation and animated designs.
For the preliminary analysis below, we generated 10-second text-to-video clips in 720p with audio.
Evaluations are early and we expect further improvements
As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.
While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes.
FLUX 3 Video is now available in Early Access here
Image
FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. The model produces a wide range of output styles (see the following samples), and is able to render high-accuracy text in multiple languages.
As with video evaluations, these are preliminary results, and we expect further improvements before release. We will open up an early access phase for FLUX 3 Image in the following weeks.
Action
FLUX 3's world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data.
For the second, mimic robotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation and production deployment.
Read our thesis on why physical AI and content creation run on the same foundation, and how it's being tested on real production tasks at Audi.
Launch Plan
Over the next few weeks and months, we will make the following capabilities available, each after an early access phase for ensuring smooth rollout, collecting feedback and rigorous safety-testing. All capabilities are built from the same underlying multimodal flow matching model. These capabilities and models include:
- Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
- Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
We will also release more technical details on the underlying approach.
Request early access here
What’s next?
We are only beginning to scratch the surface of versatile, capable, unified multimodal models, and what they will enable. From interactive image & video editing, simulation to computer use and physical AI, the frontier is wide open. While we gradually roll out these new capabilities, we are already working on the next generation models. Our goal is to unify perceptual, action and language prediction in the same unified model.
If you are interested in exploring and building with FLUX 3, get in touch here. If you are interested in contributing to our mission, join us! We are hiring in Germany and the US.
Original source - Jul 23, 2026
- Date parsed from source:Jul 23, 2026
- First seen by Releasebot:Jul 24, 2026
FLUX 3 x mimic: The Next Generation of Video-Action Models
Black Forest Labs releases FLUX 3, a new multimodal foundation model that expands beyond image generation into joint audio-visual content and robot action prediction. It also powers FLUX-mimic, a next-generation video-action model already running on robots in real factory deployments.
An early version of FLUX 3, our new multimodal foundation model, is now running on robots.
We gave mimic robotics early access to FLUX.3. Their strength in robot learning and deployment, combined with the model's world knowledge and BFL's foundation model expertise, produced FLUX-mimic: the next generation of video-action models.
FLUX 1 and FLUX 2 generate images. FLUX 3 expands into multimodality and generates audio-visual content jointly - and, at the same time, provides the foundation of FLUX-mimic: A video-action model, developed in collaboration with mimic, running robots that have been tested and deployed at Audi.
At first glance, producing convincing visual content and controlling robots seem to have little in common. One requires generating pixels, the other an understanding of how the physical world responds when you touch and manipulate it. If one model does both, it was never really only a content creation model. It is a model of how the world behaves, and content creation is one thing one can do with it.
That is what FLUX 3 is.
Video is the hard part
FLUX 3 is one model, jointly trained across images, video and audio from the beginning. The most demanding part of that training - accounting for over 95% of the total compute costs - is video prediction. To generate realistic videos, a model has no choice but to learn contact, motion, weight, cause and effect; get any of them wrong and it looks wrong. Learning to render the world accurately means learning how the world behaves.
Relatively speaking, audio is the easy modality. Low dimensional and far less detailed than video, it makes up less than 0.5% of the tokens in a 720p video with audio. Once a model has done the hard work of learning video understanding, it will learn the causal relationship between video and audio to predict speech synchronized to lip movement and audio effects synchronized to the physical events causing them.
Actions follow the same shape: a low dimensional representation of a robot's state, tightly coupled to visual observations. Actions, audio and video frames are all partial representations of a single underlying physical reality. After the model has learned about the physical processes behind video and audio, action prediction is not a new departure - it is one more view of the reality it already models.
A single backbone
If that framing is correct, teaching FLUX 3 to predict actions should not incur lasting costs: we expect a brief phase of disturbance as the model has to learn the structure of the action space and align its internal representation of the world to it, before returning to full performance. That is exactly what we observe.
In a large-scale training run, we added action prediction to the curriculum and observed the effect on video generation quality. Human ratings on text-to-video and image-to-video initially fell by up to 10% as the model started to incorporate the new action modality. After 3500 steps, the model had regained its full previous quality on video generation tasks while now also predicting actions.
Each series is normalized to its own quality before action prediction was added. Higher is better.
The model had to integrate actions into its inputs and outputs - but doing so didn't cost it capacity permanently. It merely had to learn how this new modality relates to its existing model of the world. Once this was figured out, the performance penalty on its existing capabilities was gone. Video generation and action prediction don't need separate foundations. The same backbone carries both.
This makes Physical AI a natural extension of our roadmap at Black Forest Labs rather than a change in direction. Content creation is what our multimodal FLUX 3 backbone does with image, video and audio. Physical AI is what it does with actions. One foundation model, with visual intelligence at its core, enabling two families of applications. We didn't build a separate foundation model. We focused on the hard thing: building a model that understands the world. Acting in it is what that understanding makes possible.
From lab to reality: FLUX-mimic
What happens when we point the FLUX 3 backbone at real automation tasks on real production lines?
That's the question mimic and BFL created FLUX-mimic to answer. mimic builds their own robots and brings expertise in robot learning, dexterous manipulation and production deployment; BFL builds visual foundation models and brings multimodal training and modeling expertise. Together, we built a next-generation model for general-purpose manipulation - adapted to industry requirements and integrated into mimic's full-stack deployment system. FLUX-mimic is a video-action model built on the FLUX 3 backbone.
Decoding the learned world model
Our thesis is that FLUX 3 has to learn an internal representation of the world to be able to generate videos. FLUX-mimic follows through on this thesis and decodes actions from the learned world representation of the FLUX backbone. This approach, pioneered in mimic-video, trains a lightweight action decoder on top of intermediate features extracted from the video prediction path of FLUX.
Architecture overview how FLUX-mimic is built on top of FLUX 3
The success of this approach depends on two related but different aspects: the quality of the world model learned by FLUX and the quality of the feature representation of this world model. The quality of the world model is directly related to the generation quality: if a model does not understand how the world behaves, it cannot simulate it. However, even the best world model does not help an action decoder if it is inaccessible: if the feature space keeps the causal relationships between modalities entangled nonlinearly, understanding those relationships from the feature representation remains as difficult as understanding them from the raw inputs - representation quality matters.
Generation quality and representation quality have long been studied and approached in isolation from each other. Generative approaches result in high-quality world models that enable simulations and they exhibit scaling laws for predictable returns on compute investments. However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding.
As generative models themselves rely on their own features, this divergence in their representation quality seems counter-intuitive. Improved representations within generative models should improve the quality of their world model and make them more usable for downstream tasks. In our work, Self-Flow, we demonstrated how to unify generation and representation learning in a single framework and observed exactly this reciprocal improvement: the world model improved - as measured by generation quality across video, image and audio - and its representation quality improved - as measured by success rate for robot control tasks in simulation.
Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).
Scaling the world model
Scaling laws remain true with Self-Flow, and FLUX 3 is the application of that: the scaled-up version of Self-Flow. It is trained on tens of millions of hours of general video content to learn world dynamics as broadly as possible from day one, and on hundreds of thousands of hours of video content focused on human and robot manipulation tasks to be ready as a backbone for visual intelligence. This scaling is what translates the success of Self-Flow from the lab to reality. mimic deployed FLUX-mimic in real factory use cases spanning the daily reality of production and logistics work: kitting parts into structured trays, inserting electronic control units into tight-fitting fixtures, assembling components together, and handling soft, flexible materials like seals and cables that conventional automation has never been able to touch.
Benchmarks demonstrate that the action decoder outperforms previous vision-language-action models, even with a completely frozen FLUX backbone - a setting where previous vision-language-action models fail to succeed. This highlights how scaling gives our backbone strong knowledge of the world and how to act in it, and how Self-Flow makes this knowledge readily decodable from the backbone's feature representations. When finetuning the backbone together with the action decoder, FLUX-mimic achieves state-of-the-art success rates.
Dashed line marks each model's median success rate across 20 autonomous trials. Higher is better.
From world knowledge to a working task
A backbone exposing world knowledge in decodable representations changes what it takes to teach a robot a new task. If the physics is already in the representation and readily accessible, adapting to a task is no longer a matter of teaching the model how the world works - it only has to learn how this particular task maps onto what it already knows. The expensive part is done before the robot ever moves.
This shows up directly in how much demonstration data a new task requires. In our Self-Flow experiments, action prediction reached a given success rate in half the training steps compared to a video model without Self-Flow - better representations make the world knowledge easier to extract, so less data is needed to reach the same capability. The mimic-video paper reports up to 10x sample efficiency for video-action models over vision-language-action models; FLUX-mimic combines both effects.
The same benefit shows up in behavior. FLUX-mimic naturally recovers from failure: a robot that misses a grasp corrects itself, grasps again, and completes the task. No demonstration set can cover every possible way a task can go wrong. Recovery that was never demonstrated has to come from somewhere else - from a model that already knows how the world behaves.
The backbone's predicted future (top), alongside the rollout the robot produced from the decoded actions (bottom).
Fast enough to act
Real-world deployment sets a hard constraint: the model has to act as fast as the world moves. The dominant compute cost for FLUX-mimic sits in the backbone. It is the largest component of the model and, in mimic's optimized deployment stack, its latency effectively sets the ceiling for the whole system.
This is where our methodology pays off a second time. Better representations mean more capability per parameter: a model that has learned a well-structured world model needs less capacity to reach a given level of performance than one that has not. For deployment, this translates directly into being able to run a smaller backbone - and a smaller backbone is a faster backbone.
CLIP score at 1.0M training steps; higher is better. Backbone depth is the dominant driver of deployment latency, so fewer layers means a faster model.
As a result, the backbone of FLUX-mimic can be optimized to run from input to world representation in less than 80ms on a single RTX 5090 GPU - which puts it on the same order of magnitude as human visual reaction time.
Real-world deployments require additional optimizations of the full deployment stack to avoid adding any additional latency: mimic's optimizations range from the action decoder, through cutting inter-process latency between sensors, the model and actuators, to real-time chunking such that prediction and execution overlap and keep the robots running smoothly without jitter. The end result is a self-contained robot system with reaction times of 101ms.
On the factory floor
All of this leads back to the place where automation matters: the factory. Audi runs one of the most automated production networks in the automotive industry - which gives it a precise view of where conventional automation still stops. Despite decades of robotics investment, tasks with flexible parts and fine manipulation have stayed manual, largely for economic reasons: the variant diversity of premium production makes conventionally programmed robot cells too costly to re-engineer for each case. Learning-based systems change that math.
"In partnership with mimic, Audi has been testing and deploying FLUX-mimic. We have seen these robots solve complex soft-body manipulation work that would have been simply impossible with conventional robotics. This can have a major impact in assisting our employees, increasing efficiency, and expanding flexible automation across production and logistics operations. For us, partnering with pioneering companies such as mimic and Black Forest Labs is essential in pushing the frontier of physical AI and validating these innovations in real-world production environments." — Christoph Schneider, Audi Production Lab
Closing
FLUX-mimic is a purpose-built robotics model - and a proof point for what's to come. Its sample efficiency and robustness come from the FLUX 3 backbone and the quality of the representations it exposes, not from task-specific engineering. That is what lets the approach transfer across tasks, industries, and hardware. Read more through mimic.
One model, with visual intelligence at its core, generating image, video and audio - and driving robots on a production line. Content creation and physical AI are two applications of the same foundation.
Original source - Jun 4, 2026
- Date parsed from source:Jun 4, 2026
- First seen by Releasebot:Jun 4, 2026
FLUX.2 is now on device: ASUS ProArt laptops now support Klein models
Black Forest Labs ships FLUX.2 [klein] on consumer hardware for the first time, bringing a preloaded 4B model to ASUS ProArt laptops in MuseTree with NVIDIA RTX optimization, fast on-device image generation, and no API or internet required.
For the first time a FLUX model will ship on consumer hardware. In partnership with ASUS and NVIDIA, creators picking up a new ASUS ProArt laptop will find FLUX.2 [klein] optimized for the device.
High-quality image generation, for a long time, usually means an API call. You need a connection, a backend, and latency. As of today, FLUX.2 [klein] 4B will ship preloaded in ASUS's MuseTree app on their next-generation ProArt laptop with RTX GPU lineup, which has debuted at Computex 2026 in Taipei.
Our target:
sub-5-second image generation on an 8GB VRAM laptop, while other pro apps are running in the background. No API or internet required; just the model, the hardware, and your creativity!
Why on-device, why now
Cloud inference is fast and flexible, but on-device changes what's possible for certain workflows. Privacy-sensitive use cases, offline work, low-latency iteration during a live session, are real constraints that real creators hit. ASUS's ProArt lineup is built for exactly that audience: photographers, video editors, designers who need compute that keeps pace with them.
FLUX.2 [klein] was already our fastest model by design, built for speed without compromising output quality. The latest advances in GPU performance open this up to use cases that weren't possible on consumer hardware before, and running it natively on RTX with NVIDIA's optimizations makes that real.
We're expanding our co-development partnerships to bring on-device FLUX to new form factors and workflows. If you're interested in shipping FLUX models optimized for your specific hardware or device, reach out to our team - let's build something together.
Partner with us
Original source Similar to Black Forest Labs with recent updates:
- Zed release notes166 release notes · Latest Sep 4, 2026
- Vivaldi release notes16 release notes · Latest Sep 3, 2026
- Deepseek release notes21 release notes · Latest Aug 21, 2026
- Qwen release notes31 release notes · Latest Sep 2, 2026
- Freshrss release notes28 release notes · Latest May 10, 2026
- Beeper release notes81 release notes · Latest Sep 3, 2026
- May 28, 2026
- Date parsed from source:May 28, 2026
- First seen by Releasebot:May 29, 2026
FLUX VTO: Virtual Try-On at scale
Black Forest Labs launches FLUX VTO, a public virtual try-on model for apparel that delivers fast, catalog-scale try-ons with strong identity and garment fidelity, supports multiple garments and layering, and is available in the FLUX MCP and a BFL Shop demo.
What if every shopper could see themselves in an outfit before buying? What if that experience ran fast, at catalog scale, with the garment rendered exactly as it exists?
That's what FLUX VTO is built for. Apparel runs on emotional selling, and try-on is the digital version of that moment when a shopper pictures themselves in the piece and decides they want it.
Why try-on has stayed stuck in the demo phase
Virtual try-on has been promised for years, but production deployments are rare. The reasons are familiar to anyone who has evaluated the category. Models drift between generations, so identity, hair and pose shift in ways that make outputs unusable on a live product page. Garments fare no better: logos disappear, stitching degrades, prints render incorrectly and buttons vanish. Even when the geometry is right, the look is often wrong, with outputs that don't match a brand's contrast, highlights or aesthetic.
And underneath all of that sits the economics problem. Existing try-on models take 10 to 30 seconds per generation, which is too slow for interactive shopping and too expensive to run across a full catalog.
Any one of these is enough to keep try-on out of production. Together they explain why so few catalogs have shipped it.
What FLUX VTO does differently
FLUX VTO is engineered for real shopping experiences, not demos. It clears the bar on the three dimensions retailers actually care about.
Speed and cost that work at catalog scale. Generations complete in under four seconds, fast enough to feel interactive in a consumer flow. The model is also cheaper to run than comparable systems, which is what makes try-on viable across thousands of SKUs rather than a curated handful.
Fidelity on both sides of the image. Identity is preserved across generations, and garments come through with their logos, prints, stitching and hardware intact. The output looks like the person wearing the actual product, not a reinterpretation of it.
Styling flexibility. Apply up to four garments to a single model at once, transfer full outfits or individual pieces between models, and layer items such as a shirt under a jacket with correct interaction between them.
What to know before you build with it
See the look, explore the fit. FLUX VTO shows how a garment looks on a person, making it easy to explore styling, silhouette, and overall fit. Precise body and garment sizing is still evolving, so outputs are best used as visual styling guidance.
Default moderation is on. Swimwear and lingerie are not supported under the default policy. Uploads must also follow standard safety guidelines: no adult-oriented content, no child sexual abuse imagery, no non-consensual or sexually explicit content and no dangerous, derogatory or shocking material.
Use rights matter. Only upload photos of yourself or photos where you have explicit rights to use the subject's likeness.
Need to run this even faster? You can also self-host the model and run the try-on in sub-second response time!
Go give it a try!
FLUX VTO is available publicly now. Check out the docs and try on some BFL Merch in our interactive and free to use BFL Shop Demo.
Also available in our FLUX MCP now!
Original source - May 21, 2026
- Date parsed from source:May 21, 2026
- First seen by Releasebot:May 22, 2026
FLUX Erase: Remove anything, leave no trace
Black Forest Labs launches FLUX Erase, an API-powered image removal model that erases masked objects, people, text and watermarks while reconstructing the scene behind them. It promises cleaner results, optional edge expansion, and fast, lower-cost performance.
A stray person in a product shot. A cable cutting through a landscape. Text baked into a scene. These are the types of details that can break an otherwise usable image, and correcting them manually takes a lot of time.
FLUX Erase removes whatever you mask - including its traces like shadows and subtle parts missed by the mask - and reconstructs the scene behind it coherently.
The problem with existing approaches
- Visible artifacts: most removal tools leave halos, smearing, or inconsistent texture at the edges of the removed area
- Incomplete reconstruction: the background fill is generated without understanding the full scene context, producing results that need manual correction
- Limited scope: tools trained narrowly on object removal struggle with text, people, watermarks, or compositionally complex scenes
FLUX Erase
Pass an image and a binary mask. The model erases whatever you've marked and reconstructs the scene behind it, matching lighting, texture, and background, no prompt required.
It works across objects, people, text, and anything else the mask defines. An optional edge expansion setting lets you expand the mask slightly for cleaner results on complex or soft-edged subjects.
FLUX Erase matches the quality of frontier object-removal models at a fraction of the price and latency.
We evaluated FLUX Erase against other state-of-the-art models on a held-out benchmark of 198 mask-based object-removal test images. FLUX Erase wins decisively against GPT Image-2 (68.5%) and Finegrain Eraser Standard (63.2%), ties Nano Banana 2 (49.5%), and lands closely behind Nano Banana Pro (47.3%) - putting it on par with the current frontier of mask-based object-removal while running lightning fast and at a substantially lower cost.
Get access
Try the public demo.
FLUX Erase is available via the BFL API.
View docs.
View Pricing.
Original source - May 14, 2026
- Date parsed from source:May 14, 2026
- First seen by Releasebot:May 15, 2026
FLUX Outpainting: Extend any image, in any direction
Black Forest Labs launches FLUX Outpainting, a purpose-built image expansion endpoint that helps create seamless, photorealistic outpainting with flexible canvas control and up to 4MP output. It is available now via the BFL API with a public demo and API docs.
Expanding an image beyond its original frame is harder than it should be.
Most outpainting tools today still give you visible seams, broken lighting, or loss of context. We built FLUX Outpainting to solve this, and you can try it now.
The problem with existing approaches
- Seams and artifacts: most outpainting tools produce visible boundaries where the generated content meets the original (inconsistent lighting, broken edges, mismatched texture)
- Prompt dependency: models that require detailed text prompts to expand a scene introduce extra steps and unpredictable outputs
- Rigid formats: changing an image's aspect ratio means rebuilding content rather than extending it
FLUX Outpainting
FLUX Outpainting is a purpose-built expansion endpoint. Pass an image, define your target canvas size and placement, and get back a seamlessly extended result: coherent, photorealistic, and ready to use.
What makes it different:
- Natural scene extension: the model is optimized for visually coherent continuation, carrying lighting, texture, depth, and composition through without instruction
- Flexible canvas control: define the full output dimensions and image placement coordinates directly, maps cleanly to canvas-based UIs and integrates straight into the API
- Up to 4MP output: full-resolution results, ready for production
Get access
FLUX Outpainting is available via the BFL API.
Try the public demo.
View API Docs
View Pricing
Original source - Feb 24, 2026
- Date parsed from source:Feb 24, 2026
- First seen by Releasebot:May 5, 2026
Capable, Open, and Safe: Combating AI Misuse
Black Forest Labs releases FLUX.2 open-weight image models and shares early safety results showing stronger mitigations against NCII and CSAM risks, with third-party testing reporting far fewer vulnerabilities and improved safeguards across pre-training, post-training, and deployment.
Black Forest Labs is pushing the frontier of visual intelligence. Our team has released some of the most capable and most popular AI models globally. As we grow, we're taking steps to combat misuse, protect the community, and show that high performance, open innovation, and sensible safeguards go hand in hand. Today, we're excited to share early results that help validate our mitigations for emerging risks—including synthetic non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM)—and will help us strengthen these safeguards in the future.
- On third-party evaluations, our latest FLUX.2 model family demonstrates strong >10 times fewer vulnerabilities for serious risks than other leading open-weight image models, including those from large technology firms.
- Our targeted post-training mitigations help to reduce vulnerabilities by 77-98% prior to release.
- Alongside other safeguards, these mitigations can meaningfully reduce the risk of widespread misuse. For example, industry-standard moderation practices during deployment can eliminate most, if not all, residual vulnerabilities for NCII and CSAM.
Black Forest Labs is committed to open innovation
Black Forest Labs is committed to open research and development as the bedrock for competition, innovation, and security in AI. By sharing our research breakthroughs openly, we can help to accelerate the discovery of new techniques, methods, and architectures. By sharing our models openly, we enable developers around the world to build new tools that we can scarcely imagine today—from kitchen table startups to the world’s largest enterprises. To date, our founding team has contributed three of the four most popular open-weight AI models on Hugging Face, totaling over 400 million downloads. Today, our FLUX family leads the most capable open-weight image models.
We face unique challenges
We face unique challenges We take our responsibility to mitigate emerging risks seriously. The properties that make open-weight models useful can also pose a unique challenge for risk management. These models can be deployed independently without oversight from the original developer and without appropriate safeguards, in some cases using consumer hardware. They can be modified or integrated with other systems for unauthorized purposes. If a vulnerability is discovered after release, it is not possible to fully withdraw all copies of the affected model.
We respond with layers of mitigation
We respond with layers of mitigation While there are no silver bullets to prevent all misuse, layers of mitigation can help to prevent widespread misuse. Before each release, we evaluate a number of risks, including the production of unlawful content, with a focus on synthetic NCII and CSAM. We implement a series of pre-release mitigations in our models to help prevent misuse, with other post-release safeguards to address residual vulnerabilities. These mitigations incorporate best practices outlined by nonprofit organizations such as Thorn as well as agencies such as the US National Institute of Standards and Technology and the UK Office of Communications.
- Pre-training. We filter pre-training data for multiple categories of nude and pornographic material and known CSAM. Limiting exposure to this data in the first place can help prevent a user generating unlawful content, whether by eliciting harmful features in the data, or by combining lawful features into an unlawful composite image. We have partnered with the Internet Watch Foundation, an independent nonprofit organization dedicated to preventing online abuse, to filter known CSAM from the training data.
- Post-training. We undertake multiple rounds of targeted fine-tuning to provide additional mitigation against potential abuse, spanning both text-to-image (T2I) and image-to-image (I2I) attacks. By suppressing certain concepts in the trained model, these techniques can help prevent a user generating synthetic NCII or CSAM from a text prompt, or transforming an uploaded image into synthetic NCII or CSAM.
- Deployment. We release our most capable open-weight models with enforceable licenses that prohibit unlawful or infringing misuse, and require the use of filters during inference. With our open-weight models, we provide filters to help deployers detect violative or infringing activity. On our hosted services, we implement filters for a range of content types—including sexual content, hate, violence and gore, and representations of self-harm—and maintain a reporting relationship with the U.S. National Center for Missing and Exploited Children. Additionally, we apply content provenance metadata on our hosted services to help platforms and viewers identify AI-generated content once it is shared online. We include links to the Coalition for Content Provenance and Authenticity (C2PA) in our open-weight repositories to help other developers implement this metadata.
- After deployment. We subsequently monitor for patterns of violative use in both our hosted services and the open developer community. We issue and escalate takedown requests to websites, services, or businesses that misuse our models. Additionally, we may ban users or developers we detect violating our policies. We provide a dedicated email hotline to solicit feedback from the community, and welcome ongoing engagement with authorities, developers, and researchers to share intelligence about emerging risks and effective mitigations.
Throughout the development lifecycle, we conduct multiple internal and external evaluations to identify further opportunities for mitigation. For our latest open-weight model family, FLUX.2, we partnered with Cinder to conduct third-party red-teaming prior to each of our five open-weight model releases. These included FLUX.2 [dev]—a 32 billion parameter model based on rectified flow transformer architecture that enables high-quality image generation and editing—and FLUX.2 [klein], a derivative series of four size-distilled and step-distilled models, ranging from 4 to 9 billion parameters, optimized for local deployment, faster inference, and improved photorealism.
Third-party testing informed our release decision
Third-party testing informed our release decision Prior to release, we tasked Cinder to evaluate these models throughout their development lifecycle, including early, intermediate, and final checkpoints. These evaluations focused on identifying CSAM and NCII vulnerabilities across a range of T2I and I2I attacks. By observing the models’ behavior before, during, and after fine-tuning and distillation, we could better refine our mitigation strategy and make a considered release decision. We also instructed Cinder to run the same evaluation on leading open-weight models from other firms to help assess the marginal risk of our releases compared to the baseline.
Attacks included prompts that:
- Directly attempt to elicit violative content;
- Obscure a violative request in an otherwise benign context;
- Obfuscate intent through indirect language, “l33t”, or scrambling;
- Attempt to construct a violative image by assembling features that are individually nonviolative; and
- Request analogous visual features or substitutes in place of violative terms.
For I2I evaluations, which included one or more input images, attacks included requests to:
- Undress, simulate, or reimagine an otherwise clothed individual;
- De-age an individual;
- Splice together multiple individuals to produce a violative composite figure; and
- Merge multiple images into a violative scene.
Human labelers were instructed to classify outputs as potential NCII based on a range of factors, including whether an individual in the output image was depicted in a state of undress and whether they were still identifiable from the prompt, input images, or general knowledge. They were instructed to classify outputs as potential CSAM based on age, nudity anywhere in frame, and other sexual, suggestive, abusive, or obscene features, consistent with legal definitions of CSAM. Additionally, labelers were asked to characterize the model’s defensive response—such as whether the model ignored the prompt, cropped the image, or obscured violative features—to help refine our fine-tuning approach.
Our models demonstrate 10x fewer vulnerabilities
Our models demonstrate 10x fewer vulnerabilities Totaling nearly 4,000 prompts, these evaluations yielded a rich picture of comparative risk that helped inform our release decision:
- Comparative risk. Our five models demonstrated over 10 times fewer vulnerabilities than other popular open-weight models, indicating a higher robustness to misuse. These include models recently developed or funded by large technology firms with substantial resources, such as Alibaba, Tencent, and ByteDance.
- Progressive mitigation. Our post-training mitigations yielded a 77-98 percent reduction in vulnerabilities compared to earlier checkpoints. Importantly, our most lightweight and efficient [klein] models—those likely to experience the widest adoption for local inference—demonstrated the fewest vulnerabilities.
- Residual vulnerabilities. Subsequent evaluations with Cinder suggested that residual vulnerabilities can be nearly eliminated in deployment through the adoption of industry-standard moderation practices.
Above. Violative rates across popular open-weight models compared to the most vulnerable model evaluated (Hunyuan Image 3.0). Released Black Forest Labs models in dark green. n≈3,800 prompts, covering both T2I and I2I, NCII and CSAM attacks. Outputs were classified by human labelers. Models or model families capable of T2I and I2I were evaluated for both types of attack, else they were evaluated on a single modality only. Relative performance indicated by text-to-image Elo ratings (a measure of general performance; February 2026).
We subsequently decided to release FLUX.2 [dev], FLUX.2 [klein] 9B Base, and FLUX.2 [klein] 9B under a non-commercial open-weight license that permits free use for personal and research applications, and FLUX.2 [klein] 4B Base and FLUX.2 [klein] 4B under an Apache 2.0 license.
Limitations
Limitations This evaluation could not directly measure robustness to adversarial modification (e.g. via fine-tuning or low rank adapters (LoRAs). However, we expect that circumventing our embedded mitigations will be more challenging than with other models. These safeguards should raise the expertise, data, and compute barrier to a malicious actor introducing unsafe behaviors to the model. Additionally, popular platforms like Hugging Face and CivitAI continue to improve their moderation of unlawful or violative repositories, such as LoRAs intended to produce NCII. Together, we expect that embedded mitigations in our models coupled with robust downstream moderation will help to significantly reduce the distribution of malicious derivative models and unlawful content online.
This is just the beginning, and we are constantly improving
This is just the beginning, and we are constantly improving We are a small team with global impact, and the only European and American lab releasing frontier open-weight models for visual generation. Our models compete with China’s largest technology firms. Yet despite significant risk mitigation, our models continue to rank among the most capable and most popular. Through our collaboration with Cinder, we have shown that performance, openness, and safety are not mutually exclusive.
These are early days, and we are constantly improving. We welcome ongoing dialogue with researchers, authorities, and developers as we continue to refine our approach to these risks. Please reach out to us at [email protected] with feedback!
—Black Forest Labs
Original source - Jan 15, 2026
- Date parsed from source:Jan 15, 2026
- First seen by Releasebot:May 5, 2026
FLUX.2 [klein]: Towards Interactive Visual Intelligence
Black Forest Labs releases the FLUX.2 [klein] model family, its fastest image models yet, unifying generation and editing with sub-second inference, consumer-hardware support, open weights, and new quantized versions for faster, more efficient deployment.
Today, we release the FLUX.2 [klein] model family, our fastest image models to date.
FLUX.2 [klein] unifies generation and editing in a single compact architecture, delivering state-of-the-art quality with end-to-end inference as low as under a second. Built for applications that require real-time image generation without sacrificing quality, and runs on consumer hardware with as little as 13GB VRAM.
Try it now for free here
Demo showing editing with FLUX.2 [klein]
Why go [klein]?
Visual Intelligence is entering a new era. As AI agents become more capable, they need visual generation that can keep up; models that respond in real-time, iterate quickly, and run efficiently on accessible hardware.
The klein name comes from the German word for "small", reflecting both the compact model size and the minimal latency. But FLUX.2 [klein] is anything but limited. These models deliver exceptional performance in text-to-image generation, image editing and multi-reference generation, typically reserved for much larger models.
What's New
- Sub-second inference. Generate or edit images in under 0.5s on modern hardware.
- Photorealistic outputs and high diversity, especially in the base variants.
- Unified generation and editing. Text-to-image, image editing, and multi-reference support in a single model while delivering frontier performance.
- Runs on consumer GPUs. The 4B model fits in ~13GB VRAM (RTX 3090/4070 and above).
- Developer-friendly & Accessible: Apache 2.0 on 4B models, open weights for 9B models. Full open weights for customization and fine-tuning.
- API and open weights. Production-ready API or run locally with full weights.
Note: The “FLUX [dev] Non-Commercial License” has been renamed to “FLUX Non-Commercial License” and will apply to the 9B Klein models. No material changes have been made to the license.
Text to Image collage using FLUX.2 [klein]
The FLUX.2 [klein] Model Family
FLUX.2 [klein] 9B
Our flagship small model. Defines the Pareto frontier for quality vs. latency across text-to-image, single-reference editing, and multi-reference generation. Matches or exceeds models 5x its size - in under half a second. Built on a 9B flow model with 8B Qwen3 text embedder, step-distilled to 4 inference steps.
Combine multiple input images, blend concepts, and iterate on complex compositions - all at sub-second speed with frontier-level quality. No model this fast has ever done this well.
License: FLUX NCL
Imagine editing collage using FLUX.2 [klein]
FLUX.2 [klein] 4B:
Fully open under Apache 2.0. Our most accessible model, it runs on consumer GPUs like the RTX 3090/4070. Compact but capable: supports T2I, I2I, and multi-reference at quality that punches above its size. Built for local development and edge deployment.
License: Apache 2.0
FLUX.2 [klein] Base 9B / 4B:
The full-capacity foundation models. Undistilled, preserving complete training signal for maximum flexibility. Ideal for fine-tuning, LoRA training, research, and custom pipelines where control matters more than speed. Higher output diversity than the distilled models.
License: 4B Base under Apache 2.0, 9B Base under FLUX NCL
Output Diversity using FLUX.2 [klein]
Quantized versions
We are also releasing FP8 and NVFP4 versions of all [klein] variants, developed in collaboration with NVIDIA for optimized inference on RTX GPUs. Same capabilities, smaller footprint - compatible with even more hardware.
- FP8: Up to 1.6x faster, up to 40% less VRAM
- NVFP4: Up to 2.7x faster, up to 55% less VRAM
Benchmarks on RTX 5080/5090, T2I at 1024×1024
Same licenses apply: Apache 2.0 for 4B variants, FLUX NCL for 9B.
Performance Analysis
FLUX.2 [klein] Elo vs Latency (top) and VRAM (bottom) across Text-to-Image, Image-to-Image Single Reference, and Multi-Reference tasks.
FLUX.2 [klein] matches or exceeds Qwen's quality at a fraction of the latency and VRAM, and outperforms Z-Image while supporting both text-to-image generation and (multi-reference) image editing in a unified model. The base variants trade some speed for full customizability and fine-tuning, making them better suited for research and adaptation to specific use cases. Speed is measured on a GB200 in bf16.
Into the New
FLUX.2 [klein] is more than a faster model. It's a step toward our vision of interactive visual intelligence. We believe the future belongs to creators and developers with AI that can see, create, and iterate in real-time. Systems that enable new categories of applications: real-time design tools, agentic visual reasoning, interactive content creation.
Resources
Try it
- Demo
- Playground
- HF Space for [klein] 9B, HF Space for [klein] 4B
Build with it
- Documentation
- GitHub
- Model Weights
Learn more
- https://bfl.ai/models/flux-2-klein
- Nov 25, 2025
- Date parsed from source:Nov 25, 2025
- First seen by Releasebot:May 5, 2026
FLUX.2: Frontier Visual Intelligence
Black Forest Labs releases FLUX.2, a new image model family for real-world creative workflows with stronger multi-reference consistency, better prompt following, improved text rendering, higher-detail photorealism, and image editing up to 4MP. It also offers pro, flex, dev, and VAE options.
FLUX.2 is designed for real-world creative workflows, not just demos or party tricks. It generates high-quality images while maintaining character and style consistency across multiple reference images, following structured prompts, reading and writing complex text, adhering to brand guidelines, and reliably handling lighting, layouts, and logos. FLUX.2 can edit images at up to 4 megapixels while preserving detail and coherence.
Black Forest Labs: Open Core
We believe visual intelligence should be shaped by researchers, creatives, and developers everywhere, not just a few. That’s why we pair frontier capability with open research and open innovation, releasing powerful, inspectable, and composable open-weight models for the community, alongside robust, production-ready endpoints for teams that need scale, reliability, and customization.
When we launched Black Forest Labs in 2024, we set out to make open innovation sustainable, building on our experience developing some of the world’s most popular open models. We’ve combined open models like FLUX.1 [dev]—the most popular open image model globally—with professional-grade models like FLUX.1 Kontext [pro], which powers teams from Adobe to Meta and beyond. Our open core approach drives experimentation, invites scrutiny, lowers costs, and ensures that we can keep sharing open technology from the Black Forest and the Bay into the world.
From FLUX.1 to FLUX.2
Precision, efficiency, control, extreme realism - where FLUX.1 showed the potential of media models as powerful creative tools, FLUX.2 shows how frontier capability can transform production workflows. By radically changing the economics of generation, FLUX.2 will become an indispensable part of our creative infrastructure.
Output Versatility: FLUX.2 is capable of generating highly detailed, photoreal images along with infographics with complex typography, all at resolutions up to 4MP
What’s New
- Multi-Reference Support: Reference up to 10 images simultaneously with the best character / product / style consistency available today.
- Image Detail & Photorealism: Greater detail, sharper textures, and more stable lighting suitable for product shots, visualization, and photography-like use cases.
- Text Rendering: Complex typography, infographics, memes and UI mockups with legible fine text now work reliably in production.
- Enhanced Prompt Following: Improved adherence to complex, structured instructions, including multi-part prompts and compositional constraints.
- World Knowledge: Significantly more grounded in real-world knowledge, lighting, and spatial logic, resulting in more coherent scenes with expected behavior.
- Higher Resolution & Flexible Input/Output Ratios: Image editing on resolutions up to 4MP.
All variants of FLUX.2 offer image editing from text and multiple references in one model.
Available Now
The FLUX.2 family covers a spectrum of model products, from fully managed, production-ready APIs to open-weight checkpoints developers can run themselves. The overview graph below shows how FLUX.2 [pro], FLUX.2 [flex], FLUX.2 [dev], and FLUX.2 [klein] balance performance, and control
- FLUX.2 [pro]: State-of-the-art image quality that rivals the best closed models, matching other models for prompt adherence and visual fidelity while generating images faster and at lower cost. No compromise between speed and quality. → Available now at BFL Playground, the BFL API and via our launch partners.
- FLUX.2 [flex]: Take control over model parameters such as the number of steps and the guidance scale, giving developers full control over quality, prompt adherence and speed. This model excels at rendering text and fine details. → Available now at bfl.ai/play, the BFL API and via our launch partners.
- FLUX.2 [dev]: 32B open-weight model, derived from the FLUX.2 base model. The most powerful open-weight image generation and editing model available today, combining text-to-image synthesis and image editing with multiple input images in a single checkpoint. FLUX.2 [dev] weights are available on Hugging Face and can now be used locally using our reference inference code. On consumer grade GPUs like GeForce RTX GPUs you can use an optimized fp8 reference implementation of FLUX.2 [dev], created in collaboration with NVIDIA and ComfyUI. You can also sample Flux.2 [dev] via API endpoints on FAL, Replicate, Runware, Verda, TogetherAI, Cloudflare, DeepInfra. For a commercial license, visit our website.
- FLUX.2 [klein] (coming soon): Open-source, Apache 2.0 model, size-distilled from the FLUX.2 base model. More powerful & developer-friendly than comparable models of the same size trained from scratch, with many of the same capabilities as its teacher model. Join the beta
- FLUX.2 - VAE: A new variational autoencoder for latent representations that provide an optimized trade-off between learnability, quality and compression rate. This model provides the foundation for all FLUX.2 flow backbones, and an in-depth report describing its technical properties is available here. The FLUX.2 - VAE is available on HF under an Apache 2.0 license.
Generating designs with variable steps: FLUX.2 [flex] provides a “steps” parameter, trading off typography accuracy and latency. From left to right: 6 steps, 20 steps, 50 steps.
Controlling image detail with variable steps: FLUX.2 [flex] provides a “steps” parameter, trading off image detail and latency. From left to right: 6 steps, 20 steps, 50 steps.
The FLUX.2 model family delivers state-of-the-art image generation quality at extremely competitive prices, offering the best value across performance tiers.
For open-weights image models, FLUX.2 [dev] sets a new standard, achieving leading performance across text-to-image generation, single-reference editing, and multi-reference editing, consistently outperforming all open-weights alternatives by a significant margin.
Whether open or closed, we are committed to the responsible development of these models and services before, during, and after every release.
How It Works
FLUX.2 builds on a latent flow matching architecture, and combines image generation and editing in a single architecture. The model couples the Mistral-3 24B parameter vision-language model with a rectified flow transformer. The VLM brings real world knowledge and contextual understanding, while the transformer captures spatial relationships, material properties, and compositional logic that earlier architectures could not render.
FLUX.2 now provides multi-reference support, with the ability to combine up to 10 images into a novel output, an output resolution of up to 4MP, substantially better prompt adherence and world knowledge, and significantly improved typography. We re-trained the model’s latent space from scratch to achieve better learnability and higher image quality at the same time, a step towards solving the “Learnability-Quality-Compression” trilemma. Technical details can be found in the FLUX.2 VAE blog post.
More Resources:
- FLUX.2 Documentation
- FLUX.2 Prompting Guide
- FLUX.2 Open Weights / Inference Code
- FLUX Playground
Into the New
We're building foundational infrastructure for visual intelligence, technology that transforms how the world is seen and understood. FLUX.2 is a step closer to multimodal models that unify perception, generation, memory, and reasoning, in an open and transparent way.
Join us on this journey. We're hiring in Freiburg (HQ) and San Francisco. View open roles.
Original source - Sep 25, 2025
- Date parsed from source:Sep 25, 2025
- First seen by Releasebot:May 5, 2026
FLUX.1 Kontext now in Adobe Photoshop: Powering Every Pixel
Black Forest Labs adds FLUX.1 Kontext [Pro] inside Photoshop Generative Fill, giving beta users direct in-app image editing with precise, coherent results and full creative control. The model is available from September 25 and is free for a limited time during beta.
Until now, testing different generative models meant juggling apps, exporting files, and piecing results together. With FLUX.1 Kontext [Pro] inside Photoshop, that friction disappears. You can select our model, simply describe the edits you want, then refine them with Photoshop’s suite of tools. That leads to precise and coherent results, while maintaining your full creative control.
Starting September 25th, Photoshop (beta) users around the world will be able to use FLUX.1 Kontext [Pro] directly inside Generative Fill, and for a limited time during beta, users can try out our model for free.
Our models are built on the frontier of generative visual AI research. FLUX.1 Kontext [Pro] doesn’t just deliver contextual accuracy and creative flexibility, it’s 3X faster than competing models, making the experience significantly more seamless. That means less waiting, more iterating, and more time spent in flow.
- Photographers: Select a subject and let FLUX.1 Kontext [Pro] generate contextually accurate backgrounds that blend seamlessly, while Photoshop’s masking keeps your subject untouched.
- Designers: Add realistic props, signage, or scenery into layouts. FLUX.1 Kontext [Pro]’s fills integrate naturally, and Photoshop’s blending modes and smart objects ensure polish.
- Creative Directors: Mock up campaign assets or product shots at speed. Generate consistent, on-brand details with FLUX.1 Kontext [Pro], then perfect them using Photoshop’s adjustments.
By pairing Photoshop’s professional editing environment with FLUX.1 Kontext [Pro]’s accuracy and unmatched speed, creators everywhere gain the freedom to push imagination further. We’re bringing our vision: to power every pixel, everywhere, directly into your daily workflow.
Get started now
Original source - Aug 5, 2025
- Date parsed from source:Aug 5, 2025
- First seen by Releasebot:May 5, 2026
FLUX Models Launch on Azure AI Foundry for Enterprise-Ready Image Generation
Black Forest Labs adds its FLUX flagship models to Azure AI Foundry, bringing FLUX.1 Kontext [pro] and FLUX1.1 [pro] to Azure for enterprise-ready text-to-image and image-to-image deployment with stronger scale, security, and easier access.
Starting today, Black Forest Labs’ flagship models are available directly from Microsoft on Azure AI Foundry.
FLUX.1 Kontext [pro] and FLUX1.1 [pro] can be accessed through Azure AI Foundry, offering customers an enterprise-ready path to deploy our state-of-the-art text-to-image and image-to-image foundation models with the scale, security, and simplicity of Azure.
BFL’s collaboration with Microsoft started from our earliest days when we partnered with Azure to build our training and inference clusters. Since then, we’ve worked hand in hand with Microsoft to optimize our model performance and make it easier for customers to access our models. Today, everything from exploration to deployment of our models becomes faster with FLUX.1 Kontext [pro] and FLUX.1.1 [pro] available through Azure’s powerful ecosystem.
Why this matters
With FLUX models now on Azure, you get:
- Microsoft-backed Service Level Agreements
- Azure-native deployment and observability
- Access via pay-as-you-go or Provisioned Throughput (fungible PTUs), meaning you can flexibly use your quota and reservations across any Direct from Azure models.
- Adherence to Microsoft's Responsible AI standards
- Enterprise security, governance, and scalability
- Zero compromise in speed and quality
What you can build
FLUX.1 Kontext [pro] offers the in-context image generation and iterative editing capabilities, allowing with a single model for you to make local edits, transfer styles, replace background, add typography - all while retaining character consistency and offering up to 8x the speed of other editing models for 1MP resolution. On KontextBench, FLUX.1 Kontext [pro] ranks #1 on text-guided editing and character-consistency.
FLUX1.1 [pro] is another lightning fast text-to-image model that offers up to 1.6MP resolution.
Whether you’re:
- An e-commerce company looking to integrate our models for product photography,
- A financial services company looking to streamline marketing materials, or
- A studio wanting to consistently build out a visual story
you can deploy our FLUX models directly from Azure AI Foundry with the convenience of working through your Azure account.
Get started on Azure AI Foundry by doing the following:
- If you don’t have an Azure subscription, you can sign up for an Azure account here.
- Navigate to Azure AI Foundry at ai.azure.com
- Search the model name, e.g. “FLUX Kontext” for image editing, or "Flux 1.1", for text to image, in the model catalog
- Open the model card in the model catalog on Azure AI Foundry.
- Click on “deploy” to obtain the inference API and key and also to access the playground.
- You should land on the deployment page that shows you the API and key in less than a minute. You can try out your prompts in the playground.
- You can use the API and key with various clients
We can’t wait to see what you create.
Original source - Jul 31, 2025
- Date parsed from source:Jul 31, 2025
- First seen by Releasebot:May 5, 2026
FLUX.1 Krea [dev]: An ‘Opinionated’ Text-to-Image Model
Black Forest Labs releases FLUX.1 Krea [dev], a new open-weights text-to-image model developed with Krea AI. It brings more photorealistic, distinctive image generation, stronger realism, and compatibility with the FLUX.1 [dev] ecosystem for customization.
An ‘Opinionated’ Text-to-Image Model
The BFL model garden just got an exciting update: We are proud to release FLUX.1 Krea [dev], developed in collaboration with Krea AI. FLUX.1 Krea [dev] is a new state-of-the-art open-weights model for text-to-image generation that overcomes the oversaturated 'AI look' to achieve new levels of photorealism with its distinctive aesthetic approach.
FLUX.1 Krea [dev] is the open weights version of Krea 1, offering strong performance with highly distinctive aesthetics and exceptional realism. It has been trained with the goal to generate more realistic and diverse images that do not contain oversaturated textures, a known issue in text-to-image generation. Due to these properties, we call FLUX.1 Krea [dev] ‘opinionated’ - a text-to-image model that offers its users pleasant surprises in the form of diverse, visually interesting images.
Despite its idiosyncrasies, FLUX.1 Krea [dev] outperforms previous open text-to-image models and is on par with closed solutions like FLUX1.1 [pro] in human preference assessments. Moreover it is architecturally compatible with the FLUX.1 [dev] ecosystem and serves as a flexible base model for customization for down-stream applications.
The weights of FLUX.1 Krea [dev] are now available in the BFL HuggingFace repository. Commercial Licenses can are available in the BFL Licensing Portal. Our partners FAL, Replicate, Runware, DataCrunch and TogetherAI provide API endpoints for easy integration.
Key Features of FLUX.1 Krea [dev]
The key features of our most recent open-weights text-to-image model are
- State-of-the-art open text-to-image generation
- Highly distinctive aesthetics that overcome the common "AI look" problem
- Exceptional realism and image quality
- Enhanced flexibility for customization
- Compatibility with the FLUX.1 [dev] architecture and ecosystem
Collaborative Model Development
This joint project between BFL and Krea demonstrates the value of collaborative model development between foundation model and applied AI labs. By providing a specialized and flexible base model tailored for down-stream finetuning, we helped Krea to achieve previously unfeasible results.
FLUX.1 Krea [dev] showcases how targeted collaboration between foundation model developers and application-focused teams can push the boundaries of open AI image generation.
We're just getting started. If you want to join us on our mission, we are actively hiring talented individuals across multiple roles. Apply here.
Original source - Jun 26, 2025
- Date parsed from source:Jun 26, 2025
- First seen by Releasebot:May 5, 2026
FLUX.1 Kontext [dev] - Open Weights for Image Editing
Black Forest Labs releases FLUX.1 Kontext [dev], an open-weight image editing model for researchers and developers that runs on consumer hardware and supports iterative, local and global edits. It also adds optimized TensorRT weights, a self-serve licensing portal, and updated license terms.
Up until today, all capable generative image editing models were only available as proprietary tools. Today, that changes. We release FLUX.1 Kontext [dev], our developer version of FLUX.1 Kontext [pro], which delivers proprietary-level image editing performance in a 12B parameter model that can run on consumer hardware.
Making model weights openly accessible is fundamental to technological innovation. FLUX.1 Kontext [dev] is now available as an open-weight model under the FLUX.1 Non-Commercial License, providing free access for research and non-commercial use. FLUX.1 Kontext [dev] is compatible with the existing FLUX.1 [dev] inference code and comes with day-0 support for popular inference frameworks like ComfyUI, HuggingFace Diffusers and TensorRT.
The model weights are available on HuggingFace. Our partners FAL, Replicate, Runware, DataCrunch and TogetherAI and ComfyUI provide ready-to-use API endpoints and code for cloud-based and/or local inference.
The technical report is available on arxiv.
Setting New Standards in Open Image Editing
FLUX.1 Kontext [dev] focuses exclusively on editing tasks. The model enables iterative editing, excels at character preservation across a diverse set of scenes and environments, and allows both precise local and global edits.
At Black Forest Labs, we remain committed to providing researchers and developers with best-in-class open tools that are competitive with existing proprietary solutions. To validate the performance of FLUX.1 Kontext [dev], we conducted extensive evaluation across multiple image editing benchmarks.
Human preference evaluations on KontextBench, our newly released image editing benchmark, demonstrate that FLUX.1 Kontext [dev] outperforms existing open image editing models, (Bytedance Bagel, HiDream-E1-Full) and closed models (Google's Gemini-Flash Image) across many categories. Independent evaluations run by Artificial Analysis confirm these findings.
Optimized for NVIDIA Blackwell Architecture
We have collaborated with NVIDIA to build optimized TensorRT weights specifically designed for the new NVIDIA Blackwell architecture which brings greatly improved inference speed and reduces memory usage while maintaining high-quality image editing performance.
Additionally to the original FLUX.1 Kontext [dev] weights, we’re making available these BF16, FP8 and FP4 TensorRT variants in our Hugging Face repository, giving developers the flexibility to balance speed, efficiency, and quality tailored to their use case. These optimized weights ensure that FLUX.1 Kontext [dev] can take full advantage of the latest hardware capabilities.
Streamlined Commercial Access: The BFL Self-Serve Portal
We are releasing a self-serve licensing portal with transparent terms and standardized commercials for simplifying commercial access to all of our open weights models. This includes the novel FLUX.1 Kontext [dev] as well as the FLUX.1 Tools [dev] and the popular text-to-image model FLUX.1 [dev].
Our self-serve portal provides transparent licensing terms that enable businesses to confidently integrate FLUX.1 models into their commercial products and services. Commercial Licenses to our open weights models can now be purchased with only a few clicks, accelerating the path from development to deployment. More information on self-serve licensing can be found at the BFL Helpdesk.
License Update
Black Forest Labs also updated the FLUX.1 [dev] Non-Commercial License with the following changes:
- Non-Commercial Purpose. We edited the definition of “Non-Commercial Purpose” to better clarify what constitutes Non-Commercial Purposes under the FLUX.1 [dev] Non-Commercial License.
- Content Filters. To prevent the creation and dissemination of unlawful or infringing content, the FLUX.1 [dev] Non-Commercial License requires content filters or manual review to be used with the FLUX.1 [dev] models. We’ve also made corresponding adjustments to the indemnification of the license.
- Content Provenance. Users of FLUX.1 [dev] Models under a FLUX.1 [dev] Non-Commercial License must follow applicable law for content provenance under the license.
- Restrictions. We made some clarifications on what are not permitted uses of FLUX.1 [dev] Models under a FLUX.1 [dev] Non-Commercial License.
Resources
- Model weights: https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev
- Code: https://github.com/black-forest-labs/flux
- API Documentation: https://docs.bfl.ai/quick_start/introduction
- Self-Serve Portal: http://bfl.ai/pricing/licensing
- Helpdesk: https://help.bfl.ai
We're just getting started. If you want to join us on our mission, we are actively hiring talented individuals across multiple roles. Apply here.
Original source
Curated by the Releasebot team
Releasebot is an aggregator of official release notes from hundreds of software vendors and thousands of sources.
Our editorial process involves the manual review and audit of release notes procured with the help of automated systems.