Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

FLUX 3 Image Lets Agents Build Images With Bounding Boxes

A tool wherein boxes are drawn upon a canvas, and within each, a description placed, that a model may render an image entire.

By mitch·5 min read
A canvas bears glowing boxes, each enclosing a fragment of a vast scene of festival revelers beneath a great dome.

The makers of Black Forest Labs have put out a new model called FLUX 3 Image. This multimodal creation allows agents to construct pictures from nothing more than bounding boxes. It was built with a grasp of how images should be arranged on a canvas, so people can set down each item within its own box, explain what belongs inside it, and leave the finishing work to the model itself.

This tool targets agents who require fast image generation and editing. It includes support for rendering in native 4K, pixel-perfect editing, and a commercial weights license for companies that wish to run image generation at scale.

How Bounding Boxes Work

A fixed grid with coordinates from 0 to 1000 forms the basis of the entire process. On it, boxes are drawn and labeled with descriptions. Each box is described using four numbers: its lowest and highest y and x values, arranged as [y_min, x_min, y_max, x_max].

Advertisement

A global caption describes the whole image in one paragraph. Below that comes an element table, a JSON array with one row per element. Each row includes an id, a bounding box, and a description. The caption cites every element by its id, like animal_1, where that element first appears.

If you’d rather not draw every box by hand, an LLM can handle the layout planning for you. You supply it with one line and an aspect ratio, and it produces a caption along with an element table. That table assigns a box and a semantic id to each item that matters, keeping every box editable so you can move the ones you don’t like. The model then generates inside the rest as the agent placed them.

Editing an Existing Image

The model also supports editing images you’ve already made. You can make several targeted edits at once, re-describing, replacing, or moving individual boxes. Everything you didn’t touch stays exactly where it was, so an image holds together across one round of edits after another.

The tool FLUX 3 leaves your drawn boxes untouched. Instead, it takes your brief input and expands it into a full caption that matches what the model was built on. The upsampler might add details around your boxes and propose extra items, yet each box you drew gets sent to the model exactly as it appeared, keeping its id and its coordinates intact. Only what the caption actually refers to will make it through.

What Kinds of Images Suit Bounding Boxes

This approach delivers its strongest results when applied to pieces built from numerous elements arranged with precise connections between them. It applies to layouts where type is placed around an image, to collages, panel grids, and editorial spreads, as well as to busy scenes where every face occupies its own distinct position.

The illustration given is a festival setting assembled from bounding boxes. The caption offers a description of the entire picture, while the item list sets out each box according to its id, position, and contents.

  • Fr_Text_1: “LE FESTIVAL DU SOLEIL” in a thin, elegant serif typeface in a light cream color
  • town_1: faint lights and small buildings of a coastal town at the foot of the hills
  • dome_1: a massive, smooth parabolic dome of pale concrete
  • swimmers_1: dozens of small, silhouetted figures scattered in the dark water, wading
  • crowd_1: a large crowd of people seated on the beach in casual, light-colored summer attire

Using the Playground and API

Users can access the system through two paths: the Playground, where they draw boxes by hand, or the BFL API, which lets them send a layout prompt straight to FLUX 3.

Companies gain the freedom to adjust and put the model into service on their own servers through the commercial weights license. Black Forest Labs backs these deployments with assistance.

Frequently Asked Questions

Below is a fast look at the answers to common questions about the model, which make up the FAQ.

Question Answer
What is FLUX 3 Image? FLUX 3 is Black Forest Labs’ multimodal model for video, audio, images and actions. FLUX 3 Image is the part that generates and edits images.
How do bounding boxes work? Drag out a box for every element that matters and describe what goes in it. Then give the scene one line that ties the elements together, and FLUX 3 renders it with every element inside its box.
What does a layout prompt look like? It has two parts: a global caption describing the whole image, followed by an element table, a JSON array with one row per element.
Do I have to draw every box myself? No. Give the agent one line and an aspect ratio, and an LLM plans the layout for you.
Can I edit an image I already made? Yes, and you can make several targeted edits at once. Each box stays editable.
Does FLUX 3 change the boxes I draw? No. A prompt upsampler turns your short request into the dense caption FLUX 3 was trained on, but every box you drew reaches the model verbatim.

The Bottom Line on FLUX 3 Image

FLUX 3 Image is a powerful tool for agents who need to generate or edit images quickly. The bounding box approach gives precise control over composition, and the pixel-perfect editing feature means you can target changes without touching the rest of the image.

The design of the system splits the work between two parts: one part learns to read image arrangement and composition, while the other handles the actual rendering. The whole purpose rests on that separation.

A license for commercial weights gives businesses permission to run image generation on a large scale. Anyone constructing visual content will find that the blend of exactness and freedom of choice is worth putting to the test.

The Playground and the BFL API are currently the two ways to access the system. The FAQ covers common questions about the model, while the example scene demonstrates how the bounding box method functions in practice.

The model responds to hands-on exploration by producing results based on what users draw and label. Boxes are drawn, captions are added, and the model’s output follows.

See the 3 image at bfl.ai.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.