Loading…
Loading…
Turn a reference image or idea into a structured, engine-ready prompt
Drag an image here, or
Stays in your browser — never uploaded. Look at it while you fill in the fields below.
Sets the "Camera / Technique" label on field 7 below and adds class-specific hint chips.
Output Quality: high resolution, 8k, sharp, clean render
A freeform sentence is easy to write and hard to render consistently — it skips lighting, skips composition, and leaves the model to guess at whatever you left out. The 11 fields above force every prompt to cover the same ground: what's in frame, what style it's rendered in, where it's set, how it's lit, how it's framed, what colors dominate, what captured or created it, what it feels like, what small details sell it, what must not appear, and what finish you want.
Two things make the difference between a prompt that renders what you meant and one that renders a generic average. First, specificity: "three gulls over a grey harbor" gives the model a count and a scene; "birds" gives it nothing to commit to. Second, one interpretation per field — "possibly a cloak, or maybe armor" forces a guess, and generators guess badly. Fill in one option, not a menu.
The Negative field does the opposite job — it's not what you want, it's what you're actively ruling out. Watermarks, extra limbs, text artifacts, and blur are the most common ones worth naming explicitly, every time.
Want to see the 11-field method applied to real generations, or a head-to-head of what each engine actually renders?
Need SEO and content templates?
Blog post frameworks, keyword research sheets, and on-page checklists built for real traffic growth. Starting at $4.
Browse SEO TemplatesMost image prompts are one run-on sentence that skips half of what a generator needs to know. This tool splits that sentence into 11 fixed fields — Subject, Style, Environment, Lighting, Composition, Color Palette, Camera/Technique, Mood, Details, Negative, and Output Quality — so nothing gets left to the model to guess, with a live word-count meter and a hedge-word check before you ever copy the result.
Pick an image class first — photograph, digital painting, 3D render, anime/manga, vector/flat, pixel art, traditional media, product shot, portrait, or architecture/interior — and two things happen: field 7 relabels to "Camera" for lens-and-aperture classes or "Technique" for brush-and-render classes, and hint chips appear under the fields where that class matters most (a photograph gets "golden hour" and "35mm lens," pixel art gets "16-px feel" and "dithering"). Click a chip to append it straight into the field.
As you type, the preview pane assembles a "Label: value" line per field, skipping anything left blank, and a word-count meter tracks the total against a 90-160 word target — amber under 90, green in range, amber past 160, red past 220. A hedge-word detector scans every field for "possibly," "maybe," and "or," and flags the field so you replace the hedge with one committed choice.
The Engine selector then reformats that same structured prompt for how each generator actually reads it. Midjourney gets --ar and --stylize flags appended, with stylize auto-bumped to 400 for painted or drawn classes. SDXL gets your Negative field split onto its own "SD negative prompt" line, matching how Stable Diffusion checkpoints separate positive and negative conditioning. Flux and DALL·E ignore flags entirely, so that mode instead compresses the 11 fields into one natural-language paragraph — no Label: prefixes, just prose.
A designer converting a rough mental image into a Midjourney prompt with the aspect ratio and stylize value already tuned for the chosen art style.
Someone reverse-engineering a reference photo — dropping it in as a local preview, then working through the 11 fields while looking at it, without ever uploading it anywhere.
A prompt engineer standardizing output across a batch of product-shot generations, where consistent lighting and camera language matter more than any single word choice.
A hobbyist testing the same structured prompt across Midjourney, SDXL, and Flux to see how differently each engine reads camera language versus natural-language prose.
Optionally drop in a reference image — it stays local for you to look at while you fill fields; nothing is uploaded.
Pick an image class (photograph, digital painting, 3D render, pixel art, and others) — the 7th field relabels to Camera or Technique and hint chips appear under the fields that matter for that class.
Fill in the 11 fields — count what is countable, name a style or movement rather than an artist, and commit to one interpretation instead of hedging with "maybe" or "or."
Watch the live prompt assemble in the preview pane and check the word-count meter — aim for the green 90-160 word band.
Pick an engine dialect (Midjourney, SDXL, or Flux/DALL·E) to get the output format that engine actually reads best, then copy it.
About the Image Prompt Builder
It is a fixed schema — Subject, Style, Environment, Lighting, Composition, Color Palette, Camera or Technique, Mood, Details, Negative, and Output Quality — that forces every prompt to cover the same ground a generator needs. A freeform sentence tends to skip lighting or composition entirely; the schema does not let you.
90 to 160 words is the sweet spot this tool targets. Under 90 words and generators fill gaps with defaults you did not choose; past 220 words later fields get diluted or ignored outright — the quality meter turns amber, then red, past those marks.
It depends on the image class. Photographs, 3D renders, product shots, portraits, and architecture use camera language — lens, focal length, aperture. Digital paintings, anime/manga, vector art, pixel art, and traditional media use technique language — brushwork, line style, rendering method. Selecting an image class relabels the field and swaps the hint chips automatically.
Generators cannot render a hedge. "A knight, maybe wearing armor, or perhaps a cloak" forces the model to guess, and it usually guesses badly. The hedge-word detector catches "possibly," "maybe," and "or" in any field so you commit to one interpretation before you generate.
Midjourney appends --ar and --stylize flags — stylize jumps to 400 for painted or drawn classes, since Midjourney reads those better at a higher value. SDXL breaks your Negative field into its own "SD negative prompt" line, since Stable Diffusion checkpoints read negative prompts separately. Flux and DALL·E ignore flags and read natural language, so that mode compresses your 11 fields into a single paragraph instead.
No. The image never leaves your browser — it is read locally with the File API purely so you have something to look at while filling in the fields. Nothing is uploaded, stored, or transmitted to any server.
Completely free, no signup. Every field you type and any reference image you drop in stays in your browser — nothing is sent to any server, and nothing is saved once you close the tab.
Fancy text in 24 Unicode styles — copy & paste anywhere
Open →SEO & ContentTurn one keyword into dozens of SEO video tags
Open →SEO & ContentReveal the exact tags any YouTube video uses
Open →SEO & ContentViews, length, age, views/day & tags for any video
Open →