Image models are search engines over a latent space of pictures. "Masterpiece/cinematic/8k" returns the AVERAGE of everything anyone called that — the mathematical definition of generic. A specific noun returns a specific picture. You are not wielding a paintbrush with adjectives; you are directing a camera with facts.
THE 5-SLOT SPINE (every prompt, any order): - SCENE: where, time of day, environment, background. - SUBJECT: who or what the image is actually about. - IMPORTANT DETAILS: materials, clothing, lighting, camera feel (50mm, shallow DOF), composition, mood — as facts. - USE CASE: editorial photo / product mockup / poster / UI / infographic / direct-response static ad. - CONSTRAINTS: no watermark, no extra text, preserve face, preserve layout, safe margins.
FIVE ANTI-SLOP RULES: 1) Visual facts beat vague praise. Kill stunning/cinematic/8k/masterpiece/award-winning/premium/high-end. The model can't render those. It CAN render "overcast daylight, brushed aluminum, chipped paint, soft bounce light off a wall." 2) Name objects, not moods. "Transit kiosk, boarding pass, wet pavement, neon reflection" beats "atmospheric urban scene." Moods are outputs; objects are inputs. 3) Treat text as typography. Quote the literal string, spell hard words letter-by-letter, specify font feel + size + placement: 'The headline reads m-a-c-h-i-n-a in tall condensed sans-serif, centered, white on black.' "Bold modern branding" is a prayer no one answers. 4) In edits, separate CHANGE from PRESERVE in two explicit columns, and repeat the PRESERVE list EVERY turn. Drift comes from assuming the model remembers the last turn — it doesn't. 5) One revision per turn. Small iterative edits compound; big rewrites reset identity and send you back to start. Prompt less per turn, run more turns.
PLAIN ENGLISH vs JSON: for a one-shot in a chat box, plain English wins — JSON is friction. For an AGENT PIPELINE that mutates one field at a time across many variations (which is what AutoAdy's creative machine IS), JSON-with-handles is correct: keep handles like scene=, subject=, details.lens=, wardrobe=, color_palette{primary,secondary,highlight,shadow}, then serialize them into the 5-slot spine for the model. Structure the information; don't perform expertise with syntax.
MODEL CHOICE: GPT-Images-2 owns legible text (incl. non-Latin scripts), identity-locked edits, 8-image batch consistency, and composition — run it at quality:high, ~2K (2560x1440), thinking mode, dimensions in multiples of 16, ultra aspects 3:1/1:3 for heroes/stories. Nano Banana wins ONLY sterile catalog/product-on-white shots — use it there. For ad headlines/CTAs that MUST be pixel-perfect, prefer rendering a text-free scene and compositing real vector text over it (AutoAdy's crisp-text renderer) instead of trusting baked text.
DON'T over-write the brief. Structure the user's intent into the spine; don't invent a different image. When asked for a creative, produce the 5-slot prompt first, then generate.