Nano Banana Prompt Generator
Build an image prompt for Google’s Nano Banana (Gemini image generation) — subject, style, and detail level, written in the plain descriptive language the model responds to.
Build an image prompt for Nano Banana (Gemini image generation) in the plain descriptive language this model actually responds to -- not a keyword-stacked prompt built for a different model's conventions.
How to use it
Written as a real sentence, not a keyword list -- this model responds better to natural language.
Both explicit choices, not left to the model's default assumption.
Produces a complete, plain-language image prompt.
Tips for better results
- Write the subject as a real sentence, not a keyword list.
- Be specific about lighting and setting if they matter to the result.
- Match detail level honestly to the subject's actual complexity.
Example output
“A cat wearing tiny reading glasses, sitting on a stack of books in a sunlit window -- photorealistic style, highly detailed -- written as a real descriptive sentence, not a keyword list.”
TL;DR
Google’s Nano Banana (Gemini’s image generation capability) responds well to clear, plain-language image descriptions with explicit style and detail-level direction, rather than dense keyword-stacked prompts built for a different kind of image model. This tool builds that plain-language structure, with style and detail level as their own clear fields alongside the subject description.
Why plain description works better here than keyword stacking
Some image models respond well to comma-separated keyword lists stacked together; Gemini’s image generation was built around natural language understanding more broadly, and tends to respond more reliably to an actual descriptive sentence than a pile of disconnected keywords. That means a prompt written the way you’d describe a scene to a person — specific, but in real sentences — tends to outperform the same information compressed into keyword-list form.
Style and detail level still matter as explicit choices even within that plain-language approach — specifying photorealistic versus illustration versus 3D render shapes the result meaningfully, and stating detail level (highly detailed versus simple and clean) avoids the model defaulting to whichever it happens to lean toward for a given subject.
Example
A cat wearing tiny reading glasses, sitting on a stack of books in a sunlit window, in photorealistic style with high detail, produces a complete plain-language description with style and detail level stated clearly — not a keyword-dense prompt built for a different model’s conventions.
Refining a result that’s close but not quite right
When a generated image is close to what you wanted but not quite there, the most effective adjustment is usually to describe the specific change directly in plain language — “make the lighting warmer” or “move the subject slightly to the left” — rather than restarting from a completely different description. This model’s natural-language responsiveness that makes the initial prompt work well also makes it well-suited to this kind of conversational refinement, where each adjustment builds on the previous result instead of starting over from scratch every time.
Describing scenes with multiple elements
For a scene with several distinct elements — a subject, a specific setting, and particular lighting conditions, for instance — structuring the plain-language description so each element gets a clear, separate clause tends to work better than compressing everything into one dense, run-on sentence. “A cat wearing tiny reading glasses” as one clear idea, “sitting on a stack of books” as another, and “in a sunlit window” as a third, joined together but each individually clear, gives the model distinct pieces of information to work with rather than one tangled description it has to parse apart itself. This is a small structural habit, but it noticeably improves consistency once a scene has more than one or two elements to juggle.
Reviewing a generated result against the original description line by line — checking whether each described element actually appears — is a quick way to spot exactly which part of a prompt needs to be made more explicit on the next attempt.
FAQ
Does this work for Gemini’s other capabilities, or just image generation?
This is specifically built for Nano Banana / Gemini’s image generation — for text-based tasks, the Gemini Prompt Generator elsewhere on this site is the better fit.
Why plain sentences instead of keyword tags like other image tools use?
Gemini’s image generation was built around natural language understanding, and tends to respond more reliably to real descriptive sentences than to keyword-stacked prompts built for different models’ conventions.
Can I request text to appear in the generated image?
Describe the exact text and where it should appear directly in your subject description — results vary by how much text is requested and how complex the rest of the scene is.
Is there a difference between “Nano Banana” and “Gemini image generation” as terms?
Nano Banana is the commonly used name for Gemini’s image generation capability — they refer to the same underlying feature.