How to Write Prompts for Nano Banana
- Nano Banana responds better to natural, conversational sentences than comma-separated keyword tags.
- For photo edits, name exactly what changes and explicitly state what should stay the same.
- Multi-turn refinement works well -- generate a first pass, then give specific follow-up feedback rather than perfecting one long prompt.
- It's built to parse descriptive language the way a person would explain an image, not a tag cloud.
Nano Banana (Gemini’s image generation and editing model) responds best to natural, conversational prompts that describe the full scene in one flowing sentence, rather than a comma-separated keyword list — it’s built to parse descriptive language the way a person would explain an image, not a tag cloud.
Why conversational prompts outperform keyword lists here
Unlike some image models trained heavily on keyword-tag datasets, Nano Banana is tuned for natural language. “A golden retriever puppy sitting in a sunlit garden, soft morning light, shallow depth of field” tends to work better as one described scene than “golden retriever, puppy, garden, sunlight, bokeh” as disconnected tags. This is a meaningful departure from how Midjourney tends to be prompted, where shorter tag-like fragments often work fine.
This distinction matters practically because prompts written for other image tools, if simply copied over, often underperform on Nano Banana even when they’d work well elsewhere. Rewriting a tag-style prompt into a flowing descriptive sentence before trying it here tends to close most of that gap.
Editing existing photos needs a different structure
For photo edits specifically, name exactly what changes and what stays the same — “change the background to a beach at sunset, keeping the subject and pose exactly the same” gives Nano Banana a targeted instruction, versus a vague “make it better” which produces unpredictable results across the whole image.
This “what changes, what stays” structure matters more here than in most editing contexts, because without it the model has genuine ambiguity about scope — a request to “add warmer lighting” could reasonably be interpreted as a subtle tonal shift or a dramatic recoloring of the entire image, and naming what should remain untouched removes that ambiguity directly rather than hoping the model guesses your intended scope correctly.
Multi-turn refinement works well
Nano Banana handles conversational follow-ups — generate a first pass, then give specific feedback (“make the lighting warmer,” “move the subject slightly left”) rather than trying to perfect one prompt upfront. This mirrors the same iteration habit that works well in text-based prompting, just applied to visual output instead of written content.
Specific, targeted feedback works better than general reactions here for the same reason it does in text editing — “the lighting feels too cold” gives the model less to act on than “make the lighting warmer, closer to golden hour.” The more precisely you can name what’s off, the more precisely the next generation can address it.
How Nano Banana handles complex scenes with multiple elements
Scenes with several distinct elements that all need specific treatment — multiple subjects, several distinct style requirements, competing lighting sources — tend to produce more reliable results when described in a clear logical order (foreground to background, or primary subject before secondary elements) rather than a scrambled list of details. The model appears to weight earlier-described elements somewhat more reliably, similar in spirit to how word order matters in other image generators, even though the underlying mechanism differs.
For genuinely complex scenes, breaking the request into a first pass covering the primary subject and composition, then a follow-up refining secondary details, tends to produce more controlled results than attempting to specify everything in one exhaustive initial prompt.
Common mistakes specific to Nano Banana
Falling back into comma-separated keyword habits from other image tools is the most common mistake, especially for anyone who’s built muscle memory prompting Midjourney or Stable Diffusion first. If your results feel oddly literal or disjointed, checking whether you’ve slipped into tag-list phrasing rather than a flowing description is often the fix.
Another common mistake specific to editing tasks is not clearly separating “what to change” from “what to keep” into distinct clauses, instead blending them into a single ambiguous sentence. Structuring the request explicitly around both halves, even if it makes the prompt slightly longer, produces more predictable edits than a shorter but ambiguous version.
Working with Nano Banana’s understanding of style references
Named artistic styles (“watercolor,” “1970s film photography,” “claymation”) work well as anchors within a descriptive sentence, the same way they do in most image tools, but Nano Banana tends to blend a named style more thoroughly with the described content rather than treating it as a separate stylistic filter applied after the fact. Describing style and content together in the same natural sentence — rather than tacking a style keyword onto the end — tends to produce more integrated results than a “content, then style” structure borrowed from other tools.
This also means style references benefit from the same specificity that helps everywhere else — “soft 1970s film photography with warm color grading and visible grain” gives more to work with than “vintage style,” even though both are technically style references.
Combining text editing and image generation in the same workflow
Since Nano Banana operates within Gemini, it’s often used alongside Gemini’s text capabilities in the same conversation — describing a concept in words first, refining that description through conversation, then generating the image once the concept is clear. This workflow tends to produce better results than jumping straight to image generation with an underdeveloped concept, since the conversational refinement happens in a medium (text) that’s faster to iterate on than regenerating images repeatedly.
For anyone working on a concept with real creative ambiguity — not just technical photo edits but something more exploratory — talking through the concept in text first, then translating the settled description into an image prompt, is a genuinely useful two-step process worth deliberately using rather than skipping straight to generation.
Handling text and typography within generated images
Like most AI image models, Nano Banana’s handling of legible text within an image is improving but still inconsistent, particularly for longer text strings or specific fonts. For short, simple text elements — a single word on a sign, a short label — results tend to be more reliable than for anything resembling body text or precise typography. If exact, reliable text rendering is essential to your use case, planning to add it separately in an editing tool after generation is often more dependable than expecting the model to render it correctly within the generation itself.
Resolution and output quality considerations
Output quality and resolution depend partly on how the request is framed — prompts that specify a clear, singular focal point tend to render with more fine detail in that area than scenes asking for many equally-weighted elements across the frame, since detail generation isn’t uniformly distributed regardless of what’s requested. If a specific element needs to be sharp and detailed — a product shot, a face, precise texture — naming it as the clear focal point of the composition, rather than one of several competing elements, tends to produce better results in that specific area.
This is worth knowing especially for practical use cases like product photography or portrait work, where fine detail in one specific area matters more than overall scene complexity.
Batch generating variations for a project
When you need several related images for a single project — a set of product angles, variations on a concept for a client to choose from — keeping the core descriptive sentence stable and varying only the specific element that should differ between generations produces a more cohesive set than rewriting each prompt from scratch. This mirrors the consistency habit that matters for sequences of AI-generated video clips, just applied to a set of stills instead.
Keeping a written reference of your exact base description, then generating variations by changing only the specific detail that should vary, is a simple practice that noticeably improves visual consistency across a batch compared to describing the same intended concept slightly differently each time from memory.
Try it yourself
[subject, described specifically], [setting], [style or mood], [specific lighting detail].
In this photo, change [specific element], keeping [what should stay the same] exactly the same.
FAQ
Should I use comma-separated tags or full sentences for Nano Banana?
Full descriptive sentences generally work better — it’s tuned for natural language, not keyword-tag lists.
How do I edit a specific part of a photo without changing the rest?
Name exactly what to change and explicitly state what should stay the same — both parts matter for a targeted edit.
Can I refine a result through follow-up messages?
Yes — specific conversational feedback on one element at a time tends to work better than trying to write a perfect prompt upfront.
How should I handle scenes with multiple distinct elements?
Describe them in a clear logical order, foreground to background, and consider breaking complex scenes into a first pass plus a refinement follow-up.
For more editing-specific technique, see the AI image prompt generator or browse the Gemini prompt library.