How to Write a Good Prompt for Midjourney

Last updated August 29, 2026 by GenPrompto

Key takeaways
  • Word order matters in Midjourney -- earlier words are weighted more heavily, so state the subject first, then style, lighting, and composition.
  • Parameters like --ar, --style raw, and --s are core controls, not optional extras -- --style raw specifically pulls away from the default painterly look.
  • Vague style words ("beautiful," "stunning") carry little usable information; specific terms like lighting type or color palette do real work.
  • A prompt that only looks good when you already like the subject probably isn't doing much with its style words.

A good Midjourney prompt names the subject first, then style, then lighting, then composition — in that order, since Midjourney weights earlier words more heavily and vague later details matter less than a clear subject stated up front.

Word order matters more here than in most AI tools

Midjourney’s model gives more weight to words that appear earlier in the prompt. This means “a red fox in a snowy forest, watercolor style” and “watercolor style, a red fox in a snowy forest” can produce meaningfully different results, even with identical words — put the actual subject first, then layer on style, lighting, and composition after. This is a genuine departure from how most text-based AI prompting works, where task-context-format order matters less strictly than it does here.

This weighting effect compounds with prompt length — the longer the prompt, the more the earliest words dominate relative to details buried near the end. For prompts over 40-50 words, consider whether everything after the subject and primary style is actually earning its position, or if it’s diluting the elements that matter most.

Parameters are not optional extras — they’re core control

Flags like –ar (aspect ratio), –style raw (less default stylization), and –s (stylize strength) genuinely change output, not just formatting. –style raw specifically pulls results away from Midjourney’s default painterly tendency, which matters a lot if you want something closer to photorealistic.

–ar controls aspect ratio and should match your actual intended use — a square image behaves differently in composition than a widescreen one, and choosing the aspect ratio after generating rather than before means starting over rather than adjusting. –s (stylize) ranges roughly from 0 to 1000 and controls how much Midjourney’s aesthetic preferences override literal interpretation of your prompt; lower values produce more literal, less “Midjourney-looking” results, which is worth knowing if your output consistently looks more stylized than you intended.

Why vague style words underperform specific ones

“Beautiful” or “stunning” carry almost no information the model can act on — they’re descriptions of a reaction, not a visual instruction. “Soft rim lighting” or “muted earth-tone palette” are specific enough to actually shape the output. If a prompt only works when you already like everything about the subject, the style words probably weren’t doing real work.

This distinction matters more than it first appears, because vague adjectives don’t just fail to help — they can actively crowd out more useful words given Midjourney’s word-order weighting. A prompt front-loaded with three or four vague adjectives before reaching the actual subject wastes the highest-weighted positions in the prompt on words carrying minimal instruction.

How Midjourney interprets style references differently than photo references

Naming an art style (“oil painting,” “risograph print,” “1970s film photography”) gives Midjourney a genre to work within, drawing on the visual conventions associated with that style across its training data. This differs from describing a specific photographic technique (“shot on 35mm, shallow depth of field, natural window light”), which asks for a photographic result with specific technical characteristics rather than an artistic interpretation.

Mixing these two modes in one prompt — naming both an illustrated style and photographic technical detail — tends to produce inconsistent results, since the model is being asked to reconcile two different visual logics at once. Pick one lane: illustrated/artistic style references, or photographic/technical references, rather than blending them in the same prompt.

Working with Midjourney’s version differences

Midjourney updates its underlying model periodically, and prompts that worked well on one version don’t always transfer perfectly to the next — a version that handled photorealism a certain way might interpret the same prompt differently after an update. If a prompt that used to work reliably starts producing unexpected results, checking which model version is currently active is worth doing before assuming the prompt itself needs a rewrite.

This version sensitivity is one of the reasons re-testing periodically matters for anyone publishing Midjourney prompts for others to use — a prompt genuinely tested and confirmed working in one version may need adjustment after a model update, even if nothing about the prompt itself changed.

Common mistakes beyond word order and vague style words

Requesting text within an image is a common frustration point — Midjourney has historically struggled with rendering legible text, though this has improved across versions. If text is essential to your image, expect to need multiple attempts or plan to add text separately in an editing tool rather than relying on the model to render it correctly the first time.

Another common mistake is over-specifying camera settings that don’t meaningfully change output in the way a photographer might expect — Midjourney interprets terms like aperture and focal length as stylistic cues toward certain visual conventions (shallow depth of field, compression) rather than simulating actual optical physics, so extremely precise technical camera language often adds little beyond what a simpler description would achieve.

Using negative prompting effectively

Midjourney’s –no parameter lets you exclude specific elements the model tends to add unprompted — –no text, –no watermark, –no blur are common examples. This works differently from simply not mentioning something; Midjourney will sometimes add elements based on strong associations within a style (fog with moody lighting, for instance), and –no is the direct way to suppress a specific unwanted addition rather than hoping the base prompt avoids it.

Negative prompting works best when targeted at something you’ve actually observed the model adding, rather than a defensive list of everything you can imagine going wrong. A long –no list dilutes focus the same way a long list of positive descriptors does, and can occasionally produce odd results if the excluded terms interact unexpectedly with the rest of the prompt.

Iterating on a Midjourney result

Midjourney’s variation and remix features let you refine a result without starting from scratch — generating variations on a specific output, or using remix mode to adjust the prompt while keeping the overall composition anchored to what you already have. This is genuinely different from text-based iteration, where you’re refining through follow-up messages in a conversation; here you’re iterating on a specific visual result directly.

A practical habit: when a result is close but not quite right, resist the urge to completely rewrite the prompt from scratch. Identify the specific element that’s off — lighting, composition, a style detail — and adjust just that piece first, since a full rewrite risks losing whatever was already working in favor of an unknown new combination.

Prompting for consistent characters or subjects across multiple images

Generating the same character or subject consistently across several images is a common need — for a story sequence, a brand mascot, a set of product shots — and Midjourney’s character reference and style reference features exist specifically for this, letting you anchor new generations to a previous image’s subject or style rather than relying purely on text description to recreate it identically each time.

Text description alone struggles with this kind of consistency because even a very detailed subject description leaves enough interpretive room that two separate generations rarely match closely. If consistency across multiple images matters for your project, building the workflow around reference features from the start saves considerable trial and error compared to trying to achieve it through prompt text alone.

For character or subject work specifically, keeping the underlying descriptive text stable between generations — changing only the specific scene or action — while relying on the reference image for visual consistency tends to produce more reliable results than varying both the text and hoping the reference alone carries enough weight.

How aspect ratio choice affects composition, not just canvas shape

Choosing –ar isn’t purely a matter of final image dimensions — it genuinely shapes how Midjourney composes the scene. A square aspect ratio tends to center a subject more directly; a wide aspect ratio invites more environmental context and a sense of scene, while a tall aspect ratio often produces more portrait-style framing with the subject filling more of the vertical space. Choosing the aspect ratio to match your intended composition, rather than only your intended final use, tends to produce results that need less cropping or rework after generation. Deciding this upfront, before writing the rest of the prompt, also helps you judge how much environmental or background detail is worth including in the description itself.

Try it yourself

Subject-Style-Light-Composition Version
[subject, described specifically], [style reference, e.g. “watercolor illustration” or “35mm film photograph”], [lighting description], [composition detail, e.g. “close-up” or “wide shot”] –ar [aspect ratio]
Photorealistic Version
[subject], shot on [camera reference, e.g. “85mm lens, f/1.8”], natural lighting, realistic skin texture and fine detail, shallow depth of field, photographed not illustrated –style raw –ar 3:2

FAQ

Does word order actually matter in Midjourney prompts?

Yes — earlier words are weighted more heavily. Put your subject first, then style, lighting, and composition after.

What does –style raw actually do?

It reduces Midjourney’s default painterly stylization, useful when you want a result closer to photorealistic rather than illustrated.

Why do vague words like “beautiful” or “epic” underperform?

They describe a reaction, not a visual instruction. Specific terms like lighting type or color palette give the model something concrete to act on.

Can I mix illustrated style references with photographic technical detail?

Generally not well — pick one visual logic per prompt, since blending the two tends to produce inconsistent results.

For more platform-specific technique, see a full styled Midjourney example or browse the Midjourney prompt library. For the same subject-style-lighting structure applied to other image tools, see how to write prompts for AI image generation or try the AI image prompt generator.

Written by GenPrompto Editorial Team

Every prompt on this site is tested against real model output before publishing. Guides follow a documented content standard for accuracy and depth. When something is wrong, it gets fixed -- not left for a reader to find first.

More about how this site works →

Related