AI EngineeringSeptember 10, 202515 min read
    SC
    Sarah Chen

    How to Write Prompts for Midjourney - A Step-by-Step Guide

    How to Write Prompts for Midjourney - A Step-by-Step Guide

    How to Write Prompts for Midjourney: A Step-by-Step Guide

    Start with a clear subject and a single focal style. In Midjourney prompts, you must lock onto one idea at the start of your message, then expand with precise details. For effective prompts, capture the core concept in a short noun phrase and place it at the beginning. Inside your first clause, set the mood and define a visual anchor so the output stays focused and predictable despite later refinements.

    Outline features and tone clearly. List the elements you want: lighting, texture, color palette, and composition layout. Use the features to keep you aligned with the final look. Put the key nouns and verbs after the subject, so you can swap details later without reworking the structure. The subject plus these keywords form a tight blueprint that minimizes drift.

    use studio setups and references to guide style. Attach a few references as contextual hints, and include an example line to illustrate the expected result. Inside the prompt, mention a window, a table, and a lamp to crystallize lighting. This approach yields a cohesive look across shots and supports quick iteration.

    Structure prompts to explore multiple compositions. Start with a base prompt, then split into variants by swapping adjectives and angles. Include a contrast element, such as high-contrast light against a dark background, to add depth without overcomplicating the concept. Inside each variant, adjust perspective and the relationship between the elements to emphasize the intended narrative.

    Prompt skeleton you can reuse. Subject + Context + Style + Details + Reference. For example: "Subject: a neon-lit table at a rainy window, table in focus; Context: a small studio corner, early morning haze; Style: cyberpunk, high texture; Details: reflections, wet surfaces, grain; Refs: URL1, URL2." Use this pattern to speed up iteration and keep prompts compact inside a single line or two lines.

    Practical tips for stable results. Use --ar 16:9 for cinematic width, --q 2 for detail, and --no blur to keep edges crisp. Limit prompt length to roughly 120-180 words for quick adjustments. Always verify output against a small set of references and adjust lighting and color balance by swapping primary colors for faster convergence. Through practice you'll learn how to steer outputs toward your vision without lengthy testing.

    Choose the Right Version, Model, and Aspect Ratio for Your Concept

    Use Version 5, the default model, and an aspect ratio of 16:9 for cinematic concepts. This baseline keeps lighting, color, and composition cohesive, making iterative tweaks straightforward and saving time for a fast cycle of drafts this hour.

    For a moodier, high-contrast look with controlled window light, select Version 4 to emphasize stronger textures and deeper shadows, while Version 5 preserves natural lighting and crisp details. If you aim for details with tactile texture, this baseline favors color fidelity and helps you communicate the desired outcome for the subject of focus. When you compare results, note how the dark atmosphere shifts with each version, and choose the one that aligns with the emotional intent, especially around parts of the scene that carry the most feeling.

    Choose the default model for most prompts to maintain consistency across scenes; switch to a more artistic or anime-like model only when the concept clearly benefits from a distinct look. This approach uses the same composition, then adjust texture through prompts instead of swapping models. In a cinematic frame with a subtle smile on a character, the model selection helps preserve facial proportions and the mood, making the feelings of the shot read clearly to the viewer. When the subject carries a strong visual conflict, keep the base model to avoid unintended distortions and maintain lighting fidelity.

    Match your aspect ratio to the concept: 1:1 for portraits or logo-like compositions; 4:3 for prints or framed artwork; 16:9 for broad, cinematic landscapes; 9:16 for vertical mobile previews; and 21:9 for expansive scenes. For depth and color control, 16:9 or 21:9 helps preserve window glow and color consistency across slide decks. When the idea requires a tighter composition that keeps the object in focus, you can force 1:1 with this constraint, then run variations to compare how the subject sits within the frame when you tweak the concept. This approach keeps the presentation cohesive and maintains a clear narrative flow from slide to slide, especially during quick iterations that test samples of mood and lighting.

    Concrete prompt cue (examples): Version 5, default model, --ar 16:9, dark window lighting, color-rich scene, subject in foreground, smile subtle, without extra details, adding clarifications about mood and contrast, to achieve the desired effect. You can layer details gradually, using words that describe feel and composition, and adjusting ratio to match the output you want. When you need to zoom into texture, details and physical texture emerge from the choice of version, model, and aspect ratio, ensuring the concept translates from ideas to an image that resonates with viewers this way. This method keeps prompts grounded in the base rules and helps you reach the outcome you envision, avoiding drift as you iterate the same concept this hour.

    Layer Styles Strategically: When and How to Blend Multiple Art Styles

    Begin with a single base style and layer a second, complementary approach through careful blending to keep face as the focal point and colors cohesive.

    1. Core setup: Define the foundation of the look and set the field where the face appears. This determines how the secondary style will interact without overshadowing the main message.

    2. Style pairing: Pick a second style that adds texture or mood without stealing attention. If the base is painterly, a subtle film grain can add depth while preserving colors. Avoid grandmother-nostalgia or clichés; aim for a fresh feel.

    3. Blend controls: Use opacity and blend modes to balance the two styles. Start around 20-30% opacity for the secondary layer, and mask edges to create smooth transitions, avoiding too busy edges that distract from the subject.

    4. Prompting strategy: In your prompts, lock the core look and describe the additional layer with precise parameters. Imagine the secondary style as a glaze that conveys the mood without altering the main subject. Include --no to exclude unwanted artifacts.

    5. Color and texture alignment: Ensure colors and tonal harmony across both layers. A film-like look can unify colors, while keeping secondary details restrained so the scene reads clearly at a glance.

    6. Validation and optimization: Use before/after checks to confirm the subject remains recognizable and the result aligns with goals. If it conveys the wrong vibe, tighten the weight on the second style or refine the mask; keep the process concise and practical, avoiding lengthy workflow.

    7. Experiment and iteration: Create three variants with different secondary styles to compare, then pick which variant achieves the desired impact. Save the best blend as a prompt template for future projects; such advice will accelerate reuse and enable expansion of capabilities

    Control Output with Prompt Weights: Balancing Elements in Multi-Prompt Prompts

    To achieve precise control, assign weights to each prompt segment. Set the main subject weight at 0.65–0.75, color and lighting at 0.15–0.25, and texture or background details at 0.05–0.15. In dynamic versions, shift 0.05–0.10 between groups to explore different emphasis. Use syntax like (subject:0.75) (color:0.25) (texture:0.15) to reinforce priorities in multi-prompt prompts.

    Weights: how to assign and tune

    Divide the task into basic elements: subject, color, texture, lighting, environment, and mood. Prompts can be considered as a set of layers, where each element receives weight. Keep the total near 1.0 and adjust per desired result: for bright light scenes, boost lighting and color by 0.1–0.2; for mysterious atmosphere, push mist and texture to 0.2–0.3. In forest or vintage styles, give forest and vintage keywords combined weight 0.4–0.6 while easing others.

    If you need to get feedback, ask colleagues to evaluate outputs and version, then refine weights in subsequent versions to tighten confidence in desired composition and viewer viewing experience. The approach remains dynamic, supporting experimentation across versions and prompts.

    Practical prompts and examples

    Practical prompts and examples

    Example 1: (forest:0.9) (vintage:0.4) (color:0.8) (texture:0.3) (mist:0.5) (photography:0.8) – aims for a bright, nostalgic vibe with clear composition and subtle texture.

    Example 2: (subject:0.7) (background:0.2) (color:0.6) (texture:0.2) (light:0.5) (gloom:0.25) – emphasizes dramatic lighting and crisp viewer direction, while keeping the overall balance stable. If the request produces overly bright elements, reduce conflicting weights by 0.05 and test in next version.

    Tune Modifiers: Stylize, Quality, Seed, and Negative Prompts for Precision

    To tune modifiers for precision, focus on Stylize, Quality, Seed, and Negative prompts. If you want to influence characteristics clearly, observe what happens when you adjust each parameter. A backlight glow from behind the subject plus front lighting sets the mood; glowing accents and depth cues guide the format and depth. The concept remains clear when you craft concise prompts, and use direct language to steer the render with confidence (time).

    Stylize and Quality: shaping look with intention

    Stylize governs how strongly the model leans into its internal aesthetics. A low value (around 50–150) keeps the wording tightly bound to your words, while a high value (300–1000) pushes toward more creative textures and qualities. If you want a precise, technical result, keep stylize modest and explicitly call out angle, proportions, and lighting described by backlight and front illumination. When you raise stylize, expect more illustrative approach (or more artistic 'look'), so align the intent with the chosen format. For scenes with mock or glowing edges, test several stylize values and compare depth and detail in the final images.

    Quality (the --quality parameter) controls fidelity versus speed. At 1.0 you get a balanced render; bumping to 2.0 yields more detail, richer textures, and a stronger sense of depth, but it lengthens render time and leads to a greater number of images. Use higher quality when you need subtle shading, accurate backlight interaction, and a clean boké. If time is limited, keep quality at 1.0 and compensate with focused prompts about light, mist, and angle. Always balance these parameters with a clear intent for depth, format, and tone.

    Seed and Negative Prompts for Reproducibility and Precision

    Seed fixes randomness so you can reproduce a result with the same prompt. Set a seed (–for example, --seed 12345) to lock in texture, depth, and the overall feel. To encourage variation, vary the seed while keeping key characteristics within the intent; to maintain consistency use the same seed on related frames. Negative prompts help exclude unwanted elements. Add explicit exclusions like "no watermark," "no blur," "no text," or "no fog" to refine the composition and prevent stray elements from creeping into the mist or boké. When you want to remove undesired traits while preserving depth and backlight balance, pair --no with phrases that cancel those traits and preserve front lighting and view.

    Example prompt: "a front-facing portrait of a dancer, backlight glow, mist swirling around, shallow depth of field, boké, different proportions of limbs, with glowing edges, strict format" with --stylize 400 --quality 2 --seed 9876 --no watermark, text, blur. This combination reinforces intent, aligns with desired angle and backlight, and yields more controlled images.

    Master Multi-Prompt Syntax: Structure, Ordering, and Real-World Examples

    Start with a focused base prompt that sets the core scene and vibe, then layer modular prompts in a fixed order. This keeps prompts concise and predictable, letting you control how prompts build toward the final render.

    Structure rests on three layers: the base scene, the modifiers, and the technicals. Use a toolbox of commands to manage tasks across layers. For each layer, aim for around 3–5 tokens and separate prompts with commas. The base defines the frame, the modifiers inject details, and the technical prompts tune style, lighting, color, and aspect. If you're crafting a forest mood, include forest and retro vibes, add girls with smile; this points to key details, while keeping the path simple. Next, keep layers crisp and predictable, and beyond details keep the balance between atmosphere and clarity.

    Ordering matters: start with subject and environment, then add wardrobe or props, mood, lighting, camera cues, and finally details and constraints. Put the most important elements upfront so the model prioritizes them when rendering. Use parentheses or brackets sparingly to emphasize crucial terms, and avoid repeating the same word too many times in a single layer. Example approach: base → subject → environment → wardrobe → mood → lighting → color → details → camera → constraints. This provides a clear flow and makes it easier to swap one element without breaking the rest. Next, test step by step: replace one part at a time to see the impact and make more precise decisions about settings.

    Real-world examples demonstrate how to apply the flow. Example A: base: "forest, retro mood"; modifiers: "girls, smile, vintage dress"; environment: "soft dawn light"; details: "crisp textures, leaf surfaces, mossy textures"; technicals: "--ar 16:9 --v 5"; outcome: clean balance between mood and clarity. This chain of signals indicates priorities and reduces noise. Example B: base: "urban retro poster"; modifiers: "models, rain-soaked street, neon reflections"; mood: "moody"; details: "high contrast, sharp edges"; technicals: "--ar 9:16 --v 5"; result: bold composition with crisp surface texture and strong visual punch. Example C: base: "portrait in forest light"; modifiers: "retro-style clothing, girls, smile"; details: "skin tones natural, eyes bright, surface texture visible"; technicals: "--q 2 --ar 3:4"; note: this approach lends a human focal point with controlled surroundings.

    Tips for consistency: keep a reusable base that covers 2–3 core elements, then append varying modifiers and technicals for new scenes. You can copy-paste current prompts into a new file and modify one layer at a time, which allows you to convey desired mood shifts without rewriting the entire structure. Follow a simple pattern consistently: base → modifiers → technicals; this simplifies model engine management and improves result repeatability. Use details from previous projects for a quick start, and remember that very precise control is achieved by consistently changing one element at a time.

    How Length and Specificity Influence Consistency Across Iterations

    Use a short, precise prompt to maximize consistency across iterations. This method anchors outputs by constraining length and demanding concrete details, preventing drift in images and across English-speaking audiences. Focus on a single subject and fixed mood, and include only the details that truly shape the look.

    Structure prompts into three parts: context, subject, constraints. In context, set the scene; in subject, define the main element; in constraints, lock tone and style, and specify details. Use a couple of repeatable patterns to guide the look for photography across English-speaking audiences, even for a couple of iterations.

    Set aspect with --width to fix size. This gives clarity to composition and reduces variation across iterations. Keep the width constant across runs to preserve consistency in the final image.

    Limit the number of variables. Use only a few key details like lighting in specified lighting areas, camera angle, and color palette. This approach keeps vision cohesive and results detailed, not vague.

    To modernize your workflow while preserving consistency, reuse a core phrase that defines style and patterns; update only one variable per iteration. This reduces variability while maintaining detailed results that feel coherent across projects. Start with a pair of prompts that look similar and differ by a single parameter, so you can compare impact without breaking consistency.

    Test across domains like architecture, fashion, and interiors, keeping the same method to preserve vision. Align with English-speaking audience expectations by repeating core terms and ensuring that tone and style stay constant. Reuse patterns across images to build a recognizable visual vocabulary for English-speaking viewers.

    Checklist for quick validation: keep length concise with only 3–5 details; use only essential information; fix --width and size; specify areas of lighting and camera angle; define feeling and vision, so outputs stay detailed in photography style. Run a couple iterations and compare; if drift appears, tighten constraints and reuse the same core expression to modernize prompts without breaking consistency.

    Navigating Common Pitfalls and Quick Fixes When Mixing Styles and Modifiers

    Start with a clear style anchor and use a couple of modifiers; keep the base vibe precise, and now test two quick variations to confirm the effect. In the area of creating prompts, there exist a handful of traps: muddled subjects, conflicting cues, and overly long lists. These pitfalls allow drift toward noise, so set a strict focus and keep attention on how each modifier shifts the overall feel. Write goals in words, not vague phrases, and describe in your own words to lock in precise expectations for colors, lighting, and composition. Ideal balance comes from small, deliberate changes and from restricting additional information to what truly matters.

    When you mix styles, remember the perspective side of the prompt: what works on one iteration may fail on another unless you describe the intent clearly. There exists a tendency to push too many ideas across one image; this long list often dilutes intent. Now is the moment to simplify: use a couple of core elements that amplify the effect and layer texture through the base style rather than by stacking add-ons. This approach gives you a reliable path to ideal outcomes without overwhelming the renderer.

    PitfallQuick Fix
    Unclear base style leads to muddy resultLock a single style anchor and limit modifiers to a couple; run two small variations now to compare their effect and choose the clearly better direction.
    Conflicting color cues (colors) disrupt moodChoose a harmonious palette, specify color targets in words (words), and keep colors within a four-hue range to preserve cohesion.
    Modifiers clash in meaningUse precise descriptors instead of vague terms; pick 2–3 modifiers that align with the anchor and avoid mixing incompatible vibes.
    Wrong order weakens emphasisPlace the most important modifier on the right side; test how shifting order changes subject effects and adjust toward clarity.
    Too many modifiers overload the promptLimit to ideal 2–4 modifiers; rely on the base style for texture, and reserve additional details for iteration rounds.
    Prompts drift over time with constant togglingDocument changes (changes) in versions (versions) and keep all additional tweaks concise; this helps tracking and scaling.

    📚 More on AI Generation & Prompts

    Ready to leverage AI for your business?

    Book a free strategy call — no strings attached.

    Get a Free Consultation