Prompting10 min read
Stop "Perfect Prompting" Veo 3.1: The 3-Layer Prompt System That Gets Cleaner Ads & Reels
A practical Veo 3.1 prompt structure for cleaner ads and reels: SHOT + ACTION + STYLE, with a worked before/after example, checklist, and FAQs.
On this page
- TL;DR
- Key takeaways
- Why most Veo prompts fail: “everything prompts” aren’t directable
- The 3-layer prompt system (copy/paste)
- The template (copy/paste)
- Layer 1: SHOT (what must be true in-frame)
- Framing + camera motion: pick one primary camera behavior
- “2–3 descriptors” rule for characters/products
- Location specificity reduces randomness
- Layer 2: ACTION (beats, timing, and continuity)
- Use 2–3 beats max
- Continuity anchors: repeat exact nouns
- Dialogue/audio: keep it simple and deliberate
- Layer 3: STYLE (guardrails, not story)
- Use one primary style
- Lighting is your legibility knob
- Add an “Avoid” list to prevent common ad failures
- Worked example: one messy prompt → one clean prompt (plus variations)
- Before: the messy “everything prompt”
- After: clean 3-layer prompt (single shot product mini-ad)
- Quick variations table (change only one layer)
- When to use First/Last Frame (continuity and transitions)
- Use first/last frame when
- What to add to your prompt
- Iteration order (so you don’t thrash)
- Checklist
- FAQ
- How do I write a Veo 3.1 prompt structure that doesn’t drift?
- How do I describe camera framing and movement for Veo?
- How do I get better lighting and product readability?
- How do I prompt dialogue and sound in Veo videos?
- How do I keep transitions consistent between two beats?
- Why is my prompt blocked or not generating what I asked?
- Build faster: turn the 3-layer system into a repeatable Veo3Gen workflow
- Start creating with Veo3Gen
TL;DR
Stop writing “everything prompts.” For Veo 3.1, you’ll get cleaner, more directable ads and reels by splitting prompts into three layers:
- SHOT = what must be true in-frame (subject, location, framing, camera motion, key objects)
- ACTION = what changes over time (2–3 timed beats + continuity anchors + optional dialogue/audio)
- STYLE = how it looks (visual style + lighting + constraints), without rewriting the scene
This matches how Veo’s own prompt guidance emphasizes shot framing/camera motion, style, lighting, character detail, location detail, and action (https://deepmind.google/models/veo/prompt-guide/) and aligns with “say what’s in frame and how it moves” advice from other video tools (https://help.runwayml.com/hc/en-us/articles/42460036199443-Text-to-Video-Prompting-Guide).
Key takeaways
- Direct like a director: lock SHOT (framing + camera motion), then choreograph ACTION, then apply STYLE guardrails (https://deepmind.google/models/veo/prompt-guide/).
- Less simultaneity = cleaner marketing footage: limit to one scene and 2–3 action beats.
- Continuity anchors prevent drift: repeat 2–3 exact nouns across iterations (product name/object, setting, wardrobe).
- Use first/last frame when continuity is critical: Veo 3.1 is referenced as having first frame/last frame capability (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1).
- Iterate surgically: change one layer at a time so you know what actually improved.
Why most Veo prompts fail: “everything prompts” aren’t directable
Most weak outputs come from trying to solve three problems in one paragraph:
- Composition: what we see + where the camera is + how it moves.
- Causality: what happens over time (clear beats, believable motion).
- Aesthetic: style + lighting + texture.
Veo’s prompt guide already names the levers you should control—framing/motion, style, lighting, character details, location details, action (https://deepmind.google/models/veo/prompt-guide/). The mistake is listing them as a single “mood board paragraph” instead of separating them into layers you can independently fix.
A good mental model: prompts should describe what appears in the frame and how it moves through the scene in direct language (https://help.runwayml.com/hc/en-us/articles/42460036199443-Text-to-Video-Prompting-Guide). That’s exactly what SHOT + ACTION forces you to do.
The 3-layer prompt system (copy/paste)
Use this as your default Veo 3.1 prompt structure.
The template (copy/paste)
SHOT (non-negotiables):
- Hero subject: [product/character + 2–3 specific descriptors]
- Location: [specific place + time-of-day + 1–2 concrete details]
- Framing: [wide / medium / close-up + angle]
- Camera motion: [static / slow push-in / pan / tilt / handheld]
- Key objects: [product + 1–2 props that must appear]
ACTION (beats + physics):
- Beat 1 (0–2s): [single clear action]
- Beat 2 (2–4s): [single clear action]
- Beat 3 (4–6s): [single clear action]
- Continuity anchors (repeat verbatim): [2–3 exact nouns]
- Audio (optional): [dialogue line OR topic + SFX/music]
STYLE (guardrails only):
- Visual style: [one style reference]
- Lighting: [one lighting instruction]
- Pace: [snappy / calm / cinematic]
- Avoid: [no text overlays, no extra characters, no scene cuts, etc.]
Why these fields map to Veo 3.1 guidance
- Veo explicitly calls out shot framing and camera motion (e.g., low-angle views, panning) (https://deepmind.google/models/veo/prompt-guide/).
- Veo explicitly calls out style (cartoon, claymation, film noir on 35mm, worn-out VHS texture) (https://deepmind.google/models/veo/prompt-guide/).
- Veo explicitly includes lighting (warm even lighting; spotlight in one area) (https://deepmind.google/models/veo/prompt-guide/).
- Veo recommends detailed character descriptions (https://deepmind.google/models/veo/prompt-guide/).
- Veo recommends thorough location description (https://deepmind.google/models/veo/prompt-guide/).
- Veo includes action as a prompt element (https://deepmind.google/models/veo/prompt-guide/).
- Veo can generate dialogue; you can provide a topic or exact lines (https://deepmind.google/models/veo/prompt-guide/).
Layer 1: SHOT (what must be true in-frame)
If SHOT is wrong, the clip is wrong.
Framing + camera motion: pick one primary camera behavior
Veo’s guide explicitly recommends specifying shot framing and camera motion (https://deepmind.google/models/veo/prompt-guide/). Do it like you’re writing a call sheet.
Do:
- “Close-up, slight low-angle, slow push-in”
Don’t:
- “Cool cinematic camera moves”
Rule: one main camera motion per clip. If you need multiple moves, you probably need multiple shots.
“2–3 descriptors” rule for characters/products
Veo’s guide shows why specificity wins (their example: “A woman in her twenties with wavy brown hair and light freckles” vs. “a brown-haired woman”) (https://deepmind.google/models/veo/prompt-guide/).
Apply that to marketing:
- “a bottle” → “a matte black bottle with a minimal white label”
- “a kitchen” → “a modern kitchen, white tile backsplash, morning light”
Location specificity reduces randomness
Veo contrasts generic locations with specific ones like “a smoky jazz club at night” or “a cyberpunk city with bright chrome and neon lights” (https://deepmind.google/models/veo/prompt-guide/).
For ads/reels, a specific location does two practical things:
- It constrains the background so the product reads.
- It reduces “anywhere energy” (the model inventing random sets).
Layer 2: ACTION (beats, timing, and continuity)
Veo supports action prompts (examples include dashing, backflips, chasing) (https://deepmind.google/models/veo/prompt-guide/). The key is making action unambiguous and sequenced.
Use 2–3 beats max
Write actions as beats with rough time windows. You’re not “locking duration”; you’re forcing priority.
Bad (too simultaneous):
- “She grabs it, smiles, runs outside, opens the app, dances, drinks, logo appears.”
Good (three beats):
- Beat 1: place product
- Beat 2: open product
- Beat 3: pour / reaction
Continuity anchors: repeat exact nouns
Continuity anchors are literal phrases you repeat across revisions:
- “matte black bottle”
- “oak table”
- “white tile backsplash”
When you change STYLE later, anchors keep the subject from morphing.
Dialogue/audio: keep it simple and deliberate
Veo can generate dialogue; you can provide a topic or exact lines (https://deepmind.google/models/veo/prompt-guide/).
For marketing, use one clean line:
- “Dialogue (whisper): ‘One sip. Reset.’”
If your prompt includes audio, don’t bury it. Put it in the ACTION block so it’s treated as a planned element.
Layer 3: STYLE (guardrails, not story)
STYLE should constrain the output—not send it to a different genre mid-shot.
Use one primary style
Veo gives examples such as cartoon, claymation, film noir shot on 35mm, worn-out VHS texture (https://deepmind.google/models/veo/prompt-guide/).
Rule: pick one primary style reference. If you add multiple (“anime + noir + watercolor”), you’re asking the model to average contradictions.
Lighting is your legibility knob
Lighting is a first-class Veo prompt element (https://deepmind.google/models/veo/prompt-guide/).
- For label readability: “spotlight on product label”
- For lifestyle softness: “warm even lighting”
Add an “Avoid” list to prevent common ad failures
This is cheap and effective:
- “Avoid: text overlays, extra products, extra hands, scene cuts, strange brand marks.”
Worked example: one messy prompt → one clean prompt (plus variations)
The goal here is a repeatable edit process: keep the idea, remove ambiguity.
Before: the messy “everything prompt”
“Make a super cinematic viral ad for my new sparkling drink. A cool girl in a modern kitchen grabs the bottle, opens it, drinks it, runs outside into a neon city, everyone looks amazed, add dramatic lighting, cool camera movements, slow motion, film look, make it high-end, add music and a tagline.”
What’s broken:
- No framing/motion specifics (Veo asks for them) (https://deepmind.google/models/veo/prompt-guide/).
- Two locations (kitchen → neon city) in one short clip.
- Too many actions for one shot.
- “Tagline” implies text overlays, which often derails product legibility.
After: clean 3-layer prompt (single shot product mini-ad)
SHOT (non-negotiables):
- Hero subject: matte black sparkling drink bottle with a minimal white label; woman in her twenties with wavy brown hair and light freckles
- Location: modern kitchen at morning, white tile backsplash, oak table
- Framing: close-up on bottle label at table height, slight low-angle
- Camera motion: slow push-in
- Key objects: matte black bottle, clear glass with ice
ACTION (beats + physics):
- Beat 1 (0–2s): her hand places the matte black bottle on the oak table beside the glass
- Beat 2 (2–4s): she twists the cap; crisp “click” and soft carbonation hiss
- Beat 3 (4–6s): she pours into the glass; bubbles rise; she smiles slightly behind the bottle
- Continuity anchors (repeat verbatim): matte black bottle, oak table, white tile backsplash
- Audio (optional): subtle carbonation SFX + minimal upbeat music bed
STYLE (guardrails only):
- Visual style: film noir shot on 35mm (modern, clean)
- Lighting: spotlight on the bottle label, soft fill on face
- Pace: calm, premium
- Avoid: no text overlays, no scene cuts, no extra characters
Why this is “Veo-ready” (mapped to the guide):
- It specifies framing + camera motion (https://deepmind.google/models/veo/prompt-guide/).
- It uses detailed character description (https://deepmind.google/models/veo/prompt-guide/).
- It uses specific location details (https://deepmind.google/models/veo/prompt-guide/).
- It uses action as clear beats (https://deepmind.google/models/veo/prompt-guide/).
- It uses style + lighting as constraints, not a second storyline (https://deepmind.google/models/veo/prompt-guide/).
Want to iterate this fast without rebuilding the prompt each time? Veo3Gen lets you generate with Google’s Veo 3.1 in three modes—Fast (quick default), Quality (max fidelity), and Lite (cheapest preview)—so you can do rough passes, then upgrade the same concept when it’s working. It also generates native synchronized audio (dialogue/SFX/music) in a single pass. (/pricing)
Quick variations table (change only one layer)
| Goal | Change SHOT | Change ACTION | Change STYLE |
|---|---|---|---|
| Make it more “Reel hook” | Medium close-up, handheld micro-movement | Beat 2: add one short spoken line | Worn-out VHS texture (light) (https://deepmind.google/models/veo/prompt-guide/) |
| Make it a clean tutorial | Top-down, static | Beat 1: point to label; Beat 3: pour to halfway | “Warm even lighting” (https://deepmind.google/models/veo/prompt-guide/) |
| Make label extra readable | Close-up tighter | Hold on label focus in Beat 2 | Spotlight on label (https://deepmind.google/models/veo/prompt-guide/) |
When to use First/Last Frame (continuity and transitions)
If you’re trying to preserve composition across a transition (same product position, same character, same scene) text-only prompts can be fragile.
Google’s Cloud post references Veo 3.1’s first frame / last frame capability (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1).
Use first/last frame when
- You need the same product/logo to remain consistent across a shot.
- You want an A→B transformation (cap closed → cap opened; empty glass → filled).
- You want a match-cut feel (same framing, new beat).
What to add to your prompt
Add one sentence inside ACTION:
- “Match the provided first frame composition; end on the provided last frame composition; keep the matte black bottle label readable.”
If you’re producing lots of variants, Veo3Gen supports text-to-video and image-to-video, plus first-and-last-frame control on Veo 3.1—useful when you want the same setup but different beats. (/api)
Iteration order (so you don’t thrash)
A practical order that keeps changes interpretable:
-
Fix SHOT first (composition)
- change only framing (wide→medium→close)
- change only angle (eye-level→low-angle)
- change only camera motion (static→slow push-in)
These are explicitly called out as controllable elements (https://deepmind.google/models/veo/prompt-guide/).
-
Fix ACTION next (clarity)
- remove a beat if it’s crowded
- make the subject explicit (“her right hand…”, “the bottle…”)
- keep continuity anchors identical
-
Fix STYLE last (contain drift)
- choose one style reference (https://deepmind.google/models/veo/prompt-guide/)
- choose one lighting instruction (https://deepmind.google/models/veo/prompt-guide/)
- expand “Avoid” if unwanted artifacts recur
Checklist
- Write SHOT / ACTION / STYLE as separate labeled blocks
- Specify framing and one camera motion (https://deepmind.google/models/veo/prompt-guide/)
- Add 2–3 descriptors for the hero subject (https://deepmind.google/models/veo/prompt-guide/)
- Make location specific: place + time + 1–2 concrete details (https://deepmind.google/models/veo/prompt-guide/)
- Limit ACTION to 2–3 beats; keep each beat single-purpose (https://deepmind.google/models/veo/prompt-guide/)
- Repeat 2–3 continuity anchors verbatim across iterations
- Pick one visual style reference (https://deepmind.google/models/veo/prompt-guide/)
- Pick one lighting instruction (https://deepmind.google/models/veo/prompt-guide/)
- Add an Avoid list (no text overlays, extra characters, scene cuts)
- If continuity is critical, consider first/last frame (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1)
FAQ
How do I write a Veo 3.1 prompt structure that doesn’t drift?
Use the SHOT / ACTION / STYLE split and repeat 2–3 exact continuity anchors. This mirrors Veo’s guidance to specify framing/motion, action, style, lighting, character detail, and location detail clearly (https://deepmind.google/models/veo/prompt-guide/).
How do I describe camera framing and movement for Veo?
State the framing (wide/medium/close-up, angle like low-angle) and one camera motion (static, pan, slow push-in). Veo’s guide explicitly calls out shot framing and camera motion (https://deepmind.google/models/veo/prompt-guide/).
How do I get better lighting and product readability?
Add one lighting instruction such as warm even lighting or a spotlight in one area of the shot (https://deepmind.google/models/veo/prompt-guide/). For ads, “spotlight on the label” is a practical default.
How do I prompt dialogue and sound in Veo videos?
Veo’s guide states it can generate dialogue, and you can either provide a topic or exact lines (https://deepmind.google/models/veo/prompt-guide/). Put dialogue/SFX/music explicitly in the ACTION block so it’s controlled.
How do I keep transitions consistent between two beats?
Use first frame / last frame when you need continuity across a transition; Veo 3.1 is referenced as having this capability (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1).
Why is my prompt blocked or not generating what I asked?
Safety filters can block prompts that violate responsible AI guidelines (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/video/video-gen-prompt-guide).
Build faster: turn the 3-layer system into a repeatable Veo3Gen workflow
The win isn’t a single perfect prompt—it’s a prompt you can revise without breaking everything.
If you’re producing ads/reels weekly, Veo3Gen is designed for that loop: it’s an affordable way to access Google’s Veo 3.1 without Google’s enterprise pricing, offers Fast / Quality / Lite modes for different stages, supports 720p/1080p/4K (4K on Fast/Quality) in 16:9 and 9:16, and generates native synchronized audio in one pass—so you can iterate visuals and audio together. New users get free credits to start, and credits you purchase don’t expire. (/pricing)
Next step: copy the template above, write one SHOT-only draft, then add ACTION beats. When that’s stable, explore STYLE variations—one at a time.
Start creating with Veo3Gen
Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.
- Generate your first video now: Get started
- Compare plans and pay-as-you-go pricing: See pricing
Try Veo 3 & Veo 3 API for Free
Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.