Prompting10 min read

Stop "Perfect Prompting" Veo 3.1: The 3-Layer Prompt System That Gets Cleaner Ads & Reels

A practical Veo 3.1 prompt structure for cleaner ads and reels: SHOT + ACTION + STYLE, with a worked before/after example, checklist, and FAQs.

TL;DR

Stop writing “everything prompts.” For Veo 3.1, you’ll get cleaner, more directable ads and reels by splitting prompts into three layers:

  • SHOT = what must be true in-frame (subject, location, framing, camera motion, key objects)
  • ACTION = what changes over time (2–3 timed beats + continuity anchors + optional dialogue/audio)
  • STYLE = how it looks (visual style + lighting + constraints), without rewriting the scene

This matches how Veo’s own prompt guidance emphasizes shot framing/camera motion, style, lighting, character detail, location detail, and action (https://deepmind.google/models/veo/prompt-guide/) and aligns with “say what’s in frame and how it moves” advice from other video tools (https://help.runwayml.com/hc/en-us/articles/42460036199443-Text-to-Video-Prompting-Guide).

Key takeaways

Why most Veo prompts fail: “everything prompts” aren’t directable

Most weak outputs come from trying to solve three problems in one paragraph:

  1. Composition: what we see + where the camera is + how it moves.
  2. Causality: what happens over time (clear beats, believable motion).
  3. Aesthetic: style + lighting + texture.

Veo’s prompt guide already names the levers you should control—framing/motion, style, lighting, character details, location details, action (https://deepmind.google/models/veo/prompt-guide/). The mistake is listing them as a single “mood board paragraph” instead of separating them into layers you can independently fix.

A good mental model: prompts should describe what appears in the frame and how it moves through the scene in direct language (https://help.runwayml.com/hc/en-us/articles/42460036199443-Text-to-Video-Prompting-Guide). That’s exactly what SHOT + ACTION forces you to do.

The 3-layer prompt system (copy/paste)

Use this as your default Veo 3.1 prompt structure.

The template (copy/paste)

SHOT (non-negotiables):
- Hero subject: [product/character + 2–3 specific descriptors]
- Location: [specific place + time-of-day + 1–2 concrete details]
- Framing: [wide / medium / close-up + angle]
- Camera motion: [static / slow push-in / pan / tilt / handheld]
- Key objects: [product + 1–2 props that must appear]

ACTION (beats + physics):
- Beat 1 (0–2s): [single clear action]
- Beat 2 (2–4s): [single clear action]
- Beat 3 (4–6s): [single clear action]
- Continuity anchors (repeat verbatim): [2–3 exact nouns]
- Audio (optional): [dialogue line OR topic + SFX/music]

STYLE (guardrails only):
- Visual style: [one style reference]
- Lighting: [one lighting instruction]
- Pace: [snappy / calm / cinematic]
- Avoid: [no text overlays, no extra characters, no scene cuts, etc.]

Why these fields map to Veo 3.1 guidance

Layer 1: SHOT (what must be true in-frame)

If SHOT is wrong, the clip is wrong.

Framing + camera motion: pick one primary camera behavior

Veo’s guide explicitly recommends specifying shot framing and camera motion (https://deepmind.google/models/veo/prompt-guide/). Do it like you’re writing a call sheet.

Do:

  • “Close-up, slight low-angle, slow push-in”

Don’t:

  • “Cool cinematic camera moves”

Rule: one main camera motion per clip. If you need multiple moves, you probably need multiple shots.

“2–3 descriptors” rule for characters/products

Veo’s guide shows why specificity wins (their example: “A woman in her twenties with wavy brown hair and light freckles” vs. “a brown-haired woman”) (https://deepmind.google/models/veo/prompt-guide/).

Apply that to marketing:

  • “a bottle” → “a matte black bottle with a minimal white label
  • “a kitchen” → “a modern kitchen, white tile backsplash, morning light

Location specificity reduces randomness

Veo contrasts generic locations with specific ones like “a smoky jazz club at night” or “a cyberpunk city with bright chrome and neon lights” (https://deepmind.google/models/veo/prompt-guide/).

For ads/reels, a specific location does two practical things:

  • It constrains the background so the product reads.
  • It reduces “anywhere energy” (the model inventing random sets).

Layer 2: ACTION (beats, timing, and continuity)

Veo supports action prompts (examples include dashing, backflips, chasing) (https://deepmind.google/models/veo/prompt-guide/). The key is making action unambiguous and sequenced.

Use 2–3 beats max

Write actions as beats with rough time windows. You’re not “locking duration”; you’re forcing priority.

Bad (too simultaneous):

  • “She grabs it, smiles, runs outside, opens the app, dances, drinks, logo appears.”

Good (three beats):

  • Beat 1: place product
  • Beat 2: open product
  • Beat 3: pour / reaction

Continuity anchors: repeat exact nouns

Continuity anchors are literal phrases you repeat across revisions:

  • “matte black bottle”
  • “oak table”
  • “white tile backsplash”

When you change STYLE later, anchors keep the subject from morphing.

Dialogue/audio: keep it simple and deliberate

Veo can generate dialogue; you can provide a topic or exact lines (https://deepmind.google/models/veo/prompt-guide/).

For marketing, use one clean line:

  • “Dialogue (whisper): ‘One sip. Reset.’”

If your prompt includes audio, don’t bury it. Put it in the ACTION block so it’s treated as a planned element.

Layer 3: STYLE (guardrails, not story)

STYLE should constrain the output—not send it to a different genre mid-shot.

Use one primary style

Veo gives examples such as cartoon, claymation, film noir shot on 35mm, worn-out VHS texture (https://deepmind.google/models/veo/prompt-guide/).

Rule: pick one primary style reference. If you add multiple (“anime + noir + watercolor”), you’re asking the model to average contradictions.

Lighting is your legibility knob

Lighting is a first-class Veo prompt element (https://deepmind.google/models/veo/prompt-guide/).

  • For label readability: “spotlight on product label”
  • For lifestyle softness: “warm even lighting”

Add an “Avoid” list to prevent common ad failures

This is cheap and effective:

  • “Avoid: text overlays, extra products, extra hands, scene cuts, strange brand marks.”

Worked example: one messy prompt → one clean prompt (plus variations)

The goal here is a repeatable edit process: keep the idea, remove ambiguity.

Before: the messy “everything prompt”

“Make a super cinematic viral ad for my new sparkling drink. A cool girl in a modern kitchen grabs the bottle, opens it, drinks it, runs outside into a neon city, everyone looks amazed, add dramatic lighting, cool camera movements, slow motion, film look, make it high-end, add music and a tagline.”

What’s broken:

  • No framing/motion specifics (Veo asks for them) (https://deepmind.google/models/veo/prompt-guide/).
  • Two locations (kitchen → neon city) in one short clip.
  • Too many actions for one shot.
  • “Tagline” implies text overlays, which often derails product legibility.

After: clean 3-layer prompt (single shot product mini-ad)

SHOT (non-negotiables):
- Hero subject: matte black sparkling drink bottle with a minimal white label; woman in her twenties with wavy brown hair and light freckles
- Location: modern kitchen at morning, white tile backsplash, oak table
- Framing: close-up on bottle label at table height, slight low-angle
- Camera motion: slow push-in
- Key objects: matte black bottle, clear glass with ice

ACTION (beats + physics):
- Beat 1 (0–2s): her hand places the matte black bottle on the oak table beside the glass
- Beat 2 (2–4s): she twists the cap; crisp “click” and soft carbonation hiss
- Beat 3 (4–6s): she pours into the glass; bubbles rise; she smiles slightly behind the bottle
- Continuity anchors (repeat verbatim): matte black bottle, oak table, white tile backsplash
- Audio (optional): subtle carbonation SFX + minimal upbeat music bed

STYLE (guardrails only):
- Visual style: film noir shot on 35mm (modern, clean)
- Lighting: spotlight on the bottle label, soft fill on face
- Pace: calm, premium
- Avoid: no text overlays, no scene cuts, no extra characters

Why this is “Veo-ready” (mapped to the guide):

Want to iterate this fast without rebuilding the prompt each time? Veo3Gen lets you generate with Google’s Veo 3.1 in three modes—Fast (quick default), Quality (max fidelity), and Lite (cheapest preview)—so you can do rough passes, then upgrade the same concept when it’s working. It also generates native synchronized audio (dialogue/SFX/music) in a single pass. (/pricing)

Quick variations table (change only one layer)

Goal Change SHOT Change ACTION Change STYLE
Make it more “Reel hook” Medium close-up, handheld micro-movement Beat 2: add one short spoken line Worn-out VHS texture (light) (https://deepmind.google/models/veo/prompt-guide/)
Make it a clean tutorial Top-down, static Beat 1: point to label; Beat 3: pour to halfway “Warm even lighting” (https://deepmind.google/models/veo/prompt-guide/)
Make label extra readable Close-up tighter Hold on label focus in Beat 2 Spotlight on label (https://deepmind.google/models/veo/prompt-guide/)

When to use First/Last Frame (continuity and transitions)

If you’re trying to preserve composition across a transition (same product position, same character, same scene) text-only prompts can be fragile.

Google’s Cloud post references Veo 3.1’s first frame / last frame capability (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1).

Use first/last frame when

  • You need the same product/logo to remain consistent across a shot.
  • You want an A→B transformation (cap closed → cap opened; empty glass → filled).
  • You want a match-cut feel (same framing, new beat).

What to add to your prompt

Add one sentence inside ACTION:

  • “Match the provided first frame composition; end on the provided last frame composition; keep the matte black bottle label readable.”

If you’re producing lots of variants, Veo3Gen supports text-to-video and image-to-video, plus first-and-last-frame control on Veo 3.1—useful when you want the same setup but different beats. (/api)

Iteration order (so you don’t thrash)

A practical order that keeps changes interpretable:

  1. Fix SHOT first (composition)

    • change only framing (wide→medium→close)
    • change only angle (eye-level→low-angle)
    • change only camera motion (static→slow push-in)

    These are explicitly called out as controllable elements (https://deepmind.google/models/veo/prompt-guide/).

  2. Fix ACTION next (clarity)

    • remove a beat if it’s crowded
    • make the subject explicit (“her right hand…”, “the bottle…”)
    • keep continuity anchors identical
  3. Fix STYLE last (contain drift)

Checklist

FAQ

How do I write a Veo 3.1 prompt structure that doesn’t drift?

Use the SHOT / ACTION / STYLE split and repeat 2–3 exact continuity anchors. This mirrors Veo’s guidance to specify framing/motion, action, style, lighting, character detail, and location detail clearly (https://deepmind.google/models/veo/prompt-guide/).

How do I describe camera framing and movement for Veo?

State the framing (wide/medium/close-up, angle like low-angle) and one camera motion (static, pan, slow push-in). Veo’s guide explicitly calls out shot framing and camera motion (https://deepmind.google/models/veo/prompt-guide/).

How do I get better lighting and product readability?

Add one lighting instruction such as warm even lighting or a spotlight in one area of the shot (https://deepmind.google/models/veo/prompt-guide/). For ads, “spotlight on the label” is a practical default.

How do I prompt dialogue and sound in Veo videos?

Veo’s guide states it can generate dialogue, and you can either provide a topic or exact lines (https://deepmind.google/models/veo/prompt-guide/). Put dialogue/SFX/music explicitly in the ACTION block so it’s controlled.

How do I keep transitions consistent between two beats?

Use first frame / last frame when you need continuity across a transition; Veo 3.1 is referenced as having this capability (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1).

Why is my prompt blocked or not generating what I asked?

Safety filters can block prompts that violate responsible AI guidelines (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/video/video-gen-prompt-guide).

Build faster: turn the 3-layer system into a repeatable Veo3Gen workflow

The win isn’t a single perfect prompt—it’s a prompt you can revise without breaking everything.

If you’re producing ads/reels weekly, Veo3Gen is designed for that loop: it’s an affordable way to access Google’s Veo 3.1 without Google’s enterprise pricing, offers Fast / Quality / Lite modes for different stages, supports 720p/1080p/4K (4K on Fast/Quality) in 16:9 and 9:16, and generates native synchronized audio in one pass—so you can iterate visuals and audio together. New users get free credits to start, and credits you purchase don’t expire. (/pricing)

Next step: copy the template above, write one SHOT-only draft, then add ACTION beats. When that’s stable, explore STYLE variations—one at a time.

Start creating with Veo3Gen

Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.

Limited Time Offer

Try Veo 3 & Veo 3 API for Free

Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.