Prompting14 min read

Sora 2's Official Prompting Guide → A Veo3Gen "Cinematographer Brief" Template (With 12 Copy-Paste Prompts for Ads & Reels) [2026]

Turn Sora 2 prompting principles into a Veo3Gen-ready “Cinematographer Brief” template plus 12 copy‑paste prompts for ads & Reels.

On this page

TL;DR

One-paragraph prompts cause drift: the model improvises subject, pacing, camera, and sound. Use a Cinematographer Brief—a structured, fill‑in prompt that locks goal → subject → action beats → camera → audio → dialogue block → constraints. It’s based on the most transferable parts of Sora 2’s official prompting guidance (beats, sectioning, separate dialogue), adapted to what you actually paste into Veo3Gen.

Key takeaways

Why a “Cinematographer Brief” beats one-paragraph prompts

Most unusable generations come from missing constraints, not a lack of “creativity.” When prompts are under-specified, common failures show up fast:

  • Goal is vague → you get pretty footage that can’t be cut into an ad.
  • Hero subject isn’t locked → the camera “chooses” a different main subject mid‑clip.
  • Action is a blob → pacing turns mushy, gestures look floaty.
  • Audio is generic → SFX and dialogue timing don’t match on-screen actions.

Sora 2 prompting advice (and high-performing community summaries of it) converges on one idea: structured prompts reduce randomness—use sections that state what happens, how it looks, and what is heard (https://wavespeed.ai/blog/posts/sora-2-prompting-tips-better-videos-2026). That maps cleanly to Veo3Gen workflows because Veo3Gen generations include native, synchronized audio in the same pass.

The Sora 2 principles worth stealing (and what to ignore)

You’re not “writing like Sora.” You’re borrowing practices that make any video model easier to direct.

1) Steal: beats, not blob text

OpenAI guidance, as summarized in a Sora 2 prompt collection, recommends describing action in beats/counts—small steps, gestures, pauses—to keep movement simple and legible (https://artificialcorner.com/p/sora-2-prompts).

Practical rule: 3–6 beats max for one short ad clip. Each beat should be observable on-screen.

Examples of strong beats:

  • “She taps the screen twice, pauses 1 second, then smiles.”
  • “The lid pops; steam rises; the liquid sloshes once, then settles.”

2) Steal: separate visuals from dialogue

Same source summarizes OpenAI’s recommendation: put dialogue below the prose description so the model distinguishes visual direction from spoken lines (https://artificialcorner.com/p/sora-2-prompts).

This is the fastest way to prevent:

  • dialogue becoming on-screen text,
  • visual instructions being spoken aloud,
  • messy voice timing.

3) Steal: request audio that syncs to the visuals

Wavespeed explicitly recommends structuring prompts to include what is heard, and requesting sound elements that sync with actions (https://wavespeed.ai/blog/posts/sora-2-prompting-tips-better-videos-2026).

If your beat is “cap twists,” your audio section should say “cap twist SFX synced.” Don’t say “add satisfying sounds” and hope.

4) Ignore in prompt text: API-only parameter talk

The official Sora 2 Prompting Guide states that some attributes are governed only by API parameters and can’t be reliably requested in prose (https://developers.openai.com/cookbook/examples/sora/sora2_prompting_guide). It also documents parameters like size formatted as {width}x{height} and a seconds parameter with supported values 4/8/12/16/20 (https://developers.openai.com/cookbook/examples/sora/sora2_prompting_guide).

Translation for creators: keep your prompt focused on cinematic intent, and use your tool’s UI/API controls for hard constraints when available.

If you’re generating at scale, note: the Sora guide also mentions a batch API for video for asynchronous jobs in larger workflows (https://developers.openai.com/cookbook/examples/sora/sora2_prompting_guide). You can apply the same production thinking to Veo3Gen too, because Veo3Gen has a developer API for generating videos programmatically.

The Veo3Gen “Cinematographer Brief” (copy-paste template)

Use this as a reusable prompt skeleton. Fill it like a creative brief, then paste into Veo3Gen.

CINEMATOGRAPHER BRIEF (Veo3Gen Prompt Template)

1) Goal (1 sentence):

  • What should the viewer do/feel/understand by the end?

2) Deliverable:

  • Platform: [Reels/TikTok/Shorts]
  • Aspect ratio: [9:16 or 16:9]
  • Resolution: [720p / 1080p / 4K]
  • Mode: [Veo 3.1 Fast / Quality / Lite]

3) Hero subject (lock it):

  • Main subject description with specific identifiers (material/color/label placement/outfit)
  • “Hero subject stays on-screen the whole time unless stated.”

4) Action beats (3–6 max):

  • Use physical verbs + timing cues (pause/hold/snap/tilt/set down)

5) Setting:

  • Location + time of day + background activity level (quiet/clean vs busy)

6) Camera (one primary behavior):

  • Shot size: [close/medium/wide]
  • Behavior: [locked-off / slow push-in / handheld micro-shake]
  • Prohibitions: “no whip pans, no random zooms, no cutaways”

7) Lighting:

  • [soft window light / warm vanity / fluorescent office] + contrast level

8) Style:

  • [UGC phone video / clean studio demo / cinematic commercial / documentary handheld]

9) Audio (native synced):

  • Voice: [calm/friendly/urgent], spoken style [voiceover/on-camera]
  • Music: [none/subtle/energetic] and how loud
  • SFX: list 3–5 effects explicitly tied to beats

10) DIALOGUE (separate block):

  • Put spoken lines here only

11) Constraints (anti-chaos):

  • No extra characters
  • No random text overlays unless specified
  • No background chatter
  • Keep motion realistic

Mid-article CTA: use this template with Veo3Gen’s modes + audio

If you want fewer reruns, run this exact brief structure in Veo3Gen and pick the right mode for the job: Veo 3.1 Fast (quick default), Veo 3.1 Quality (max fidelity), or Veo 3.1 Lite (cheapest preview). Because Veo3Gen generates native, synchronized audio with the video, the “Audio” and “DIALOGUE” blocks pay off immediately.

New users get free credits to start, and you can generate via text-to-video or image-to-video.

Worked example (with real copy): vague prompt → usable ad clip

Here’s a concrete before/after you can reuse.

The vague version (what most people type)

“Make a cool cinematic video of a woman using a skincare serum, trendy lighting, with voiceover and music.”

Why it fails: no locked hero identifiers, no beats, no camera rules, and “voiceover and music” has no sync targets.

The Cinematographer Brief version (copy-paste)

Goal: UGC-style hook + product proof for a hydrating face serum.

Deliverable: Platform Reels, 9:16, 1080p, Veo 3.1 Fast.

Hero subject (lock it): Woman (mid‑20s to early‑30s), natural look, hair in a loose bun. Clear glass dropper bottle with pale blue label facing camera. Subject and bottle stay in frame.

Action beats:

  1. Medium close-up in bathroom mirror. She holds bottle up to camera, label centered; hold 1 second.
  2. She twists cap; dropper lifts with a soft pop; one droplet clings to the tip.
  3. Close-up: she dots 3 drops on cheek; liquid moves with slight drag then settles (realistic viscosity).
  4. She gently pats twice; skin looks dewy; she smiles and nods once.

Setting: Clean bathroom vanity, neutral background, minimal clutter.

Camera: Phone-like lens, mostly locked-off with tiny handheld micro-movement. No zooms, no cutaways.

Lighting: Soft warm vanity lights, low harsh shadows.

Style: Authentic UGC ad, not glossy.

Audio (native synced): Friendly voiceover. Subtle upbeat background music low volume. SFX synced: cap twist, dropper pop, tiny liquid plop, skin pat.

DIALOGUE (voiceover):

  • “My skin gets tight after I wash it—so I do three drops.”
  • “It sinks in fast… and the glow is real.”

Constraints: No captions or on-screen text. No extra people. No background chatter.

What changed (so you can copy the method)

Element Vague prompt Brief prompt Why it matters
Hero lock “a woman” hair + age range + bottle identifiers reduces subject drift
Pacing none 4 beats + hold prevents mushy timing
Motion realism none “drag,” “settles” reduces floaty physics
Audio “voiceover and music” voice + music loudness + synced SFX improves sync (https://wavespeed.ai/blog/posts/sora-2-prompting-tips-better-videos-2026)
Dialogue mixed into idea separate block reduces confusion (https://artificialcorner.com/p/sora-2-prompts)

12 copy‑paste prompts (ads + Reels)

Each prompt follows the Brief format. Replace brackets.

1) UGC hook (problem → curiosity)

Goal: Stop-scroll hook for [product] that sets up a clear payoff.

Deliverable: Reels, 9:16, 1080p, [Veo 3.1 Fast].

Hero subject: [Creator], [distinct outfit detail], holding [product] label facing camera.

Action beats:

  1. Looks at camera, raises product into frame; hold 0.5s.
  2. Does one simple demo gesture (open/apply/pour) once.
  3. Small reaction: eyebrow raise + half-smile.

Camera: Handheld phone feel, medium close-up, no zooms.

Audio: Voiceover + light room tone. SFX synced to the demo action.

DIALOGUE (spoken to camera):

  • “If you [common pain], try this for 7 days.”
  • “Here’s what surprised me…”

Constraints: No cutaways. No background crowd noise.

2) Product hero reveal (clean studio)

Goal: Premium hero shot for [product] landing page.

Deliverable: 16:9, 4K, [Veo 3.1 Quality].

Hero subject: Single [product] on [surface], brand mark visible.

Action beats:

  1. Slow push-in.
  2. Light sweep across edges.
  3. Subtle rotation stops with logo centered.

Audio: Minimalist whoosh synced to light sweep + soft room tone. No voice.

Constraints: No extra props unless specified.

3) Before/after (same framing, no text)

Goal: Show transformation without captions.

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Hero subject: Same person, same outfit, same background.

Action beats:

  1. “Before” expression + hair/skin detail visible.
  2. Uses [product] with one simple motion.
  3. “After” look: confident posture shift + smile.

Camera: Locked-off, identical composition.

Audio: One distinct transition SFX between beat 2 and 3.

Constraints: Only the result changes; wardrobe stays consistent.

4) App demo (hands + phone)

Goal: Demonstrate one core feature of [app] clearly.

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Hero subject: Hands holding a phone; screen is the hero.

Action beats:

  1. Thumb taps [button].
  2. Screen transitions to [feature].
  3. Result appears.
  4. Thumb confirms.

Audio: Tap/click SFX synced; subtle UI blips; no loud music.

Constraints: Keep screen readable; minimal background.

5) Testimonial (talking head)

Goal: Believable testimonial for [service].

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Hero subject: Person seated, natural lighting, eye contact.

Action beats:

  1. Starts mid-thought.
  2. One small hand gesture.
  3. Smile at end.

Audio: Clean voice, minimal room echo. No music.

DIALOGUE (spoken):

  • “I tried [alternative] and it didn’t work.”
  • “With [brand], I finally [result].”

Constraints: No dramatic camera moves.

6) Offer reveal (urgency without numbers)

Goal: Create urgency without on-screen text.

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Hero subject: [product] + creator.

Action beats:

  1. Creator leans in.
  2. Holds product close to camera.
  3. Gesture “today only.”
  4. Points to blank negative space for editor CTA.

Audio: Energetic voice. One sting SFX on beat 2.

DIALOGUE (spoken):

  • “If you’ve been waiting—this is the best time.”
  • “Tap to grab it before it’s gone.”

Constraints: Leave clean space for overlays.

7) Unboxing (tactile)

Goal: Tactile unboxing for [product].

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Action beats:

  1. Tape peel (slow).
  2. Box opens with slight resistance.
  3. Protective wrap crinkles.
  4. Product placed on table, centered.

Audio: Crisp tape peel + cardboard creak + wrap crinkle synced.

Constraints: Clean table, no clutter.

8) Food/drink pour (physics-forward)

Goal: Make [drink] look irresistible.

Deliverable: 9:16, 1080p, [Veo 3.1 Quality].

Action beats:

  1. Ice drops with bounce.
  2. Pour begins slow then thickens (inertia).
  3. Foam rises then settles.
  4. Condensation beads on glass.

Audio: Ice clink + pour splash + fizz synced.

Constraints: Realistic liquid behavior.

9) Local business day-in-the-life (mini montage)

Goal: 12–20s mini-story of [shop] craftsmanship.

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Action beats:

  1. Prep: hands working close-up.
  2. Customer interaction: smile + exchange.
  3. Finished product hero.

Camera: Documentary handheld, smooth (not chaotic).

Audio: Light ambient shop sound + soft music low.

Constraints: No random cutaways to unrelated items.

10) Fitness form cue (single move)

Goal: Teach one form fix fast.

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Action beats:

  1. Incorrect rep (subtle).
  2. Reset posture.
  3. Correct rep.
  4. Thumbs-up.

Audio: Coaching voiceover + one beep cue at correction.

Constraints: One exercise only.

11) SaaS problem → solution (desk)

Goal: Show frustration then relief using [tool].

Deliverable: 16:9, 1080p, [Veo 3.1 Fast].

Action beats:

  1. Person sighs at messy tabs.
  2. Opens [tool].
  3. One click organizes.
  4. Relaxed smile.

Audio: Keyboard clicks + subtle “resolve” music swell.

Constraints: No extra people.

12) Event teaser (energy + editable space)

Goal: Hype [event] with clean space for overlays.

Deliverable: 9:16, 1080p, [Veo 3.1 Fast].

Action beats:

  1. Quick walk-in shot.
  2. Stage lights flare.
  3. Close-up high-five.
  4. End on signage area with blank negative space.

Audio: Crowd roar + bass hit synced with stage flare.

Constraints: Keep signage readable; avoid chaotic camera.

Motion & physics: 5 words that fix “floaty AI movement”

When motion looks synthetic, it’s often missing forces. Add one force cue per beat.

  • Weight: “sets it down with a heavy thud; table vibrates slightly.”
  • Inertia: “stops suddenly; the bag swings forward then back.”
  • Drag: “thick cream stretches and trails before snapping clean.”
  • Bounce: “hits once, rebounds lower on the second bounce.”
  • Splash: “droplets scatter outward, then fall back and ripple.”

This matches the beats-first approach recommended in Sora prompt guidance summaries (https://artificialcorner.com/p/sora-2-prompts).

Dialogue & audio blocks (the formatting that prevents mix-ups)

OpenAI’s recommended pattern (as summarized) is simple: dialogue goes below visuals (https://artificialcorner.com/p/sora-2-prompts). Keep it consistent.

Format A: voiceover

DIALOGUE (voiceover):

  • “Line 1…”
  • “Line 2…”

Format B: on-camera

DIALOGUE (spoken to camera):

  • “Line 1…”
  • “Line 2…”

When you do want audio, request it explicitly and tie it to beats—Wavespeed highlights this as a practical way to improve sync (https://wavespeed.ai/blog/posts/sora-2-prompting-tips-better-videos-2026).

Consistency: how to get “series” outputs without a pipeline

Two constraints you should plan for:

What works in prompt-only practice:

  1. Write one canonical “hero subject” paragraph and reuse it verbatim across prompts.
  2. Keep camera behavior constant across a series (e.g., always handheld UGC or always locked-off studio).
  3. Generate small families: keep structure identical, change one variable (lighting or setting or dialogue). This makes selection easier.

Checklist

  • Goal is one sentence (viewer takeaway + intended action)
  • Deliverable is set (platform + aspect ratio + resolution + chosen mode)
  • Hero subject is locked with specific identifiers
  • Action is 3–6 beats with at least one hold/pause
  • Camera has one primary behavior (locked/slow push/handheld), no surprises
  • Setting is simplified (minimal clutter)
  • Audio includes voice type + music loudness + 3–5 synced SFX
  • Dialogue is in a dedicated DIALOGUE block (not mixed into visuals)
  • Constraints include “no cutaways / no random text overlays” unless desired

FAQ

How do I write an AI video prompt that doesn’t drift into random scenes?

Lock the hero subject, keep the clip to 3–6 beats, and include constraints like “no cutaways” plus a single camera behavior. Drift usually comes from leaving “what must stay constant” unspecified.

Why should dialogue be in a separate block?

Because OpenAI prompting guidance (as summarized) recommends separating dialogue from visuals so the model doesn’t confuse spoken text with visual instructions (https://artificialcorner.com/p/sora-2-prompts).

How do I get audio that actually matches on-screen actions?

Request audio elements that are explicitly synced to beats (cap twist, pop, pour, pat). A Sora 2 prompting tips article recommends including what is heard and syncing it to visuals (https://wavespeed.ai/blog/posts/sora-2-prompting-tips-better-videos-2026).

Why does the same prompt produce different videos?

Variability is expected: running the same prompt multiple times can yield different outputs (https://artificialcorner.com/p/sora-2-prompts). Use robust structure (beats + constraints) and plan to pick winners from a small set.

What should I do when a parameter can’t be reliably requested in prose?

The Sora 2 official guide notes that some attributes are governed only by API parameters and must be set explicitly in the API call (https://developers.openai.com/cookbook/examples/sora/sora2_prompting_guide). Use your tool’s controls for hard constraints, and keep the prompt focused on creative direction.

Create more usable videos per credit with Veo3Gen (closing CTA)

If you adopt the Cinematographer Brief, you’ll waste fewer generations on “almost right” clips because your prompt stops leaving core decisions to improvisation.

Veo3Gen is an affordable way to access Google’s Veo 3.1 video models without Google’s enterprise pricing. It supports text-to-video and image-to-video, plus first-and-last-frame control on Veo 3.1, and generations include native, synchronized audio in a single pass.

  • Want to start immediately with free credits and choose the right mode (Fast/Quality/Lite)? https://veo3gen.com/pricing
  • Need a repeatable workflow for batches of ads or Reels? Veo3Gen has a developer API so you can generate videos programmatically: https://veo3gen.com/api

Start creating with Veo3Gen

Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.

Limited Time Offer

Try Veo 3 & Veo 3 API for Free

Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.