Prompting9 min read
Adobe Firefly's Official Video Prompt Rules (June 2026) → A Veo3Gen "Shot Spec" You Can Reuse for Ads, Reels & Tutorials
Turn generic AI prompts into a reusable “Shot Spec” for consistent ads, Reels, and tutorials—plus a worked before/after example and iteration checklist.
On this page
- TL;DR
- Key takeaways
- Why “prompt rules” matter (and what to ignore)
- The Veo3Gen “Shot Spec” template (copy/paste)
- Mid-article CTA: use this template inside Veo3Gen
- Step 1: Subject + context (what must be true in frame 1)
- Pin down the scene like a production designer
- Step 2: Action (one change) + what stays fixed
- Step 3: Camera language that reduces randomness
- Worked example (before/after) you can reuse today
- Step 4: Temporal beats (sequence without brittle timestamps)
- Step 5: An iteration loop that doesn’t waste runs
- Prompt revision table (quick fixes)
- Marketing add-on: output notes that make prompts convert
- 3 ready-to-use Shot Specs (copy/paste)
- 1) Short-form ad hook (vertical)
- 2) UGC product demo (hands-only)
- 3) Creator tutorial b-roll (editing workflow)
- Common mistakes (and the exact rewrite)
- Mistake: Overloaded scenes
- Mistake: Too many camera instructions
- Mistake: Numeric specificity
- Mistake: Asking for text overlays inside the generation
- Checklist
- FAQ
- How do I structure an AI video prompt so it’s consistent?
- How do I stop random camera moves?
- How do I prompt sequence without exact timestamps?
- What should I include for marketing videos so they’re usable in ads?
- How do I generate lots of variations without rewriting prompts?
- Generate these Shot Specs faster with Veo3Gen
- Start creating with Veo3Gen
- Sources
TL;DR
A reliable AI video prompt is less “creative writing” and more a shot specification: define what must be true in frame 1, what changes, and how the camera observes it. Use a constrained prompt structure (Subject + Action + Scene + Camera/Lighting/Style) and add simple temporal beats (“starts with / then / ends with”) so the model has a clear, shootable plan (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos). For marketing outputs, include the intent (video type, duration, features/topic, CTA) in your output notes (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos).
Key takeaways
- Treat each generation as one shot: one main subject, one main action, one camera move.
- Use FlexClip’s proven structure: Subject + Action + Scene + (Camera Movement + Lighting + Style) (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos).
- Add sequence with beat language (starts/then/ends) instead of brittle timestamps.
- Iterate like a producer: change one field at a time so you learn what fixed the output.
- For marketing prompts, include type of video, duration, features/topic, and CTA in output notes (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos).
Why “prompt rules” matter (and what to ignore)
FlexClip’s core point is practical: a well-crafted prompt dictates what the model produces (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos). InVideo says the same thing more bluntly: AI performs as well as the prompt it’s given (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos).
What that means in day-to-day creator work:
- If your prompt reads like an entire commercial (multiple locations, wardrobe changes, product closeups, punchlines), you’re asking for too many discrete events.
- If your prompt reads like a shootable single take, you get consistency.
Ignore:
- Long adjective chains (“ultra cinematic, insane, viral, breathtaking…”).
- Exact counts and measurements (numbers often introduce failure modes).
- “And then and then…” storyboards. If it’s not one take, split it into shots.
The Veo3Gen “Shot Spec” template (copy/paste)
This template maps to the structure FlexClip recommends—Subject + Action + Scene + (Camera Movement + Lighting + Style) (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos)—and adds two creator necessities: sequencing and marketing intent.
VEO3GEN SHOT SPEC (copy/paste)
Subject: (who/what is the focus; include 1–2 defining traits)
Scene/Setting: (where; foreground + background anchors; time of day)
Action: (one main action; concrete verbs; what stays still)
Camera: (shot size + angle + ONE movement)
Lighting: (mood + key light direction)
Style: (1–2 style references; keep it minimal)
Temporal beats: (starts with… then… ends with…)
Constraints: (no cuts; no extra people; no readable text/logos; realism notes)
Audio (optional): (dialogue/SFX/music vibe)
Output notes: (type of video, duration target, platform, key feature/topic, CTA)
Mid-article CTA: use this template inside Veo3Gen
If you want a fast “prompt → preview → revise one field → rerun” loop, Veo3Gen is an affordable way to access Google’s Veo 3.1 video models without Google’s enterprise pricing. It supports text-to-video and image-to-video, includes native synchronized audio in one pass, and offers three modes—Veo 3.1 Fast, Quality, and Lite—so you can preview cheaply and finish at higher fidelity once the shot logic is locked.
Step 1: Subject + context (what must be true in frame 1)
FlexClip defines Subject as what or who is the focus of the video (people, animals, plants, objects) (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos). Your job is to eliminate ambiguity with one identity + 1–2 anchors.
Use:
- Identity + anchors: “matte-black insulated bottle with flip-straw lid”
- Frame ownership: “center frame,” “foreground,” “held in one hand”
Avoid:
- Large casts unless the extras are explicitly background.
- Numeric precision (“exactly 12 ice cubes”). If you need quantity, say “a small amount” or “a few.”
Pin down the scene like a production designer
FlexClip defines Scene as where the action takes place, including foreground/background elements (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos). Don’t write a paragraph—write 2–3 anchors.
Examples:
- “sunlit home kitchen, white tile backsplash, light wood countertop”
- “gym bench area, blurred weights in background, cool gray tones”
Step 2: Action (one change) + what stays fixed
FlexClip calls Action the core of a prompt that drives the storyline and should be clear and concise (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos).
Pick one hero verb chain that can happen in one take:
- “flips lid open, takes one sip, sets bottle down upright”
- “peels backing, applies patch once, presses to secure”
Then state what must not change:
- “background stays static”
- “no wardrobe change”
- “hands remain in frame”
Step 3: Camera language that reduces randomness
FlexClip defines Camera Movement as shot/angle/movement that adds to narrative and visual appeal (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos). In prompting, camera instructions double as stability constraints.
Use a simple recipe:
- Shot size: close-up / medium / wide
- Angle: eye-level / top-down / low angle
- One move: slow push-in / pan left / gentle handheld
FlexClip notes movements can be combined (e.g., “move down and zoom out”) (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos). Combine only when it’s essential; otherwise, one move is easier to control.
Worked example (before/after) you can reuse today
Below is the same idea rewritten from “marketing wish-list” to “single-shot spec.”
Before (typical prompt that breaks):
Create a cool ad for my bottle, show the bottle in different angles, someone uses it at the gym and at work, cinematic lighting, add text with a strong CTA.
What’s wrong (in practical terms):
- Multiple locations (“gym and work”) = multiple shots.
- “Different angles” = cuts.
- “Add text” = often garbled typography; better handled in edit.
After (Veo3Gen Shot Spec):
Subject: matte-black insulated water bottle with flip-straw lid, light condensation
Scene/Setting: gym bench area, blurred weights in background, neutral gray tones
Action: athlete’s hand flips the straw lid open, takes one sip, sets bottle down upright
Camera: medium close-up, eye level, slow push-in
Lighting: cool overhead gym lighting with a soft rim light on the bottle
Style: clean realistic product commercial
Temporal beats: starts with bottle centered on bench; then hand enters and flips lid + sip; ends with bottle upright, label area blank
Constraints: single continuous shot, no cuts; no on-screen text; no readable logos; no extra props appearing
Audio (optional): subtle gym ambience + soft lid click, no music
Output notes: short-form ad hook, vertical 9:16, highlight “flip-straw convenience,” CTA for edit: “Tap to learn more”
Step 4: Temporal beats (sequence without brittle timestamps)
Instead of “at 2s do X,” use:
- Starts with = frame-1 truth
- Then = the single change
- Ends with = final composition you want to cut on
This plays nicely with marketing intent: keep the CTA in Output notes for your edit, or request spoken audio if you want the line delivered in-generation.
Step 5: An iteration loop that doesn’t waste runs
InVideo’s guidance (“AI only performs as well as the prompt it’s given”) becomes useful when you turn it into an operating rule: change one variable per iteration (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos).
Prompt revision table (quick fixes)
| Failure you see | Likely cause | Edit only this field |
|---|---|---|
| Subject drifts / changes identity | Subject too vague | Subject: add 1–2 anchors (material/color), “foreground/center frame” |
| Random props appear | Scene under-specified | Scene/Setting: add 2–3 anchors; Constraints: “no extra objects introduced” |
| Unwanted cuts / angle jumps | Camera not constrained | Camera: specify one shot size + one move; Constraints: “single continuous shot, no cuts” |
| Weird hand/object motion | Action too abstract | Action: replace “uses/shows” with physical verbs; add one realism note |
| Camera does something wild | Too many camera adjectives | Camera: delete extras; keep one move; remove “dynamic/crazy/energetic” |
Marketing add-on: output notes that make prompts convert
InVideo lists five marketing prompt elements: type of video, duration, brand website, features/topic, and CTA (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos). You don’t need all five every time, but you should consistently include:
- Type of video (ad hook, UGC demo, tutorial b-roll)
- Duration target (range is fine)
- Features/topic (one promise)
- CTA (what the viewer should do next)
Put these in Output notes so your shot stays simple while your intent stays clear.
3 ready-to-use Shot Specs (copy/paste)
Each example stays inside “one shot, one action, one move.”
1) Short-form ad hook (vertical)
Subject: frosty iced coffee in a clear glass, visible condensation
Scene/Setting: bright morning kitchen counter, minimal clutter, kettle blurred in background
Action: hand drops an ice cube in; coffee swirls; hand slides glass slightly toward camera
Camera: close-up, slightly top-down, slow push-in
Lighting: warm morning light from the left
Style: crisp modern realistic food commercial
Temporal beats: starts with still glass centered; then ice drop + swirl; ends with glass closer to camera as swirl settles
Constraints: single continuous shot; no text overlays; no readable logos; no extra objects appearing
Audio (optional): ice clink + subtle kitchen ambience
Output notes: ad hook, 6–10s, 9:16, feature “refreshing morning ritual,” CTA: “Tap to try the recipe”
2) UGC product demo (hands-only)
Subject: hands holding a compact face moisturizer jar, white lid, label area blank
Scene/Setting: bathroom vanity, mirror blurred, minimal everyday items in background
Action: hands open jar, scoop a small amount, apply once to back of hand, hold hand still to show sheen
Camera: medium close-up, eye level, gentle handheld
Lighting: soft bathroom lighting with slight backlight for texture
Style: natural UGC smartphone look, realistic
Temporal beats: starts with jar in hands; then lid opens + product shown; ends with hand held still for clear texture view
Constraints: no cuts; no extra people; no readable brand text; realistic skin motion/texture
Audio (optional): quiet room tone + soft lid twist; optional whispered line: “Feels lightweight and smooth.”
Output notes: UGC demo, 10–15s, 9:16, feature “lightweight texture,” CTA: “Tap to see ingredients”
3) Creator tutorial b-roll (editing workflow)
Subject: laptop keyboard and hands, editing timeline visible but unreadable
Scene/Setting: desk with a small lamp and a plant blurred in background
Action: hand taps trackpad, drags one clip in the timeline, then hits play
Camera: over-the-shoulder medium shot, slow pan right
Lighting: warm desk lamp + cool screen glow
Style: clean creator workspace, realistic
Temporal beats: starts with hands hovering; then drag-and-drop; ends with play initiated and timeline moving
Constraints: single continuous shot; no pop-ups; no legible text; no jump cuts
Audio (optional): soft clicks + subtle lo-fi bed
Output notes: tutorial b-roll, 5–8s, 16:9 or 9:16, topic “simple editing workflow,” CTA: “Watch the full tutorial”
Common mistakes (and the exact rewrite)
Mistake: Overloaded scenes
Symptom: the model ignores half your locations.
Rewrite: one Shot Spec = one location. Generate 3–5 shots and edit a montage.
Mistake: Too many camera instructions
Symptom: jittery framing, angle jumps.
Rewrite: keep one move. FlexClip allows combining moves (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos), but treat combos as advanced.
Mistake: Numeric specificity
Symptom: wrong counts, strange UI text.
Rewrite: replace numbers with “a small amount / a few / brief moment” and add “no legible text” in constraints.
Mistake: Asking for text overlays inside the generation
Symptom: garbled typography.
Rewrite: put CTA in Output notes for your edit. If you want the CTA delivered, request short spoken audio instead.
Checklist
- One main Subject with 1–2 anchors (material/color/role)
- One main Action using physical verbs (open, pour, place, point)
- Scene/Setting includes 2–3 background anchors
- Camera = shot size + angle + one movement
- Lighting specifies mood + direction (warm morning, backlighting, spotlight)
- Temporal beats: starts with / then / ends with
- Constraints: single continuous shot, no cuts; no extra props/people; no readable text/logos
- Output notes include type, duration target, feature/topic, CTA (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos)
FAQ
How do I structure an AI video prompt so it’s consistent?
Use Subject + Action + Scene + (Camera Movement + Lighting + Style) and keep it to one shot with one action (https://help.flexclip.com/en/articles/10326783-how-to-write-effective-text-prompts-to-generate-ai-videos).
How do I stop random camera moves?
Specify one shot size and one camera movement, then add the constraint: “single continuous shot, no cuts.”
How do I prompt sequence without exact timestamps?
Use beat language: “starts with… then… ends with…” to describe order without brittle timecodes.
What should I include for marketing videos so they’re usable in ads?
InVideo recommends including elements like video type, duration, features/topic, and CTA (https://invideo.io/blog/ultimate-ai-prompting-guide-for-marketing-videos). Put them in Output notes so the shot stays simple.
How do I generate lots of variations without rewriting prompts?
Turn the Shot Spec into a template and swap fields (setting, style, feature, CTA). If you’re generating programmatically, Veo3Gen has a developer API.
Generate these Shot Specs faster with Veo3Gen
If you’re adopting a Shot Spec workflow, your biggest leverage is iteration speed: preview, fix one failure mode, rerun. Veo3Gen is an affordable way to access Google’s Veo 3.1 video models without Google’s enterprise pricing, supports text-to-video and image-to-video, and can generate native synchronized audio (dialogue, SFX, music) in a single pass.
When you’re ready to go from preview to final, Veo3Gen offers three modes—Veo 3.1 Lite (cheapest, preview), Veo 3.1 Fast (quick, great default), and Veo 3.1 Quality (max fidelity)—plus supported resolutions 720p, 1080p, and 4K (4K on Fast/Quality) and aspect ratios 16:9 and 9:16.
Start creating with Veo3Gen
Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.
- Generate your first video now: Get started
- Compare plans and pay-as-you-go pricing: See pricing
Sources
Try Veo 3 & Veo 3 API for Free
Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.