Video AI Prompting10 min read
Veo 3.1 First+Last Frame (2026): A Creator's Guide to Seamless Transitions, Outfit Changes & Product Reveals
Learn Veo 3.1 first+last frame with a 5-line prompt template, alignment workflow, and worked examples for seamless transitions, outfit swaps, and product reveal
On this page
- TL;DR
- Key takeaways
- What “First + Last Frame” is (and when it beats text-only)
- The part creators miss: you must choreograph the “bridge”
- Choose a transition that hides artifacts
- Outfit swaps
- Room makeovers / set changes
- Product reveals
- Prep your two images (the part that determines 80% of success)
- Do the 60-second overlay test
- Match camera language (framing + perspective)
- Match lighting direction and softness
- Remove “temptation details”
- The 5-line prompt template (copy/paste)
- Camera constraint lines that reduce drift
- Audio: keep it tight and timed
- Worked example (with a concrete “before → after” improvement)
- Inputs
- Before (typical vague prompt)
- After (template-driven prompt)
- Scale this workflow in Veo3Gen (mid-article CTA)
- 3 copy-paste prompts (creator + small business)
- 1) Countertop product reveal (wipe)
- 2) Room makeover (foreground wipe + gentle push)
- 3) Makeup before/after (hand-to-lens)
- Troubleshooting table: symptom → cause → exact edits
- Posting quality: pass the “first-second + one-motion” test
- Aspect ratio + resolution (make the choice upfront)
- Trim intentionally
- When first+last frame is the wrong tool
- Checklist
- FAQ
- How do I use Veo 3.1 first and last frame for a smooth transition?
- How do I stop random zooms, pans, or cuts?
- What prompt format works best for outfit changes?
- Why doesn’t my last frame match even when I provide it?
- Can Veo generate audio that matches the transition beat?
- How can I batch-generate variations efficiently?
- Ship your first bookend transition (closing CTA)
- Start creating with Veo3Gen
TL;DR
Use Veo 3.1 first and last frame as bookends: you provide a start image and an end image, then you direct one continuous transition beat that connects them (spin, wipe, whip pan) while you lock camera + identity and forbid cuts. The practical formula:
- Align the two frames (same crop, horizon, lens feel, lighting direction).
- Write a “bridge,” not just endpoints (one motion that plausibly hides artifacts).
- Constrain everything else (locked camera, single shot, no extra props/actions).
- Time audio to the change moment (one SFX + optional short line).
Key takeaways
- Bookending (first+last frame) is the fastest way to get controlled before→after (outfit swaps, makeovers, product reveals) with less camera drift than text-only.
- Your images do most of the steering: matching framing/perspective and lighting beats “more adjectives.”
- The most common failure is not “bad prompting”—it’s misaligned bookends and a missing single transition beat.
- A repeatable 5-line template (Subject lock → Start → Motion → End → Audio) makes results predictable and scalable.
- For short-form, prioritize the “first-second + one-motion” rule: recognizable subject fast, one clear transition, end before degradation (https://www.vidu.com/blog/social-media-video-ai).
What “First + Last Frame” is (and when it beats text-only)
First + Last Frame means you supply a start image and an end image, and the model generates the motion between them. Google Cloud references Veo 3.1’s “first frame, last frame” capability (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1). Replicate also describes first/last frame input as a tool used in Veo 3.1 workflows (https://replicate.com/blog/veo-3-1).
Use it when you need continuity more than novelty:
- Outfit change / match cut (same pose, different wardrobe)
- Before/after (cleaning, makeup, renovation, plating food)
- Product reveal (hidden → hero shot)
- Brand-safe content (same subject, same framing, controlled environment)
Text-only prompts still win for vibe-first scenes where you don’t care if the camera moves or the subject changes.
The part creators miss: you must choreograph the “bridge”
Two anchors aren’t enough. If you only label endpoints (“change outfit from A to B”), the middle becomes a free-for-all: extra gestures, random camera moves, unasked-for props. Your job is to specify:
- What stays invariant (identity, camera position, framing)
- The one motion that causes the change (spin / wipe / whip pan)
- What is forbidden (cuts, scene changes, additional actions)
Choose a transition that hides artifacts
Pick a transition where motion blur or occlusion is expected.
Outfit swaps
Best transitions:
- 360° spin (motion blur masks fabric morphing)
- Hand covers lens (hard cut without calling it a cut)
- Whip pan (blur masks identity drift)
- Snap + flash (instantaneous change you can time to audio)
Room makeovers / set changes
Harder because geometry must stay consistent. Prefer:
- Foreground wipe (plant/curtain passes near lens)
- Slow push-in (subtle motion; fewer hallucinated background changes)
- Doorway pan (a wall edge hides the “swap”)
Product reveals
Often the most reliable if you keep composition simple. A useful reality check: even for a basic “product rotating against a gradient,” Vidu reports only 3 out of 5 text-prompt tests were smooth enough to post (https://www.vidu.com/blog/social-media-video-ai). Bookends reduce variance by pinning composition at both ends.
Prep your two images (the part that determines 80% of success)
If your bookends don’t look like the same shot, the model will fight you.
Do the 60-second overlay test
- Put both images in any editor.
- Set the top layer opacity to ~50%.
- Nudge/scale to align key points:
- Face: pupils, nose tip, jawline
- Products: corners, logo center
- Rooms: horizon line, countertop edges, door frames
Pass condition: you can align without obvious warping.
Match camera language (framing + perspective)
DeepMind’s prompt guide notes prompts can specify shot framing and camera motion (https://deepmind.google/models/veo/prompt-guide/). That helps—but only if your images already agree.
Avoid mixing:
- Close-up vs medium shot
- Wide-angle distortion vs telephoto compression
- Different camera heights (horizon line mismatch)
Match lighting direction and softness
DeepMind also calls out lighting control in prompts (https://deepmind.google/models/veo/prompt-guide/). If your start is soft frontal light and your end is hard side light, you often get “flicker” as the model tries to reconcile shadows.
Fixes:
- Re-shoot one bookend to match
- Or lightly relight one image so shadows fall the same way
Remove “temptation details”
Busy backgrounds create unintended story cues.
- Remove readable text unless it’s essential
- Avoid extra people/limbs in the background
- Keep the set consistent and uncluttered
The 5-line prompt template (copy/paste)
Google Cloud frames Veo 3.1 prompting as a shift from simple generation toward creative control (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1). Use this template to make that control repeatable.
1) SUBJECT LOCK: [identity constraints; what must remain the same]
2) START FRAME: Match the provided first frame exactly: [pose, framing, lighting, environment]
3) TRANSITION MOTION: [one continuous action], single continuous shot, no cuts
4) END FRAME: Arrive at the provided last frame exactly: [final state], same camera position
5) AUDIO: [one SFX timed to the change] + [optional short dialogue line]
Camera constraint lines that reduce drift
Because prompts can specify camera motion (https://deepmind.google/models/veo/prompt-guide/), explicitly state when you want no motion:
- “Locked camera, tripod-stable, no zoom/pan/tilt.”
- “Single continuous shot, no cuts, no scene change.”
- “Keep framing identical from start to end.”
Audio: keep it tight and timed
DeepMind notes Veo can generate dialogue and prompts can include specific lines (https://deepmind.google/models/veo/prompt-guide/). Google Cloud also describes rich synchronous audio for Veo 3.1 (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1).
For transitions, audio should be a cue, not a second storyline:
- One SFX (whoosh/snap/pop)
- Optional 3–7 word line
- Avoid multi-speaker dialogue or complex music direction
Worked example (with a concrete “before → after” improvement)
Goal: an outfit change that lands exactly on the last frame without camera drift.
Inputs
- First frame: creator in casual outfit, centered medium shot
- Last frame: same creator, same pose/framing, glam outfit
Before (typical vague prompt)
Use my first and last frame. Make her outfit change from casual to glam. Add cool music.
Problems:
- “Her” is not an identity lock
- No camera constraints
- No single bridge motion
- “Cool music” is undefined (timing will float)
After (template-driven prompt)
1) SUBJECT LOCK: Same adult woman throughout; keep face, hairstyle, and body proportions consistent.
2) START FRAME: Match the provided first frame exactly: centered medium shot, warm even lighting, same background.
3) TRANSITION MOTION: One smooth 360-degree spin in place; motion blur during the spin; single continuous shot; locked camera; no zoom; no cuts.
4) END FRAME: Complete the spin and match the provided last frame exactly: same framing and pose; glam outfit.
5) AUDIO: One clean whoosh during the spin + a crisp snap exactly at the moment the outfit changes; no dialogue.
What you changed (the actual levers):
- Identity lock (prevents “new person” drift)
- Camera lock (prevents invented zoom/orbit)
- One motion bridge (forces a coherent middle)
- Audio tied to the beat (forces timing)
Scale this workflow in Veo3Gen (mid-article CTA)
If you want to produce a set (e.g., 8 outfits, 5 colorways, 3 “before/after” tiers), Veo3Gen makes the bookend workflow practical because it supports image-to-video plus first-and-last-frame control on Veo 3.1, and generations include native synchronized audio in a single pass (no separate audio step). You can start with free credits and scale later with pay-as-you-go credits (credits don’t expire) or optional monthly plans, and there’s a developer API for programmatic generation.
3 copy-paste prompts (creator + small business)
Each assumes you upload two images: Start and End.
1) Countertop product reveal (wipe)
1) SUBJECT LOCK: Single hero product only; keep packaging colors consistent; do not add extra objects.
2) START FRAME: Match first frame exactly: empty countertop center, soft daylight from left.
3) TRANSITION MOTION: A neutral cloth wipes left-to-right across the frame; continuous shot; locked camera; no cuts.
4) END FRAME: Match last frame exactly: cloth exits right; product now centered and upright.
5) AUDIO: Soft fabric swipe + subtle “pop” at the reveal; no dialogue.
2) Room makeover (foreground wipe + gentle push)
1) SUBJECT LOCK: Same room architecture and layout; no new doors/windows; keep camera height consistent.
2) START FRAME: Match first frame exactly: wide shot, neutral lighting, slightly messy.
3) TRANSITION MOTION: Gentle steady push-in; a foreground plant leaf passes close to lens as a wipe; single continuous shot; no cuts.
4) END FRAME: Match last frame exactly: same angle/framing; room styled and tidy.
5) AUDIO: Airy whoosh during the leaf wipe + soft chime at reveal; no dialogue.
3) Makeup before/after (hand-to-lens)
1) SUBJECT LOCK: Same adult woman; keep skin tone and facial features consistent; same background.
2) START FRAME: Match first frame exactly: close-up, centered, soft front light.
3) TRANSITION MOTION: She raises her hand to cover the lens fully, then lowers it; locked camera; single continuous shot; no cuts.
4) END FRAME: Match last frame exactly: after makeup look, same framing.
5) AUDIO: Quick whoosh as hand covers lens + light “ta-da” chime when revealed; no dialogue.
Troubleshooting table: symptom → cause → exact edits
| Symptom | Likely cause | Prompt edits that actually help |
|---|---|---|
| Identity drift (face/outfit morphs) | No hard subject lock; competing descriptors | Add: “Same person throughout; keep face and hairstyle identical.” Remove extra style adjectives. Restate: “match first/last frame exactly.” |
| Random props/actions appear | You asked for too much, or no bridge motion | Replace with one transition beat. Add: “no extra props, no additional actions.” |
| Camera zooms/pans/orbits | Camera not constrained | Add: “locked camera, tripod-stable, no zoom/pan/tilt; single continuous shot; no cuts.” |
| Timing feels off (change early/late) | No specified moment; audio not anchored | Add: “change happens at peak of spin / at exact moment hand uncovers lens.” Tie SFX to that moment. |
| Hands/cloth look weird | Intricate interactions | Simplify: use spin/whip pan; or specify “neutral cloth, smooth wipe, hand mostly off-screen.” |
| End frame doesn’t match composition | Bookends misaligned; prompt didn’t demand exact match | Re-crop using overlay test. Add: “arrive at last frame exactly, same framing and position.” Remove camera motion. |
Posting quality: pass the “first-second + one-motion” test
Vidu’s tests found effective social clips shared three traits: the subject is recognizable within the first second, there’s one clear motion/transition, and the clip ends before anything degrades (https://www.vidu.com/blog/social-media-video-ai). Use that as your go/no-go filter.
Aspect ratio + resolution (make the choice upfront)
Pick based on destination:
- Reels/TikTok/Shorts-style vertical: 9:16
- YouTube/landing pages: 16:9
Veo3Gen supports 16:9 and 9:16 and 720p, 1080p, and 4K (4K on Veo 3.1 Fast/Quality). Start lower while you iterate, then re-generate at higher resolution once the transition is approved.
Trim intentionally
Treat the reveal as the ending.
- Hit the transition by ~1.5–2.5 seconds (for short-form)
- Hold the final state briefly
- Cut before the model “adds ideas”
When first+last frame is the wrong tool
Use bookends when the concept is truly one beat: “A turns into B via X.”
Switch approaches when you need multiple beats (approach → pick up → rotate → place → reveal). Two options:
- Last-frame chaining: generate shot 1, then use its last frame as shot 2’s first frame.
- More anchors: Replicate notes a “reference-to-video” capability that can combine up to three reference images into a coherent scene guided by text (https://replicate.com/blog/veo-3-1). Use that when start/end isn’t enough.
Checklist
- Start + end images pass the overlay test (same crop, horizon, key points aligned)
- Perspective matches (no wide vs tele mismatch; consistent camera height)
- Lighting direction/softness matches (avoid soft→hard swaps)
- Prompt includes a hard SUBJECT LOCK (same person/product throughout)
- Prompt defines one transition motion (spin/wipe/whip pan) and forbids extra actions
- Prompt forbids edits: “single continuous shot, no cuts”
- Camera constraints are explicit (locked camera; no zoom/pan/tilt unless intentional)
- Audio is minimal and timed (one SFX + optional short line)
- Subject is recognizable in the first second; you trim before degradation
FAQ
How do I use Veo 3.1 first and last frame for a smooth transition?
Upload a start and end image, then describe one continuous motion that connects them while locking camera and identity. Don’t just label endpoints—direct the bridge.
How do I stop random zooms, pans, or cuts?
Add explicit negatives: “locked camera, tripod-stable, no zoom/pan/tilt” and “single continuous shot, no cuts.” Prompts can specify camera motion (https://deepmind.google/models/veo/prompt-guide/), so you must specify stillness when you want a match cut.
What prompt format works best for outfit changes?
Use the 5-line template: Subject lock → Start frame match → Transition motion → End frame match → Audio. It reduces mid-clip invention by forcing a single job at a time.
Why doesn’t my last frame match even when I provide it?
Usually the two images don’t align (different crop/perspective), or you didn’t demand an exact arrival. Re-crop via the overlay test, then restate “arrive at the provided last frame exactly, same framing.”
Can Veo generate audio that matches the transition beat?
DeepMind notes Veo can generate dialogue and include specific lines (https://deepmind.google/models/veo/prompt-guide/). Google Cloud describes rich synchronous audio for Veo 3.1 (https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1). Keep it simple: one SFX timed to the change.
How can I batch-generate variations efficiently?
Standardize bookends and swap only the end-state line (outfit colorway, product SKU, styled vs unstyled). Veo3Gen offers a developer API for programmatic generation, which is ideal for batching.
Ship your first bookend transition (closing CTA)
Build a small library of aligned start/end frames (your “plates”), then reuse the 5-line template for every variation. If you want an affordable way to access Google’s Veo 3.1 models without enterprise pricing, Veo3Gen offers three modes—Veo 3.1 Fast (quick default), Quality (max fidelity), and Lite (cheapest preview)—plus first-and-last-frame control, text-to-video and image-to-video, and native synchronized audio in one pass. Start with the free credits, then scale with pay-as-you-go credits (non-expiring) or an optional monthly plan when you’re ready.
Start creating with Veo3Gen
Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.
- Generate your first video now: Get started
- Compare plans and pay-as-you-go pricing: See pricing
Try Veo 3 & Veo 3 API for Free
Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.