Creator Workflow9 min read
Vertical Veo 3.1 Videos in Flow (2026): A Creator Workflow for Reels/TikToks That Don't Crop Weird
A vertical-first Veo 3.1 workflow for Flow that prevents weird crops on Reels/TikTok—safe zones, shot cards, and a worked prompt rewrite.
On this page
- TL;DR
- Key takeaways
- Why vertical-first beats “generate wide, crop later”
- Decision rule: native 9:16 vs 16:9 for flexibility
- What Flow / Veo 3.1 supports for vertical (grounded summary)
- The 9:16 framing system (the part most prompts skip)
- 1) Safe zones: think “bands + column”
- 2) Headroom + shot size (the two most common vertical failures)
- 3) Motion rule: micro-motion, not lateral motion
- Copy-paste template: the Vertical Shot Card (use this every time)
- Worked example: fix “cropped head + missing product” with a single rewrite
- Before (sounds fine, fails often)
- After (same concept, shot-card constrained)
- Build vertical as 3 clips (Hook → Proof → CTA)
- Clip 1: Hook (0–2 seconds)
- Clip 2: Proof (2–7 seconds)
- Clip 3: CTA (7–12 seconds)
- Mid-article CTA: iterate vertical variants without enterprise pricing
- Common vertical failures (and the specific fix to add)
- Failure: forehead clipped / head cut off
- Failure: hands disappear during the demo
- Failure: product lands where captions go
- Failure: subject drifts to the edge
- Failure: inconsistent character/object across clips
- Checklist
- FAQ
- How do I generate true 9:16 vertical AI video with Veo 3.1?
- How do I stop Flow outputs from cropping off heads in Reels?
- How do I keep the product visible when someone gestures?
- How do I get consistent characters/objects across multiple vertical clips?
- Does Veo generate audio natively or do I need a separate audio pass?
- Closing CTA: ship vertical variants faster with a consistent shot card
- Start creating with Veo3Gen
TL;DR
Generate native 9:16 instead of “wide then crop” whenever faces, hands, or a product must stay on-screen. Most “weird crop” problems are really composition + safe-zone problems you can solve by specifying shot size, headroom, subject position, and reserved caption space.
Use a repeatable workflow:
- Lock framing with a Vertical Shot Card prompt.
- Keep the subject inside a center-safe column.
- Reserve top/bottom bands for platform UI + captions.
- Add any critical on-screen text in post (don’t rely on in-model text).
Key takeaways
- Prefer native 9:16 when the story depends on faces/hands/product staying visible; crop-later is mainly for scenic B-roll.
- Prompt for shot size + headroom + subject position + caption space (these four remove most “cropped weird” failures).
- If you need consistency across clips, use reference/ingredient images—Veo 3.1 Ingredients to Video is explicitly designed to create videos from reference images and supports portrait mode. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
- Build vertical content as 3 clips (Hook → Proof → CTA) with consistent framing; add captions/graphics in edit.
Why vertical-first beats “generate wide, crop later”
“Weird crop” usually comes from one of these:
- The model composed for 16:9, not 9:16.
- The subject moves side-to-side, so a vertical crop can’t follow cleanly.
- Key details (hands, product, face) live near the edges.
When you crop a wide frame into 9:16, you’re forcing a new composition the model never planned for. Vertical-first generation flips that: the model composes for portrait from the start.
Google explicitly states that Veo 3.1 Ingredients to Video supports vertical video generation (portrait mode). (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
Decision rule: native 9:16 vs 16:9 for flexibility
Use this in production:
Generate native 9:16 when you care about:
- Talking-head hooks
- UGC testimonials
- Product-in-hand demos (hands must stay visible)
- Anything caption-heavy
Generate 16:9 only when:
- It’s scenic/establishing B-roll
- You truly need multiple crops (9:16 + 1:1 + 16:9)
- The subject is centered with lots of negative space
What Flow / Veo 3.1 supports for vertical (grounded summary)
Here’s what’s supported based on Google’s public notes:
- Flow is an AI filmmaking tool powered by Veo. (https://blog.google/innovation-and-ai/products/veo-updates-flow/)
- Veo 3.1 Ingredients to Video lets you create videos based on reference (ingredient) images, and Google highlights improved character identity consistency and improved background/object consistency. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
- The same post states portrait mode (vertical generation) for Ingredients to Video. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
- Google’s Flow notes say Flow is powered by various Veo models (plus Gemini models), and Flow will notify you if you select a feature that isn’t supported by Veo 3.1. (https://support.google.com/flow/answer/16352836?hl=en)
- Veo can generate audio natively (dialogue, SFX, ambient noise) as part of the generation. (https://deepmind.google/models/veo/)
Practical implication: when a control isn’t behaving, don’t guess—Flow may warn you that the feature isn’t supported for the chosen model. (https://support.google.com/flow/answer/16352836?hl=en)
The 9:16 framing system (the part most prompts skip)
Vertical success comes from giving the model layout constraints—the same constraints a human DP uses.
1) Safe zones: think “bands + column”
You don’t need pixel math to get most of the benefit. Use three zones:
- Top reserved band: keep foreheads, hats, and raised hands out of the extreme top.
- Bottom reserved band: reserve space for captions and any overlay you’ll add in edit.
- Center-safe column: keep face/product primarily in the middle so minor reframing doesn’t cut essentials.
Prompt language that maps to this:
- “Subject centered”
- “Leave generous empty space at the bottom for captions”
- “Avoid important details near the very top”
2) Headroom + shot size (the two most common vertical failures)
If you only change two things in your prompts, change these:
- Shot size: medium / medium close-up / medium-wide
- Headroom rule: “comfortable headroom” + “eyes on the upper third”
This prevents:
- accidental extreme close-ups (forehead clipped)
- hands/product falling out of frame
3) Motion rule: micro-motion, not lateral motion
Centered doesn’t have to mean static. The trick is to allow motion that doesn’t break framing:
- Prefer slow push-in over pans.
- If handheld: “subtle handheld feel” but no drifting off-center.
- Keep actions at chest height (especially product demos).
Copy-paste template: the Vertical Shot Card (use this every time)
Treat this like a shot spec.
9:16 vertical video, social media reel. SHOT SIZE: [medium / medium close-up / medium-wide]. Subject: [WHO/WHAT] in [LOCATION]. Composition: subject centered, eyes on upper third, comfortable headroom. Hands/Key object: keep [HANDS/PRODUCT] fully visible. Safe zones: leave generous empty space at the bottom for captions; avoid key details near the very top. Action (one sentence): [WHAT HAPPENS]. Camera: [static tripod / slow push-in], minimal lateral movement. Style/lighting: [UGC realistic / cinematic], [lighting]. Audio: [dialogue/SFX/ambient] (optional). Do not include on-screen text (captions added in post).
Worked example: fix “cropped head + missing product” with a single rewrite
The fastest way to improve outputs is to rewrite prompts so they force framing.
Before (sounds fine, fails often)
“A woman talks about a new vitamin drink while holding the bottle, vertical video, realistic.”
Typical failure pattern:
- the model goes too tight → forehead clipped
- bottle drops into the bottom area where captions will sit
- hands drift out of frame during gestures
After (same concept, shot-card constrained)
9:16 vertical video, social media reel. Medium shot. Subject: a woman speaking to camera in a bright home kitchen, holding a vitamin drink bottle. Composition: subject centered, eyes on upper third, comfortable headroom; both hands and the full bottle visible at all times. Safe zones: leave generous empty space at the bottom for captions; avoid key details near the very top. Blocking: bottle held at chest height, label facing camera. Action: she delivers one enthusiastic line, then lifts the bottle slightly and takes a small sip. Camera: static tripod, no lateral movement. Style/lighting: natural window light, realistic UGC. Audio: clean dialogue + subtle sip SFX. No on-screen text.
What changed (the actual levers):
- Medium shot + headroom prevents accidental close-ups.
- Chest-height rule keeps the product out of the caption band.
- No lateral movement reduces edge drift that makes vertical feel “auto-cropped.”
Build vertical as 3 clips (Hook → Proof → CTA)
Don’t try to generate “one perfect 12-second take.” Generate three short shots with the same framing rules.
Clip 1: Hook (0–2 seconds)
Goal: earn the pause.
- One subject, one action.
- Tighter shot than the rest (but keep headroom).
- Bottom space reserved for the first caption line.
Prompt move: specify one motion only (e.g., “raises product into frame and smiles”).
Clip 2: Proof (2–7 seconds)
Goal: show the claim.
- Demonstrate with hands.
- Favor a medium shot so hands + object stay visible.
- If you need consistency across multiple proof clips, use reference/ingredient images—Veo 3.1 Ingredients to Video is designed for creating videos from reference images and highlights improved consistency. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
Prompt move: “hands fully visible the entire time” + “label facing camera.”
Clip 3: CTA (7–12 seconds)
Goal: one clear next step.
- Return to talking head for clarity.
- Keep gestures contained so captions remain readable.
Prompt move: “static tripod” + “minimal lateral movement.”
Mid-article CTA: iterate vertical variants without enterprise pricing
If you want to run this exact vertical shot-card workflow outside Google’s enterprise pricing, Veo3Gen is an affordable way to access Google’s Veo 3.1 video models. It supports text-to-video and image-to-video, 9:16 and 16:9, and generations include native, synchronized audio in a single pass (dialogue/SFX/music)—no separate audio step. New users get free credits, and there’s a developer API for programmatic generation.
Common vertical failures (and the specific fix to add)
Failure: forehead clipped / head cut off
Add:
- “medium close-up” (or wider)
- “comfortable headroom”
- “eyes on upper third”
Failure: hands disappear during the demo
Add:
- “medium shot”
- “hands fully visible the entire time”
- “action kept at chest height”
Failure: product lands where captions go
Add:
- “leave generous empty space at the bottom for captions”
- “product held at chest height”
Failure: subject drifts to the edge
Add:
- “subject centered”
- “minimal lateral movement”
- “static tripod” or “slow push-in only”
Failure: inconsistent character/object across clips
Use image-to-video with ingredient images; Google positions Ingredients to Video as a reference-image workflow with improved identity and object/setting consistency, and it supports portrait mode. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
Checklist
- Decide: native 9:16 vs 16:9 then crop (use the decision rule).
- Choose shot size first (medium / medium close-up / medium-wide).
- Add comfortable headroom + eyes on upper third for talking heads.
- Reserve a bottom caption band (“leave generous empty space at the bottom”).
- If demoing, require hands + key object fully visible throughout.
- Constrain motion: slow push-in / micro-motion, avoid pans.
- Ask for no on-screen text; add type/captions in post.
FAQ
How do I generate true 9:16 vertical AI video with Veo 3.1?
State “9:16 vertical video” and include composition constraints (shot size, headroom, caption space). Veo 3.1 Ingredients to Video supports portrait mode. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
How do I stop Flow outputs from cropping off heads in Reels?
Go wider than you think: “medium close-up,” “comfortable headroom,” “eyes on upper third,” and keep the subject centered with minimal lateral movement.
How do I keep the product visible when someone gestures?
Use a medium shot, require “hands fully visible the entire time,” and constrain the action to chest height so it doesn’t collide with captions.
How do I get consistent characters/objects across multiple vertical clips?
Use reference/ingredient images with Ingredients to Video; Google describes this workflow as creating videos based on reference images with improved identity and object/setting consistency, and it supports portrait mode. (https://blog.google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/)
Does Veo generate audio natively or do I need a separate audio pass?
Veo can generate sound effects, ambient noise, and dialogue natively as part of creation. (https://deepmind.google/models/veo/)
Closing CTA: ship vertical variants faster with a consistent shot card
If you adopt two habits—native 9:16 generation and a reusable Vertical Shot Card—you’ll eliminate most unusable takes caused by cropping and drifting composition.
When you’re ready to turn that into repeatable output, generate in Veo3Gen: pick Veo 3.1 Fast/Quality/Lite depending on speed vs fidelity, keep audio native in one pass, and scale variants via the developer API. Pricing is pay-as-you-go credits with optional monthly plans, and purchased credits do not expire. New users can start with free credits.
Start creating with Veo3Gen
Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.
- Generate your first video now: Get started
- Compare plans and pay-as-you-go pricing: See pricing
Try Veo 3 & Veo 3 API for Free
Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.