AI Video11 min read

Kling's "Subject + Movement + Scene" Prompt Structure → A Veo3Gen Shot Brief Template (With 10 Copy-Paste Examples)

Turn Kling’s Subject→Movement→Scene formula into a Veo3Gen shot-brief template, with a worked rewrite, troubleshooting table, checklist, and 10 examples.

On this page

TL;DR

Use Kling’s prompt spine—Subject → Movement → Scene—as a shot brief inside Veo3Gen, then optionally add Camera Language + Lighting + Atmosphere for control. The practical trick is a motion budget: one subject action + (optional) one camera move per clip so the model doesn’t improvise extra actions, props, or framing.

Key takeaways

  • Kling’s reference formula is: Subject (description) + Subject Movement + Scene (description) + (Camera Language + Lighting + Atmosphere). Treat it as a reusable shot brief. (https://kling.ai/quickstart/text-to-video-prompt-guide)
  • Keep movement straightforward for short clips; the guide explicitly recommends movement suited for a ~5-second video—i.e., one clear beat. (https://kling.ai/quickstart/text-to-video-prompt-guide)
  • Separate subject motion from camera motion. If both are vague, you’ll get random framing.
  • Prevent “random props” by constraining the scene: location + 1–3 props + what must not appear.
  • Iterate like a producer: change one line at a time (Subject or Movement or Scene).

Why “Subject → Movement → Scene” works (even outside Kling)

Kling’s text-to-video guide states that the prompt directly dictates the content of the video and offers a clean reference formula: Subject + Subject Movement + Scene + optional Camera/Lighting/Atmosphere. (https://kling.ai/quickstart/text-to-video-prompt-guide)

This structure fixes the most common failure in creator prompts: they mix identity, action, environment, and style into one paragraph. When that happens, the model has too many degrees of freedom.

Think of it as three production questions:

  1. Subject: What is the hero?
  2. Movement: What changes over time?
  3. Scene: What stays stable around the hero?

When you answer those separately, you get outputs that are easier to troubleshoot—because you know which line to edit.

The Veo3Gen Shot Brief Template (copy-paste)

Kling’s formula is already close to a shot brief; below is the same idea, formatted so you can reuse it across ads, UGC, and social cutdowns. (https://kling.ai/quickstart/text-to-video-prompt-guide)

Mid-article CTA: If you want to turn these shot briefs into repeatable production, Veo3Gen gives you an affordable way to access Google’s Veo 3.1 models (without Google’s enterprise pricing) and includes a developer API for programmatic generation when you’re ready to batch variants.

Copy-paste template

SUBJECT (identity anchors):

  • Hero:
  • 2–4 anchors (type, material/color, one distinctive detail, framing anchor):

MOVEMENT (motion budget):

  • Subject action (one verb + end state):
  • Camera motion (static OR one move):

SCENE (stable constraints):

  • Location + foreground/background:
  • Props (limit 1–3):
  • Constraints (continuity + exclusions):

OPTIONAL ADD-ONS (only after the core works):

  • Camera language (shot type / composition):
  • Lighting (one clear source):
  • Atmosphere/mood:

A filled-in mini example (so you can see the “shape”)

  • SUBJECT: matte black insulated bottle, subtle logo, centered on light wood table
  • MOVEMENT: bottle slides in from left and stops center; camera static
  • SCENE: bright kitchen softly blurred; only one bottle on table; single continuous shot; no text
  • ADD-ONS: medium close-up; soft morning daylight; clean minimal mood

Step 1 — Subject: lock identity with 2–4 anchors

Kling defines the subject as the main focus (people, animals, plants, objects). (https://kling.ai/quickstart/text-to-video-prompt-guide)

Use 2–4 anchors. More words aren’t more stable—clear anchors are.

Good identity anchors

  • Type: “ceramic mug”, “hand model”, “skincare serum bottle”
  • Material + color: “frosted glass”, “brushed aluminum”, “matte black”
  • Distinctive detail: “silver dropper cap”, “rounded square lid”, “small logo on handle”
  • Framing anchor: “center frame”, “waist-up”, “top-down on tabletop”

Avoid

  • Long stacks of adjectives (“gorgeous, stunning, cinematic…”)—they don’t define the subject.
  • Sneaking actions into the subject line (“holding, pouring, walking…”)—keep action in Movement.

Step 2 — Movement: one clear beat (and name the camera move)

Kling’s guide explicitly calls out that subject movement should be straightforward and suitable for short duration (e.g., ~5 seconds). (https://kling.ai/quickstart/text-to-video-prompt-guide)

The motion budget rule

For a single clip, budget:

  • One subject action (one primary verb)
  • Optional: one camera move (or keep it static)

Write subject action like a storyboard beat

Format: verb + end state

  • “hand rotates the bottle 30 degrees, then stops
  • “ribbon loosens and falls flat on the box”
  • “wrist turns to reveal the watch face, then holds

If you catch yourself writing “and then” twice, you’ve described multiple beats.

Choose exactly one camera behavior

Examples:

  • “static tripod”
  • “slow push-in”
  • “slow pan left-to-right”

Don’t ask for orbit + zoom + rack focus unless you’re willing to accept unpredictable framing.

Step 3 — Scene: specify environment + what must NOT change

Kling defines the scene as the environment, including foreground, background, and other elements. (https://kling.ai/quickstart/text-to-video-prompt-guide)

A reliable scene line has three parts:

  1. Where (location)
  2. What’s in it (1–3 props)
  3. Constraints (continuity + exclusions)

Scene constraints that reduce “random additions”

  • “single continuous shot, no cuts”
  • “only the listed props”
  • “clean background, no extra objects”
  • “no text, no signage, no labels”

If you need text, write the exact string

Luma’s best practices note you can request text by specifying it (example: a poster with text that reads “Dream Machine”). (https://lumalabs.ai/learning-hub/best-practices)

Actionable rule: if you want text, include:

  • the exact words
  • where it appears (poster/label/sign)
  • keep it short

Optional add-ons: Camera Language, Lighting, Atmosphere

Kling positions Camera Language + Lighting + Atmosphere as optional add-ons. Keep them optional until your core shot is stable. (https://kling.ai/quickstart/text-to-video-prompt-guide)

Camera language (use shot terms, not “cinematic”)

  • “macro close-up on product texture”
  • “medium shot, eye-level, centered composition”
  • “top-down tabletop shot”

Lighting (one clear source beats five adjectives)

  • “soft window daylight from camera-left”
  • “hard rim light, dark background”

Atmosphere/mood (pick one lane)

  • clean/minimal (DTC)
  • cozy/warm (food/home)
  • energetic (fitness)

Worked example: messy prompt → predictable shot brief

Here’s a true before/after using the exact structure.

Before (too many degrees of freedom)

“Cinematic video of a skincare serum, luxury vibes, beautiful bathroom, sparkling water, model picks it up and smiles, camera moves around, dramatic lighting, text on screen saying ‘Glow Now’.”

Why it fails

  • Competing subjects (product + model) with no identity anchors
  • Multiple actions (pick up, smile, text overlay)
  • Vague camera instruction (“moves around”)
  • Scene invites extra props (“beautiful bathroom”, “sparkling water”)

After (shot brief)

SUBJECT:

  • Frosted-glass skincare serum bottle with a silver dropper cap, centered on a clean marble countertop.

MOVEMENT (motion budget):

  • Subject action: a hand enters from the right and rotates the bottle 30 degrees, then stops.
  • Camera motion: slow push-in.

SCENE (stable constraints):

  • Minimal modern bathroom background, softly blurred.
  • Only the serum bottle and the hand; no extra objects on the counter.
  • Single continuous shot, no cuts. No on-screen text.

OPTIONAL ADD-ONS:

  • Camera language: product close-up, eye-level, centered composition.
  • Lighting: soft diffused daylight.
  • Atmosphere: clean, premium, calm.

10 copy-paste Veo3Gen shot briefs (ads, creators, socials)

Each example follows Kling’s structure: Subject + Movement + Scene + optional Camera/Lighting/Atmosphere. (https://kling.ai/quickstart/text-to-video-prompt-guide)

1) DTC product hero (tabletop reveal)

  • SUBJECT: matte black insulated water bottle, subtle logo, centered on light wood table
  • MOVEMENT: bottle slides in from left and stops center; camera static
  • SCENE: bright kitchen softly blurred; only one bottle; single continuous shot, no cuts; no text
  • ADD-ONS: medium close-up; soft morning daylight; crisp minimal mood

2) Food pour (one satisfying action)

  • SUBJECT: clear glass with ice cubes on a coaster, centered
  • MOVEMENT: amber soda pours into the glass in a steady stream; slow push-in
  • SCENE: simple countertop; neutral background; no hands visible; single continuous shot; no text
  • ADD-ONS: close-up; bright backlight; refreshing mood

3) Phone “tap” (simple UI moment)

  • SUBJECT: modern smartphone held in one hand, screen facing camera
  • MOVEMENT: thumb taps once near center of screen; camera static
  • SCENE: neutral desk setup softly blurred; no extra devices; single continuous shot
  • ADD-ONS: medium close-up; soft diffused light; clean tech mood

4) UGC creator intro (product raise)

  • SUBJECT: creator, waist-up, facing camera; product in their hand
  • MOVEMENT: raises product into frame and holds it still; gentle handheld sway
  • SCENE: home office background softly blurred; no other people; single continuous shot
  • ADD-ONS: medium shot; warm indoor lighting; casual mood

5) Beauty texture macro (swatch)

  • SUBJECT: fingertip with glossy lip oil swatch
  • MOVEMENT: fingertip slowly turns to catch light, then stops; camera static
  • SCENE: clean neutral background; no extra objects; single continuous shot
  • ADD-ONS: macro close-up; soft key light; premium mood

6) Fitness gear (lace tighten)

  • SUBJECT: training shoes on gym floor, centered
  • MOVEMENT: hands pull laces tight once, ends taut; slow push-in
  • SCENE: gym background blurred; no other people visible; single continuous shot
  • ADD-ONS: close-up; hard overhead lighting; energetic mood

7) Packaging ASMR (ribbon loosen)

  • SUBJECT: small rigid gift box with ribbon, centered on table
  • MOVEMENT: ribbon is pulled and loosens; camera static
  • SCENE: minimal tabletop; only the box; single continuous shot; no confetti; no text
  • ADD-ONS: top-down shot; soft warm light; cozy mood

8) B2B “planning” (sticky note place)

  • SUBJECT: blank sticky notes on a wall, aligned
  • MOVEMENT: a hand places one sticky note and presses it flat; slow pan right
  • SCENE: simple office wall; no logos; single continuous shot
  • ADD-ONS: medium close-up; soft office lighting; focused mood

9) Pet product context (kibble pour, no pet)

  • SUBJECT: stainless steel dog bowl on kitchen floor
  • MOVEMENT: kibble pours into the bowl in a steady stream; camera static
  • SCENE: clean floor; no pets visible; single continuous shot; no text
  • ADD-ONS: close-up; bright natural light; wholesome mood

10) Fashion accessory detail (watch turn)

  • SUBJECT: stainless steel watch on a wrist, centered
  • MOVEMENT: wrist turns slowly to show the watch face, then stops; slow push-in
  • SCENE: neutral background; no extra jewelry; single continuous shot
  • ADD-ONS: close-up; soft directional key light; refined mood

Common failure modes (diagnose → fix one line)

Edit the line that matches the symptom—don’t rewrite everything.

Symptom Likely cause One-line fix
Random props/clutter appear Scene under-specified SCENE: “Only the listed props; clean surface; no extra objects.”
Subject changes (shape/wardrobe) Not enough identity anchors SUBJECT: add material + color + one distinctive detail.
Feels like multiple shots No continuity constraint SCENE: “Single continuous shot, no cuts, no transitions.”
Everything moves Too many actions/camera moves MOVEMENT: “One action only; camera static.”
Framing drifts off hero Vague camera language ADD-ONS: “Centered composition; camera stays locked on subject.”
Unwanted text/signage Scene invites labels SCENE: “No text, no labels, no signage.”

A tight testing loop (so you don’t thrash prompts)

  1. Base pass: write only Subject + Movement + Scene.
  2. Stabilize: if it’s wrong, fix one line (Subject or Movement or Scene).
  3. Only then add one optional layer (camera language or lighting or atmosphere), matching Kling’s “optional add-ons” framing. (https://kling.ai/quickstart/text-to-video-prompt-guide)

If you’re building lots of variants for social, Veo3Gen supports text-to-video and image-to-video, plus first-and-last-frame control on Veo 3.1 when you need tighter continuity. It also supports 720p, 1080p, and 4K output (4K on Veo 3.1 Fast/Quality) with 16:9 and 9:16 aspect ratios.

Checklist

  • Write SUBJECT with 2–4 identity anchors (type, material/color, distinctive detail, framing).
  • Apply a motion budget: one subject action + one camera move (or none).
  • Write MOVEMENT as verb + end state (one beat).
  • Constrain SCENE: location + 1–3 props + exclusions (“only listed props”, “no extra objects”).
  • Add “single continuous shot, no cuts” when you need ad-ready continuity.
  • Add one add-on at a time (camera language or lighting or atmosphere).
  • Iterate by changing one line only per test.

FAQ

How do I write an AI video prompt structure that doesn’t drift?

Use Subject → Movement → Scene, then optionally add Camera Language + Lighting + Atmosphere—this mirrors Kling’s reference prompt formula. (https://kling.ai/quickstart/text-to-video-prompt-guide)

How do I stop random objects from appearing?

Strengthen the Scene: name the location, limit props to 1–3, and add exclusions like “only the listed props” and “no extra objects.”

How do I keep movement clean in a short clip?

Follow Kling’s guidance to keep movement straightforward for short duration: write one subject action with a clear end state, and avoid chaining multiple beats. (https://kling.ai/quickstart/text-to-video-prompt-guide)

How do I stop weird camera moves?

Name exactly one camera behavior: “static tripod” or one move (push-in/pan). If you don’t specify, the model may invent one.

How can I scale ad variants once I have a stable shot brief?

Batch variations by changing only one field (e.g., lighting or scene constraints). In Veo3Gen, you can also automate generations programmatically using the developer API.

Ready to turn these shot briefs into production?

Veo3Gen is an affordable way to access Google’s Veo 3.1 video models without Google’s enterprise pricing. It offers three modesVeo 3.1 Fast (quick default), Veo 3.1 Quality (max fidelity), and Veo 3.1 Lite (cheapest preview)—and generates native, synchronized audio (dialogue, SFX, music) in a single pass.

Start with the free credits for new users, lock your best-performing shot brief, then scale into variants via pay-as-you-go credits (purchased credits don’t expire) or an optional monthly plan.

Start creating with Veo3Gen

Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.

Sources

Limited Time Offer

Try Veo 3 & Veo 3 API for Free

Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.