AI Video Prompting8 min read

JSON Prompting for Repeatable AI Video Ads (2026): The "Shot Config" Template Creators Can Reuse in Veo3Gen (Even Without a JSON UI)

A reusable “Shot Config” JSON-style template for repeatable AI video ads in Veo3Gen—plus a worked example, debugging table, checklist, and FAQ.

TL;DR

JSON prompting isn’t “writing curly braces.” It’s writing your prompt like a config so you can lock what must stay consistent (subject, camera, lighting, style, constraints) and only change what you’re testing (hook, offer, action, pacing). This post gives you a copy/paste “Shot Config” template you can use in Veo3Gen as plain text to make AI video ads more repeatable—even without a JSON UI.

Key takeaways

  • Treat your prompt like a structured spec: lock SUBJECT / ENVIRONMENT / CAMERA / LIGHTING / STYLE / CONSTRAINTS, then iterate one field at a time.
  • Use a campaign “Core Config” so wardrobe, location, camera language, and grading don’t drift between variants.
  • Put negative prompts in a dedicated CONSTRAINTS block so they’re consistent across every version (negative prompts help reduce issues like blurry motion and harsh lighting) (https://www.designyourway.net/blog/ai-video-prompting-made-simple-tips-for-better-results).
  • Debug faster: random zooms are usually CAMERA, outfit drift is SUBJECT, unreadable product/UI is ACTION + PACING + CAMERA.
  • Build a small, versioned library of Shot Configs so you can ship next week’s ads without “re-briefing” the model from scratch.

Why structured prompting beats rewriting prompts from scratch

If you publish ads, reels, or promos weekly, your main problem isn’t “lack of creativity.” It’s prompt drift:

  • you change the offer and the model changes the set
  • you change the hook and the camera becomes chaotic
  • you tweak a vibe adjective and your brand look disappears

Structured prompting solves the operational side of creative.

The VidAU comparison frames JSON prompting as critical for repeatable pipelines, multi-shot sequences, and tool-integrated workflows (including custom orchestration layers) (https://www.vidau.ai/json-prompting-ai-video-comparison). Even if your tool doesn’t expose literal JSON fields, you can still get most of the benefit by writing your prompt in consistent sections.

What you gain immediately:

  1. Faster iteration: you edit one block instead of rewriting a paragraph.
  2. Cleaner QA: when something breaks, you know where to fix it.

What “JSON prompting” means in practice (for ad creators)

VidAU defines JSON prompting as a way to separate scene description from camera motion (https://www.vidau.ai/json-prompting-ai-video-comparison). That single separation is the backbone of ad consistency.

VidAU also lists typical structured elements used in Veo-style inputs: scene description, subject descriptors, camera parameters, lighting model, motion intensity, duration, and style constraints (https://www.vidau.ai/json-prompting-ai-video-comparison). That’s basically a producer’s shot brief—just more explicit.

And while you may not get a raw JSON panel in every UI, VidAU notes Veo’s internal architecture is built around structured input blocks even if public UI access doesn’t expose them directly (https://www.vidau.ai/json-prompting-ai-video-comparison). The practical takeaway: clear sections map well to how these models respond.

The “five layers” that keep showing up

DesignYourWay reduces good prompts to five layers: subject, environment, action, camera, style (https://www.designyourway.net/blog/ai-video-prompting-made-simple-tips-for-better-results). We’ll keep those, then add ad-specific fields you’ll actually use in production: objective, pacing, text-safe area, constraints, audio notes, output.

The Veo3Gen “Shot Config” template (copy/paste)

Use this as plain text. Don’t obsess over syntax—obsess over stable headings.

Veo3Gen context you can plan around: it provides access to Google’s Veo 3.1 video models with three modes—Veo 3.1 Fast, Veo 3.1 Quality, and Veo 3.1 Lite (cheapest/preview). Generations include native synchronized audio (dialogue, SFX, music) in a single pass. It supports text-to-video and image-to-video, plus first-and-last-frame control on Veo 3.1. Supported resolutions include 720p/1080p/4K (4K on Fast/Quality), aspect ratios 16:9 and 9:16. Pricing is pay-as-you-go credits with optional monthly plans, and purchased credits do not expire; new users get free credits; there’s also a developer API (Veo3Gen facts).

Shot Config (Template)

SHOT CONFIG v1.0

OBJECTIVE:
- What this clip must achieve (e.g., “stop-scroll hook + show benefit in 8s”).
- Viewer + platform (e.g., “TikTok 9:16, cold audience”).

SUBJECT (LOCKS):
- Hero: who/what must stay the same across versions.
- Fixed descriptors: wardrobe, hair, props, brand colors, defining features.
- Continuity locks: “same outfit throughout,” “same product color,” “label always facing camera.”

ENVIRONMENT (LOCKS):
- Location + time of day.
- 2–3 anchors that must not change (mirror type, shelf, table material).
- Brand safety notes (no logos/signage except ours).

ACTION (VARIABLE):
- 1–3 observable beats in order (no abstract verbs).

CAMERA (LOCKS):
- Shot size (CU/MS/WS) + camera height.
- Movement: pick ONE (locked-off / slow push-in / slow lateral slide / gentle orbit).
- Framing rules: where the product must live in frame.

LIGHTING (LOCKS):
- Lighting style (soft daylight, studio high-key, neon night).
- Contrast intent (e.g., “soft shadows, not flat”).

STYLE (LOCKS):
- Ad style (UGC / commercial / cinematic / minimal).
- Texture + grade notes.

PACING (VARIABLE):
- Duration target.
- At least one explicit hold (e.g., “hold on label 0.7s”).

TEXT/OVERLAYS SAFE AREA (LOCKS):
- Reserve region for captions/offers (e.g., “bottom 20% clear”).

CONSTRAINTS (LOCKS):
- Negative prompts / avoid list (distorted hands, unreadable text, random zooms, harsh flat lighting, etc.).

AUDIO NOTES (VARIABLE):
- Audio goal: VO tone, SFX, music vibe.
- Any required words (if dialogue) + when they occur.

OUTPUT NOTES:
- Aspect ratio: 16:9 or 9:16.
- Resolution: 720p / 1080p / 4K (where supported).
- If using first/last frame control: describe first frame + last frame.

CTA (mid-article): If you want to keep your ad workflow consolidated, Veo3Gen can generate video plus synchronized audio in one pass and lets you choose Fast/Quality/Lite depending on whether you’re previewing or pushing fidelity (Veo3Gen facts).

Worked example: from “good but drifty” to repeatable

Below is a concrete before/after you can reuse. The goal is not better prose—it’s controlled variance.

Before (one-off paragraph)

Make a cool video ad of a woman showing a skincare serum in a bright bathroom. Cinematic lighting, smooth camera, premium vibe. Add text: “Glass Skin in 7 Days”.

Typical failure modes:

  • “premium vibe” mutates into different aesthetics every run
  • camera invents zooms
  • bathroom and wardrobe change
  • product label becomes unreadable
  • text placement collides with face/product

After (repeatable Shot Config)

SHOT CONFIG v1.0

OBJECTIVE:
- 9:16 paid social ad, 8–10s: demonstrate serum + end on clear label hero moment.

SUBJECT (LOCKS):
- One woman, late 20s–early 30s, natural makeup.
- Wardrobe: white robe. Hair: neat bun.
- Product: clear glass dropper bottle, minimal white label.
- Continuity locks: same robe + bun; bottle label always facing camera.

ENVIRONMENT (LOCKS):
- Bright modern bathroom, white tile.
- Anchors: round mirror + oak floating shelf + white countertop.
- No visible brand logos or signage.

ACTION (VARIABLE):
- Beat 1: hold bottle close, label-forward.
- Beat 2: apply 2–3 drops to cheekbones.
- Beat 3: final hero hold of bottle, label-forward.

CAMERA (LOCKS):
- Medium close-up, eye level.
- Movement: slow push-in only (no zooms).
- Framing: center face + bottle; keep bottom 20% clear.

LIGHTING (LOCKS):
- Soft daylight key, gentle contrast, highlights on glass.

STYLE (LOCKS):
- Premium commercial-UGC hybrid; natural skin texture; clean neutral grade.

PACING (VARIABLE):
- Single take feel.
- Hold on label for 0.7s at the end.

TEXT/OVERLAYS SAFE AREA (LOCKS):
- Bottom 20% clear for captions/offer.

CONSTRAINTS (LOCKS):
- Avoid: distorted hands/fingers, unreadable label, random zooms, jittery camera, harsh flat lighting.

AUDIO NOTES (VARIABLE):
- Soft modern beat + subtle glass/dispense SFX.
- Optional whisper VO: “Two drops. All glow.”

OUTPUT NOTES:
- 9:16, 1080p.

Why this works:

Workflow: generate 10 ad variants without losing the “look”

Use a two-file approach: a locked core and a variants sheet.

Step-by-step iteration loop (one variable at a time)

  1. Create your CORE (do not change): SUBJECT, ENVIRONMENT, CAMERA, LIGHTING, STYLE, TEXT SAFE AREA, CONSTRAINTS.
  2. Generate Baseline A (save the exact text you used).
  3. Pick one variable to test (usually ACTION first).
  4. Duplicate and produce A1–A5 where only that field changes.
  5. Promote the winner into the core (or label as a winning variant).
  6. Repeat with the next variable (PACING → AUDIO NOTES → STYLE only if necessary).

This matches why structured prompting matters: building repeatable pipelines and shot-level consistency (https://www.vidau.ai/json-prompting-ai-video-comparison).

When to switch modes (preview vs final)

Because Veo3Gen offers Fast / Quality / Lite modes (Veo3Gen facts), you can align mode choice to your iteration phase:

  • Lite: cheapest preview when you’re still deciding ACTION beats
  • Fast: default when you want quick, solid iterations
  • Quality: final polish when the config is already stable

Debugging map: symptom → field → specific fix

Stop rewriting the whole prompt. Fix the block that governs the failure.

Symptom Likely cause Fix this field first Specific change
Random zooms / surprise camera moves camera language isn’t constrained CAMERA + CONSTRAINTS Choose ONE movement (“slow push-in only”); add “no zooms”
Outfit/hair drift between variants subject not locked SUBJECT Add wardrobe/hair locks + “same throughout”
Location changes each run environment lacks anchors ENVIRONMENT Add 2–3 anchors (mirror type, shelf, countertop)
Product label / UI not readable no hero hold + too wide framing PACING + CAMERA + ACTION Add 0.7–1.0s hold; tighter shot; “label-forward” action
Lighting becomes flat/harsh lighting intent missing LIGHTING + CONSTRAINTS Specify soft key + contrast; add “avoid harsh flat lighting”
Captions cover the face/product no safe zone TEXT/OVERLAYS SAFE AREA Reserve bottom/top % and enforce framing

Mini library: 3 reusable ad cores (with swappable variants)

These aren’t “magic prompts.” They’re starting cores you can version and reuse.

Core A: Physical product demo (9:16)

SHOT CONFIG v1.0

OBJECTIVE:
- 7–9s: show the product working in one clear action.

SUBJECT (LOCKS):
- Hands-only demo; clean nails.
- Product: handheld fabric steamer, matte white.
- Continuity: same steamer model/color.

ENVIRONMENT (LOCKS):
- Minimal laundry room, neutral tones.
- Anchors: white wall + simple clothing rack + light wood table.

CAMERA (LOCKS):
- Close-up at chest height.
- Movement: slow lateral slide only.

LIGHTING (LOCKS):
- Soft daylight, gentle contrast.

STYLE (LOCKS):
- Clean product demo, crisp details.

TEXT/OVERLAYS SAFE AREA (LOCKS):
- Top 15% clear.

CONSTRAINTS (LOCKS):
- Avoid: random zoom, motion blur, warped fabric texture, extra fingers.

ACTION (VARIABLE):
- Beat 1: wrinkled shirt on hanger.
- Beat 2: one downward glide.
- Beat 3: hold on smooth “after” fabric.

PACING (VARIABLE):
- Hold on “after” for 0.6s.

AUDIO NOTES (VARIABLE):
- Light whoosh + subtle steam SFX; minimal upbeat music.

OUTPUT NOTES:
- 9:16, 1080p.

Core B: UGC testimonial (caption-friendly)

SHOT CONFIG v1.0

OBJECTIVE:
- 10–12s: hook + one benefit + CTA.

SUBJECT (LOCKS):
- One creator, casual sweatshirt; friendly tone.

ENVIRONMENT (LOCKS):
- Cozy living room.
- Anchors: neutral sofa + plant + plain wall.

CAMERA (LOCKS):
- Medium close-up, eye level.
- Movement: stable handheld feel (no zooms).

LIGHTING (LOCKS):
- Soft window light, warm fill.

STYLE (LOCKS):
- Authentic UGC (not glossy).

TEXT/OVERLAYS SAFE AREA (LOCKS):
- Bottom 25% clear.

CONSTRAINTS (LOCKS):
- Avoid: over-smoothing skin, uncanny mouth movement, harsh flat lighting.

ACTION (VARIABLE):
- Single take speaking to camera; brief product hold near face.

PACING (VARIABLE):
- Include 2 short pauses for captions.

AUDIO NOTES (VARIABLE):
- Conversational VO; low music or none.

OUTPUT NOTES:
- 9:16, 1080p.

Core C: App screen reveal (legibility-first)

SHOT CONFIG v1.0

OBJECTIVE:
- 6–8s: show one key app screen clearly.

SUBJECT (LOCKS):
- Hand holding modern smartphone; screen must remain legible.

ENVIRONMENT (LOCKS):
- Clean desk.
- Anchors: light desk surface + minimal notebook + mug.

CAMERA (LOCKS):
- Close-up, slight top-down.
- Movement: slow push-in only.

LIGHTING (LOCKS):
- Soft studio light; reduce reflections.

STYLE (LOCKS):
- Modern tech ad, minimal.

TEXT/OVERLAYS SAFE AREA (LOCKS):
- Left 20% clear for callouts.

CONSTRAINTS (LOCKS):
- Avoid: illegible UI, screen warping, flicker, random zoom.

ACTION (VARIABLE):
- Beat 1: phone lifts into frame.
- Beat 2: one tap.
- Beat 3: hold on key screen.

PACING (VARIABLE):
- Hold key screen for 1.0s.

AUDIO NOTES (VARIABLE):
- Subtle tap SFX + light synth bed.

OUTPUT NOTES:
- 9:16, 1080p.

Checklist

  • Write a Core Config (SUBJECT, ENVIRONMENT, CAMERA, LIGHTING, STYLE, TEXT SAFE AREA, CONSTRAINTS) and don’t touch it during testing.
  • Define ACTION as 1–3 observable beats.
  • Define PACING with a duration target and at least one explicit hold (0.6–1.0s).
  • Put all negative prompts in CONSTRAINTS (keep a single campaign blacklist) (https://www.designyourway.net/blog/ai-video-prompting-made-simple-tips-for-better-results).
  • Iterate one field at a time (ACTION first, then PACING, then AUDIO).
  • Save winners with a version name: FORMAT__OBJECTIVE__STYLE__v#.
  • If you need automation later, plan your fields so they can become API variables (Veo3Gen has a developer API) (Veo3Gen facts).

FAQ

How do I do JSON prompting in Veo3Gen if there’s no JSON UI?

Use JSON structure, not JSON syntax: paste a labeled Shot Config block so each instruction is isolated and editable. VidAU’s point is that structured prompting separates scene description from camera motion (https://www.vidau.ai/json-prompting-ai-video-comparison).

How do I stop random camera moves and zooms?

Fix CAMERA first: choose one movement type (locked-off / slow push-in / slow lateral slide) and ban zooms in CONSTRAINTS. Don’t rely on “cinematic” to imply stability.

How do I keep the same outfit and product look across variants?

Over-specify SUBJECT locks: wardrobe, hair, props, and explicit continuity (“same throughout”). Treat that block as untouchable while you test hooks.

How do I make pacing feel intentional instead of messy?

Put timing into PACING: total duration + at least one hold beat (0.6–1.0s) on the label/UI. If it’s still unclear, reduce ACTION to 1–2 beats.

How do I reduce artifacts like weird hands or harsh flat lighting?

Maintain a dedicated CONSTRAINTS blacklist with negative prompts; negative prompts are recommended to reduce issues such as blurry motion and harsh lighting (https://www.designyourway.net/blog/ai-video-prompting-made-simple-tips-for-better-results).

When should I use Veo 3.1 Lite vs Fast vs Quality?

Use Lite for cheapest previews, Fast as a quick strong default, and Quality when you’re pushing maximum fidelity. Veo3Gen offers all three modes (Veo3Gen facts).

Create a repeatable ad pipeline in Veo3Gen

Once you have 2–3 stable Core Configs, you can produce variants by swapping only OBJECTIVE / ACTION / PACING / AUDIO—and keep your camera language and brand look consistent.

Veo3Gen is designed for practical production: it supports text-to-video and image-to-video, includes native synchronized audio in one generation pass, and offers Fast/Quality/Lite modes so you can preview cheaply and finalize with higher fidelity (Veo3Gen facts).

CTA (closing): Start by generating one baseline with your Core Config, then iterate one field at a time. When you’re ready to scale from copy/paste to programmatic batch generation, use Veo3Gen’s developer API (Veo3Gen facts): /api.

Start creating with Veo3Gen

Veo3Gen gives you affordable Veo 3.1 video generation with native audio, up to 4K, and credits that never expire — with free credits to start.

Sources

Limited Time Offer

Try Veo 3 & Veo 3 API for Free

Experience cinematic AI video generation at the industry's lowest price point. No credit card required to start.