LayerV · Academy Video Pipeline Operator Guide · August 2026

Writing prompts for generated video

The better the prompt, the better the animation

A strong video prompt reads like a compact directing brief. It says what happens, what the camera does, and what must stay identical from the first frame to the last. A weak prompt describes a pretty frame and leaves everything else to the model. On this pipeline, every generated second is paid for. The quality of the animation is decided in the prompt, before any money is spent. This guide is how we write them.

10 seconds, generated from one approved still

This clip is ten seconds of generated animation. One approved still went in. The hands open, a cloud gathers, two labelled cards condense out of it, the camera drifts in. Nothing was animated by hand. Every move was a written beat before anything was rendered. That is the standard this guide protects: the animation is only as good as the sentences that order it.

Section 01Write a directing brief, not a caption

The model executes decisions. Where the prompt makes no decision, it makes one for you. So build every prompt like a brief, in a fixed order. The model weighs early tokens hardest. Put the important decisions first.

  1. Subject and event: who is in the shot, and the one thing that happens.
  2. Setting: where, when, in what light.
  3. Visual treatment: palette, texture, medium.
  4. Camera: one sentence, one behavior.
  5. Timeline: the beats, in order.
  6. Continuity rules: what must not drift.
  7. Negatives: grouped at the end, targeted.
The skeleton

[Subject] performs [one specific action] in [setting]. The image uses [light, texture, palette]. The camera uses [shot size, angle, one movement]. The clip is silent. Every section below develops one part of this skeleton.

Section 02Weak vs strong: the same shot, twice

Weak
A holographic figure stands in a dark digital space, cinematic, dramatic lighting, high quality.
Strong
A luminous figure in a white and silver body suit stands alone at the center of a glowing grid floor, in a deep navy volumetric space crossed by a faint constellation mesh. He watches a single point of light descend toward his open palm. Begin in a locked wide shot, then make one restrained push toward him as the floor panels light up in sequence beneath the falling point. Deep navy and pale blue only. The clip is silent. No text, no logos, no cuts.

Why the second one works: it makes decisions. An event: the descending light. A camera behavior: locked wide, then one push. An attention shift: his gaze, the floor responding. A palette rule. The weak version only asks for a mood.

Section 03Describe motion, not photographs

Static wording produces static shots. Convert every attribute into a change someone could watch happen. Name momentum, weight, drag, contact and secondary motion when they matter.

Static wording Video direction
"His cape is flowing" "As he turns, the fabric lags behind the movement, then settles with visible weight"
"He looks focused" "His eyes track the rising bar, his chin lifts slightly, and his hand closes as the figure locks in"
"The space is alive" "Motes drift across the foreground, and the constellation mesh pulses once behind him"
"Dramatic camera" "The camera pushes in at walking speed, then stops when he stops"
"Energetic scene" "The grid lights ripple outward from his feet, one ring per step"

Section 04References: one narrow job each

Each reference gets one named role and explicit exclusions. "Image 1 defines the character" is too broad. The model will import the background, the pose and the lighting along with the face.

Too broad
@Image 1 defines the character.
Narrow role
@Image 1 defines the character's face, build, and the white and silver body suit only. Ignore its background, pose, and lighting.

Multiple views are one object

Several angles of one object must be declared as one object. Otherwise you get several of it:

@Image 1 shows the front of the character. @Image 2 shows his left side. Both images define one character. Only one character appears in the video.

When references conflict, this is the priority order

  1. Character identity
  2. Essential prop or object geometry
  3. Wardrobe
  4. Location and spatial layout
  5. Lighting and texture
  6. Motion and camera style
House rule: the three-reference stack

On Academy shots, image_urls carries three files, in a fixed order: the two character references, then the approved still of the episode. The still anchors composition and color. Two references hold identity in plain motion. They do not hold it against a text-bearing prompt. The third reference is what stabilized our keyframes. More references are not more control: past a handful, stability drops.

Section 05Image-to-video is a different contract

This is the most expensive lesson in the pipeline. A text-to-video prompt must describe subject, setting and style. There is no input image, so the description is the scene. An image-to-video prompt already has the scene: the approved still. Copy the scene blocks into it and the model re-renders what they describe instead of animating the still. The character comes back re-scaled. The layout shifts. The lettering is redrawn from frame zero.

The contract line: verbatim, first line of every i2v prompt

"Animate the start frame. Every visual property comes from the reference frames and only from them: nothing is re-rendered, re-framed, re-scaled or re-styled."

Scene blocks never appear in an image-to-video prompt. The beats, the camera line, the appearance rule and the grouped negatives do.

Section 06The beat sheet: directing time

A ten-second shot written as one paragraph is one undirected instruction. The model resolves everything whenever it likes. On one paid test, a headline and two figures snapped into frame between one third of a second and the next. Nothing had told them when to arrive. The fix is a timed beat sheet. Every couple of seconds is directed. No stretch of the clip is left to the model.

Camera: one slow continuous push in, no cuts, no shake.
Appearance rule: nothing pops. Everything that arrives rises from
nothing to full brightness over roughly a second.
00:00 to 00:02  he raises his right hand; the first bar rises from the floor with it
00:02 to 00:05  the bar reaches full height; the ring at its base brightens
00:05 to 00:08  he turns his palm; the second, shorter bar rises beside the first
00:08 to 00:10  he lowers his hand; both bars hold; the floor ripple fades
Audio: none, the clip is silent.
Goal: one sentence on what the shot teaches.

What makes this structure work:

Long clips are continuity systems. Each stage states its initial state, one main event, the camera behavior, and a visible end state. The end state is the handoff to the next stage. Every stage after the first opens with the same line: continue the same identity, wardrobe, layout and light.

The clock is a request, not an edit

The beat sheet does two jobs. It orders the arrivals, and that part works. It also carries a clock, and the clock is read loosely. The model vendor's own documentation says exact segment timing is unstable. Keep the timecodes: they order the beats, and the captions are cut against them. Never re-prompt to move a beat by half a second. Never pay a regeneration for pacing. speed on the shot retimes the footage locally, for free, predictably.

Section 07Camera: one motivated move

The model knows the standard vocabulary: wide shot, medium close-up, low angle, overhead, push in, pull out, pan, track, orbit, tilt, handheld, dolly zoom, rack focus. A camera instruction states five things. What the camera follows. Where it begins. Where it ends. What motivates the move. What must stay readable.

Vague
Dynamic orbit around the character.
Directed
Begin at knee height behind the character's right shoulder. As he raises the bar, rise into one slow clockwise orbit and finish at eye level in a medium shot. His face stays readable throughout. The camera slows when his body becomes still. No collision, no weightless floating, no speed ramp.

One motivated move beats a pile of moves. A technical term comes with its visible result: "rack focus: the foreground motes soften while the figure behind them becomes sharp."

Section 08Performance: behavior, not emotion labels

Emotion labels produce theatrical mime. "He looks determined" gives the model nothing to stage. Direct the performance as a four-part beat: trigger, immediate reaction, gradual change, held expression.

When the second figure locks into place, his eyes stay on it for one beat. His shoulders drop slightly. He exhales, turns his head toward the gap between the bars, and lets one measured nod land as if the number confirms what he already suspected. He holds that posture without looking into the lens.

Two to four physical signals per change. More destabilizes the face or produces exaggerated acting. Direct gaze, breath, posture, hands and mouth.

Section 09Audio: our clips are silent by design

Current video models generate their own audio. The grammar marks music in ( ), sound effects in < >, dialogue in { }, on-screen subtitles in 【 】. Learn the markers so you never trigger them by accident.

House rule

Academy episodes carry a voiceover and sound design added in post, cut against the real recorded audio. Generated audio is never used. Every motion prompt ends with "Audio: none, the clip is silent". No dialogue, music or subtitle markers ever appear in a prompt. Unmarked quoted speech is a trap: the model may render it as on-screen text.

Section 10Positives first, negatives grouped

State what must hold before what must not appear. Positive invariants first:

Keep the camera locked at chest height. Preserve one character and two bars throughout. Deep navy and pale blue only.

Then the exclusions, grouped at the end:

No text, no logos, no duplicated character, no new objects, no cuts, no camera shake.

Long generic negative lists create contradictions. Banning blur while asking for shallow depth of field is a contradiction. They also spend attention. On one paid test, tripling the length of one clause made the model obey that clause and violate the unchanged clause after it. Adding a constraint costs one somewhere else. Keep every clause to one sentence. Keep the most critical clause last, and never touch it.

Never name a shape you do not want drawn

A beat said an area "takes on a cold rim of light". The model heard rim as an object and drew a rectangle around the whole area. Describe light as a verb on an existing object: "the interval brightens", "the floor lights under him". Never give an effect a noun that is also furniture: rim, edge, border, outline, frame, panel, plane, box.

Rank the constraints that can collide

Some constraints cannot all hold at once. Say which one gives way, or the model decides for you:

First priority: the lettering in the reference frames stays exactly as drawn.
Second priority: the right two thirds of the frame stay empty.
Sacrifice the smoothness of the camera move before either priority.

Section 11Text on screen: never the model's job

Models cannot letter. Every rule in this section was bought on a real clip.

Section 12The prompt template

Copy this and fill it in. Delete the blocks the shot does not need. The [Contract] block is image-to-video only, [World] is text-to-video only, and a prompt never carries both.

[Goal]
Generate a [duration]-second [format] in which [the one main event].

[Contract]        image-to-video only, verbatim
Animate the start frame. Every visual property comes from the reference frames
and only from them: nothing is re-rendered, re-framed, re-scaled or re-styled.

[World]           text-to-video only
Location, time, light, palette, texture, and the physical behaviour that matters.

[Reference roles]
@Image 1 defines [feature] only. Ignore [what to exclude].
@Image 2 defines [feature] only. Ignore [what to exclude].
All views of one object define one object. Only one appears.

[Camera]
One sentence: what it follows, where it begins, where it ends, what motivates
the move, what must stay readable.

[Appearance rule]
Nothing pops. Everything that arrives rises from nothing to full brightness over
roughly a second. Lettering visible in the reference frames is reproduced exactly
as drawn there, never redrawn, never restyled.

[Timeline]
00:00 to 00:0X   initial state. one main event. what arrives. end state.
00:0X to 00:0Y   continue all invariants. one main event. end state.

[Performance]
The trigger, then two to four visible changes in eyes, breath, hands, posture
or gaze.

[Audio]
Audio: none, the clip is silent.

[Invariants]
Identity, wardrobe, prop count, screen direction, spatial layout, light.

[Priority]
First priority: ...
Second priority: ...
Sacrifice ... before either priority.

[Negatives]
Grouped, targeted, critical clause last and unchanged.

[Goal line]
One sentence on what the shot is for.

Section 13Worked examples

A · Academy shot, image-to-video with beat sheet (our bread and butter)

Animate the start frame. Every visual property comes from the reference frames and only from them: nothing is re-rendered, re-framed, re-scaled or re-styled.

Camera: locked, no movement, no cuts.
Appearance rule: nothing pops. Everything that arrives rises from nothing to full brightness over roughly a second. All lettering visible in the reference frames is sacred: reproduce it exactly as drawn there, never redraw it, never restyle it.

00:00 to 00:02  he raises his right hand, palm up
00:02 to 00:05  the tall bar rises from the grid floor to full height
00:05 to 00:08  he turns toward it; the ground under the bar brightens
00:08 to 00:10  the lettering beneath the bar from the end reference is revealed, exactly as drawn there

Audio: none, the clip is silent.
Goal: the viewer's eye lands on the bar, then on its label, in that order.
No object appears that is not in the reference frames: no outline, no border, no frame, no box, no panel, no glass plane anywhere in the picture at any moment.

B · Multi-reference product film

@Image 1 defines the watch: brushed steel case, deep blue dial, and worn leather strap only. Ignore the background and the lighting.
@Image 2 defines the jeweller's bench: scarred walnut surface, brass tools, and the single angled bench lamp. Ignore any hands in the image.
@Image 3 defines the loupe, its knurled barrel, and its size against the case.

Generate a 12-second one-take. 0-5s: a low macro shot across the bench as two hands set the watch face-up beside the loupe; the strap settles with visible weight and one coil stays raised. 5-9s: the camera makes one slow push toward the dial as a hand lifts the loupe into the light; the bench lamp draws a moving highlight across the steel. 9-12s: the hand withdraws. The dial holds still, centred, and fully readable.

Keep one watch, one loupe, the same bench geography, and one continuous camera axis. Warm tungsten light and deep shadow only. Audio: none, the clip is silent. No cuts, no second watch, no text, no logos, no polished studio lighting.

Why it works: each reference has one job. The product detail becomes visible at a planned moment. The negatives protect the count and the character of the light.

C · 30-second scene with irreversible continuity

[0-8s - Sew]
A bookbinder pulls the last stitch through a single grey notebook block and cuts the thread. Static medium shot. End state: the block sewn, thread cut, scissors laid to frame right.

[8-17s - Press]
Continue the same hands, apron, bench and window light. Slow push in as the binder slides the block into the press and turns the handle twice. End state: the block clamped, handle horizontal, offcuts swept to the left.

[17-25s - Wrap]
Hard cut to a high three-quarter shot. The binder lifts the same block out and folds brown paper around it once. End state: the notebook wrapped, one corner still open, twine uncut on the bench.

[25-30s - Hand off]
Eye level. The binder ties the twine once and slides the parcel across the bench. End on the empty bench surface.

The notebook never returns to an earlier state. Preserve hand identity, apron, prop count, bench orientation, window direction and dusk light. Audio: none, the clip is silent. No dialogue, no music, no second notebook, no text.

Why it works: each stage changes the object once and hands a clean arrangement to the next. A stage with two events becomes two stages.

D · First-and-last-frame transition

@Image 1 is the first frame: the dark grid floor, locked wide composition, the character standing at center, floor unlit.
@Image 2 is the last frame: the same composition with the full constellation mesh lit above him and the grid glowing.

Generate one continuous 8-second locked-off transformation from @Image 1 to @Image 2. The change happens through visible physical progression, not a dissolve: the grid lights spread outward from his feet, the mesh points ignite from low to high, and the rim light reaches his shoulders last. The character, camera, horizon and floor geometry remain fixed. Audio: none, the clip is silent. No camera movement, no crossfade, no text, no sudden object replacement.

Why it works: both frame roles are named. The shared geometry is protected. The transition is one continuous event. Both frames share the same aspect ratio.

E · Edit one part of an existing clip

Edit @Video 1 only from 4 to 7 seconds. Change the cool blue glow on the floor ring to a warmer pale gold.

@Video 1 is the sole editing master. It defines the character, space, action, composition, camera, and event order.

Keep identity, wardrobe, pose, floor geometry, camera and timing unchanged. Allow nearby surfaces to respond naturally to the warmer light. No other edit.

F · Extend a clip forward

Extend @Video 1 forward. The first frame of the extension directly continues its last frame. Preserve the locked wide shot, the character's position, the lit bar, the grid floor, and the palette. The floor ripple completes its outward travel and fades; the character holds his position. Do not reintroduce anything that has already faded.

Section 14It came back wrong: repair table

A clip is on the table and it is wrong. Read the cost column first. Most repairs are free, local and predictable. A regeneration on a free row pays for the episode twice. Classify before regenerating: only a prompt fault, a reference fault or model variance justifies a new submission.

Symptom Likely cause First repair Cost
The acting reads too slow or too fast The model reads the beat clock loosely speed on the shot (1.0–1.35), rebuild from trim free
Captions wrong or mistimed Timings cut against an estimate, not the audio Edit the voiceover line, rebuild subtitles locally free
Two shots butt together and read as a glitch Separately generated shots share no continuity A transition in the manifest, rebuild from concat free
A cut must land on an exact frame Timestamps are not frame-accurate Split into two shots, join locally free
Character re-rendered / re-scaled from frame zero The i2v prompt opened on a scene block Contract line first, verbatim; scene blocks removed
Lettering redrawn, mangled, or reads an invented value A beat had writing arrive (appear, form, fade up) Time a reveal of end-frame lettering, "exactly as drawn there"
An uninvited frame, border, panel or plane An effect was given a noun that is also furniture Describe light as a verb on an existing object
The reserved zone comes back occupied Body clause lengthened, or zone clause moved off the end Restore both clauses to exact length and order
Character changes between shots Identity reference too broad or unnamed One identity reference, named role, background and pose excluded
Duplicate object or prop Several views read as several objects State all views define one object; only one appears
Character's colors are wrong A locked character or world block was paraphrased Restore the block verbatim from the reference set
Events happen out of order Long clip has no stages or end states Consecutive ranges, one event and one visible handoff per stage
Camera feels random Prompt says "cinematic" or stacks movements One movement with subject, start point, end frame
Several negatives ignored at once A cheaper model tier was used The standard endpoint; do not reopen on a price list

Section 15What the model cannot guarantee

Section 16Endpoints and cost discipline

The pipeline generates through fal.ai: image-to-video for most shots, reference-to-video when needed, and the cheaper tiers of both. The grammar in this guide is not tied to one model family, but the endpoint you pick is a quality decision. Two rules, both paid for:

Spend gate

Every generation run is priced from the manifest before anything is submitted. It waits for a written authorization on that exact amount. A prompt costs nothing to rewrite. A submission always costs.

Section 17Pre-generation checklist

Before any paid submission, verify:

The best prompt is rarely the longest one. It gives every element a role, every event enough time, and every stage a clean handoff to the next. Whatever the prompt leaves open, the model decides. Whatever the model decides, someone reviews, rejects, and pays to regenerate.