Writing prompts for generated video
A strong video prompt reads like a compact directing brief. It says what happens, what the camera does, and what must stay identical from the first frame to the last. A weak prompt describes a pretty frame and leaves everything else to the model. On this pipeline, every generated second is paid for. The quality of the animation is decided in the prompt, before any money is spent. This guide is how we write them.
This clip is ten seconds of generated animation. One approved still went in. The hands open, a cloud gathers, two labelled cards condense out of it, the camera drifts in. Nothing was animated by hand. Every move was a written beat before anything was rendered. That is the standard this guide protects: the animation is only as good as the sentences that order it.
The model executes decisions. Where the prompt makes no decision, it makes one for you. So build every prompt like a brief, in a fixed order. The model weighs early tokens hardest. Put the important decisions first.
[Subject] performs [one specific action] in [setting]. The image uses [light, texture, palette]. The camera uses [shot size, angle, one movement]. The clip is silent. Every section below develops one part of this skeleton.
A holographic figure stands in a dark digital space, cinematic, dramatic lighting, high quality.
A luminous figure in a white and silver body suit stands alone at the center of a glowing grid floor, in a deep navy volumetric space crossed by a faint constellation mesh. He watches a single point of light descend toward his open palm. Begin in a locked wide shot, then make one restrained push toward him as the floor panels light up in sequence beneath the falling point. Deep navy and pale blue only. The clip is silent. No text, no logos, no cuts.
Why the second one works: it makes decisions. An event: the descending light. A camera behavior: locked wide, then one push. An attention shift: his gaze, the floor responding. A palette rule. The weak version only asks for a mood.
Static wording produces static shots. Convert every attribute into a change someone could watch happen. Name momentum, weight, drag, contact and secondary motion when they matter.
| Static wording | Video direction |
|---|---|
| "His cape is flowing" | "As he turns, the fabric lags behind the movement, then settles with visible weight" |
| "He looks focused" | "His eyes track the rising bar, his chin lifts slightly, and his hand closes as the figure locks in" |
| "The space is alive" | "Motes drift across the foreground, and the constellation mesh pulses once behind him" |
| "Dramatic camera" | "The camera pushes in at walking speed, then stops when he stops" |
| "Energetic scene" | "The grid lights ripple outward from his feet, one ring per step" |
Each reference gets one named role and explicit exclusions. "Image 1 defines the character" is too broad. The model will import the background, the pose and the lighting along with the face.
@Image 1 defines the character.
@Image 1 defines the character's face, build, and the white and silver body suit only. Ignore its background, pose, and lighting.
Several angles of one object must be declared as one object. Otherwise you get several of it:
@Image 1 shows the front of the character. @Image 2 shows his left side. Both images define one character. Only one character appears in the video.
On Academy shots, image_urls carries three files, in a fixed
order: the two character references, then the approved still of the
episode. The still anchors composition and color. Two references hold
identity in plain motion. They do not hold it against a text-bearing
prompt. The third reference is what stabilized our keyframes. More
references are not more control: past a handful, stability drops.
This is the most expensive lesson in the pipeline. A text-to-video prompt must describe subject, setting and style. There is no input image, so the description is the scene. An image-to-video prompt already has the scene: the approved still. Copy the scene blocks into it and the model re-renders what they describe instead of animating the still. The character comes back re-scaled. The layout shifts. The lettering is redrawn from frame zero.
"Animate the start frame. Every visual property comes from the reference frames and only from them: nothing is re-rendered, re-framed, re-scaled or re-styled."
Scene blocks never appear in an image-to-video prompt. The beats, the camera line, the appearance rule and the grouped negatives do.
A ten-second shot written as one paragraph is one undirected instruction. The model resolves everything whenever it likes. On one paid test, a headline and two figures snapped into frame between one third of a second and the next. Nothing had told them when to arrive. The fix is a timed beat sheet. Every couple of seconds is directed. No stretch of the clip is left to the model.
Camera: one slow continuous push in, no cuts, no shake. Appearance rule: nothing pops. Everything that arrives rises from nothing to full brightness over roughly a second. 00:00 to 00:02 he raises his right hand; the first bar rises from the floor with it 00:02 to 00:05 the bar reaches full height; the ring at its base brightens 00:05 to 00:08 he turns his palm; the second, shorter bar rises beside the first 00:08 to 00:10 he lowers his hand; both bars hold; the floor ripple fades Audio: none, the clip is silent. Goal: one sentence on what the shot teaches.
What makes this structure work:
Long clips are continuity systems. Each stage states its initial state, one main event, the camera behavior, and a visible end state. The end state is the handoff to the next stage. Every stage after the first opens with the same line: continue the same identity, wardrobe, layout and light.
The beat sheet does two jobs. It orders the arrivals, and that part works.
It also carries a clock, and the clock is read loosely. The model vendor's
own documentation says exact segment timing is unstable. Keep the
timecodes: they order the beats, and the captions are cut against them.
Never re-prompt to move a beat by half a second. Never pay a regeneration
for pacing. speed on the shot retimes the footage locally,
for free, predictably.
The model knows the standard vocabulary: wide shot, medium close-up, low angle, overhead, push in, pull out, pan, track, orbit, tilt, handheld, dolly zoom, rack focus. A camera instruction states five things. What the camera follows. Where it begins. Where it ends. What motivates the move. What must stay readable.
Dynamic orbit around the character.
Begin at knee height behind the character's right shoulder. As he raises the bar, rise into one slow clockwise orbit and finish at eye level in a medium shot. His face stays readable throughout. The camera slows when his body becomes still. No collision, no weightless floating, no speed ramp.
One motivated move beats a pile of moves. A technical term comes with its visible result: "rack focus: the foreground motes soften while the figure behind them becomes sharp."
Emotion labels produce theatrical mime. "He looks determined" gives the model nothing to stage. Direct the performance as a four-part beat: trigger, immediate reaction, gradual change, held expression.
When the second figure locks into place, his eyes stay on it for one beat. His shoulders drop slightly. He exhales, turns his head toward the gap between the bars, and lets one measured nod land as if the number confirms what he already suspected. He holds that posture without looking into the lens.
Two to four physical signals per change. More destabilizes the face or produces exaggerated acting. Direct gaze, breath, posture, hands and mouth.
Current video models generate their own audio. The grammar marks music in
( ), sound effects in < >, dialogue in
{ }, on-screen subtitles in 【 】. Learn the
markers so you never trigger them by accident.
Academy episodes carry a voiceover and sound design added in post, cut against the real recorded audio. Generated audio is never used. Every motion prompt ends with "Audio: none, the clip is silent". No dialogue, music or subtitle markers ever appear in a prompt. Unmarked quoted speech is a trap: the model may render it as on-screen text.
State what must hold before what must not appear. Positive invariants first:
Keep the camera locked at chest height. Preserve one character and two bars throughout. Deep navy and pale blue only.
Then the exclusions, grouped at the end:
No text, no logos, no duplicated character, no new objects, no cuts, no camera shake.
Long generic negative lists create contradictions. Banning blur while asking for shallow depth of field is a contradiction. They also spend attention. On one paid test, tripling the length of one clause made the model obey that clause and violate the unchanged clause after it. Adding a constraint costs one somewhere else. Keep every clause to one sentence. Keep the most critical clause last, and never touch it.
A beat said an area "takes on a cold rim of light". The model heard rim as an object and drew a rectangle around the whole area. Describe light as a verb on an existing object: "the interval brightens", "the floor lights under him". Never give an effect a noun that is also furniture: rim, edge, border, outline, frame, panel, plane, box.
Some constraints cannot all hold at once. Say which one gives way, or the model decides for you:
First priority: the lettering in the reference frames stays exactly as drawn. Second priority: the right two thirds of the frame stay empty. Sacrifice the smoothness of the camera move before either priority.
Models cannot letter. Every rule in this section was bought on a real clip.
Copy this and fill it in. Delete the blocks the shot does not need. The [Contract] block is image-to-video only, [World] is text-to-video only, and a prompt never carries both.
[Goal] Generate a [duration]-second [format] in which [the one main event]. [Contract] image-to-video only, verbatim Animate the start frame. Every visual property comes from the reference frames and only from them: nothing is re-rendered, re-framed, re-scaled or re-styled. [World] text-to-video only Location, time, light, palette, texture, and the physical behaviour that matters. [Reference roles] @Image 1 defines [feature] only. Ignore [what to exclude]. @Image 2 defines [feature] only. Ignore [what to exclude]. All views of one object define one object. Only one appears. [Camera] One sentence: what it follows, where it begins, where it ends, what motivates the move, what must stay readable. [Appearance rule] Nothing pops. Everything that arrives rises from nothing to full brightness over roughly a second. Lettering visible in the reference frames is reproduced exactly as drawn there, never redrawn, never restyled. [Timeline] 00:00 to 00:0X initial state. one main event. what arrives. end state. 00:0X to 00:0Y continue all invariants. one main event. end state. [Performance] The trigger, then two to four visible changes in eyes, breath, hands, posture or gaze. [Audio] Audio: none, the clip is silent. [Invariants] Identity, wardrobe, prop count, screen direction, spatial layout, light. [Priority] First priority: ... Second priority: ... Sacrifice ... before either priority. [Negatives] Grouped, targeted, critical clause last and unchanged. [Goal line] One sentence on what the shot is for.
Animate the start frame. Every visual property comes from the reference frames and only from them: nothing is re-rendered, re-framed, re-scaled or re-styled. Camera: locked, no movement, no cuts. Appearance rule: nothing pops. Everything that arrives rises from nothing to full brightness over roughly a second. All lettering visible in the reference frames is sacred: reproduce it exactly as drawn there, never redraw it, never restyle it. 00:00 to 00:02 he raises his right hand, palm up 00:02 to 00:05 the tall bar rises from the grid floor to full height 00:05 to 00:08 he turns toward it; the ground under the bar brightens 00:08 to 00:10 the lettering beneath the bar from the end reference is revealed, exactly as drawn there Audio: none, the clip is silent. Goal: the viewer's eye lands on the bar, then on its label, in that order. No object appears that is not in the reference frames: no outline, no border, no frame, no box, no panel, no glass plane anywhere in the picture at any moment.
@Image 1 defines the watch: brushed steel case, deep blue dial, and worn leather strap only. Ignore the background and the lighting. @Image 2 defines the jeweller's bench: scarred walnut surface, brass tools, and the single angled bench lamp. Ignore any hands in the image. @Image 3 defines the loupe, its knurled barrel, and its size against the case. Generate a 12-second one-take. 0-5s: a low macro shot across the bench as two hands set the watch face-up beside the loupe; the strap settles with visible weight and one coil stays raised. 5-9s: the camera makes one slow push toward the dial as a hand lifts the loupe into the light; the bench lamp draws a moving highlight across the steel. 9-12s: the hand withdraws. The dial holds still, centred, and fully readable. Keep one watch, one loupe, the same bench geography, and one continuous camera axis. Warm tungsten light and deep shadow only. Audio: none, the clip is silent. No cuts, no second watch, no text, no logos, no polished studio lighting.
Why it works: each reference has one job. The product detail becomes visible at a planned moment. The negatives protect the count and the character of the light.
[0-8s - Sew] A bookbinder pulls the last stitch through a single grey notebook block and cuts the thread. Static medium shot. End state: the block sewn, thread cut, scissors laid to frame right. [8-17s - Press] Continue the same hands, apron, bench and window light. Slow push in as the binder slides the block into the press and turns the handle twice. End state: the block clamped, handle horizontal, offcuts swept to the left. [17-25s - Wrap] Hard cut to a high three-quarter shot. The binder lifts the same block out and folds brown paper around it once. End state: the notebook wrapped, one corner still open, twine uncut on the bench. [25-30s - Hand off] Eye level. The binder ties the twine once and slides the parcel across the bench. End on the empty bench surface. The notebook never returns to an earlier state. Preserve hand identity, apron, prop count, bench orientation, window direction and dusk light. Audio: none, the clip is silent. No dialogue, no music, no second notebook, no text.
Why it works: each stage changes the object once and hands a clean arrangement to the next. A stage with two events becomes two stages.
@Image 1 is the first frame: the dark grid floor, locked wide composition, the character standing at center, floor unlit. @Image 2 is the last frame: the same composition with the full constellation mesh lit above him and the grid glowing. Generate one continuous 8-second locked-off transformation from @Image 1 to @Image 2. The change happens through visible physical progression, not a dissolve: the grid lights spread outward from his feet, the mesh points ignite from low to high, and the rim light reaches his shoulders last. The character, camera, horizon and floor geometry remain fixed. Audio: none, the clip is silent. No camera movement, no crossfade, no text, no sudden object replacement.
Why it works: both frame roles are named. The shared geometry is protected. The transition is one continuous event. Both frames share the same aspect ratio.
Edit @Video 1 only from 4 to 7 seconds. Change the cool blue glow on the floor ring to a warmer pale gold. @Video 1 is the sole editing master. It defines the character, space, action, composition, camera, and event order. Keep identity, wardrobe, pose, floor geometry, camera and timing unchanged. Allow nearby surfaces to respond naturally to the warmer light. No other edit.
Extend @Video 1 forward. The first frame of the extension directly continues its last frame. Preserve the locked wide shot, the character's position, the lit bar, the grid floor, and the palette. The floor ripple completes its outward travel and fades; the character holds his position. Do not reintroduce anything that has already faded.
A clip is on the table and it is wrong. Read the cost column first. Most repairs are free, local and predictable. A regeneration on a free row pays for the episode twice. Classify before regenerating: only a prompt fault, a reference fault or model variance justifies a new submission.
| Symptom | Likely cause | First repair | Cost |
|---|---|---|---|
| The acting reads too slow or too fast | The model reads the beat clock loosely | speed on the shot (1.0–1.35), rebuild from trim |
free |
| Captions wrong or mistimed | Timings cut against an estimate, not the audio | Edit the voiceover line, rebuild subtitles locally | free |
| Two shots butt together and read as a glitch | Separately generated shots share no continuity | A transition in the manifest, rebuild from concat | free |
| A cut must land on an exact frame | Timestamps are not frame-accurate | Split into two shots, join locally | free |
| Character re-rendered / re-scaled from frame zero | The i2v prompt opened on a scene block | Contract line first, verbatim; scene blocks removed | regenerate |
| Lettering redrawn, mangled, or reads an invented value | A beat had writing arrive (appear, form, fade up) | Time a reveal of end-frame lettering, "exactly as drawn there" | regenerate |
| An uninvited frame, border, panel or plane | An effect was given a noun that is also furniture | Describe light as a verb on an existing object | regenerate |
| The reserved zone comes back occupied | Body clause lengthened, or zone clause moved off the end | Restore both clauses to exact length and order | regenerate |
| Character changes between shots | Identity reference too broad or unnamed | One identity reference, named role, background and pose excluded | regenerate |
| Duplicate object or prop | Several views read as several objects | State all views define one object; only one appears | regenerate |
| Character's colors are wrong | A locked character or world block was paraphrased | Restore the block verbatim from the reference set | regenerate |
| Events happen out of order | Long clip has no stages or end states | Consecutive ranges, one event and one visible handoff per stage | regenerate |
| Camera feels random | Prompt says "cinematic" or stacks movements | One movement with subject, start point, end frame | regenerate |
| Several negatives ignored at once | A cheaper model tier was used | The standard endpoint; do not reopen on a price list | regenerate |
The pipeline generates through fal.ai: image-to-video for most
shots, reference-to-video when needed, and the cheaper tiers of
both. The grammar in this guide is not tied to one model family, but the
endpoint you pick is a quality decision. Two rules, both paid for:
Every generation run is priced from the manifest before anything is submitted. It waits for a written authorization on that exact amount. A prompt costs nothing to rewrite. A submission always costs.
Before any paid submission, verify: