Director scene examples
Assign separate roles to blocking, motion, identity and environment references. These adaptable prompts follow the methods in the teaching examples. Replace characters, dialogue and timing for your project. The videos illustrate the method; they are not claimed to be outputs of the rewritten prompts below.
Example 1: Two-person blocking and a 30-second emotional scene
Prepare: Image 1 for the woman, Image 2 for the man, Image 3 for the seaside, Image 4 for their Director blocking and Image 5 for her close-up composition. Add voice references if needed. Establish her on the left and him on the right, with connected eyelines.
Task: 30 seconds, 16:9, a consistent detailed 2D animation style.
Plan three shots, cutting at 12 and 22 seconds.
@Image1 and @Image2 define identity, clothes and materials, not reference poses.
@Image3 defines seaside rocks, afternoon light and space.
@Image4 defines initial blocking: woman left, man right, about 1.5 m apart on one axis.
No additional people. Preserve light direction and clothes.
0–12 s: Medium shot, slow push in.
She looks at the bouquet, extends it with both hands and softly says, “This is for you.”
Her fingers show slight tension; she looks up after speaking. He pauses before accepting.
12–22 s: Cut to her close-up using @Image5, camera locked.
She keeps offering the flowers, breath changes slightly and fingers tighten as she waits.
No new dialogue or exaggerated tears.
22–30 s: Return to the two-shot and slowly pull back.
He accepts the bouquet with both hands. Their fingers briefly touch; both relax and smile.
Do not add a hug, cheering or unrelated action.
Audio: clear dialogue, soft surf, wind and clothing. No music, narration, subtitles or text.
Retain the final expressions without cutting abruptly to black.Check: one bouquet changes hands without duplication; the close-up keeps the axis; emotion develops through action. Timestamps guide rhythm, but actual cuts must be reviewed.
Example 2: Three-person relationships and an exit
Focus on the excluded person's reaction instead of giving all three large simultaneous actions. Prepare three camera views, then provide independent identity and setting references.
The labels are previs editing marks and should not appear in the generated scene. Check the actual reference colors: in these images, blue represents the man, yellow Woman A and pink Woman B. Do not reuse a different previs clip's color mapping blindly.
Prepare the references
Reference the three blocking views as Images 5, 6 and 7. Identity images do not prescribe acting; blocking images do not prescribe final materials.
Task: 30 seconds, 16:9, live-action cinematic realism.
Plan three shots, cutting at 12 and 22 seconds.
@Image1 defines the man, @Image2 Woman A and @Image3 Woman B.
@Image4 defines the campus road, trees, steps and natural light direction.
@Image5, @Image6 and @Image7 define only their respective shot compositions and blocking.
Blue=man; yellow=Woman A; pink=Woman B. Do not render proxy colors, labels or grids.
Keep identity, clothing and age stable. No additional main characters or axis crossing.
0–12 s: Use @Image5, medium shot with a slow push.
B looks at the man, grips then releases her coat edge, waiting to be recognized.
He briefly looks puzzled, shifts his weight slightly toward A and says, “Were you calling me?”
B pauses after hearing him, then breathes in softly. A notices their eyelines without taking focus.
12–22 s: Use @Image6, locked medium close-up.
B looks at him and quietly asks, “You don't remember me?”
He waits before looking at A and saying, “Let's go.”
A nods lightly; they prepare to turn. B raises her hand, then stops after hearing the line.
22–30 s: Use @Image7 in one continuous shot, slowly retreating.
The pair leave in the planned direction. B stays, softly saying, “Wait.”
She takes half a step, settles back and watches them recede before lowering her gaze.
Do not add a fourth shot, sudden zoom, flashback or slow motion.
Each line belongs to its designated speaker; listeners do not move their mouths in sync.
Use only dialogue, light wind, leaves, steps and clothing. No score, dramatic stings or subtitles.Check: correct speaking/listening roles, matching backgrounds and B remaining in place after the pair leave. The final eight seconds form one shot. A three-shot request must not introduce a fourth cut in its timeline.
Example 3: Dynamic proxies become natural performances
Prepare: one complete previs clip, three identity images and a setting image. The following color convention is specific to this example and must match your own previs. It differs from the campus images above.
@Video1 defines starts, key waypoints, stops, action order, final positions and camera path.
In this clip: red=Lin @Image1; blue=Gu @Image2; yellow=Su @Image3.
@Image4 defines the banquet hall, doorway, stage, lights and floor; objects stay fixed.
Images define identity and clothes; proxies guide space and timing only.
Movement:
Retain main routes, crossing order, arrival times and occlusions.
Start with gaze or weight shifts; swing arms naturally; decelerate before stopping to speak.
Turn through gaze, head, shoulders and footwork rather than rotating mechanically in place.
Stationary people breathe, blink and shift weight slightly. Reactions follow dialogue.
Smooth paths between waypoints without changing key stops or final positions.
Events:
Lin follows her path to the stage, stops, turns to Gu and says, “Now we can talk.”
Gu approaches on his assigned path, stops at a reasonable distance and says, “I'm listening.”
Su follows the third path to join them, stops and reacts with her gaze, without a collision.
The camera then approaches Lin in the previs direction and settles.
Camera:
Keep start/end positions, height, main direction and followed subject.
Accelerate and decelerate gently, with some tracking lag rather than game-style target locking.
No added orbit, whip pan, abrupt zoom or horizontal mirroring.
Sound:
Assign each line and voice to its speaker. Others listen silently.
Retain only intended dialogue, footsteps and ambience, not irrelevant previs audio.
Remove colored proxies, paths, axes, camera cones and interface marks.
No fourth person, duplicate character, interpenetration or foot sliding.If the selected duration cannot hold the dialogue, shorten it or split the sequence instead of forcing rushed speech. Preserve key blocking and camera intent while regenerating natural motion between key positions.
Example 4: Re-render a detailed model
A structurally complete model already contains motion. Preserve structure, action, spatial layout, camera moves and cuts; use images to define final people, materials, color and environment. Remove paths, controllers and camera aids before input. See the detailed-model template.
