Complete storyboard rule template
Use these rules as a starting point for custom parsing. Adapt them to the genre and shot duration, test one episode, then have an administrator publish. They govern how a script is split, rather than describe a single video to generate.
Decide three things first
- Choose the model and shot-duration cap, allowing time for speech, action and reactions.
- Choose one visual style and permitted shot sizes. Retain medium-wide shots when space must be established. If only tighter framing is permitted, revise every large-scene rule accordingly.
- Decide whether narration, ambience, music and subtitles are needed. Avoid conflicting audio instructions.
For creation, testing and publication, see Custom storyboard parsing. Select the relevant scene patterns below; not every scene needs five shots.
Role and task
text
Convert the supplied script into producible shots.
Retain plot, exact dialogue, relationships and event order. Do not invent facts.
Give each shot one primary action and one visual focus.
Split two important details if one framing cannot show both clearly.
Time dialogue at a natural pace; retain every word when splitting a long line.
Identify each speaker and leave room for the listener's reaction.
Describe people, positions, action, shot size, viewpoint, movement, light, sound and ending state.
Keep asset names consistent. An outfit change does not create a new identity.These are creative instructions, not mandatory API fields or a promise of a fixed JSON output.
1. Dialogue
text
Speaker conveys information → listener is affected → relationship or action changes.
Prefer over-shoulder shots, compact two-shots, singles and close-ups.
Give key lines to the speaker, significant effects to the listener and distance changes to a relationship shot.
Use an insert for evidence or a prop without crossing the dialogue axis.
Default to a locked camera; use one tracking movement for walking dialogue.
Push in slowly only when pressure, realization or a decision intensifies.
Do not replace reactions with orbiting, abrupt zooms or repeated pans.2. Action and combat
text
Establish space and opponents → preparation → one core action → contact/dodge/block → new positions.
State who is where, facing which direction and trying to reach what destination.
Do not pack preparation, impact, collapse, recovery and counterattack into one short clip.
Use medium-wide framing for paths, medium shots for upper-body interaction and close-ups for intent.
Hand or prop inserts provide brief emphasis, not the complete action sequence.
After each contact, record position, orientation, posture and prop state.
Keep pursuit, strikes and falls consistent in screen direction.
Choose one principal camera move: tracking, a pan, brief handheld motion or a locked view.3. Narrative processes
text
Goal or start → meaningful step A → step B or a key detail → result → consequence.
Use for investigation, preparation, training, work, travel or compressed time.
Every process shot must add a new state, rather than repeat the same behavior.
If the process is unimportant, show only the start and result.
Use medium-wide views for routes, medium framing for tasks and inserts for proof of completion.
Track along a route; pan or tilt between an operator and an adjacent tool or result.
Do not introduce a different camera movement in every shot merely to suggest montage.4. Emotion and decisions
text
Initial state → stimulus → physical or facial change → decision → external action.
Express emotion through breath, gaze, mouth, fingers, shoulders, balance or distance.
Establish the initial state in medium framing, use close-ups for the decision and inserts sparingly.
Default to a locked camera. Push in slowly when emotion becomes irreversible.
Pull back when the result is isolation or loss.
Do not substitute shaking, orbiting or repeated focus changes for acting.
A listener reacts after hearing the stimulus, not before it.5. Discovery and revelation
text
Hidden state → unusual sign → attention or approach → revealed information → reaction.
Establish the question before showing the answer. Do not expose evidence during the hidden stage.
Use a relationship shot for the search, a close-up for attention and an insert for the evidence.
Choose one reveal based on space:
Rack focus for depth; pan or tilt for adjacent space; push toward evidence; pull back for a larger context.
Do not combine panning, pushing and rack focus in one reveal.6. Entrances, exits and transitions
text
Entrance: stable space → doorway or sound cue → clear path → reactions → new blocking.
Exit: decision or pressure → turn/rise/collect prop → established exit direction → remaining reaction.
Track when the path matters; stay locked for a crossing; pan between adjacent doorway and people.
Reserve brief handheld motion for a sudden intrusion or pursuit.
Update positions after entrance. Remove departed people from the on-screen cast.
Keep doorway location, movement direction and held objects consistent with the previous shot.7. Powers and spectacle
text
Everyday scale reference → sign or preparation → one core power → environmental result → reaction or cost.
Demonstrate scale through a visible comparison rather than adjective lists.
Use medium-wide views for scale, medium framing for activation and inserts for one sign or cost.
A locked view supports comparison; a pan or tilt follows one path; a pullback shows spreading effects.
Do not pack activation, multiple people flying, a collapsing building and every reaction into one short clip.Global constraints
text
Build rhythm gradually and retain reaction time at important beats.
Every shot advances plot, conveys essential emotion or establishes necessary space.
Merge or remove shots that add no information.
Preserve identity, clothing, prop ownership, light, left/right relationships and event order.
Do not request both an uninterrupted take and a hard cut in the same shot.
Do not invent dialogue or make a silent character speak.
Use one coherent visual style without mutually exclusive instructions.Test one episode
| Check | Passing result | Rule to revise if it fails |
|---|---|---|
| Story and dialogue | No omissions or invented facts; correct speakers | Task priorities and dialogue splitting |
| Feasibility | Actions and lines fit the selected duration | Shot granularity and reaction time |
| Space | Positions remain traceable after entrances/exits | Blocking and axis rules |
| Acting | Stimulus precedes visible reaction | Emotion and action descriptions |
| Camera | Clear focus and purposeful movement | Framing and viewpoint preferences |
| Consistency | Stable people, outfits and prop ownership | Asset naming and global rules |
Save the tested rules and output, then compare revisions using the same episode. Before publication, remove placeholders and conflicts such as banning wide shots while requesting a wide city establishing shot.
