User Guide
Working with Scenes
Explore, edit, and regenerate individual scenes in your project
The Scenes tab is the primary workspace for reviewing and refining your generated project. It provides a split-panel interface with a scene list, a media player, and detailed editing tabs.
Layout
The scenes view has three main areas:
- Scene List (left sidebar on desktop, bottom drawer on mobile) — Scrollable list of all scene thumbnails with status indicators
- Scene Player (center) — Large preview showing the selected scene's image or video, sized to your project's aspect ratio
- Detail Tabs (below the player) — Tabs for inspecting and editing each scene
Scene List
Each scene in the sidebar shows:
- Thumbnail — The generated image (or a placeholder if still generating)
- Status indicators — Visual badges showing image and motion generation status
- Scene number — Position in the project
Click any scene to select it. On mobile, scenes appear in a bottom drawer that can be pulled up.
Batch Motion Generation
At the top of the scene list, a Generate Motion button lets you start video generation for all eligible scenes at once (scenes with completed images but no video). You can optionally include music generation in the batch.
Scene Player
The central player displays:
- Still image — When no video exists yet
- Video playback — When motion has been generated, with play/pause controls
- Progress overlay — During generation, shows the current phase name
- Variant preview — When browsing image variants, shows the alternate image with a "Click 'Use saved AI image' to apply" badge
The player automatically sizes to match your project's aspect ratio (16:9, 9:16, or 1:1).
Detail Tabs
Below the player, the tabs run roughly in production order — what the scene says, who and where it is, what it looks like, how it moves, and what happens to the finished clip.
| Tab | What it does |
|---|---|
| Story | The script extract for this scene, its dialogue, and duration |
| Actors | Characters appearing in the scene |
| Location | The scene's location and reference image |
| Props | Element and prop references attached to the scene |
| Image | The still frame — prompt, model, generation |
| Avatar | Render the scene as a talking avatar instead of B-roll |
| Variants | Alternate images to choose between |
| Audio | Narration voice, per-scene VO, and who speaks |
| Motion | How the scene moves — the action outline and choreography |
| Post | Operations that run on the finished clip |
| Graphics | Lower thirds, titles, captions, logos |
| Result | The rendered clip, render history, and trimming |
| Debug | The exact request that would be submitted — nothing is sent |
Story Tab
Shows the original script extract for this scene — the text from your screenplay that this scene was built from — along with its parsed dialogue lines and the scene duration.
The extract seeds the image and the motion direction. Once you have edited those downstream, the extract is history rather than the source of truth, so it is fine for it to drift.
Actors Tab
Shows which characters appear in this scene. Each character links to their detail page where you can view their full profile and recast them.
Location Tab
Shows the location for this scene with its reference image and details. Links to the location's detail page for updating references.
Props Tab
Attach element and prop references — a product, a logo, a specific object — so they stay consistent when the scene is generated.
Image Tab
Full control over the scene's image generation:
- Editable prompt — The full visual prompt used to generate the image. Edit it to refine the composition, lighting, or details.
- Character count — Displayed in real-time as you edit
- Model selector — Switch between image models for this specific scene
- Shorten Prompt — AI-powered prompt compression that preserves intent while reducing length
- Generate Image / Regenerate Image — Create a new image with the current prompt and model
- Use saved AI image — Re-apply a previously generated AI image as the scene's frame, with no new generation — handy after uploading your own image or switching models
- Upload your own image — Use a still you already have as the starting frame, skipping generation. The AI image is preserved so you can switch back for free
- Use current image as End Frame — Promote the scene's still to be the frame the clip ends on, rather than the one it starts from
Avatar Tab
Render the scene as a talking avatar rather than generated B-roll. Pick a look, and the avatar speaks the scene's dialogue.
Avatars can also be composited as a presenter overlay — a talking head in a corner inset over the scene's B-roll, speaking its voiceover. The overlay's corner, size, and shape are set once for the whole project in Render.
Overlays can be rendered with a transparent background, so the presenter appears as a cut-out figure over your footage instead of a video inside a box.
Variants Tab
Generate multiple visual interpretations of a scene. Useful for A/B testing different compositions or finding the best framing.
- Generate Scene Variants — Creates a grid of variant images using the selected image model
- Click any variant to set it as the scene's primary image
- Choose a different model before generating to compare outputs across providers
Audio Tab
Controls the voice for this scene:
- Voice and Language — Which narrator voice speaks this scene's lines
- Model and Voice settings — Fine-tune delivery
- Delay — Push the voiceover start later within the clip
- Generated audio — Preview what will be used
The Audio tab sits before Motion because who is speaking affects how the scene should be directed.
Who actually speaks
Scene.Video splits speech by who says it:
- Narrator / voiceover lines are synthesized with your chosen voice and played over the clip.
- Named on-camera characters are voiced by the video model itself, using the real words, so a visibly-talking mouth matches what is heard.
The video model speaks switch overrides this and sends every line — narrator included — through the video model instead. Turn it on when you want the clip to carry all of its own audio.
When the switch is on, the lines being sent are listed underneath it, so you can see exactly what the model was asked to say.
If a character is mumbling gibberish, it is almost always because the model was given a talking mouth but no words. Check the Audio tab to confirm the lines are reaching the model, and the Debug tab to see the assembled prompt.
Motion Tab
The Motion tab is built in layers. You describe what happens in plain English; the app converts that into timestamped direction the video model follows.
1 · Action outline
One action per row, in order, in plain English — "He throws the puck hard, like a baseball pitch."
- Press Enter to add the next action
- Use the up/down arrows to reorder rows without retyping them
- The camera wish field below is optional — say what you want the camera to do and it is honoured; leave it blank and the director style decides
Your outline saves as you type, so it survives navigating away. A Saving… / Saved marker next to the action count tells you when it is safe to leave.
2 · Choreography
Press Generate choreography from outline and the direction skill converts your rows into timestamped beats — [0s], [3s], [6s] — written in the grammar the chosen video model actually follows.
Every action you wrote survives into the timeline, but the beat budget governs how many beats they land in (roughly one beat per three seconds). If you write more actions than the duration can hold, adjacent ones are merged rather than dropped — an overloaded timeline makes video models simplify or silently skip things.
The generated choreography stays hand-editable. Your edits are what gets sent, until you generate from the outline again.
3 · Director notes
One camera move, one line. Stacking moves causes jitter, so a single continuous treatment reads best.
3.5 · Ambient & SFX
The ambient bed and sound effects the video model adds. Clear it to drop model-generated sound entirely.
This field does not follow your action edits. If you rewrite the outline, check that the ambient description still matches — a stale bed describing things that no longer happen is a common source of odd audio.
4 · Final prompt
The exact text submitted to the model, assembled from the sections above plus the scene's dialogue. Read it before spending credits — everything the model sees is here.
Naming look-alikes
When several characters in frame look alike — clones, twins, a crowd, matching uniforms — a wardrobe detail like "the one with glasses" cannot be resolved and the model will pick at random.
The direction skill handles this by anchoring on position instead: it identifies the character by their side of frame plus a neighbour ("the Sue at frame right, standing between the maroon-shirt clone and the one in the denim jacket") and reuses a short form of that anchor in every later beat.
Post Tab
Operations that run on a finished clip, as opposed to Motion and Avatar which produce one. Each requires a rendered video.
- Generate SFX — Watches the clip and replaces its audio with matched sound effects
- Cinematic Lipsync — Re-voices on-camera dialogue with a chosen voice and matching lip movement
- Revert to original render — Restores the clip as it came out of the render engine
Post operations are mutually exclusive by design: each derives from the original render, so revert first to switch between them.
Graphics Tab
Lower thirds, titles, captions, and logos composited over the finished clip. Nothing here re-renders the scene or costs credits.
See Graphics for the full guide.
Result Tab
The rendered clip for this scene, plus:
- History — Every previous render of this scene, with one-click revert. If a regeneration comes out worse than what you had, you can go back to the earlier take rather than paying to try again.
- Trim — Adjust the clip's in and out points
- Transition — How this scene cuts into the next, including match cut, which conditions the clip to end on the next scene's opening frame for a seamless join
Debug Tab
A dry run of the exact request Generate image or Generate video would submit — the final assembled prompt, the conditioning and reference images in submission order, the resolved spoken lines, and the raw request body. Nothing is sent and nothing is charged.
This is the fastest way to answer "why did it do that?" before spending anything.
Pre-flight check
Before you spend credits on a render, Re-analyze runs an AI review of the pending request — the exact prompt and the starting image — and flags problems that would make the generation come out wrong.
It catches things like an effect whose precondition is missing from the still, more action than the duration can show, stacked camera moves, or dialogue that is too long for the clip.
It is advisory, not a gate. Findings are ranked by whether they break the shot, and text you wrote yourself is treated as a deliberate choice rather than a defect.
Smart Retry
If any images or videos fail to generate, a failure summary banner appears at the top of the scenes view with:
- A count of failed items
- Smart Retry — Automatically retries only the failed items
- Regenerate All — Takes you back to the script tab for a full regeneration
Retries are deduplicated, so clicking twice cannot submit — or bill you for — the same render twice.
Real-Time Progress
During generation, progress banners show:
- Generation Progress — For the initial pipeline (script analysis, image generation, etc.) with phase indicators
- Motion Progress — For batch motion generation, tracking individual scene completion
- Estimated time — Based on the number of scenes and selected model
Progress updates arrive in real-time. If the connection drops, the UI falls back to polling.