How it works
Narration
The voiceover lands on the right frame.
Write or edit the script and Clueso works out which moment of the recording each line belongs to, then times the narration to it. No dragging audio along a timeline.
The words are right. The timing has nothing to do with what is on screen.
Type the narration you want, or change three words in the script you already have.
What Clueso does
Matches lines to frames
It pulls keyframes from the clip and a vision model pins each part of the script to the moment it describes.
What you get
Timing you didn't do by hand
The narration follows the footage, and re-running after an edit re-derives it.
What it actually does
The controls an evaluator asks about, not a restatement of the headline.
Matched to frames, not to a clock
Clueso extracts keyframes from the clip and decides which frame each part of the script is about — so the timing survives a re-cut.
Runs on every clip in a build
When the agent builds a project, every video clip is synced, whatever the footage came from.
Up to 20 sync points per clip
Past that, the clip wants splitting — and the editor says so rather than failing quietly.
Clips up to five minutes
A longer clip is split first, with a message that tells you why.
Three ways in
The sync button in the editor, Clueso Agents, or MCP — all three share one queue.
Source audio is respected
A clip set to keep its own audio is skipped rather than overridden.
Re-sync after a script edit
Change the words and re-run; the timings re-derive against the same footage.
Without it
With Auto Sync
Without it
Changing one line of narration meant re-timing everything after it by hand.
With Auto Sync
Change the words. The timing re-derives itself against the footage.