Caption first
YouTube captions already contain a timeline. Mono uses that timeline as the timing source, then gives Gemini enough surrounding context to translate spoken units naturally.
Rebuild before translating
Automatic captions often split one sentence across several tiny fragments. Mono groups those fragments into meaningful speech units, preserves speaker changes where possible, then translates the reconstructed unit.
Lock finished blocks
Once a block is translated and checked, Mono treats it as Final and moves forward. The viewer can start watching while later blocks continue to finish.
Why this matters for seeking and playback speed
Because the final subtitle cues keep the original timestamps, the same translated cue can be found again after pause, seek, replay, or a playback-rate change. Timing is a property of the video timeline, not the time the AI response arrived.