Timed speech recognition
Generate, correct and style automatic captions in the browser
Timeline Studio uses Whisper small q8 in a browser WASM worker to create editable caption segments. Recognition output is aligned toward nearby speech energy, then displayed on the same timeline used by preview and export.
Captions remain editable
Automatic recognition is a starting point rather than a burned-in result. Caption text, segment boundaries and timing can be reviewed on the timeline. Styling and placement are resolved from the same project state during preview and deterministic export.
Audio-aware timing
Coarse recognition timestamps are adjusted conservatively toward nearby waveform energy. Chinese cleanup is limited to high-confidence contextual corrections instead of rewriting the transcript broadly.
Caption and voice relationships
Captions can remember their source audio association. AI voice generation leaves captions detached by default, while the project-wide link control can restore still-valid remembered associations without reconnecting deleted audio.
What to review
Always check names, numbers, technical terms, punctuation and speech with heavy noise or overlapping speakers. Device performance and source-audio quality affect recognition speed and accuracy.