Editable browser narration
Generate AI voiceovers directly inside the video timeline
Timeline Studio creates multilingual narration in the browser and stores each result as an editable audio asset. Generated speech can be timed, trimmed, faded, linked to captions and mixed with the rest of the project instead of replacing earlier results.
Supported voice routes
Chinese and mixed Chinese/English synthesis uses the built-in Hojo TTS Light 80M speakers 晴岚 and 若溪. English also includes Kokoro voices, while additional browser voices cover German, Spanish, French, Italian and Brazilian Portuguese. Actual availability depends on the selected language and the model assets available to the browser.
Authorized voice conversion
OpenVoice V2 is used as a second-stage Chinese and English voice converter, not as a standalone text-to-speech engine. A user must explicitly authorize enrollment of a reference voice. Saved profiles remain browser-local and completed conversions go to My assets; they do not replace timeline audio automatically.
Two visible stages, one user action
When a saved target voice is selected, Timeline Studio first generates the base language voice and then converts it to the authorized target timbre. Both stages retain independent progress and retry information while presenting the converted result as the final asset.
Editing after generation
Voiceover clips support timing, gain, fades and caption association. Explicit replacement of an existing clip preserves its timing and links while keeping the original restorable.