A voice over generator
that knows when each line lands.
Paste the script. Every sentence becomes a line with a timecode. Each line is rendered on its own and timed from the real audio. Listen with the line highlighted, export MP3 plus SRT, VTT or CSV.
mp3 · srt · vtt · csv
Built for the slow part of editing
The slow part of a voice over is rarely the recording. It is lining up thirty sentences against thirty shots and then retyping them as captions.
Each sentence is rendered on its own, so its start time and duration come from the real audio, not a guess.
The playing line lights up. Click any line to jump to it. Change the pause and the track re-times instantly.
MP3 for the audio track, SRT or VTT for captions, CSV for the timeline. All from the same timing.
What an AI voice over generator should do
Say the words well, obviously. But a voice over generator for video has a second job: tell you where every line starts, so you can cut to it. That is why the timeline is on the page and not hidden in a download. It is also why pauses are a setting: narration over quick cuts wants almost none, narration over long shots wants room to breathe.
Ten voices cover the usual needs: an upbeat read for Shorts and Reels, an even explainer voice for tutorials and product videos, a slow deep voice for documentary narration, a soft voice for long reads. All of them are multilingual, so a video dubbed into Spanish or Japanese can keep its narrator.
Where it stops
- No voice cloning: we will not imitate you, your client or anyone else.
- No music or sound effects library. Bring your own bed.
- One narrator per render; multi-speaker dialogue is not offered.
- Pace is an instruction to the voice, not an exact multiplier; the timeline always reflects the real audio.
Also on the desk
Recording your own voice instead? Transcribe it to get captions. Need the sound out of a reference clip? Video to audio.
Questions
What makes this a voice over generator rather than text to speech?
Timing. A voice over has to land on picture. This tool splits your script into lines, gives each line a start time and a duration, highlights the line that is playing, and exports the same timing as SRT and VTT subtitles and as a CSV you can read into an editor. Text to speech gives you a file; this gives you a file plus a map of it.
Can I export subtitles that match the voice?
Yes. Every line is rendered separately and the timings are measured from the finished audio, so SRT and VTT captions stay in sync with the MP3. The CSV export has line number, start, end and text for spreadsheets or scripting. Exports unlock after the first render; before that the timeline is only an estimate.
Can I change the pause between lines?
Yes, the gap slider adds silence between lines (0 to 2 seconds). Use a longer gap when each line sits over its own shot, a shorter one for continuous narration.
Which voice should I use for Shorts and Reels?
Fast, bright voices read well over quick cuts: try Puck or Zephyr at a brisk pace. For explainers and tutorials, Kore or Zephyr at a natural pace. For documentary-style narration, Charon at a relaxed pace. Listen on the voices page.
Can I voice a video in several languages?
Yes. Paste the translated script, switch the language, keep the same voice: every voice here is multilingual, so the narrator stays consistent across language versions.
Does the AI voice over count as commercial use?
Monetised channels, client work and ads need the Creator or Studio plan. The free plan is for personal and non-monetised projects. See terms.