AI · July 8, 2026 · 6 min read
You don't need a mic, a quiet room, or a good speaking day to narrate a video anymore. Type your script, pick a voice, and an AI voiceover drops onto your timeline as an audio clip you can trim and mix. Here's how to add a text-to-speech narration to any video, right in the browser — no separate app, no export-and-reimport dance.
Why type your narration instead of recording it
Recorded voiceovers are a pain. You need decent audio gear, a room without an echo, and the patience to re-record every line you fumble. Then you clean up the breaths and the "ums." For a lot of videos — faceless tutorials, explainers, product walkthroughs, screen recordings where nobody's on camera — a clean synthetic voice is faster and honestly sounds more consistent than a tired human take at 11pm.
Text-to-speech has also gotten genuinely good. The voices here don't sound like a GPS unit; they have natural rhythm and warmth. And because you're typing, you can rewrite a line and regenerate in seconds instead of setting up the mic again.
What you need
- A recent Chrome, Edge or other Chromium browser
- A recording or clip open in the editor — a screen capture, a slideshow, some b-roll, anything you want narrated
- Your script, or a rough draft of it — up to 4000 characters per voiceover
- An EZ Web Streams Pro plan (AI Voiceover is a Pro feature)
Step 1 — Open the AI Voiceover wizard
In the editor, open the Insert menu and choose AI Voiceover…. A wizard opens with a big script box and a row of voices. That's the whole tool — no plugin to enable, no account to link somewhere else.
Step 2 — Write your script
Type or paste your narration into the box. You get up to 4000 characters, which covers several minutes of speech. A few things that make synthetic voices sound better:
- Write the way you talk — short sentences, contractions, plain words
- Use commas and periods to control the pacing; punctuation becomes pauses
- Spell tricky names phonetically if the voice trips on them
- Break a long script into logical chunks so you can place each part precisely later
Step 3 — Pick one of the 6 voices
Choose the voice that matches your video's tone:
- Alloy — neutral and clear, a safe default for tutorials
- Echo — warm, good for friendly walkthroughs
- Fable — expressive, nice for storytelling and ads
- Onyx — deep, authoritative for explainers
- Nova — bright and upbeat, works for quick promos
- Shimmer — soft and calm, great for relaxed voice-overs
Not sure? Generate with one, listen, and regenerate with another — each render costs about a penny, so trying two or three is basically free.
Step 4 — Generate
Click Generate. Under the hood, OpenAI's tts-1 model renders your script in the voice you picked and the finished narration drops onto the timeline as an audio clip. It's saved with your project, so it's still there when you reload the tab.
Step 5 — Trim, fade and mix it in
Because the voiceover is a normal audio clip, you can shape it like any other track. Drag it to line up with the right moment on screen. Trim the ends, split it to insert a pause, and add a short fade in and out so it doesn't start abruptly. If you've got background music from the royalty-free library, lower the music under the narration — or use the "duck under speech" option so the music dips automatically whenever the voice is talking.
When it sounds right, export to MP4 with the platform preset you need and the voiceover is baked into the file.
Narrate your next video without a mic
Type a script, pick a voice, and the AI voiceover lands on your timeline.
Start free →How it compares
The best-known text-to-speech tools are ElevenLabs and Play.ht. Both make great voices, but they're separate products with their own paid subscriptions, and they're built to hand you an audio file. So the workflow is: write your script over there, generate, download the WAV or MP3, then come back to your editor and import it. Every time you tweak a line, you repeat the whole round-trip. Descript's Overdub keeps things in one app, but Descript is a desktop install you have to download and manage.
In EZ Web Streams the voiceover is generated inside the same editor where you're already cutting the video. There's no export, no re-import, no second subscription — you type in the wizard and the clip appears on the timeline next to your footage and music. Rewrite a sentence and regenerate without leaving the tab. And it runs in the browser on any OS, not just a Mac or a downloaded app.
FAQ
How many AI voices can I choose from?
Six. Alloy is neutral, Echo is warm, Fable is expressive, Onyx is deep, Nova is bright, and Shimmer is soft. You preview and pick one in the AI Voiceover wizard before you generate, and you can generate again with a different voice if the first isn't the right fit.
Can I edit the voiceover after it's generated?
Yes. The narration lands on your timeline as a normal audio clip, so you can trim it, split it, move it, fade it in and out, and duck it under background music. It's not a locked render — it behaves like any other audio track, and it persists when you reload the project.
How much does an AI voiceover cost?
AI Voiceover is a Pro feature. Each generation costs roughly pennies because it uses OpenAI's tts-1 text-to-speech model. There's no separate voice subscription to buy — it's included with Pro at $7.99/mo, and the free plan still records up to 5 minutes at 720p, unlimited, forever.
How long can the script be?
Up to 4000 characters per voiceover, which is roughly 600–700 spoken words or several minutes of narration. For longer videos you can generate multiple voiceover clips — one per section — and place each one where it belongs on the timeline.