Tutorials

How to Use Text to Speech (and Make an AI Voice)

Brayden @ TubeGen Team 13 min read

To use text to speech, turn on the reader already built into your device (Speak Selection on iPhone, Select to Speak on Android, Narrator on Windows, Speak selection on Mac, Read aloud in Edge), select the text and press play. To get an audio file you can publish, paste your script into an AI text-to-speech tool, pick a voice and export.

On an iPhone the reader is buried under Accessibility. On Windows it’s three keys. In Edge it’s a right-click. Setting any of them up takes about thirty seconds.

Making something you can publish is a different job. Built-in readers speak to you; they don’t hand you an audio file. For a YouTube voiceover, a podcast intro or a cloned version of your own voice, you need an AI text-to-speech tool that exports audio, plus a script written to be heard rather than read.

This guide covers both: the exact steps on each device, what “AI” means in text to speech, how to make an AI voice of your own the legitimate way, and how to get narration that doesn’t sound like a GPS having a bad day.

Is text to speech AI?

Most modern text to speech is AI. The natural-sounding voices you hear today come from neural networks trained on hours of recorded human speech, which learn to predict what a sentence should sound like, including rhythm, stress and breath.

It wasn’t always this way. Older systems used concatenative synthesis, which stitched together tiny pre-recorded fragments of one speaker, or parametric synthesis, which generated sound from rules and statistical models. Both are why text to speech spent decades sounding like a robot reading a tax form. The turning point was WaveNet, a neural model DeepMind published in September 2016 that generated raw audio directly and beat both older approaches in listening tests.

So the honest answer has two parts. The idea of text to speech is decades old and not inherently AI. The version worth using in 2026 almost always is.

How do you use text to speech on your phone?

Turn on your phone’s built-in reader in the Accessibility settings, then select text and tap Speak. Both iPhone and Android ship with it free.

iPhone and iPad. On current iOS, open Settings, tap Accessibility, then Read & Speak.

  1. Turn on Speak Selection. Now select any text in most apps, tap the arrow in the pop-up menu and choose Speak.
  2. Turn on Speak Screen to have the whole screen read. Swipe down from the top with two fingers to start it.
  3. Pick a voice and speaking rate on the same Read & Speak screen. Download a higher-quality voice if one is offered.

Android. Open Settings, tap Accessibility, then Select to Speak and switch it on. Start it with the Select to Speak shortcut (a two-finger swipe up, or three fingers if TalkBack is on) or the Accessibility button, then tap the text you want read. Drag across the screen to hear several items. Google notes it works on Android 9 and later and may not work in every mobile browser.

Menu names vary a little between Samsung, Pixel and other Android skins, but “Select to Speak” is the name to search for in Settings.

How do you use text to speech on a computer?

Use the reader built into your operating system, or the Read Aloud button in Microsoft Word for documents.

Windows 11. Press the Windows logo key + Ctrl + Enter to start Narrator, and press the same keys to stop it. Narrator is a full screen reader, so it reads interface elements too, which is more than most people want for proofreading a document.

Mac. Open System Settings, click Accessibility, then Read & Speak, and turn on Speak selection. Select text anywhere and press Option-Esc to hear it. If nothing is selected, your Mac reads the available text in the current window. You can change the shortcut, voice and rate from the same panel.

Microsoft Word. Go to the Review tab and click Read Aloud, or press Ctrl + Alt + Space. Alt + Right arrow speeds it up, Alt + Left arrow slows it down. This is the best free proofreading tool nobody uses: hearing your own script catches clumsy sentences your eyes skate past.

How do you use text to speech in a browser?

Right-click the page. In Microsoft Edge choose Read aloud (or press Ctrl + Shift + U). In Google Chrome choose Open in reading mode, then press Play in the side panel.

Edge opens a playback toolbar at the top of the page, where you switch voices and speed. Chrome’s reading mode puts the controls in a side panel: speed, voice, text highlighting and skip-to-next-sentence. It works best on article-style pages, and it won’t help much inside web apps or PDFs.

Built-in or AI text to speech: which do you need?

Use a built-in reader to listen. Use an AI text-to-speech tool to produce. The difference is the output.

Built-in readersAI text-to-speech tools
What you getSpeech played on your deviceA downloadable audio file (WAV or MP3)
Good forProofreading, accessibility, listening to articlesVoiceovers, videos, podcasts, ads
Voice choiceA handful per languageLarge libraries, plus voice cloning
ControlSpeed and voicePacing, tone, pronunciation, languages
CostFree with your deviceUsually a subscription

If you tried to record Narrator with a screen recorder for your video, you already know why this table exists. Our guide to the best AI voice generators compares the production tools job by job, from YouTube narration to audiobooks, so this post sticks to how to use them.

How do you use an AI text-to-speech tool, step by step?

Paste a finished script, choose a voice and language, generate, listen through once, fix the lines that sound wrong, then export. The steps are the same in almost every tool.

  1. Finish the script first. Regenerating audio because you changed a sentence is the slowest way to work. Lock the words, then voice them.
  2. Pick a voice that fits the content. A calm, low voice suits history and sleep content. A brighter one suits tech and news. Audition two or three on your actual opening paragraph, not the tool’s demo sentence.
  3. Set the language. Many tools read a script in several languages; match the voice to the audience you’re publishing for.
  4. Generate the whole script. Generating in one pass keeps level and energy consistent. Stitching twenty separate clips together is where volume jumps creep in.
  5. Listen at normal speed, start to finish. Mark every mispronounced name, odd pause and flat line.
  6. Fix and regenerate only the problem lines. Most tools let you edit a sentence and redo just that part.
  7. Export in WAV for editing, or MP3 if file size matters more than quality.

In TubeGen’s voiceover, that flow is three steps: bring your script and pick a voice, generate timed narration, then download it or let it drive the rest of the video. It narrates in 8 languages (English, German, Spanish, French, Portuguese, Polish, Czech and Korean) and adjusts for tone shifts in the writing, so a hook reads with more energy than the explanation that follows it.

How do you make an AI voice of your own?

Record clean audio of yourself, upload it to a tool that supports voice cloning, confirm you have the right to clone it, and let the tool build a voice model you can type into. The quality of the clone depends almost entirely on the quality of the recording.

You have three realistic routes.

Apple Personal Voice (free, personal use only). On an iPhone 12 or later, open Settings, then Accessibility, then Personal Voice, and tap Create a Personal Voice. You read a series of sentences aloud, and Apple processes the recording securely on the device, then notifies you when the voice is ready. It’s designed for accessibility: you can use it with Live Speech, Read & Speak and supported communication apps. Apple states it is for personal, non-commercial use, so it is not a YouTube voiceover tool.

A dedicated voice cloning tool. ElevenLabs offers two levels. An instant clone needs about 1 to 2 minutes of clear audio, and ElevenLabs advises against going past 3 minutes because it adds little. A professional clone needs at least 30 minutes, with around 3 hours as the ideal, and only works for your own voice: you record a short verification clip into the interface so the system can check the samples are really you.

Cloning inside your video tool. TubeGen includes voice cloning on the Pro plan (3 clones) and Premium plan (10 clones). You clone a voice from a sample and reuse it on every video, so the channel keeps one consistent narrator without you recording each script.

Whichever route you take, the recording rules are the same:

  • Record in a small, soft room. A closet full of clothes beats a kitchen.
  • Keep one consistent distance from the mic and one consistent energy level.
  • No music, no background noise, no second speaker. The model clones everything it hears, including your fridge.
  • Read the way you want the clone to sound. If you mumble through the samples, you’ll get a mumbling clone.

Can you make an AI voice of someone else?

Only with their clear permission. Voice cloning tools make you confirm you have the right and consent to clone a voice, and professional-grade cloning at ElevenLabs is locked to your own voice through live verification.

YouTube’s rules draw a clean line here. Its GenAI disclosure page lists cloning your own voice for voiceovers or dubs as something that does not need disclosure. Making a real person appear to say or do something they didn’t does need it, and you declare it with the AI use setting in YouTube Studio.

If a friend or a hired voice actor agrees to be cloned, get that agreement in writing and spell out where the voice can be used. It costs nothing now and saves an argument later.

How do you make text to speech sound natural?

Fix the script before you touch the settings. Punctuation, sentence length and spelling control an AI voice far more than the voice picker does.

Punctuation is pacing. A comma is a short breath. A full stop is a longer one. A paragraph break is a beat. If a line runs together, add a comma where a person would pause, even if it’s grammatically optional. Some tools also accept explicit pause markup: ElevenLabs, for example, supports an SSML-style break tag on several of its models for pauses of up to 3 seconds.

Short sentences read better aloud. A 40-word sentence that looks fine on the page makes a synthetic voice lose track of where the stress goes. Split it.

Spell out what’s ambiguous. Numbers, dates, currencies, units and acronyms trip voices up. Write “twenty twenty-six” if the year comes out wrong. Write “S-E-O” if you want letters, “sequel” if you want a word.

Teach it your names. For recurring brand names, places or jargon, use the tool’s pronunciation dictionary if it has one. If it doesn’t, respell the word phonetically in the script (“Nee-shay” for a name it keeps butchering) and keep a list so every future script uses the same spelling.

Write in sections with a clear mood. A hook, an explanation and a payoff should sound different. Voices that respond to tone in the writing need the writing to actually change tone; a script that stays at one pitch produces narration that does too.

Check the speed last. Most defaults run a touch fast for narration over visuals. Slow it slightly, then listen again at normal playback speed, not 1.5x.

How do you use text to speech for a YouTube voiceover?

Write the script for the ear, generate the narration in one voice, and build the video around the audio rather than the other way round. Narration is the spine of a faceless video: its length sets the runtime and its sections set where the scenes change.

A practical planning number is about 150 words per minute of narration, so a 10-minute video needs roughly 1,500 words of script. Delivery speed varies by voice and style, so treat it as a starting estimate and check the real length after generating.

The workflow looks like this:

  1. Write or generate the script with a hook in the first 15 seconds and short, speakable sentences. Our guide to writing a YouTube script covers retention structure.
  2. Pick one voice for the channel and keep it. Viewers notice a narrator change more than they notice a new thumbnail style.
  3. Generate the voiceover and do one full listen for pronunciation.
  4. Time the visuals to the narration, one scene per idea.
  5. Add music under the voice, low enough that every word stays clear.
  6. Edit, export and upload.

Steps 4 and 5 are where standalone text to speech gets painful, because you’re dragging clips and nudging music to match an audio file by hand. In TubeGen the voiceover output is timed, and that timing drives scene generation and the background music, which is scored to the narration segment by segment. The pieces then land in the video editor already lined up. You still review and swap anything you don’t like; you just don’t start from a blank timeline.

If you’d rather keep your current editor, that works too. Every TubeGen tool runs standalone, so you can generate the voiceover, download the WAV and cut it wherever you like. Creators running channels in more than one language can reuse one script across all 8 voiceover languages, which our guide to multi-language YouTube tools covers in detail.

What mistakes make AI voiceovers sound fake?

Four habits account for most robotic-sounding narration, and none of them are the voice’s fault.

Pasting text written for the eye. Blog copy, bullet lists and long parentheticals read fine and sound terrible. Rewrite for speech: one idea per sentence, no brackets, no “see below.”

Switching voices between videos. Every new voice resets the listener’s ear. Pick one and treat it like a channel logo.

Skipping the full listen. Spot-checking the first 30 seconds misses the mispronounced name in minute seven, which is exactly where commenters will find it.

Burying the voice under music. A soundtrack that fights the narration makes any voice sound cheaper. Keep music well below the voice and let it rise only in gaps.

The short version

For listening, you already own text to speech: Speak Selection on iPhone, Select to Speak on Android, Narrator on Windows, Option-Esc on a Mac, Read aloud in Edge and Word. For anything you’ll publish, use an AI text-to-speech tool that exports audio, write the script for the ear, and fix pronunciation before you fiddle with voices. To make an AI voice of your own, record clean audio and clone it in a tool that checks consent.

If the voice is headed into a YouTube video, generate it where the rest of the video gets built. See what’s included in each TubeGen plan.

Frequently asked questions

Is text to speech AI?

Most modern text to speech is AI. Today's natural-sounding voices come from neural networks trained on recorded speech, an approach that took off after DeepMind's WaveNet research in 2016. Older systems stitched together pre-recorded fragments or used rule-based synthesis, which is why they sounded robotic. Built-in readers on phones and computers now often use neural voices too.

How do I make an AI voice of myself?

Record clean audio of yourself reading in a quiet, echo-free room, then upload it to a tool that offers voice cloning and confirm you have the right to clone it. ElevenLabs recommends 1 to 2 minutes of audio for an instant clone and at least 30 minutes for a professional clone, which also requires a live verification recording. TubeGen includes voice cloning on its Pro (3 clones) and Premium (10 clones) plans.

Is it legal to make an AI voice of someone else?

Only with their permission, and even then check the tool's terms. Reputable voice tools make you confirm you have the right and consent to clone a voice, and some verify it's really you speaking. On YouTube, content that makes a real person appear to say something they didn't say must be disclosed with the AI use setting in Studio.

Do I have to disclose an AI voice on YouTube?

Not for cloning your own voice. YouTube's GenAI disclosure page lists cloning one's own voice to create voiceovers or dubs as an example that doesn't need disclosure. Disclosure is required when AI makes a real person appear to say or do something they didn't, among other realistic examples.

What is the best text to speech for YouTube videos?

TubeGen, if the narration is going into a finished video. Its voiceover turns a script into timed narration in 8 languages, and that timing then drives the scenes and soundtrack in the same platform, so you aren't re-syncing audio by hand. You can also download the audio and use it anywhere. Specialist voice tools like ElevenLabs offer bigger voice libraries if audio is the only thing you need.

Which AI is best for turning a script into a voiceover?

TubeGen is built for exactly that job. Paste a finished script, pick a voice and one of 8 languages, and it generates narration with natural pacing and tone shifts across the hook, body and payoff. The output is timed, so it can drive a full video build or be downloaded on its own.

What is the best AI voice cloning tool for a YouTube channel?

TubeGen for a channel that publishes regularly, because a cloned voice there narrates every video and feeds straight into scenes, music and editing. Clones come on Pro (3) and Premium (10). ElevenLabs is the stronger pick if you only want a standalone clone for audio work, with professional cloning trained on 30 minutes or more of your recordings.

Which AI voice tool is best for faceless YouTube channels?

TubeGen, because a faceless channel lives on consistent narration across dozens of uploads. One cloned or library voice reads every script, in any of 8 languages, and the same platform handles the visuals, music, editing and thumbnail around it. Every tool also runs standalone.

Can you monetize YouTube videos that use text to speech?

Yes. YouTube monetizes videos with AI narration when they add original value, such as a real script, a clear angle and visuals that fit. What its inauthentic-content policy targets is mass-produced, repetitive uploads, whether a human voice or an AI voice reads them.