Skip to the tool

English to IPA, and then hear it said

free, no account, read by people who grew up speaking it

Schwa

noun

The weak vowel in about, taken, pencil, supply and sofa. Five spellings, one sound, and no way to tell from the page.

30 / 500 words
British and American, side by side.

Upload a clip of 3 to 8 seconds. The voice appears in the list above straight away and is already selected. Only you can use it.

The voice is yours alone. Its code stays in this browser. Details

Paste any English text. Schwa writes it in the phonetic alphabet, in the two spellings dictionaries use, then reads it aloud in 25 voices from 9 places where English is a first language. You also get 64 other languages, each read by someone who speaks it.

25
native voices
9
accents, each one measured
71%
same-spelling words read right
$0
always, no account

English spelling does not tell you the sound

Choir is said “kwire”. Colonel is said “kernel”. The letters a, e, i, o and u all spell the same weak vowel in about, taken, pencil, supply and sofa. That weak vowel is the schwa, and it is the most common sound in the language. The phonetic alphabet writes the sound instead of the spelling, so there is nothing left to guess.

  • 8,639 of 20,000 common words differ
  • 78.2% agreement with CMUdict
  • 71% on same-spelling words

The choir sang about the harbour.

British/ðə ˈkwaɪə ˈsæŋ əˌbaʊt ðə ˈhɑːbə/American/ðə ˈkwaɪɚ ˈsæŋ əˌbaʊt ðə ˈhɑːrbɚ/

Two differences, and both are the letter r. It is the single biggest split in spoken English, and it is the reason the accent you pick changes the transcription as well as the voice.

Same spelling, two pronunciations

Nothing inside these words is different. Only the rest of the sentence is, which is why a dictionary lookup cannot tell them apart and why this page reads the whole sentence instead.

  • readstill gets one of these wrong
    1. 1Please read it now. (present tense)
    2. 2I read the book yesterday. (past tense)
  • recordstill gets one of these wrong
    1. 1The record was set in 1968. (noun)
    2. 2Record it before you forget. (verb)
  • presentboth read right
    1. 1She opened the present. (noun)
    2. 2He will present the findings. (verb)
  • windstill gets one of these wrong
    1. 1The wind blew all night. (moving air)
    2. 2Wind the clock before bed. (to turn)
  • liveboth read right
    1. 1They live near the river. (verb)
    2. 2Do not touch a live wire. (adjective)
  • desertboth read right
    1. 1Nothing grows in the desert. (dry place)
    2. 2Do not desert your post. (to abandon)
  • objectboth read right
    1. 1She picked up the object. (noun)
    2. 2They object to the plan. (verb)
  • closestill gets one of these wrong
    1. 1Please close the door. (verb)
    2. 2The station is close to us. (near)
  • leadstill gets one of these wrong
    1. 1She will lead the group. (to go first)
    2. 2The pipes are made of lead. (the metal)
  • bassstill gets one of these wrong
    1. 1He plays the bass guitar. (low sound)
    2. 2They caught a bass in the lake. (the fish)
  • tearboth read right
    1. 1A tear ran down her cheek. (from crying)
    2. 2Do not tear the paper. (to rip)
  • dovestill gets one of these wrong
    1. 1A dove landed on the roof. (the bird)
    2. 2He dove into the water. (past of dive)

Nine tools on the first page of Google turn English into phonetic script. The one that ranks first does not say a single word out loud. How the transcription works, with the measurements.

25 voices from 25 different people

Grouped by where their English comes from, and every group was checked against the recording rather than trusted from a label. The waveform beside each name is measured on the exact file you can play, not drawn. The number in Hz is the median pitch of that same file.

  • 11 women · 14 men
  • 9 accents
  • Common Voice · CC0 1.0

United States

  • Kaylee

    Woman 212 Hz · clear

  • Cheyenne

    Woman 184 Hz · full

  • Wyatt

    Man 122 Hz · full

  • Dashiell

    Man 112 Hz · full

Canada

  • Alanis

    Woman 201 Hz · middle

  • Geneviève

    Woman 184 Hz · full

  • Gilles

    Man 153 Hz · middle

  • Grayson

    Man 128 Hz · full

England

  • Marjoriechosen

    Woman 180 Hz · full

  • Ralph

    Man 121 Hz · full

Scotland

  • Mhairi

    Woman 197 Hz · middle

  • Eilidh

    Woman 191 Hz · middle

  • Hamish

    Man 122 Hz · full

  • Alasdair

    Man 110 Hz · full

Wales

  • Rhys

    Man 104 Hz · deep

Ireland

  • Eoin

    Man 124 Hz · full

Australia

  • Bronte

    Woman 216 Hz · clear

  • Lachlan

    Man 148 Hz · middle

New Zealand

  • Aroha

    Woman 189 Hz · middle

  • Ngaire

    Woman 181 Hz · full

  • Manaia

    Man 164 Hz · middle

  • Tamati

    Man 117 Hz · full

South Africa

  • Thandi

    Woman 165 Hz · full

  • Sipho

    Man 131 Hz · middle

  • Jacques

    Man 111 Hz · full

Every recording carries a speaker id, so we know each voice is its own person rather than one voice at a different pitch. Source: Mozilla Common Voice 17, licensed CC0 1.0. The nine accents, and how each one was checked.

Eight seconds of your own voice is enough

Useful if you teach: students hear the model in a voice they already know.

Recordings up to 2 MB, at most 5 voices a day from one IP address. Your voice gets a random code that is stored only in this browser: Schwa keeps no shared list of voices, no page shows the code, and there is no route by which your voice could become public. The other side of the same coin: if you clear your browser data or move to another computer, the voice is gone.

  1. 01

    The upload box sits directly under the voice picker at the top of the page. Give it a clip of three to eight seconds: one person, clear speech, no music behind.

  2. 02

    The machine learns the voice in about ten seconds. After that it stays in your list and is already selected.

  3. 03

    Use it like any other voice, including for other languages. The words come out right, but your own accent comes through.

  4. 04

    Press Delete whenever you want. The voice leaves this computer, and the server is told to forget the recording.

Four steps, under a minute

  1. 01

    Paste your text

    Type or paste into the box above. Up to 500 words at a time, as many times as you like. Numbers, dates and abbreviations are read out in full.

  2. 02

    Read the transcription

    You get both spellings side by side: the British one used by Cambridge and Oxford, and the American one. Words that change between the two are marked, so the difference is visible rather than described.

  3. 03

    Pick an accent

    Every voice belongs to its own person, and every accent group was checked against the recording rather than trusted from a label. You can listen to all of them before you type anything.

  4. 04

    Download the file

    Five hundred words takes about seventeen seconds. You get an MP3 with no watermark: for a lesson, for a video, for paid work.

64 more languages, each read by someone who speaks it

Every one of these 64 languages has a voice built from a recording of a person who grew up speaking it, and a speech recogniser checked each one separately before it got here. A language with no native-speaker voice has no listen button: an English voice reading Hausa tells you nothing about the result. The sentence under each name is from the Universal Declaration of Human Rights.

Se hele listen med resultatet for hvert sprog

Beside the other free tools

Numbers read off each provider’s own pages on 5 September 2026.

FeatureSchwatoPhoneticsIPA ReaderPhoneticGen
Writes English in phonetic scriptyes, both spellingsyes, both spellingsno, you supply the scriptyes
Says it out loudyesnoyesyour browser voice
Whose voice you hear25 people, one recording eachno voice at alla speech enginewhatever your device has
Accent stated, and checkedstated and measuredno voice at allnono
Uses the sentence to pick the right readingyesno, it lists the optionsnot applicableno
Free in one go500 words (about 3,100 characters)no stated limitno stated limitno stated limit
Account needednononosign-in offered
Where the voices come fromCommon Voice, CC0 1.0no voice at allnot statednot stated

Questions people ask

Which phonetic spelling do you use?

Both of the ones dictionaries use. The British column follows Received Pronunciation, the spelling in Cambridge and Oxford: bath is /bɑːθ/, car is /kɑː/. The American column follows General American: bath is /bæθ/, car is /kɑːr/. Stress marks sit at the start of the syllable, the way every dictionary prints them, not next to the vowel.

Does it get words like “read” and “record” right?

Usually. Those words are spelled one way and said two ways, and the only thing that separates them is the rest of the sentence. Schwa reads the whole sentence rather than looking words up one at a time. On 24 test sentences built in matched pairs it picked the right reading 71% of the time. It is not perfect, and the page lists the pairs it still gets wrong instead of claiming it does not miss any.

How many words are said differently in British and American English?

Out of the 20,000 most common English words, 8,639 come out differently in the two spellings, which is 43.2%. Most of that is one thing: the letter r after a vowel, which is said in America and Canada and dropped in most of England, Australia and New Zealand. The rest is a smaller set of real differences, like the vowel in bath and the first syllable of schedule.

How do you know a voice really has the accent you say it has?

We measure it instead of trusting the label. Every speaker in Common Voice states their own accent, but a voice built from a recording can lose it, or pick up one that was never there. So each voice reads a sentence full of words with r after a vowel and a control sentence with none, and both are measured for the resonance that drops when an English r is curled. 13 voices came out r-coloured and 15 did not. 5 of the 31 read against their own accent, mostly by acquiring an American r the speaker never had, and those are not on the site.

How many voices are there, and where are they from?

25 English voices, 11 women and 14 men, from 25 different people across 9 places where English is a first language. They are built from Mozilla Common Voice recordings, which are CC0 1.0, public domain, with no conditions attached. They are not one voice at different pitches.

Why is Indian English not on the list?

Not for lack of recordings. Indian and South Asian English is the third largest group in the whole English corpus, ahead of six of the accents that are on the site. The site draws its line at communities where English is the first language of the community itself, because that is the only line here that is not arbitrary, and because the phonetic spelling the page prints is British or American. Drawing it elsewhere would be a different site, and a good one.

Is it really free?

Yes. Up to 500 words (3,900 characters) at a time, with 5 seconds between two goes. No account, no card, no watermark on the MP3. You can use it in work you are paid for.

Does it read other languages?

Yes, 64 more, and each one is read by a person who speaks it, not by an English voice with a foreign accent. 30 of them have their own page with a sample you can hear before typing anything.

What happens to my text?

It is not stored. The server keeps only a checksum, so that if the same text is asked for twice with the same voice it can hand back the finished file instead of using the graphics card again. Audio files are deleted after 24 hours.

Updated 2026-09-05