Text to speech
Type something and hear it read aloud, in one of 28 voices.
What should it say
It stays on this device. Nothing is sent anywhere.
Settings
What makes this one different
Every free text-to-speech site works the same way: you paste your script into a box, it goes to their server, and you get back a watermarked clip, a character limit, and an invitation to subscribe. The script itself, which is often a video narration, a lesson or a message you have not published yet, sits in their logs.
This runs the voice model in your browser. The text is never sent anywhere, there is no character-per-month allowance, and the audio comes out clean.
The voices
Twenty-eight of them, American and British, male and female. They come from Kokoro, an open model released under the Apache licence, which means the audio you generate is yours to use, including commercially. It is worth trying several: they differ more in pacing and warmth than the names suggest, and the right one for a tutorial is rarely the right one for an audiobook.
Speed and length
Generation runs at roughly the length of the audio itself, so a minute of speech takes about a minute, on one CPU thread. That is fine for a paragraph and tedious for a chapter, so this is built for scripts rather than for books. The model is a 326 MB download the first time and is cached afterwards.
Getting a natural read
Punctuation is what the model uses for rhythm. Commas and full stops give it places to breathe, and a wall of text without them comes out rushed. Numbers, dates and abbreviations are read the way they are written, so if you want "Dr." spoken as "Doctor", write it that way. Spelling out an acronym with spaces or full stops usually gets you the letters rather than a mangled word.
Common questions
Can I use the audio commercially?
Yes. The model is Apache-licensed and generates the audio on your machine, so there is no service in the middle claiming rights over the output. Nothing here is watermarked.
Is there a character limit?
The box holds a few thousand characters at a time, which is a practical limit rather than a commercial one: it all has to be generated on your own processor. For something long, do it a section at a time and join the results.
Does my text get sent anywhere?
No. The model weights come down to you once, then the synthesis happens in the page. You can watch the network panel while it runs.
Can it clone a voice?
No, and it is not trying to. You get the 28 voices the model ships with.
It mispronounced a word.
Spell it phonetically and try again. The model reads what is written, so an unusual name often does better spelled the way it sounds.