Most text to speech engines read your words. This one performs them — because the mood is a control you set, not something a model infers from punctuation and gets wrong.
100,000 characters a month on the free plan. No card, no trial clock.
Other engines read the punctuation and infer a mood — which is why the same sentence comes back cheerful when you needed it grave. Here the mood is a setting, and the same line in two moods is two different performances.
Our own Prosody model for emotion and voice cloning, OpenAI for the widest language coverage, Gemini for studio-grade regional voices. Pick per project rather than per subscription.
Numbers, dates, currencies and abbreviations are rewritten into speakable words before anything is generated, so "24/7" is not read as a fraction.
Or add a line of direction instead. The character cost is on screen before you commit, counted against the right pool.
MP3, WAV or FLAC by plan. Or keep going — lay it over video in Studio, or hand it to your own systems through the API.
Not a demo with a countdown. A working allowance that comes back on the first of every month.
| Free | Paid plans | |
|---|---|---|
| Characters a month | 100,000 — about two hours of speech | Up to 2,000,000 |
| Voice clones | 1 | 1 to 30 by plan |
| Built-in voices | All of them, unlimited use | All of them, unlimited use |
| Premium engines | — | OpenAI + Gemini, shared allowance |
| Export | MP3 | MP3, WAV, FLAC by plan |
| Commercial licence | — | Included |
The premium allowance is one shared pool across OpenAI and Gemini, so you can spend all of it on whichever engine suits the project rather than watching two counters.
Generate the same line calm, then excited, and the point makes itself. A hundred thousand characters a month, no card, and the allowance returns every month for as long as you use it.