Every other text-to-speech tool hands you an MP3 and stops. Then you open a video editor, line the audio up by hand, and export it again. Studio is that second half, in the place the voice was made.
100,000 characters a month on the free plan. No card, no trial clock.
All of them on footage you already have, and none of them needing a second application.
Replace a clip's audio with your generated voice, or keep the original underneath at a level you choose. Drag the voice along the timeline to land it on the right frame.
Up to twenty stills, crossfading, with an optional slow push on each. The voice sets the length and the slides divide it evenly.
Up to ten, end to end, with a straight cut or a crossfade. Different sizes are letterboxed to fit rather than cropped, so nothing loses its edges.
Cut a section out and keep it. Where nothing needs re-encoding it copies the streams instead, which is why it finishes almost immediately.
Cue timings are recovered from the rendered audio — the pauses are found and the words distributed across the speech between them, so a dramatic silence does not stretch the caption in front of it.
Any section, with the width and smoothness you want. A GIF stores every frame whole, so the screen tells you what that costs before you commit.
Any video in, an MP3 out. Useful when the recording you need is trapped inside footage someone sent you.
Landscape, portrait or square, up to 4K. Keeping the original size skips re-encoding entirely — and nothing is ever cropped to fit.
Drop in footage, pick a voice you already generated, and drag it along the timeline until it sits where it should. The original sound can go, or stay underneath at a level you choose.
Up to twenty stills become one video. The voice sets the length and the slides divide it evenly, crossfading, with an optional slow push on each so nothing sits still.
Generate it on the Text to Speech screen, in whichever engine and mood suits the piece. Studio picks it up from your history.
A video, or the stills you want narrated. Studio reads the file itself rather than trusting its name, so a mislabelled upload fails immediately with a reason.
Move the voice to the frame it belongs on, decide what happens to the original sound, choose the framing, and let it run.
Studio is editing, not synthesis. It runs on processing time we already pay for.
| Detail | |
|---|---|
| Character quota | Untouched — Studio never consumes it |
| Upload size | 100 MB on free, 500 MB on paid plans |
| Source length | Five minutes on free, twenty on paid plans |
| Slides in a slideshow | Up to 20 |
| Clips in a join | Up to 10 |
| GIF length | Up to 15 seconds |
| Keeping your renders | Free renders are removed after a few hours; paid renders are kept |
A render made on the free plan is there to try the feature, not to store it. Upgrade before the window closes and the file stays — the plan is checked again at deletion, not only when the render finished.
The voice, the picture and the subtitles in one place, with nothing to line up by hand. Studio is on every plan, and it never touches your character allowance.