graphic_eq
ProsodyAI
menu
movie_editStudio

The voice is half of it.
Studio is the other half.

Every other text-to-speech tool hands you an MP3 and stops. Then you open a video editor, line the audio up by hand, and export it again. Studio is that second half, in the place the voice was made.

1,000 premium characters to try, on the house. No card, no trial clock.

Eight things Studio does to a video

All of them on footage you already have, and none of them needing a second application.

movie

Voice over video

Replace a clip's audio with your generated voice, or keep the original underneath at a level you choose. Drag the voice along the timeline to land it on the right frame.

image

Photographs into a narrated video

Up to twenty stills, crossfading, with an optional slow push on each. The voice sets the length and the slides divide it evenly.

library_add

Join clips

Up to ten, end to end, with a straight cut or a crossfade. Different sizes are letterboxed to fit rather than cropped, so nothing loses its edges.

content_cut

Trim

Cut a section out and keep it. Where nothing needs re-encoding it copies the streams instead, which is why it finishes almost immediately.

subtitles

Burn in subtitles

Cue timings are recovered from the rendered audio — the pauses are found and the words distributed across the speech between them, so a dramatic silence does not stretch the caption in front of it.

gif_box

Export a GIF

Any section, with the width and smoothness you want. A GIF stores every frame whole, so the screen tells you what that costs before you commit.

graphic_eq

Pull the audio out

Any video in, an MP3 out. Useful when the recording you need is trapped inside footage someone sent you.

aspect_ratio

Framed for where it is going

Landscape, portrait or square, up to 4K. Keeping the original size skips re-encoding entirely — and nothing is ever cropped to fit.

The finishing work, in the same place

What a video usually goes to a second editor for before it is published.

content_cut

Cut the silences

Long pauses are found and removed automatically. Choose how tight — light, medium or tight — and the rest of the take is left alone.

music_note

A music bed

Lay a track under the voice at the level you set, so the narration stays on top.

branding_watermark

Your logo in the corner

A watermark in any corner, at the size and opacity you choose.

smart_display

Presets for where it is going

YouTube: 16:9, 1080p, loudness at −14 LUFS. Shorts, Reels and TikTok: 9:16. Podcast: −16 LUFS. A preset only fills the fields; you can change any of them.

closed_caption

Subtitles as a file too

Burn them into the picture in one of three styles, or take the .srt alongside the video and upload it to the platform yourself.

photo_camera

A still for the thumbnail

Grab any frame as an image — the starting point for a thumbnail.

Voice over video

Land the line on the frame it belongs on

Drop in footage, pick a voice you already generated, and drag it along the timeline until it sits where it should. The original sound can go, or stay underneath at a level you choose.

  • check_circleThe waveform is drawn from the real audio, so you place it by eye
  • check_circleKeep the music and effects, quieter, or replace them outright
  • check_circleNothing is cropped — a portrait clip stays portrait
VIDEOVOICEdrag it to the frame it belongs on
Image to video

Photographs, narrated, with the timing worked out for you

Up to twenty stills become one video. The voice sets the length and the slides divide it evenly, crossfading, with an optional slow push on each so nothing sits still.

  • check_circleCrossfades sized to the clip so the last slide is never clipped
  • check_circleSubtitles burned in from the audio itself, pauses and all
  • check_circleExport to landscape, portrait or square, up to 4K
one narrated video

How to add a voice over to a video

01

Make the voice

Generate it on the Text to Speech screen, in whichever engine and voice suits the piece. Studio picks it up from your history.

02

Bring the picture

A video, or the stills you want narrated. Studio reads the file itself rather than trusting its name, so a mislabelled upload fails immediately with a reason.

03

Place it and render

Move the voice to the frame it belongs on, decide what happens to the original sound, choose the framing, and let it run.

What it costs you

Studio is editing, not synthesis. It runs on processing time we already pay for.

Detail
Character quotaUntouched — Studio never consumes it
Upload size100 MB on free, 500 MB on paid plans
Source lengthFive minutes on free, twenty on paid plans
Slides in a slideshowUp to 20
Clips in a joinUp to 10
GIF lengthUp to 15 seconds
Keeping your rendersFree renders are removed after a few hours; paid renders are kept

A render made on the free plan is there to try the feature, not to store it. Upgrade before the window closes and the file stays — the plan is checked again at deletion, not only when the render finished.

What Studio is not

It is built for voiced videos: narration, explainers, slideshows, Shorts. For some work you still want a full editor.

In Studio
Layered timelineNo — one picture track, with the voices, a music bed and the original sound beneath it
Keyframes, effects, colour gradingNo
TransitionsA straight cut or a crossfade
Longest sourceTwenty minutes on paid plans, five on free

If a project needs multicam, motion graphics or grading, cut it in a full editor and bring the voice and subtitles from here.

Stop exporting into a second app

The voice, the picture and the subtitles in one place, with nothing to line up by hand. Studio is on every plan, and it never touches your character allowance.