What happens to your file
An MP4 is a container that holds video, audio and other data. The speech is in the audio track, so converting MP4 to text means reading that track and recognising the words in it.
On this page your browser reads the audio track and cuts it into parts of about a minute. Each part is sent to ElevenLabs’ Scribe v2 model, which the site reaches through fal.ai, and the model returns the text. The browser then joins the parts in order and builds the paragraphs. The video itself stays on your computer. How the converter works goes through the stages in more detail.
Get a better transcript
The recording decides most of the quality.
- Use the original file when you have it. A copy that was compressed again has less detail for the model to work with.
- Clear speech with little background noise gives the best text. A recording of a distant room, music under the voices or several people talking at once gives more mistakes.
- Check names, numbers and technical terms. They are the words the model gets wrong most often.
- If a recording has long silences or music, look at those stretches. The model can add words that were never said.
- If a file mixes languages, read those passages closely.
When you need timecodes or subtitles
Timecodes are off by default. Turn on [HH:MM:SS] timecodes, which are free, and each paragraph starts with the time it begins in the video. A pack adds HH:MM:SS.mmm and HH:MM:SS:FF timecodes, a start timecode and a timecode every 10 to 60 seconds.
For subtitles, download SRT or VTT. Each subtitle has a start and end time, at most two lines of 42 characters and at most 7 seconds. These downloads come with a pack; see pricing.
Using the text for something important
An automatic transcript is a first draft. For court, medical or other records where accuracy matters, have a person check it against the recording. If you need your money back for minutes you did not use, the refund policy covers it.