Skip to content
Cutshed
Menu

Subtitles & text · 22

Subtitle generator

Get a first pass of captions from a talking-head video or a voice memo, fix the lines, and leave with an SRT and a VTT. Whisper tiny runs locally.

First run downloads Whisper tiny quantized (about 69 MB including the runtime) from this site’s copy of the ONNX Runtime and from the Hugging Face hub for the weights. The weights are cached in this browser. The recording is transcribed locally.

How to get captions you can publish

Record or export the cleanest dialogue you have. A lav or a boom with the music ducked will beat a camera mic in a cafe. If you already separated a vocal or a dialogue stem, transcribe that file instead of the full mix. Music under speech is the fastest way to invent words.

Choose the spoken language before you start. Whisper tiny is multilingual and MIT licensed (OpenAI). The quantized weights are about 45 MB, plus about 27 MB for the ONNX Runtime build this page uses. The status line shows the download. After that, the Cache API keeps the files for the next clip.

Playback and the cue list stay in sync. Click “Play from here” to hear a line in context, then rewrite names, brands, and numbers. Aim for a line a viewer can read in the time it is on screen: roughly 42 characters, one or two lines, and at least about a second of hold. Split a cue that runs across a cut.

Video files are decoded with ffmpeg.wasm inside the tab, down to 16 kHz mono, which is what Whisper expects. Audio files try the browser’s own decoder first. Either way the samples do not leave the device.

Tips

  • Fix spelling before you touch timestamps. Most timing from a clean take is already close enough.
  • YouTube wants SRT. A self-hosted player usually wants VTT. Export both from the same edit.
  • If a whole sentence is wrong, the language menu is the first thing to change, not the spelling.
  • Delete filler cues the model invented during silence, or viewers will read “[silence]” moments that are not there.

FAQ

Which Whisper model is this?

Xenova/whisper-tiny, quantized to 8-bit. It is the smallest multilingual Whisper build that still gives usable timestamps. It will miss names, jargon, and noisy dialogue. The editor is there because of that.

Is the recording uploaded to transcribe it?

No. The weights download into the browser cache. Audio is resampled to 16 kHz mono in the tab and passed to the local model. Video pictures are only used for the preview player.

Should I pick a language or leave it on auto?

Pick the language when you know it. Auto-detect is weaker on short clips, accents, and anything with music under the voice. The list covers the languages this tiny model can be pointed at.

SRT or VTT?

Upload SRT to YouTube as a caption file. Use VTT for an HTML5 video track element. Both exports use the lines you edited, not the raw model text.

Can I burn the captions into the picture?

Not in this version. You get a sidecar file. Your editor or YouTube can place it on the picture.

How long a file can I run?

Twenty minutes or 200 MB. Longer interviews should be cut into reels first. Whisper itself works in 30 second windows with a short overlap so a long file does not arrive as one block.