Subtitles & text · 25
Transcript to post
Whisper tiny writes the words. A small set of rules turns the cues into Markdown: a title from the first sentence, paragraphs where the speaker paused, and optional filler removal. No chat API.
How to use it
Pick the language, the way you do on the subtitle page. The same Whisper tiny model runs locally, in 30 second windows. The first run downloads the weights, about 45 MB, from the Hugging Face hub into the browser cache.
The formatter drops “um”, “uh”, “er”, “ah”, “you know”, and “i mean” when the box is checked. It starts a new paragraph when two cues are more than 1.4 seconds apart or the paragraph would pass about 480 characters. The first sentence, if it is under 90 characters, becomes the heading.
Read it. Tiny Whisper mishears names, and the formatter does not know your argument. It only groups time and strips a few fillers. Fix the Markdown before you publish.
Tips
- A clean voiceover becomes a better post than a noisy room. Denoise first if the room is the problem.
- Leave fillers in when the piece is a quote and the hesitation matters.
- For captions, use the subtitle generator. This page is for a post, and it throws away the timestamps.
FAQ
Does a cloud model write the post?
No. Transcription is Whisper tiny in the browser. The post shape is rules in the page: pauses, length, and a short filler list.
Which Whisper model?
Xenova/whisper-tiny, the same 8-bit model as the subtitle generator. It is fast and it makes mistakes. Edit the result.
Is the recording uploaded?
No. The weight download is the model, not your audio.
License
Whisper is MIT (OpenAI), loaded through Transformers.js (Apache-2.0) and ONNX Runtime (MIT). The Markdown formatter is original code in this site. There is no paid language-model API.