Audio/Video Transcriber
Turn speech in audio or video into text with timestamps, in Hindi, English and more. Nothing is uploaded.
Fastest. Use it for English speech only.
The model downloads once from huggingface.co and is then kept in this browser. Your audio never leaves this device.
Runs in your browser. Long files need time and memory: a 10-minute recording takes a few minutes on a laptop.
Transcript
Audio/Video Transcriber FAQ
Can it transcribe Hindi?
Yes. Choose Hindi + 90 languages and set Spoken language to Hindi (or Detect automatically). The text comes out in Devanagari. It works best with clear speech; check the result, because noisy recordings and Hindi-English mixing cause mistakes.
Is my audio uploaded to a server?
No. Your file is decoded and transcribed inside your browser and never leaves your device. The only downloads are the speech model from huggingface.co, the first time you use it, and the speech engine from this site.
How do I make subtitles for a video?
Drop the MP4 or WebM video, click Transcribe, then click SRT or VTT. Most video editors and YouTube accept SRT files; VTT is the format for web players.
Why is the first transcription slow to start?
The speech model has to download once: 41 MB for English or 76 MB for the multilingual model. After that it loads from your browser's cache unless you clear this site's data.
What is the difference between the two models?
English only (Whisper tiny.en) is smaller and faster but understands only English. Hindi + 90 languages (Whisper base) is larger and slower but handles Hindi, other Indian languages and many more, and can detect the language for you.
Which file formats can I transcribe?
Anything your browser can play: MP3, WAV, M4A/AAC, OGG/Opus, WebM and the sound in MP4 videos. MKV and AVI often fail; convert them to MP3 first with the Universal Conversion Suite.