Audio/Video Transcriber

Downloads a speech model once (41–76 MB)Productivity#Audio#Local AI

Turn speech in audio or video into text with timestamps, in Hindi, English and more. Nothing is uploaded.

Fastest. Use it for English speech only.

The model downloads once from huggingface.co and is then kept in this browser. Your audio never leaves this device.

Runs in your browser. Long files need time and memory: a 10-minute recording takes a few minutes on a laptop.

Transcript

The transcript appears here, with a time for every line. Click a time to play from there.
How to use, limits & privacy

About Audio/Video Transcriber

Turn speech in an audio or video file into text with timestamps, using the Whisper speech model in your browser. Choose the fast English-only model or the multilingual model, which understands Hindi and about 90 other languages. Your recording is never uploaded; you can copy the text or download it as TXT, or as SRT and VTT subtitles.

How to Use

1

Pick a model

Choose English only (41 MB) for English speech, or Hindi + 90 languages (76 MB) and set Spoken language, or leave it on Detect automatically.

2

Add your file

Drop or choose an MP3, WAV, M4A, OGG, WebM or MP4 file of up to 60 minutes and 500 MB.

3

Transcribe

Click Transcribe. The first time, the model downloads from huggingface.co and is then kept in this browser. Lines appear as they are transcribed.

4

Check and save

Click a time to play that part, then use Copy, TXT, SRT or VTT to save the transcript.

Privacy & Processing

  • Mode: local
  • Files Leave Browser: Local tool processing; review details below
  • Max Input Size: Device memory limits
  • Account Required: No
  • Data Stored Locally: The downloaded models are kept in this browser's cache (Cache Storage, "transformers-cache"); clear this site's data to remove them. Transcripts are not saved.
  • Network Processing: Assets or models may require an initial download

Your audio and the transcript stay on your device. The speech model is downloaded from huggingface.co (Hugging Face) on first use, and the speech engine from this site; nothing about your file is sent.

Rules & Limitations

  • The first use of each model downloads it from huggingface.co (41 MB English, 76 MB multilingual) plus a 10 MB speech engine from this site. Later visits load it from the browser cache.
  • Up to 60 minutes and 500 MB per file. The audio is decoded into memory (about 4 MB per minute), so long files can fail on phones; cut them with the Audio Trimmer first.
  • Transcription runs on your device, so it is slower on older laptops and phones. Expect a few minutes for a 10-minute recording on a laptop, longer with the multilingual model.
  • Accuracy depends on the audio. Hindi and other Indian languages make more mistakes than English, especially with noise, music or Hindi and English mixed in one sentence.
  • It does not label speakers. Music and silence may show up as short notes like [Music].

Top Suggestions

  • Typing up lecture and class recordings for notes
  • Making subtitles (SRT or VTT) for YouTube and Instagram videos
  • Writing down interviews and meeting recordings
  • Getting the text of a WhatsApp voice note

Audio/Video Transcriber FAQ

Can it transcribe Hindi?

Yes. Choose Hindi + 90 languages and set Spoken language to Hindi (or Detect automatically). The text comes out in Devanagari. It works best with clear speech; check the result, because noisy recordings and Hindi-English mixing cause mistakes.

Is my audio uploaded to a server?

No. Your file is decoded and transcribed inside your browser and never leaves your device. The only downloads are the speech model from huggingface.co, the first time you use it, and the speech engine from this site.

How do I make subtitles for a video?

Drop the MP4 or WebM video, click Transcribe, then click SRT or VTT. Most video editors and YouTube accept SRT files; VTT is the format for web players.

Why is the first transcription slow to start?

The speech model has to download once: 41 MB for English or 76 MB for the multilingual model. After that it loads from your browser's cache unless you clear this site's data.

What is the difference between the two models?

English only (Whisper tiny.en) is smaller and faster but understands only English. Hindi + 90 languages (Whisper base) is larger and slower but handles Hindi, other Indian languages and many more, and can detect the language for you.

Which file formats can I transcribe?

Anything your browser can play: MP3, WAV, M4A/AAC, OGG/Opus, WebM and the sound in MP4 videos. MKV and AVI often fail; convert them to MP3 first with the Universal Conversion Suite.