Transcribe Audio to Text Online — Private Whisper AI
Convert speech in audio and video files to text with Whisper AI running in your browser. No uploads, free, with copy and TXT download.
🔒 Runs entirely in your browser — nothing is uploadedSpeech to text without uploading your recordings
This transcriber turns spoken words in an audio or video file into plain text using OpenAI's Whisper speech recognition model. Unlike most online transcription services, it does not send your file anywhere. The model runs inside your browser through Transformers.js and ONNX Runtime, using your graphics card via WebGPU when it is available and falling back to WebAssembly on the CPU otherwise. Interviews, meeting recordings, lectures, voice memos and private videos stay on your device the whole time.
The only network download is the model itself. On the first run the tool fetches the Whisper weights from the Hugging Face model hub, which the browser then caches, so the next transcription starts almost immediately. Nothing about your audio is included in that request.
How the transcription works
When you press Transcribe, the audio track is decoded and resampled to 16 kHz mono, the format Whisper expects. Long recordings are split into windows of up to 30 seconds, and each cut is placed at the quietest moment near the end of the window so words are rarely chopped in half. Silent windows are skipped because Whisper tends to invent phrases such as "Thank you" when it hears nothing. Each window is transcribed in a background worker so the page stays responsive, and the text appears as soon as each piece is finished.
Whisper Tiny is the fastest option and works well for clear speech. Whisper Base is noticeably more accurate with accents and noisy audio at the cost of a larger download and slower processing. If you know the spoken language, select it; auto-detect listens to the first part of the recording and picks the most likely language. The Translate option produces English text from speech in another language.
Tips for better results
Speed depends on your hardware: with WebGPU a few minutes of audio often takes well under a minute, while CPU-only processing can take close to real time. Recordings with one speaker close to the microphone give the best text. Background music, crosstalk and heavy compression reduce accuracy. The output has no speaker labels or timestamps, so if you need captions with timing, use the auto subtitle generator instead. Always proofread names, numbers and specialist vocabulary before publishing or quoting the transcript. Very long files need a lot of memory; if a tab runs out of memory, split the recording into shorter parts first.
How to use
- Add a fileDrop an audio recording or a video with speech into the box.
- Pick model and languageChoose Tiny for speed or Base for accuracy, and set the spoken language or leave Auto-detect.
- TranscribePress Transcribe and watch the text appear window by window as the model works.
- Copy or downloadProofread the transcript, then copy it or save it as a .txt file.