Auto Subtitle Generator — Free SRT & VTT with Whisper AI
Automatically generate subtitles from speech with Whisper AI in your browser. Edit cues, preview them and download SRT or VTT. No uploads.
🔒 Runs entirely in your browser — nothing is uploadedAutomatic captions that stay on your device
Captions make videos accessible to deaf and hard-of-hearing viewers, help people who watch with the sound off and improve how search engines understand your content. Writing them by hand is slow, so this tool listens to the speech in your video and creates timed subtitle cues automatically. It uses the Whisper speech recognition model with timestamps, running locally in your browser through Transformers.js with WebGPU acceleration when your device supports it and WebAssembly otherwise.
Your video is never uploaded. The only download is the model itself from the Hugging Face hub, which your browser caches after the first run. That makes the tool suitable for unreleased footage, client work, internal training videos and anything else you would rather not hand to a third-party captioning service.
From speech to timed SRT and VTT cues
The audio track is decoded to 16 kHz mono and split into windows of up to 30 seconds at quiet moments. Whisper returns short phrases with start and end times for each window, which are shifted to their position in the full video. Phrases longer than your chosen character limit are split at word boundaries with the time shared proportionally, and long lines are broken into two balanced lines for comfortable reading. A limit of 84 characters, about two lines of 42, matches common broadcast and streaming guidelines.
The cue list is fully editable. Change the times in HH:MM:SS.mmm format, correct misheard words, delete false captions or add new ones at the current preview position. The preview player shows your latest edits as live captions so you can check the sync before exporting. SRT is the most widely supported format for editors and upload platforms; VTT is the native format for web video players.
Getting accurate subtitles
Choose the spoken language when you know it instead of relying on auto-detect, and prefer the Base model for interviews, accents or noisy audio. Music-heavy passages and overlapping voices are the most common source of errors, and Whisper timestamps can drift by a fraction of a second, so a quick review pass is always worthwhile. To display captions permanently, for example on social media where subtitle files are not supported, load the finished SRT into the burn subtitles tool.
How to use
- Add your videoDrop a video or audio file; a preview player appears so you can check the result.
- Choose optionsPick the model, the spoken language (or auto-detect) and how many characters each cue may hold.
- GeneratePress Generate subtitles and cues fill in as each 30-second window is recognized.
- Edit and exportReview the captions in the preview, fix any mistakes and download .srt or .vtt.