Description
Whisper Web is a browser-based speech recognition platform powered by OpenAI's Whisper AI model. It processes audio locally in your browser using WebGPU and WebAssembly, providing speech-to-text transcription without downloads, installations, or server uploads. In Free mode, your audio never leaves your device, and once the model loads on your first visit, transcription works completely offline.
The tool supports more than 100 languages with automatic language detection, so no manual selection is needed. You can transcribe from microphone input, file upload, or a URL, and export results to TXT, JSON, SRT, and VTT. Because processing is local, Whisper Web requires no account and collects no data.
A free plan lets you transcribe files up to 200 MB and 20 minutes with no limit on the number of transcriptions. Whisper Web Unlimited adds an optional secure cloud workflow for larger files up to 10 hours and 5 GB, batch uploads of 50 files at once, and transcript syncing across devices, at US$20/month or US$10/month billed yearly. Cloud files are deleted after transcription.
Whisper Web's Core Features
Browser-based speech-to-text powered by OpenAI Whisper
Privacy-first local processing with no uploads in Free mode
100+ languages with automatic detection
WebGPU and WebAssembly acceleration
Microphone, file upload, and URL input
Export to TXT, JSON, SRT, and VTT
Offline transcription after the model loads
Optional cloud mode for large files and batch uploads
How to use Whisper Web?
Open Whisper Web: Load the site in your browser, no account required.
Select a model: Use the default Base model, or Tiny for slower devices.
Load your audio: Drag and drop an MP3, WAV, M4A, or video file, or use the mic or a URL.
Transcribe: Click start to run the AI model locally in your browser.
Export: Download the transcript or subtitles as TXT, JSON, SRT, or VTT.
Whisper Web's Use Cases
- Private transcription
- Subtitle creation
- Quick clips
- Large-file transcription







