Speech to Text
Local model or browser speech serviceTranscribe your microphone or audio files with a local model or browser recognition
Choose the language that is spoken. The model does not detect languages automatically.
Audio Source
Audio and transcript stay in this browser. Only public model and runtime files are downloaded, which shares normal request details such as your IP address with those hosts.
Local model download
Checking browser cache…
Whisper Tiny (multilingual, q8), pinned revision ff4177021cc4. About 41.3 MB of model files from huggingface.co (large files may come from its CDN, us.aws.cdn.hf.co), plus about 20.6 MB of speech runtime from this site.
Files are kept in this browser’s cache for reuse and count toward site storage. Without a cached copy, the first load needs a network connection. Transcription uses your CPU and memory. A minute of audio can take a minute or more on slower devices. Model license: Apache-2.0 (model card); original Whisper code and weights MIT.
Drop an audio file here
Up to 25 MB and 5 minutes. This browser can decode: WAV.
You can also paste from the clipboard.
For other formats or codecs, convert to WAV first.
Load the local model first.
Ready
Transcript
Record or add audio, then transcribe. The transcript appears here.
Automatic transcript: words and punctuation may be wrong, and speakers are not identified.
CodingTool.dev does not upload, store or log your audio or transcript.
Related Tools
Speech to Text
Turn microphone speech or an audio file into text with timestamps, then copy it or download TXT, SRT or VTT. Choose a local model that runs in your browser or your browser’s own speech recognition. Pick Microphone or Audio File, the engine and the spoken language. For Local Model, review the download details and load the model once. Record or add a file, then transcribe. Edit segment text if needed and export plain text or subtitles.
Engines
Local Model runs Whisper Tiny (multilingual, quantized) in a background worker in your browser for both recordings and files. Browser Recognition uses the Web Speech API built into some browsers; it works with the microphone only, and the browser vendor may process audio on its servers.
Privacy
CodingTool.dev never receives, stores or logs your audio or transcript. With Local Model, audio stays in this browser; only public model and runtime files are downloaded. With Browser Recognition, your browser sends microphone audio to its speech service. Only your source, engine, language and view choices are saved.
Limits and accuracy
Files up to 25 MB and recordings or files up to 5 minutes. Transcripts are automatic and can contain mistakes, missing words or imperfect punctuation. Speakers are not identified. Browser Recognition times are approximate.
