# Transcribe a podcast to text

> Turn any podcast or audio recording into a clean transcript with AI — plain text, SRT, or VTT, all in your browser. Nothing is uploaded.

Auto-generate accurate SRT, VTT, or plain-text subtitles from any video or audio using OpenAI's Whisper speech-recognition model running in the browser (WebGPU or WebAssembly), across roughly 100 languages — or upload your own SRT/VTT. From there you can download the file, burn styled TikTok-style captions permanently into the video, or add a soft subtitle track viewers can toggle on and off. Nothing is uploaded; handles up to 1 GB / 2 hours of audio.

## Podcast transcripts that stay on your machine

A transcript makes a podcast searchable, quotable, and accessible — and it feeds show notes and blog posts. This tool runs Whisper locally in your browser, so even long recordings are transcribed on your own hardware with nothing uploaded to a server.

Export a plain-text transcript for editing, or a timestamped SRT/VTT if you're captioning a video version of the episode. Clear speech transcribes very well; running a noisy recording through the free noise remover first noticeably improves accuracy.

## Frequently asked questions

### Does it work with audio files, not just video?

Yes. Drop an MP3, WAV, M4A, FLAC, or other audio file and it transcribes the same way. For audio you get the text outputs — plain-text transcript, SRT, or VTT.

### Can I get a plain-text transcript for show notes?

Yes — choose the TXT output for a clean transcript with no timestamps, ideal for show notes, blog posts, or meeting minutes, and copy it to your clipboard in one click. SRT and VTT with timestamps are available too.

### Can I use subtitles I already have?

Yes. Upload an existing SRT or VTT file to skip transcription — then download it in another format, burn the captions into the video, or add them as a soft track.

### Do you upload my video?

No. The Whisper AI transcription runs locally in your browser via WebGPU or WebAssembly, and if you burn captions in, that re-encode runs on your device too. Your file never leaves your computer — once the page has loaded it even works offline.

### What languages are supported?

Whisper was trained on about 100 languages, including English, Spanish, Catalan, French, German, Chinese, Japanese, Arabic, and Hindi. Pick the spoken language from the grid before transcribing (English is the default) so it transcribes accurately.

### Why does the first file take longer?

The Whisper speech model (~77–150 MB depending on your browser) downloads the first time you use the tool and is then cached. After that, transcription starts immediately.

### Is it really free?

Yes. The transcript generator is completely free, with no signup, watermarks, or limits. We make money from our video translation service, not from this tool.

---

_Free · No sign-up · No watermark · Runs entirely in your browser — your file never leaves your device._

- Open this tool: https://easydub.video/tools/subtitles/podcast-transcript
- All free EasyDub tools: https://easydub.video/tools
