What Whisper Does and How to Access It
OpenAI Whisper is a speech-to-text tool that converts audio files into written text. You speak or play an audio recording, and Whisper transcribes what it hears into a document you can read, edit, and use. It works with audio files in formats like MP3, WAV, M4A, and FLAC, and it handles multiple languages.
You can use Whisper in three ways: through OpenAI's web interface (the simplest route if you just need occasional transcription), through their API (if you want to build Whisper into your own software), or by downloading the open-source version to run on your own computer. Each route has different costs and technical requirements.
The web interface requires an OpenAI account and costs money per minute of audio transcribed. The API works the same way but lets you automate the process. The open-source version is free to download but requires some technical setup and runs on your machine rather than OpenAI's servers.
Key Takeaways
- Whisper transcribes audio files to text and works with common formats like MP3, WAV, and M4A across dozens of languages.
- The web interface is the easiest starting point and charges per minute of audio, with pricing visible before you upload.
- The API lets you automate transcription and integrate it into your own tools, but requires basic programming knowledge.
- The open-source version runs free on your computer but needs technical setup and works best on machines with good processing power.
- Whisper works better with clear audio and struggles with background noise, accents it hasn't seen much training on, and technical jargon.
Using Whisper Through the Web Interface
Start at platform.openai.com and sign in with your OpenAI account (create one if you don't have one). Click on the Whisper section or navigate to the audio transcription tool. Upload your audio file by clicking the upload button or dragging the file into the window.
Before you upload, OpenAI shows you the estimated cost based on the file size and length. Audio costs roughly $0.02 per minute, though this can vary. Once you confirm, Whisper processes the file and returns the transcript. You can then copy the text, download it as a file, or edit it directly in the interface.
The whole process usually takes a few minutes depending on the audio length and how busy OpenAI's servers are. Longer files (over an hour) may take longer to process. You can upload multiple files in sequence without starting over each time.
Using Whisper Through the API for Automation
If you want to transcribe many files or integrate transcription into software you're building, the API is the route. You'll need a basic understanding of how APIs work and some ability to write or copy code. OpenAI provides code examples in Python, JavaScript, and other languages on their documentation site.
Set up an API key in your OpenAI account (under API keys in account settings), then use that key to send audio files to Whisper programmatically. The pricing is the same as the web interface — roughly $0.02 per minute — but you only pay for what you actually transcribe. You can set spending limits in your account to avoid surprise bills.
The API is useful if you're transcribing podcasts regularly, processing customer support calls, or building a tool that needs transcription as part of a larger workflow. If you're just transcribing a few files by hand, the web interface is simpler and doesn't require coding.
Running Whisper Locally on Your Computer
OpenAI released Whisper as open-source software, meaning you can download it and run it on your own machine for free. This route has no per-minute cost, but it requires technical setup and a computer with decent processing power (a modern laptop usually works, but older machines may be slow).
You'll need to install Python and a few software libraries, then download Whisper from GitHub. The OpenAI GitHub repository includes step-by-step instructions for Windows, Mac, and Linux. Once installed, you run Whisper from the command line by typing a command that points to your audio file.
The local version is useful if you transcribe large amounts of audio regularly and want to avoid API costs, or if you need to keep audio files on your computer rather than uploading them to OpenAI's servers. It's slower than the cloud version and requires more technical knowledge to set up, but it gives you full control and privacy.
What Whisper Transcribes Well and What It Struggles With
Whisper performs best with clear, high-quality audio where one person speaks at a time. It handles accents, background music, and some background noise, but accuracy drops as noise increases. If you're transcribing a podcast recorded in a quiet studio, expect very accurate results. If you're transcribing a phone call with static or a meeting with multiple people talking over each other, expect more errors.
Whisper also struggles with technical jargon, brand names, and specialized vocabulary it hasn't seen much training on. A medical transcription might have errors in drug names or procedures. A tech podcast might misheard acronyms. You'll usually need to review and correct the transcript by hand, especially for formal documents or technical content.
Timestamps are included in the transcript, so you can see exactly when each part of the audio was spoken. This is useful for editing videos or finding specific moments in a long recording. The transcript also includes confidence scores in some formats, telling you how sure Whisper is about each part.
Costs and Choosing Between the Three Routes
The web interface and API both cost roughly $0.02 per minute of audio. A one-hour file costs about $1.20. There's no monthly subscription — you only pay for what you transcribe. OpenAI charges your account monthly based on usage.
The open-source version costs nothing to run, but you need a computer powerful enough to process audio without taking hours. A modern laptop can transcribe an hour of audio in 10 to 30 minutes, depending on the machine. Older computers or those with limited RAM may be much slower.
Choose the web interface if you transcribe occasionally and want the simplest experience. Choose the API if you're building software or transcribing many files and want to automate the process. Choose the local version if you transcribe large amounts regularly and want to avoid API costs, or if privacy is a concern and you don't want to send audio to OpenAI's servers.
Tips for Better Transcription Results
Record or source audio that's as clear as possible. Use a good microphone if you're recording yourself, and minimize background noise. If you're transcribing existing audio, try to use the highest-quality version available — MP3s at 128 kbps will produce worse results than 320 kbps versions of the same file.
Review the transcript after Whisper finishes. Even with good audio, you'll usually find a few errors to fix, especially in names, numbers, or technical terms. For formal documents, budget time for a full read-through. For casual notes or rough transcripts, the output is often good enough to use as-is.
If you're transcribing multiple languages or switching between languages in one file, tell Whisper which language to expect. This improves accuracy. You can specify the language in the web interface or in the API call. Whisper supports dozens of languages, though accuracy is best for English and other widely-spoken languages.
Frequently Asked Questions
Can Whisper transcribe video files directly?
No, Whisper only works with audio files. If you have a video, you'll need to extract the audio first using a tool like FFmpeg (free, command-line) or a video editor. Once you have an audio file, you can send it to Whisper.
Does Whisper work with live audio or only recorded files?
Whisper works only with files you upload. It doesn't transcribe live speech in real-time. If you need live transcription, you'll need a different tool. Some services like Otter.ai or Google Meet offer real-time transcription as a feature.
What languages does Whisper support?
Whisper supports 99 languages, including all major ones. Accuracy is highest for English and other widely-spoken languages. Less common languages may have more errors. You can specify the language when you upload to improve results.
Is my audio private when I use the web interface or API?
Audio you upload to OpenAI's servers is processed by their systems. OpenAI states they don't use audio from the API for training, but the file does pass through their infrastructure. If privacy is critical, use the local open-source version, which runs entirely on your computer.
How long does transcription take?
Through the web interface or API, transcription usually takes a few minutes for files under an hour. Longer files may take longer depending on server load. The local version is slower — typically 10 to 30 minutes per hour of audio on a modern laptop, depending on your computer's power.