What WhisperX does and why you might want it

WhisperX is a free, open-source tool that transcribes audio and video files into text. It uses artificial intelligence to convert speech to text, and it runs on your own computer rather than uploading your files to someone else's server. For people who are deaf or hard of hearing, WhisperX can generate captions for video. For anyone who needs a written record of what was said — in meetings, lectures, interviews, or personal recordings — it does that work without sending your files to a cloud service.

WhisperX is more accurate than some free transcription tools because it aligns the text with the exact timing of the audio, which makes it useful for creating properly timed captions. It also handles multiple speakers better than the basic version of OpenAI's Whisper, which WhisperX builds on.

The trade-off is that WhisperX requires some technical setup. You will need to install software on your computer and run it from a command line rather than clicking a button in a web browser. This guide walks through that setup step by step.

Key Takeaways

  • WhisperX requires Python, Git, and PyTorch to be installed on your computer before you can use it.
  • The installation process differs slightly depending on whether you use Windows, Mac, or Linux, and whether your computer has an Nvidia GPU.
  • After installation, you run WhisperX by typing commands in a terminal or command prompt, not through a graphical interface.
  • The first time you run WhisperX, it downloads a language model (about 3 GB), which takes time but only happens once.

Check your computer's specifications before you start

WhisperX runs faster if your computer has a dedicated graphics card (GPU), but it will work without one — it will just take longer. If you have an Nvidia GPU, the setup is slightly different than if you don't.

To check what you have on Windows, right-click on your desktop, look for "Nvidia Control Panel" or open Settings and search for "Device Manager." On Mac, click the Apple menu, choose "About This Mac," and look at the Graphics line. On Linux, open a terminal and type lspci | grep -i nvidia or lspci | grep -i amd.

If you find an Nvidia card, note the model number. If you have an AMD card or no dedicated GPU, that is fine — WhisperX will still work, just more slowly. You do not need to upgrade your hardware to use it.

Install Python and Git

WhisperX runs on Python, a programming language. You need Python version 3.9 or newer. Go to python.org, click Downloads, and choose the installer for your operating system. Run the installer. On Windows, make sure to check the box that says "Add Python to PATH" before you click Install — this lets your computer find Python from the command line.

You also need Git, which is version control software that lets you download WhisperX from its online repository. Go to git-scm.com, download the installer for your operating system, and run it. Accept the default options unless you have a reason to change them.

After both installations finish, restart your computer. Then open a terminal (on Mac or Linux) or Command Prompt (on Windows). Type python --version and press Enter. You should see a version number like "Python 3.11.5". Type git --version and press Enter. You should see a version number like "git version 2.42.0". If you see either command is not recognized, restart your computer again and try once more.

Download WhisperX and install its dependencies

Open your terminal or Command Prompt again. Choose a folder where you want WhisperX to live — your Documents folder or Desktop works fine. Navigate to that folder by typing cd followed by the path. For example, on Windows you might type cd Documents. On Mac or Linux, you might type cd ~/Documents.

Then type this command and press Enter: git clone https://github.com/m-bain/whisperx.git. Git will download WhisperX into a new folder called whisperx. When it finishes, type cd whisperx to enter that folder.

Now you need to install the libraries that WhisperX depends on. Type pip install -e . (note the period at the end). This tells Python to install WhisperX and everything it needs. The download will take a few minutes. You will see text scrolling past as packages install. When it finishes, you should see a line that says "Successfully installed" followed by a list of packages.

Install PyTorch with GPU support (if you have an Nvidia GPU)

PyTorch is the machine learning library that WhisperX uses. If you have an Nvidia GPU, you want the GPU version of PyTorch so transcription runs faster. Go to pytorch.org. Under "Start Locally," select your operating system, Python, and CUDA (not CPU). CUDA is Nvidia's parallel computing platform. The page will show you a command to copy.

Go back to your terminal, make sure you are still in the whisperx folder, and paste that command. It will look something like pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118. Press Enter and let it download. This step takes several minutes because PyTorch is large.

If you do not have an Nvidia GPU, skip this step. The standard PyTorch installation that came with WhisperX will work fine, just slower.

Download the language model and test the installation

The first time you run WhisperX, it downloads a language model — the AI that actually does the transcription. This is about 3 GB and only happens once. After that, WhisperX uses the model you already have.

To test that everything is installed correctly, type this command: whisperx --help. You should see a list of options and flags. If you see "command not found" or "is not recognized," close your terminal, reopen it, and try again. Sometimes the terminal needs to restart to find newly installed programs.

Now transcribe a short test file. If you have an audio or video file on your computer, use that. Otherwise, you can find sample audio files online. Type a command like this: whisperx /path/to/your/file.mp3. Replace /path/to/your/file.mp3 with the actual location of your file. On Windows, you can drag the file into the terminal window and it will paste the path for you.

WhisperX will start. The first run downloads the language model, which can take 10 to 20 minutes depending on your internet speed. After that, transcription begins. When it finishes, you will see a new file in the same folder as your audio file, with a name like file.vtt or file.srt. Open it in any text editor — the file contains your transcript with timestamps.

Common problems and how to fix them

If you get an error about CUDA or GPU, WhisperX is trying to use your graphics card but something is not set up correctly. The simplest fix is to run WhisperX without GPU support by adding --device cpu to your command. It will be slower, but it will work: whisperx /path/to/your/file.mp3 --device cpu.

If the command whisperx is not recognized, Python did not add WhisperX to your system path. Go back to the whisperx folder in your terminal and run the command as python -m whisperx /path/to/your/file.mp3 instead. This tells Python to run WhisperX directly.

If transcription is very slow even on a short file, your computer is using CPU instead of GPU. Check that PyTorch installed correctly by typing python -c "import torch; print(torch.cuda.is_available())". If it prints False, the GPU version of PyTorch did not install. Go back to the PyTorch website, copy the command for your GPU, and run it again.

Frequently Asked Questions

Do I have to use the command line every time?

Yes, WhisperX does not have a graphical interface. You open a terminal or Command Prompt and type a command each time you want to transcribe a file. Some people create a simple script or batch file to make this easier, but the basic workflow always involves typing a command.

What file formats does WhisperX accept?

WhisperX works with most common audio and video formats: MP3, WAV, M4A, FLAC, OGG for audio, and MP4, MKV, WebM, AVI for video. If you have a file in an unusual format, you may need to convert it first using free tools like FFmpeg.

Can WhisperX create captions for video files?

Yes. WhisperX creates .vtt and .srt files, which are standard caption formats. You can load these into video editing software or upload them to video platforms like YouTube. The timestamps in the file tell the video player when to show each line of text.

How long does transcription take?

On a computer with an Nvidia GPU, WhisperX typically transcribes audio faster than real time — a 10-minute audio file might take 2 to 5 minutes. On CPU alone, it is slower, often taking as long as the audio itself or longer. Very long files (over an hour) may take several hours on CPU.

Is my audio file private when I use WhisperX?

Yes. WhisperX runs entirely on your computer. Your audio files never leave your machine, and no data is sent to WhisperX servers or any third party. This is different from cloud-based transcription services, which upload your files to their servers.