Stays on your device. WebGPU (or WebAssembly fallback) runs Whisper in your browser. No upload, no signup, unlimited use.
Interview transcription for recordings that cannot be shared
Journalists, qualitative researchers, oral historians, and lawyers all have the same problem with transcription services: the recording is the thing they are least allowed to hand over. A source spoke on condition of anonymity. A participant signed a consent form that names exactly who may hear the audio. An interview is under embargo until a publication date. In each case the upload step, not the transcript, is the compliance event.
The tool above never takes that step. OpenAI's Whisper model is downloaded into your browser once and then runs on your own machine, so the recording is read off your disk and transcribed in place. No account, no upload, no limit on hours, and no vendor that has to be named in an ethics application.
A workflow that holds up
- Drop the recording above. Voice memo M4A files, handheld recorder WAV files, and video interviews all work.
- Let the transcript build, then export plain text for coding and quoting, or JSON if you are importing into analysis software and want the segment timings preserved.
- Work from the transcript to find your material, and go back to the audio at the timestamp before you publish any direct quote. The transcript locates; the recording confirms.
- Keep the audio where it already is. Nothing about this workflow creates a second copy somewhere you would have to account for later.
What this replaces, and what it does not
It replaces the per minute transcription bill, which for a research project with dozens of hours of tape is the single largest line item after travel. It also replaces the awkward conversation about which vendor is allowed to hold the audio, because none of them is.
It does not replace a human transcriptionist where the standard is a verbatim record for legal or archival use, and it does not do speaker labels. Being clear about that is more useful than overselling it: this is a fast, free, private first pass that gets you from hours of audio to searchable text, and the last mile is still yours.
Frequently asked questions
- Is this safe for embargoed or confidential material?
- The audio is never transmitted, so there is no third party copy to leak, subpoena, or breach. The model runs inside your browser on your own hardware. If your obligation is that the recording must not be disclosed to anyone, a tool with no server in the loop is structurally different from a cloud service with a good privacy policy, and the difference is the part that survives the vendor changing its terms.
- Does it label the interviewer and the interviewee?
- No, it does not do speaker diarization. You get one accurate timestamped transcript. In practice a two person interview is easy to split afterwards because the questions are obvious, and the timestamps let you jump straight back to any moment in the audio to confirm attribution before you quote it.
- How accurate does it need to be before I can quote from it?
- Treat any machine transcript as a finding aid rather than a citable record. It is very good on clear speech and it will still miss things: proper nouns, technical terms, and anything said over the top of somebody else. The professional habit is to use the transcript to locate the quote and then listen to that timestamp before you publish it.
- Can it handle an interview in another language?
- Whisper recognizes 99 languages with automatic detection, and it can also translate a non English interview into an English transcript in one pass. For fieldwork that means a usable working transcript without hiring a translator for the first read through.
- Is there a limit on how much I can transcribe?
- None. No monthly minutes, no per file cap, no account. Files up to 500 MB each, with no restriction on how many you run. A research project with forty hours of interviews costs exactly nothing, which is not true of any per minute service.
- What if the recording is noisy or the room was bad?
- Accuracy will drop, the same as it would for any speech model and for a human transcriptionist. Distance from the microphone hurts most, followed by crosstalk and background music. If you have any control over the recording, a cheap lapel microphone on the person being interviewed improves the transcript more than any software choice.