Estimate time to transcribe audio.
Transcription time is quoted as a real-time ratio — how many minutes of work per minute of audio. Even a fast typist on clean single-speaker audio runs about 3:1, because you cannot type continuously while listening; you play, type, rewind and confirm. Three factors multiply that baseline. Poor audio nearly doubles it since passages need repeated playback. Each additional speaker adds roughly 12 percent for diarisation and cross-talk. Full verbatim, capturing every filler and false start, adds about 35 percent over clean verbatim. A separate review pass at half the audio length should be budgeted on top.
Real-time ratio
Ratio = 3 x quality factor x speaker factor x verbatim factor x typing factor
Total time
Transcription minutes = audio minutes x ratio, plus a review pass at 0.5x audio length
Because typing speed is not the constraint — listening is. You play a phrase, type it, rewind to check, and repeat. A 90 words-per-minute typist and a 65 words-per-minute typist differ far less in transcription throughput than in raw typing, because both spend most of the time on playback control.
Usually yes, but it does not remove the work. Automated output on good audio needs a correction pass of roughly 1:1 to 1.5:1 against the audio length — much better than 4:1 from scratch, but far from free, and poor audio can make correction slower than transcribing fresh.