Hi all, I am looking for a decent STT model that can transcribe audio from a podcast. I’ve listed some details about the content of the audio file, in case that helps:
- Straightforward, English-speaking, two-person-conversation audio that is approximately 1 hour in length.
- The conversation between the podcast speakers is casual, so it does contain typical filled pauses, discourse markers, disfluencies, etc.
- Timestamps aren’t necessary.
- Identifying the active speaker and what they have (generally) said is necessary, but the precision with which the transcript reports this is not.
- For example, suppose Jack and Jill are the speakers. Jill is explaining something, and while she is speaking, Jack interjects with a filler line like “Mmm, right, that makes sen-”, only for Jack to pause so that Jill’s initial speaking line can continue uninterrupted. If the STT were to transcribe this precisely, it would need to very suddenly realise that Jill is no longer speaking, create a new speaker line for Jack, only to return back to Jill again. I can see lots of issues with that in a more casual and non-turn-based conversation, such as inaccuracies of who said what exactly, and in distinguishing a speaker’s intent with what is literally said.