When the transcript says [inaudible]
There is one moment of perfect honesty in the modern meeting stack, and it wears square brackets. Somewhere in the transcript of your Tuesday planning call, between a decision about hiring and a decision about scope, sits [inaudible] — a machine putting up its hand, in the official record, to say: I did not catch that.
We'd like to say a word in its defense, because [inaudible] is the best behaved thing in your transcript. It failed loudly. Everything else that went wrong that hour failed quietly.
Here's the mechanics of it. A speech recognizer is a probability machine: given the audio, it outputs the words it considers most likely. When the audio is clean, the most likely words are the ones you said. When a chair scrapes or two people overlap or someone drifts back from their microphone — the usual suspects — the evidence blurs, and the machine has two options. If its confidence collapses entirely, you get the bracket. But confidence rarely collapses entirely. Far more often it dips just a little, and the machine does what probability machines do: it picks the most plausible wrong thing, and writes it down with the same serene typography as everything else.
That's the part that should keep minute-keepers up at night. For every honest [inaudible], statistics suggests a family of confident near-misses living undetected in the same paragraph. "Can't commit to Q2" and "can commit to Q2" are one soft consonant apart in bad audio. The bracket you can see is the bycatch; the guesses you can't see are the fish.
And then the summary layer — the part everyone actually reads — launders the whole business. Large language models write fluent prose over rough transcripts; that's their nature. The [inaudible] disappears, the near-misses get paraphrased into confident bullet points, and the action items arrive wearing suits. Fluency is not fidelity. A summary never says "we're not sure what was decided here"; it says something plausible, which is worse.
So when a transcript hands you the bracket, treat it as a gift and as a message. The gift: an honest flag, at the exact timestamp, telling you which sixty seconds of audio to replay before you forward the minutes. The message: your capture chain dropped below the machine's hearing threshold at least once — which means it spent longer than that in the gray zone where guessing happens.
The fixes are unglamorous, which is how you know they work. Get the microphone closer to the mouths — geometry beats gear. Let sentences finish; overlap is the diarizer's kryptonite and the recognizer's too. Silence the noises you own. And put a noise-cancellation layer in front of the whole chain so the recognizer never hears the espresso machine in the first place — that's Krisp's standing trick, and the reason its note-taker gets handed cleaner evidence than anything else on our board. (Disclosure: that's an affiliate relationship, as is our habit to say in the same breath.)
None of this will make [inaudible] extinct, and honestly, we'd miss it. Every archive needs one entry that tells the truth about the archive. When yours shows up, don't sigh at the machine's failure. Replay the minute, fix the room, and appreciate the only line in the minutes that admitted what it didn't know.
Further filing: why accuracy is an audio problem · the pipeline, step by step · the other margin note