Meeting transcription that happens while you talk
Two audio streams go in — your microphone and the sound of the call — and one transcript comes out, with every line attributed. It is written as the conversation runs, not reconstructed from a recording afterwards.
60 free minutes a month · live calls only
The answers below are an example. Your own base connects in the app.
…
…
Let's start with Q3 — how did conversion move?
Conversion is up 14% — let me pull up the funnel.
And the cost per lead?
Cost per lead is €12.40 — 9% under plan. Most of the gain came from organic.
Not a video — this is Whisperer itself, and the panel really answers. The microphone starts only when you press it.
Mechanism
Why the speaker labels are reliable
Most tools guess who spoke by comparing voices in a single mixed track. Whisperer does not have to guess: the two streams never meet until the transcript is assembled.
- Capture Two separate sources Your voice arrives from the microphone. Everyone else arrives from system audio — the same sound your speakers play.
- Recognition Short chunks, sent as they come Audio goes out in short fragments while you are still speaking, each one already carrying the label of the stream it came from.
- Result A running transcript Lines appear in the window above your screen and, at the same time, become the context every AI answer is built from.
Language
Set the language before you start, not after
Recognition runs on Whisper, which supports more than 90 languages. The session language is a setting you choose before the call: on a mixed call, pick the one that will be spoken most, and the model handles the switches around it.
If the transcript comes back in the wrong alphabet, this is almost always why — not a fault in recognition.
Per service
The same mechanism, whatever you are calling on
There is no integration with any conferencing service, which is why there is no list of supported ones. If the app plays the other person through your speakers, the transcript will have them.
Limits
Where this does not apply
Both limits follow from the same design decision — that recognition is a stream — and neither can be configured away.
A finished recording can be uploaded: the Meetings section takes one file up to 25 MB — mp3, m4a, wav, ogg, flac, aac, plus mp4 and webm, from which the audio track is used. What is missing is bulk upload of a whole archive; and it spends the same minutes a live call does.
System audio is where other voices come from. Around a table with no online meeting, your laptop microphone hears everyone at once and the labels stop meaning anything.
Questions
What people ask first
Can Whisperer transcribe a meeting in real time?
Yes, and only in real time. Audio is sent for recognition in short chunks while the conversation runs, and the line appears in the window within moments of being spoken.
What happens to the transcript after the meeting?
It is saved in your account history with speaker labels and timestamps, and it becomes the basis for the meeting map. You can copy it or export it as plain text or Markdown. In no-logs mode nothing is stored after the session ends.
Can I search previous transcripts?
Yes. History has a keyword search across transcripts and AI answers, and the assistant can answer questions about past calls in plain language with the source attached.
Do I need permission from the host?
Not from the conferencing service — Whisperer is not using its recording feature and does not join the call. Whether you need permission from the people in the room is a separate question, and a legal one in some places; that decision stays with you.
Why is only my own voice in the transcript?
System audio access is missing. On macOS grant the Screen Recording permission, which is how the other side is captured. On Windows no extra permission is needed, but the call has to be playing through the default output device.
Check it on a call that does not matter
Fifteen minutes with a colleague is enough. Watch for one thing: that both labels show up. If they do, everything downstream — notes, search, translation — has what it needs.