All posts
3 min read

Why listnr never mixes your microphone with the call audio

Mixing the mic and the call into one waveform makes Whisper invent confident, wrong sentences. How listnr transcribes two separate lanes on-device instead.

By Rokibul Hasan

My team decides things on calls. Every call ended one of two ways: somebody took notes and stopped taking part, or nobody took notes and we rebuilt the decision from memory two days later, badly. I wanted a transcript, but I was not willing to put a bot in the call or send those conversations to somebody else's server. So I built listnr, a local meeting transcriber for Apple Silicon Macs.

The first version failed in a useful way

My first attempt was the obvious one: mix the microphone and the call audio into one waveform and hand it to Whisper. The result looked fine at a glance. It was plausible text with the speakers scrambled, invented sentences whenever two people spoke at once, and no way to tell which lines to trust. A transcript that is confidently wrong is worse than no transcript, because you cannot tell where it is wrong.

Two lanes, never mixed

listnr captures your microphone and your system audio as two separate lanes and transcribes each one on its own. Your voice is identified by which wire it arrived on, so it can never be attributed to someone else. The remote voices share one wire, so they go through speaker diarization after the session to be split into Speaker 1, Speaker 2 and so on.

It is more work than mixing. It is also the only approach I found that produces a transcript I actually believe.

Everything runs on your Mac

  • Transcription uses WhisperKit and Core ML, so the models run on the Apple Silicon neural engine and GPU.
  • Audio and transcripts never leave the machine. The only network request is the one-time download of model weights.
  • There is no bot in the call, no account and no subscription. It works with whatever you use for the call, because it listens to the Mac, not the meeting platform.
  • It handles English, Bangla, Hindi, Spanish, French, German, Japanese and Chinese, with optional translation to English.

Installing it

listnr is a command-line tool for now, because that is what I needed first. It installs through Homebrew from the repository, which is its own tap:

brew tap rokib16x/listnr https://github.com/rokib16x/listnr
brew trust --tap rokib16x/listnr
brew install listnr
listnr setup

# quit your terminal completely, reopen it, then:
listnr doctor && listnr

Speech models download on first use, not during install: about 139 MB for English and up to roughly 1.5 GB for the multilingual models. You can fetch one ahead of a call with listnr models download whisper-base.en. After that, the whole thing works offline, and you can confirm it by pulling your network cable.

What it does not do yet

  • Diarization runs after the session. While a call is running, remote audio is labeled "Others", and the Speaker 1 to N split is computed when you stop.
  • Memory grows with session length, roughly 460 MB per hour across both lanes, because the whole session is kept in RAM.
  • Non-English transcription is much weaker than English. Whisper itself is far better at English at every model size, and listnr tunes its thresholds per writing system but cannot close the gap.
  • Translation only targets English, and only one way.
  • There is no menu bar app yet. It is on the roadmap.

I list these in the README on purpose. A tool that records conversations should be honest about where it is unreliable.

A bug that shipped for five releases

The most useful lesson came from a failure. For five releases, the speaker lane could be silent without any error, because the Screen Recording permission on macOS attaches to your terminal application and only takes effect for processes started afterwards. Unit tests passed the whole time, since they cover the pure logic and not a live capture stream. I added a listnr doctor command that confirms both permissions took effect before you spend a meeting finding out, and I now say plainly in the README that the capture layer is partly verified by hand.

listnr is MIT licensed and installs through Homebrew. It is still beta, and the README lists the known limits: diarization runs after the session, memory grows with session length, and non-English accuracy is below English. The code and install steps are on GitHub.

About this project

Listnr — On-Device Meeting Transcription for macOS

More posts