How do I take AI meeting minutes without the audio leaving my computer, and check every point?
Deep dive into the Voxtrama project
Taking meeting minutes with AI while keeping the audio at home takes three steps on your own machine: transcription, speaker separation, and a summary written by a local model. Trusting those minutes takes a fourth: every point in the summary has to point back to something that was actually said, and the system has to tell you when it cannot find it. Voxtrama, the open source project I published on GitHub on 28 September 2026, does all four in Docker.
The name is easy to mix up with Voxtral, Mistral's speech model. They are unrelated projects.
Why should the audio of a meeting stay on your computer?
Meeting audio holds personal data and, often, confidential information. A cloud transcription service sends that audio to a third party's server, and whoever made the recording is answerable for that processing too.
That is where Voxtrama started. During an important work meeting I realised I could not say which server the recording would end up on, or under what licence the tools I had would handle it. I wanted something that worked without an internet connection, with everything Whisper needs packed into a container. A public office, a law or accounting firm and a research group all face the same question, and under the GDPR they have to answer it.
What happens to the audio in Voxtrama?
The audio never leaves your computer. A recording uploaded to Voxtrama follows this path, from upload to deletion.
| Moment | Where the audio is | What goes over the network |
|---|---|---|
| Upload | copied into the data folder on the host, ~/Voxtrama/recordings |
nothing |
| Clean-up, transcription, speaker separation | read by the machine's processor | nothing |
| Summary | stays on disk: the model receives only the transcript text | the text, to the Ollama server you configured; if Ollama runs on the same machine, as it does by default, nothing |
| Storage | stays in the data folder until you delete it | nothing |
| Deletion | Delete › Audio on the job, Clean up, retention, or rm -rf ~/Voxtrama |
nothing |
Retention is off by default, so a recording stays on disk until you delete it or set a limit of 7, 30, 90 or 365 days. Clean up removes on its own the recordings no job uses, once they are at least a week old.
How does Voxtrama transcribe and summarise locally?
Voxtrama is a web application made of four Docker containers: the Redis queue, a database migration that runs once, the web server, and the worker that runs the jobs. A recording goes through three stages.
| Stage | Component | Where it runs |
|---|---|---|
| Transcription | faster-whisper, small, medium and large-v3 profiles in int8 | the machine's processor |
| Speaker separation | SpeechBrain ECAPA-TDNN plus clustering | the machine's processor |
| Summary | a language model served by Ollama | Ollama on the same machine, or a server you pick |
Transcription runs on the processor even when a graphics card is present, because Docker on macOS keeps the GPU out of the container. The summary model does use a GPU when its Ollama server has one.
How does Voxtrama check that a sentence in the summary was really said?
Every point the model writes has to carry a quote from the transcript, and Voxtrama searches the transcript for it. The match is exact: before comparing, it only evens out capital letters and spaces, and punctuation stays as it is. The timestamp the model gives is thrown away too, since the only position that counts is where the sentence actually appears.
| What Voxtrama finds | What it shows |
|---|---|
| the quote appears once | the point, with the speaker and the minute; one click plays the passage |
| the quote appears in several places | the point flagged Needs review, because the right passage is ambiguous |
| the quote is missing from the transcript | the point flagged Needs review |
I chose an exact match over an approximate one for a specific reason, which is also written in the code: a looser match would accept a quote the model had paraphrased, and paraphrase is exactly where a summary drifts away from what was said.
What do you do when a point is flagged Needs review?
Needs review marks a point to check again: the model either paraphrased, or wrote something the recording does not contain. The point stays on screen and the job carries on. From the Transcript tab you listen to the passage and choose one of two fixes: rewrite the point, or tie it by hand to the right turn of speech. From then on the point is labelled Edited, and the PDF, HTML and DOCX exports use the corrected text. What the model originally wrote is kept in a separate file.
Which workflows are there, and how do you add one?
A workflow tells Voxtrama what the summary is for: what to pull out of a recording and in what shape. The Community edition ships three, plus plain transcription.
| Workflow | What it produces |
|---|---|
meeting-decisions |
summary and decisions, each with an owner, a deadline and a quote |
lesson-companion |
summary and key concepts with their definitions |
research-interview |
summary and themes, each with a quote |
transcribe-only |
the transcript with its speakers |
A workflow is a YAML file in the workflows/ folder, and adding one takes no code. It behaves like a skill: the same call becomes a list of decisions or a map of themes depending on the workflow you pick and the model that runs it. The Community edition lets you add one workflow of your own next to the three it ships with.
How does Voxtrama summarise hours of audio with a small model?
Voxtrama splits the transcript into windows of about 12,000 characters, moving each cut by up to 15% so it lands at a natural break, and gives every window a stretch of the one before and the one after. That lets a 7-billion-parameter model with a short context work through recordings several hours long. When Ollama is started with OLLAMA_NUM_PARALLEL, windows are processed in parallel: the VOXTRAMA_PARALLEL_WINDOWS variable sets how many, and the default is 2.
Each call to the model is capped at 1,800 seconds. The figure comes from a measurement recorded in the code: summarising a 59-minute call took 33 minutes of processor time.
What do you need to install Voxtrama?
Docker, and for summaries an Ollama server with at least one model pulled. Transcription alone works without Ollama.
| Requirement | Value |
|---|---|
| Docker | Docker Desktop on macOS or Windows, or Docker Engine with the Compose plugin on Linux |
| Memory | 8 GB for transcription with the smallest model, about 16 GB with a summary model alongside |
| Disk | about 0.6 GB for the smallest transcription model and the speaker model, up to 3 GB more for the most accurate one |
| Network | one free local port; make up starts from 8000 |
Installation is four commands:
git clone https://github.com/sjpagan/voxtrama.git
cd voxtrama
cp .env.example .env
make upmake up starts the containers in order, waits for the web server to answer and prints its address, for example http://127.0.0.1:8000. To connect a summary model:
ollama pull qwen3:4b
echo 'VOXTRAMA_OLLAMA_MODEL=qwen3:4b' >> .env
make down && make upOn first start Voxtrama measures the machine and proposes a transcription profile and how many cores to give each chunk of audio; nothing is applied until you confirm it. The command docker compose exec web voxtrama doctor checks the installation, and it is the first thing to run when something goes wrong.
How long does it take?
Thirty minutes of audio took 10 minutes and 40 seconds from start to finish, about a third of the recording's length. I measured it on 30 September 2026 on giorgioserver, my development server: an AMD Ryzen 9 9950X with 32 cores and 60 GB of RAM, no graphics card, running Voxtrama 0.1.0-alpha2 with the large-v3 transcription profile. The audio was the first 30 minutes of act two of Goldoni's La Locandiera, a multi-voice public domain recording from LibriVox.
| Step | Time | Relative to audio length |
|---|---|---|
| Transcription, large-v3 | 6 min 30 s | 0.22 |
| Speaker separation | 13 s | 0.01 |
| Summary, qwen2.5:7b | 1 min 42 s | 0.06 |
| Decision extraction, qwen2.5:7b | 2 min 11 s | 0.07 |
The worker peaked at 3.1 GB of RAM, not counting Ollama. Transcription used 8 of the 32 available threads, which is the default; giving it more cores brings the time down further.
On my 2017 iMac Pro, with an Intel processor, the same work on a call takes roughly as long as the audio itself. That figure comes from timing a handful of recordings by hand.
The summary produced 8 points and 16 decisions, and 11 of those 24 came out flagged Needs review. An eighteenth-century comedy, full of quips and quick exchanges, is a hard test for a 7-billion-parameter model, so the number says more about that text than about a work meeting. It does show the mechanism at work: almost half of the lines the model attributed to the characters did not appear in that form in the transcript, and Voxtrama flagged them instead of presenting them as fact.
How does it compare with Speakr, Scriberr and Meetily?
Speakr, Scriberr and Meetily are the open source projects closest to Voxtrama. The table compares what their own pages stated on 30 September 2026.
| Project | Licence | Speakers | LLM summary | Points tied to a verified quote |
|---|---|---|---|---|
| Voxtrama | Apache 2.0 | yes, ECAPA-TDNN | yes, Ollama | yes, with Needs review |
| Speakr | AGPLv3 or commercial | yes | yes | timestamps in Q&A answers, no documented check |
| Scriberr | MIT | yes, pyannote | yes, Ollama | not in the documentation |
| Meetily | MIT | in the PRO edition | yes | not in the documentation |
The comparison measures documentation, and a feature can exist in the code even when the pages leave it out. Meetily is a desktop app; the other three run in Docker.
What can't Voxtrama do yet?
Voxtrama is an alpha, 0.1.0-alpha2 from 29 September 2026, and it has clear limits.
- Speaker separation gives you the context of who said what, and it gets things wrong. On the
theatre audio used in the test it counted 32 voices in a scene with five or six characters. Turns can be corrected by hand, and voices are only matched within a single recording.
- The interface is in English; the code is ready for translations.
- I have tested it with Docker on macOS with an Intel processor. Other systems come next.
- The automated measurements that would stop a release from lowering quality are not in place yet.
- Until 1.0, an update may change the format of workflows and stored jobs.
How do I try Voxtrama?
The code is on GitHub (opens in a new tab): clone it, run make up, and the voxtrama demo command runs a sample that ships with the package. I am collecting the first reports from real use to decide on languages and priorities, and the repository's issues are the place to leave them.
Sources
All sources were checked on 30 September 2026.
- Voxtrama, repository (opens in a new tab), GitHub, tag 0.1.0-alpha2. README, requirements, installation, workflows. Primary source.
- Voxtrama, quote matching code (opens in a new tab): normalisation rule and the reason for exact matching. Primary source.
- Voxtrama, transcript windows (opens in a new tab): window size and overlap. Primary source.
- Voxtrama, job page guide (opens in a new tab): Needs review, correcting a point, exports. Primary source.
- Voxtrama, generic tuning profile (opens in a new tab): 33 minutes of processor time to summarise a 59-minute call, 1,800-second cap per call. Primary source.
- Voxtrama, speaker embeddings (opens in a new tab): 4.04 times real time in batches of 32, 0.51 times one at a time, measured in the worker container. Primary source.
- Voxtrama, releases (opens in a new tab): 0.1.0-alpha1 of 28/09/2026, 0.1.0-alpha2 of 29/09/2026. Primary source.
- Speakr (opens in a new tab), GitHub. Licence, features, Q&A with timestamps. Primary source.
- Scriberr (opens in a new tab), GitHub. Licence, speaker separation with pyannote, summaries with Ollama. Primary source.
- Meetily (opens in a new tab), GitHub. Licence, desktop app, speaker separation in the PRO edition. Primary source.
- Self-hosted AI notetakers (opens in a new tab), Anarlog, 20/04/2026. Overview of self-hosted projects.
- La Locandiera, act two (MP3) (opens in a new tab), LibriVox, public domain. The audio used in the 30/09/2026 test.
- Voxtral (opens in a new tab), Mistral AI. The speech model with a similar name.
This text was translated by AI.
How was AI used?
Translated from the Italian original with AI assistance, then read and corrected by a person, who holds editorial responsibility.