Voxtrama
Voxtrama turns the recordings you already have, meetings, lessons and interviews, into a transcript with the speakers told apart and a summary you can check. It runs in Docker on your own computer, and the audio stays there. I wrote it from scratch and published it on GitHub on 28 September 2026 under the Apache 2.0 licence.
Who builds Voxtrama, and since when
What Voxtrama is built with
What problem Voxtrama solves
Anyone who records a meeting to write up the minutes usually uploads the audio to an outside service, and rarely knows which server it lands on or how long it stays there. The summary that comes back is written by a language model, which can put words in someone's mouth.
How Voxtrama solves the problem
Voxtrama brings transcription, speaker separation and summary onto the machine of whoever made the recording. Every point of the summary has to quote the transcript: Voxtrama looks the quote up word for word, links the point to its minute in the audio, and flags as Needs review whatever it finds worded differently.
What Voxtrama achieved
First public release on GitHub on 28 September 2026, 0.1.0-alpha1; 0.1.0-alpha2 followed the next day. Apache 2.0 licence. It is an alpha: I use it on my own work calls and I am looking for the first people to try it.
What happens to the audio
The audio never leaves your computer. A recording uploaded to Voxtrama follows this path, from upload to deletion.
| Moment | Where the audio is | What goes over the network |
|---|---|---|
| Upload | copied into the data folder on the host, ~/Voxtrama/recordings |
nothing |
| Clean-up, transcription, speaker separation | read by the machine's processor | nothing |
| Summary | stays on disk: the model receives only the transcript text | the text, to the Ollama server you configured; if Ollama runs on the same machine, as it does by default, nothing |
| Storage | stays in the data folder until you delete it | nothing |
| Deletion | Delete › Audio on the job, Clean up, retention, or rm -rf ~/Voxtrama |
nothing |
Retention is off by default, so a recording stays on disk until you delete it or set a limit of 7, 30, 90 or 365 days. Clean up removes on its own the recordings no job uses, once they are at least a week old.
Why I wrote Voxtrama
The idea came during an important work meeting that I was recording to write up the minutes. The tools I knew asked me to upload the audio to their service, and their pages made it hard to tell which server it would end up on and under what licence. I wanted something that worked without an internet connection, with everything Whisper needs packed into a container.
Today I use Voxtrama on my own work calls. I follow the meeting without taking notes, and when the call ends the key points and decisions are already there. The goal from the start was specific: summarise recordings several hours long with small models, around 7 billion parameters, that run on an ordinary computer.
Voxtrama is for people who record conversations they have to act on or study from and would rather keep the audio at home: a team keeping minutes, researchers working with interviews, a student going back over a lesson. For a public office or a professional firm the GDPR turns the same question into an obligation.
From audio to minutes: what happens to a recording
A recording goes through four steps, all on the same machine.
- Audio clean-up. A high-pass filter, noise reduction and even loudness prepare the file for
transcription.
- Transcription. faster-whisper transcribes on the processor, with three profiles: small,
medium and large-v3.
- Speaker separation. SpeechBrain's ECAPA-TDNN model tells the voices apart within the
recording. In the Speakers tab you name each voice, listen to its first minute, add a speaker the model missed, and move a single sentence to whoever actually said it.
- Summary. A language model served by Ollama, on the same machine or on a server you pick,
writes the summary according to the chosen workflow.
It accepts WAV, FLAC, MP3, M4A and OGG files. Several files can be joined one after the other, or treated as separate microphones recorded in the same meeting.
Why every point of the summary quotes the recording
Every point the model writes has to carry a quote from the transcript, and Voxtrama looks it up word for word, evening out only capital letters and spaces. If it finds the quote exactly once, it links the point to the speaker and the minute: one click and you hear the passage. If the quote is missing, or appears in more than one place, the point stays on screen with the label Needs review.
I chose an exact match because a looser one would accept a paraphrase from the model, and paraphrase is where a summary drifts away from what was said. A flagged point can be rewritten by hand or tied to the right turn of speech, and from then on it carries the label Edited. The model's original text is kept separately. The article on verifiable meeting minutes explains the mechanism in full.
Workflows: what the summary is for
A workflow tells Voxtrama what to pull out of a recording. The Community edition ships three, plus plain transcription:
| Workflow | What it produces |
|---|---|
meeting-decisions |
summary and decisions, each with an owner, a deadline and a quote |
lesson-companion |
summary and key concepts with their definitions |
research-interview |
summary and themes, each with a quote |
transcribe-only |
the transcript with its speakers, nothing else |
A workflow is a YAML file, and it behaves like a skill: the same call becomes a list of decisions or a map of themes depending on the workflow and the model that runs it. Adding one means writing the file, with no code to touch. The Community edition lets you define one of your own next to the three it ships with, for instance minutes laid out the way your office requires.
Hours of audio with a small model
Voxtrama reads a long transcript in windows of about 12,000 characters, two at a time, and gives each one a stretch of the window before and after so the thread of the conversation holds. Answers the model has already given are kept: if a call fails, the retry asks only for the missing windows. Each call is capped at 1,800 seconds, a figure set from a real measurement: 33 minutes of processor time to summarise a 59-minute call.
Every job leaves a record
For every job Voxtrama writes a manifest.json file with the models used, the settings and the source file. From there the job can be regenerated with another model or other settings, and steps whose inputs have not changed are reused instead of run again. The voxtrama compare command sets two jobs side by side and reports only the differences that matter: another model, other settings, another file.
The summary exports to PDF, HTML and DOCX. The downloaded manifest replaces the context and the server addresses with [removed], while the copy in the data folder keeps them.
What goes over the network, and when
No part of the code sends the audio file off the machine. Only two things go over the network, at two specific moments.
| Connection | When | What travels |
|---|---|---|
| huggingface.co | only when Voxtrama downloads a model it does not have yet: at first start, or on the first job that needs it | the request for the model files; no audio and no text |
| the Ollama server for the summary | on every summary | the transcript text, to the server you configured: by default it is on the same machine |
To keep everything in, run Ollama on the same machine. If you configure a remote server, that server receives the transcript text, never the audio; the Settings › Data & privacy page shows, step by step, where each workflow runs, and a step marked local_only refuses remote servers.
I checked this on 30 September 2026 in the logs of three jobs run on giorgioserver. The first one downloaded both models and contacted huggingface.co three times, all for the download. The other two jobs, with the models already on disk, contacted no outside server; the summary went to the Ollama instance on the same machine.
The code also records a fault that has since been fixed: the speaker model, although already on disk, queried huggingface.co four times on every job to check its own files. No audio left, but the fact that this machine was running that model did. The model now loads in offline mode, and a test makes sure it stays that way.
The rest of the protection lives in the code:
- the application answers only on
127.0.0.1and refuses requests carrying another host name or
origin, which also covers DNS rebinding;
- a credential for the model server travels only over
httpsor to the same machine; - the versions of downloaded models are pinned, because SpeechBrain builds objects from the model's
configuration file;
- retention deletes finished jobs after 7, 30, 90 or 365 days.
Speaker separation works on one recording at a time: in the next call, the same person is once again an unnamed voice. That is deliberate. A voice print is biometric data, and keeping one would turn a transcription tool into an identification system. For the same reason Voxtrama works only on files you already have, and leaves video platforms' content to them.
Two choices I made from the start
Voxtrama comes without a default summary model. A model picked by me would in effect become a recommendation, and the right choice depends on the machine and on the language of the recordings. On first start Voxtrama looks for Ollama on the machine, lists the models it finds, and you choose.
In the same way, Voxtrama measures the machine and proposes a transcription profile and how many cores to give each chunk of audio, but applies the proposal only once you confirm it. A wrong profile makes the product look slow, and whoever knows the machine makes a better call than an estimate.
How it is built
| Part | Technology |
|---|---|
| Web application and API | Python 3.11, FastAPI, Jinja2 |
| Job queue | RQ on Redis |
| Database | SQLite with Alembic migrations |
| Transcription | faster-whisper (CTranslate2), int8 on the processor |
| Speaker separation | SpeechBrain ECAPA-TDNN, clustering with scikit-learn |
| Summary | Ollama, local or remote |
| Command line | Typer: doctor, demo, run, compare, retention |
| Distribution | Docker Compose, four containers |
The containers are the Redis queue, a database migration that runs once, the web server and the worker that runs the jobs. Everything Voxtrama keeps lives in one folder on the host, ~/Voxtrama by default: recordings, results, models, database and settings.
How I released it
On 28 September 2026 I published 0.1.0-alpha1, the first version meant to be installed by someone other than me. 0.1.0-alpha2 came out the next day, with four fixes found by running Voxtrama on real recordings. The most significant one was in transcription: on a machine with twenty processors, a job set to four cores and two chunks decoded on just four threads, because the second chunk was prepared and then never ran.
Branches and tags follow the drupal.org project model, the same one I use for my Drupal modules: one development branch per series, 0.1.x, and tags that state the stability, from alpha through rc to the release.
Project status
Voxtrama is an alpha. I have tested it with Docker on a 2017 iMac Pro with an Intel processor and, on 30 September 2026, on a Linux server with a Ryzen 9 9950X, where 30 minutes of audio took 10 minutes and 40 seconds from transcription with large-v3 to the extracted decisions. The next step is running it on other operating systems. Until 1.0 an update may change the format of workflows and stored jobs, and the release notes say so when it does. The automated measurements that would stop a release from lowering quality are not in place yet. The interface is in English and the code is ready for translations: I am waiting for the first reports from real use before choosing languages.
If you try it, the repository's issues are where your feedback helps most.
Sources
All sources were checked on 30 September 2026.
- Voxtrama on GitHub (opens in a new tab): README, licence, requirements, workflows, branches and tags. Primary source.
- Voxtrama changelog (opens in a new tab): contents of 0.1.0-alpha1 and the four fixes in 0.1.0-alpha2. Primary source.
- Voxtrama releases (opens in a new tab): 0.1.0-alpha1 of 28/09/2026 and 0.1.0-alpha2 of 29/09/2026. Primary source.
- Job page guide (opens in a new tab): Needs review, correcting a point, exports. Primary source.
- Command line (opens in a new tab):
doctor,run,compare,retention. Primary source. - Generic tuning profile (opens in a new tab): 33 minutes of processor time for a 59-minute call, 1,800-second cap. Primary source.
- Test on Hugging Face requests (opens in a new tab): the four requests per job recorded before the fix, and the check that the speaker model stays offline. Primary source.
- Speaker model loading (opens in a new tab): offline mode forced once the model is downloaded. Primary source.
- Logs of the jobs run on giorgioserver on 30/09/2026: requests to huggingface.co only in the job that downloaded the models. Author's measurement.
- SECURITY.md (opens in a new tab): listening on 127.0.0.1, host and origin checks, credentials. Primary source.
This text was translated by AI.
How was AI used?
Translated from the Italian original with AI assistance, then read and corrected by a person, who holds editorial responsibility.