Voxtrama

In short
Voxtrama is an open source application that runs in Docker on your own computer, and the audio never leaves it. It transcribes recordings with Whisper, separates the speakers and, with a model served by Ollama, writes a summary, decisions, themes or concepts. Every point carries the sentence from the recording and the minute it was said; when the sentence is missing from the transcript, the point is flagged Needs review.

Voxtrama turns the recordings you already have, meetings, lessons and interviews, into a transcript with the speakers told apart and a summary you can check. It runs in Docker on your own computer, and the audio stays there. I wrote it from scratch and published it on GitHub on 28 September 2026 under the Apache 2.0 licence.

View on GitHub sjpagan/voxtrama (opens in a new tab)

Who builds Voxtrama, and since when

Client
Personal project
Sector
Data confidentiality GDPR public administration
Role
Author and maintainer
Year
2026

What Voxtrama is built with

Python FastAPI Docker Ollama Redis SQLite

What problem Voxtrama solves

Anyone who records a meeting to write up the minutes usually uploads the audio to an outside service, and rarely knows which server it lands on or how long it stays there. The summary that comes back is written by a language model, which can put words in someone's mouth.

How Voxtrama solves the problem

Voxtrama brings transcription, speaker separation and summary onto the machine of whoever made the recording. Every point of the summary has to quote the transcript: Voxtrama looks the quote up word for word, links the point to its minute in the audio, and flags as Needs review whatever it finds worded differently.

What Voxtrama achieved

First public release on GitHub on 28 September 2026, 0.1.0-alpha1; 0.1.0-alpha2 followed the next day. Apache 2.0 licence. It is an alpha: I use it on my own work calls and I am looking for the first people to try it.

What happens to the audio

The audio never leaves your computer. A recording uploaded to Voxtrama follows this path, from upload to deletion.

Moment Where the audio is What goes over the network
Upload copied into the data folder on the host, ~/Voxtrama/recordings nothing
Clean-up, transcription, speaker separation read by the machine's processor nothing
Summary stays on disk: the model receives only the transcript text the text, to the Ollama server you configured; if Ollama runs on the same machine, as it does by default, nothing
Storage stays in the data folder until you delete it nothing
Deletion Delete › Audio on the job, Clean up, retention, or rm -rf ~/Voxtrama nothing

Retention is off by default, so a recording stays on disk until you delete it or set a limit of 7, 30, 90 or 365 days. Clean up removes on its own the recordings no job uses, once they are at least a week old.


Why I wrote Voxtrama

The idea came during an important work meeting that I was recording to write up the minutes. The tools I knew asked me to upload the audio to their service, and their pages made it hard to tell which server it would end up on and under what licence. I wanted something that worked without an internet connection, with everything Whisper needs packed into a container.

Today I use Voxtrama on my own work calls. I follow the meeting without taking notes, and when the call ends the key points and decisions are already there. The goal from the start was specific: summarise recordings several hours long with small models, around 7 billion parameters, that run on an ordinary computer.

Voxtrama is for people who record conversations they have to act on or study from and would rather keep the audio at home: a team keeping minutes, researchers working with interviews, a student going back over a lesson. For a public office or a professional firm the GDPR turns the same question into an obligation.


From audio to minutes: what happens to a recording

A recording goes through four steps, all on the same machine.

  1. Audio clean-up. A high-pass filter, noise reduction and even loudness prepare the file for

transcription.

  1. Transcription. faster-whisper transcribes on the processor, with three profiles: small,

medium and large-v3.

  1. Speaker separation. SpeechBrain's ECAPA-TDNN model tells the voices apart within the

recording. In the Speakers tab you name each voice, listen to its first minute, add a speaker the model missed, and move a single sentence to whoever actually said it.

  1. Summary. A language model served by Ollama, on the same machine or on a server you pick,

writes the summary according to the chosen workflow.

It accepts WAV, FLAC, MP3, M4A and OGG files. Several files can be joined one after the other, or treated as separate microphones recorded in the same meeting.


Why every point of the summary quotes the recording

Every point the model writes has to carry a quote from the transcript, and Voxtrama looks it up word for word, evening out only capital letters and spaces. If it finds the quote exactly once, it links the point to the speaker and the minute: one click and you hear the passage. If the quote is missing, or appears in more than one place, the point stays on screen with the label Needs review.

I chose an exact match because a looser one would accept a paraphrase from the model, and paraphrase is where a summary drifts away from what was said. A flagged point can be rewritten by hand or tied to the right turn of speech, and from then on it carries the label Edited. The model's original text is kept separately. The article on verifiable meeting minutes explains the mechanism in full.


Workflows: what the summary is for

A workflow tells Voxtrama what to pull out of a recording. The Community edition ships three, plus plain transcription:

Workflow What it produces
meeting-decisions summary and decisions, each with an owner, a deadline and a quote
lesson-companion summary and key concepts with their definitions
research-interview summary and themes, each with a quote
transcribe-only the transcript with its speakers, nothing else

A workflow is a YAML file, and it behaves like a skill: the same call becomes a list of decisions or a map of themes depending on the workflow and the model that runs it. Adding one means writing the file, with no code to touch. The Community edition lets you define one of your own next to the three it ships with, for instance minutes laid out the way your office requires.


Hours of audio with a small model

Voxtrama reads a long transcript in windows of about 12,000 characters, two at a time, and gives each one a stretch of the window before and after so the thread of the conversation holds. Answers the model has already given are kept: if a call fails, the retry asks only for the missing windows. Each call is capped at 1,800 seconds, a figure set from a real measurement: 33 minutes of processor time to summarise a 59-minute call.


Every job leaves a record

For every job Voxtrama writes a manifest.json file with the models used, the settings and the source file. From there the job can be regenerated with another model or other settings, and steps whose inputs have not changed are reused instead of run again. The voxtrama compare command sets two jobs side by side and reports only the differences that matter: another model, other settings, another file.

The summary exports to PDF, HTML and DOCX. The downloaded manifest replaces the context and the server addresses with [removed], while the copy in the data folder keeps them.


What goes over the network, and when

No part of the code sends the audio file off the machine. Only two things go over the network, at two specific moments.

Connection When What travels
huggingface.co only when Voxtrama downloads a model it does not have yet: at first start, or on the first job that needs it the request for the model files; no audio and no text
the Ollama server for the summary on every summary the transcript text, to the server you configured: by default it is on the same machine

To keep everything in, run Ollama on the same machine. If you configure a remote server, that server receives the transcript text, never the audio; the Settings › Data & privacy page shows, step by step, where each workflow runs, and a step marked local_only refuses remote servers.

I checked this on 30 September 2026 in the logs of three jobs run on giorgioserver. The first one downloaded both models and contacted huggingface.co three times, all for the download. The other two jobs, with the models already on disk, contacted no outside server; the summary went to the Ollama instance on the same machine.

The code also records a fault that has since been fixed: the speaker model, although already on disk, queried huggingface.co four times on every job to check its own files. No audio left, but the fact that this machine was running that model did. The model now loads in offline mode, and a test makes sure it stays that way.

The rest of the protection lives in the code:

  • the application answers only on 127.0.0.1 and refuses requests carrying another host name or

origin, which also covers DNS rebinding;

  • a credential for the model server travels only over https or to the same machine;
  • the versions of downloaded models are pinned, because SpeechBrain builds objects from the model's

configuration file;

  • retention deletes finished jobs after 7, 30, 90 or 365 days.

Speaker separation works on one recording at a time: in the next call, the same person is once again an unnamed voice. That is deliberate. A voice print is biometric data, and keeping one would turn a transcription tool into an identification system. For the same reason Voxtrama works only on files you already have, and leaves video platforms' content to them.


Two choices I made from the start

Voxtrama comes without a default summary model. A model picked by me would in effect become a recommendation, and the right choice depends on the machine and on the language of the recordings. On first start Voxtrama looks for Ollama on the machine, lists the models it finds, and you choose.

In the same way, Voxtrama measures the machine and proposes a transcription profile and how many cores to give each chunk of audio, but applies the proposal only once you confirm it. A wrong profile makes the product look slow, and whoever knows the machine makes a better call than an estimate.


How it is built

Part Technology
Web application and API Python 3.11, FastAPI, Jinja2
Job queue RQ on Redis
Database SQLite with Alembic migrations
Transcription faster-whisper (CTranslate2), int8 on the processor
Speaker separation SpeechBrain ECAPA-TDNN, clustering with scikit-learn
Summary Ollama, local or remote
Command line Typer: doctor, demo, run, compare, retention
Distribution Docker Compose, four containers

The containers are the Redis queue, a database migration that runs once, the web server and the worker that runs the jobs. Everything Voxtrama keeps lives in one folder on the host, ~/Voxtrama by default: recordings, results, models, database and settings.


How I released it

On 28 September 2026 I published 0.1.0-alpha1, the first version meant to be installed by someone other than me. 0.1.0-alpha2 came out the next day, with four fixes found by running Voxtrama on real recordings. The most significant one was in transcription: on a machine with twenty processors, a job set to four cores and two chunks decoded on just four threads, because the second chunk was prepared and then never ran.

Branches and tags follow the drupal.org project model, the same one I use for my Drupal modules: one development branch per series, 0.1.x, and tags that state the stability, from alpha through rc to the release.


Project status

Voxtrama is an alpha. I have tested it with Docker on a 2017 iMac Pro with an Intel processor and, on 30 September 2026, on a Linux server with a Ryzen 9 9950X, where 30 minutes of audio took 10 minutes and 40 seconds from transcription with large-v3 to the extracted decisions. The next step is running it on other operating systems. Until 1.0 an update may change the format of workflows and stored jobs, and the release notes say so when it does. The automated measurements that would stop a release from lowering quality are not in place yet. The interface is in English and the code is ready for translations: I am waiting for the first reports from real use before choosing languages.

If you try it, the repository's issues are where your feedback helps most.


Sources

All sources were checked on 30 September 2026.

Giorgio Alfredo Pagano
AI modified

This text was translated by AI.

How was AI used?

Translated from the Italian original with AI assistance, then read and corrected by a person, who holds editorial responsibility.