How I work with AI in development: my method

I'm Giorgio Alfredo Pagano, and I work with AI by following a common thread: a method of my own that decides what to ask, which model to ask, and how to check the answer. For seven months, with Claude Code, I analyse the problem, break it into atomic tasks and hand them to a server that routes the work between Opus, Sonnet and two Qwen3 models, starts environments on demand and keeps the decisions shared across sessions in ctx-memory. I review every delivery myself.

I work with AI by following a common thread. In my experience, asking Claude or ChatGPT for one thing at a time gives you pieces that each work and never add up to a whole. My method decides what to ask, which model to ask, with what context, and how to check the result.

AI has changed the way I work as a freelance Drupal developer, and the method is still mine. Every job goes through four steps I decide on: I analyse the problem, break it down, delegate, and review. The delegated work lands on a server that acts as a remote core: it routes tasks between models, starts environments on demand and remembers the decisions I've made. This is how it's built, and what the measurements say.

What is the common thread when I work with AI?

The common thread is four steps, which I always follow in the same order:

  1. I analyse the problem. I gather everything in a brief: what needs doing, why, and under which

constraints.

  1. I define the flow by breaking the problem into atomic tasks. I write the full list of steps

and use AI to check it: covering every step, finding the functional gaps the list leaves open, keeping track from one session to the next. The work starts from a GitLab issue, code goes through an OpenSpec change that pins down the specification, and architecture choices end up in an ADR.

  1. I delegate to the system. Each task goes to the right model, with a precise specification.
  2. I review and decide. I read everything that comes back before pushing. The final call is mine.

AI works inside these steps and takes the volume off my hands: writing, reading logs, checking texts. I set the direction.

What role does the server play in the flow?

Everything in the flow goes through the server, which over time has become a set of microservices that support development and centralise shared work. Today it runs 55 containers across 21 Docker Compose projects. The same machine hosts GitLab with its CI, Ollama with the local models, ctx-memory with the decision memory, the skill registry and the development environments.

Centralising keeps everything aligned. Skills live in a registry on the server, and the system pushes updates to every machine I work from automatically. Some of them run on the server itself, next to Ollama, and hand Claude only the result: that's how the text checks and part of the SEO pipeline work. The system also keeps sensitive data on a separate path, and condenses text with Caveman, which shortens messages before they reach Claude.

The hardware is an AMD Ryzen 9 9950X with 32 threads and 60 GB of RAM, and no graphics card: for now, CPU and plenty of RAM are enough.

How does the server route the work?

Through a table of rules written into Claude Code's global instructions, which applies to every session without being repeated. An MCP server receives the delegated work and sends it on:

The work is… It goes to
a new, self-contained file of 40 lines or more Qwen3-Coder on Ollama, which saves the file to disk
a mechanical change to an existing file Qwen3-Coder, which sends back only the diff
a long log or output to make sense of Qwen3 instruct, which summarises it
code that needs repository context, several files at once, tests Sonnet
a change of a few lines Opus, directly
architecture, synthesis, review, decisions Opus

Opus, the model that drives the session, designs and reviews, and the others write the bulk of the code. The body of a delegated file stays out of Opus's context: it sends the specification, gets back the file path or the diff, and reads the content only when it reviews it.

One rule covers Ollama when it's asleep. Waking it up takes time, once per session, and the system treats that as a price worth paying, since it covers all the work that follows. Delegation is skipped only when Ollama can't be reached.

How do environments start only when they're needed?

Environments start on demand thanks to Traefik and Sablier. Traefik receives the requests for every service on the server, and Sablier keeps containers stopped until someone calls them, starts them on the first request and stops them after a period of inactivity, usually fifteen minutes. Of the 55 containers, 34 work this way, in 13 groups: the site environments, ctx-memory, the AI tools, even the browser I use to generate images. While I was writing this article, 4 of those 34 were running.

Each environment is ready when I open it, and the machine leaves its memory to Ollama while the environments sleep.

How do I avoid explaining the context again in every session?

With ctx-memory, my RAG system for tasks and decisions, built on Drupal 11 and Typesense. Issues record what was done, while ctx-memory keeps the why, with the evidence next to it, and shares it across Claude sessions. A new session starts from context that is already defined, instead of working out tasks and decisions from scratch.

I always use it from Claude. An MCP server exposes three tools, search, read and capture, and a dedicated skill teaches the agent how to use them: when I open a session on a project, the agent retrieves the decisions that concern it and works with that context in front of it. When I decide something, the same skill writes it into the graph. Two hooks capture automatically: every proposal.md of an OpenSpec change goes in as evidence along with the decision it states, and every committed ADR goes in with the permalink to its commit.

Underneath, Drupal acts as the graph database: contexts, decisions, evidence and relations are Drupal entities on PostgreSQL, with revisions and permissions. Typesense, connected through Search API, indexes them for search; the index lives in RAM and can be rebuilt from scratch from the database. The graph holds about 180 decisions and 440 pieces of evidence today. Four rules keep it reliable:

  1. a captured decision always starts as «proposed», and only I can confirm it;
  2. a decision is confirmed only with at least one piece of evidence pointing to a stable source;
  3. the graph is append-only: a decision changes only through a new one that supersedes it;
  4. every record states who wrote it, a person or a model.

How much work goes through the local model?

The system has logged every delegation since 8 July, with the tokens the local model reads and the ones it writes. Counting only the work with Claude, meaning skills and MCP tools, and leaving out the other applications that use Ollama on the server:

Measure Since 8 July Last 30 days
delegations 555 303
tokens read by Qwen3 about 1,691,000 about 1,029,000
tokens written by Qwen3 about 517,000 about 293,000
total processed locally about 2.2 million about 1.3 million

The tokens read are the logs, files and texts to check that Claude would otherwise have read, along with the instructions the local model needs. The tokens written are the output Claude would otherwise have produced. In the same 30 days I also launched 201 Claude executors, 170 of them Sonnet.

The real saving is larger, because every file Claude reads stays in its context for the turns that follow. The session in which I wrote this article, which also covered other content and four merge requests, produced about 578,000 output tokens in the main thread and read about 316 million, mostly from the cache. The number that matters to me is practical: with delegation I reach the end of the week, where before I ran out halfway.

How well does a local model perform on a CPU?

Qwen3 on a CPU finishes a job in under a minute: a skill run on the server takes 49 seconds on average, writing a file 63.

The guides you find online advise against Qwen3-Coder on a CPU, with answers that take minutes. The difference is the model: both Qwen3 models I use are 30B-A3B, mixture-of-experts models with 30 billion parameters, of which about 3 billion are active for each token, quantised to 4 bits in 18 GB. That's why they answer at the speed of a much smaller model. These times hold for this machine and for short jobs, such as one file or a log summary.

Reliability holds up too: out of more than 800 delegations logged since 8 July, 7 attempts came back empty, under 1%. The system counts them separately, outside the savings figures.

How do I review what comes back from delegation?

I read every delivery before pushing, because delegating the writing leaves the responsibility for the result with me. On one working day I noted the defects of every delegated delivery, and each one had at least one: a --quiet option that let the output through, a comm run on unsorted files, sample comments copied from a core module, a verdict stated without any evidence behind it. They are small defects, and that's exactly why they slip through. Skills and pipeline scripts run automatic checks before merging: English identifiers, translations, exported configuration.

Texts follow the same rule. Every text on the site goes through llm-detect, a skill that measures how many fingerprints of a language model a text carries: calques, contrived antitheses, stock phrases. A red result blocks publication, and green is only the minimum bar: in August a bilingual human reviewer spotted, within minutes, a text the pipeline had passed, and that led to new markers and four checks done by hand. This article went through them too.

Sources

All measurements were read on 1 October 2026 on the server where the system runs.

Sablier configuration, Claude Code global instructions, skill registry, ctx-memory contexts and decisions. Author's own measurements.

Giorgio Alfredo Pagano
AI modified

This text was translated by AI.

How was AI used?

Translated from the Italian original with AI assistance, then read and corrected by a person, who holds editorial responsibility.