← BACK TO ENGINEERING
Runtime 12 min read

Agent Teams That Never Talk to Each Other: Running a 24/7 Lead Pipeline on pi

Picture a workshop where nobody speaks.

The scribe finishes a page and hangs it on a hook. The illuminator takes it down, paints the margin, and hangs it one hook further along. The binder takes it from there. The master never walks the floor. He reads the ledger, chalks tomorrow's orders on a slate by the door, and deals with whatever lands in the basket marked problems.

Nobody needs to know who else is working. Nobody waits for a reply. If a scribe falls ill mid-page, the page goes back on its hook and someone else picks it up.

That's the architecture of Oak Prospector's agent teams. It's a pipeline that finds local businesses with websites that are hurting them, gathers evidence, builds each one a preview of a better site, and handles the outreach. It's designed to run 24/7 on teams of pi coding agents, and the central rule is that the agents never talk to each other.

This article is the design. Some of it is built and running in production; the core of the runner isn't built yet. I'll be precise about which is which.


I – Why Agents Shouldn't Chat

The first version of this idea, in a previous article about OSM lead generation, was a batch script. Download a country, parse it, crawl, score, output a CSV. It worked. It also stopped at the CSV, because everything after the CSV (judging a website, writing to a real person, following up) needs judgment.

Judgment is what language models are for. The tempting next step is a "crew": an assessor agent that talks to a copywriter agent that talks to a reviewer agent, coordinated by a manager agent that reads everyone's messages.

I don't want that, for three reasons:

  • Messages get lost and nobody notices. If the copywriter misses a note from the assessor, the lead just silently gets worse copy.
  • Waiting is stalling. A 24/7 pipeline where agent A blocks on a reply from agent B is a pipeline that stops at 3 a.m.
  • You can't audit a conversation. When a prospect asks "why did you send me this?", I want a record with a timestamp, an actor and a cost, not a transcript to reconstruct.

So the rule is: every handoff is a write to a record. The records live in one API, the prospector vibe, and that API is the only thing any agent can touch.


II – Two Loops, Only One of Them Intelligent

There are two loops in the system, and only the inner one uses a model.

The outer loop is plain code, identical for every worker:

while (running) {
  const batch = await api.claim({ phase: 'assess', limit: 20 }); // lease: 15 min, heartbeat every 60 s
  if (!batch.leads.length) {
    await api.waitForEvent('lead.captured');
    continue;
  }
  const result = await runPhase(batch);        // code, pi, or both
  await api.complete(batch.leaseId, result);   // validated; moves each lead to its next state
}

A lease means nobody else gets those twenty leads for fifteen minutes. The worker heartbeats while it works. If it crashes, the lease expires and the leads go back in the queue. Every step also carries an idempotency key such as lead:123:assess:v3, so running it twice changes nothing.

The inner loop is pi: read the instructions and inputs, think, call a tool, look at the result, repeat, stop. It runs only where a decision can't be written as an if.

That split matters more than anything else in the design, because most of the pipeline doesn't need a model at all:

Phase Who does it Why
Discover Code OpenStreetMap queries are deterministic
Enrich Code Fetch pages, check DNS, SSL and redirects, detect the CMS
Capture Code Desktop and mobile screenshots
Assess Code builds an evidence pack → pi scores → code validates "Is this site costing them customers?" is a judgment
Preview pi builds a site in a sandbox Design and copy are judgments
Outreach copy pi writes → a deterministic gate approves or rejects Writing is a judgment; sending is a rule
Replies pi classifies → code applies the rule "Is this a no or a not-now?" is a judgment
Supervisor pi reads metrics → proposes → code applies within bounds Deciding to pause a team is a judgment

The cheapest thing that works wins. Language models make judgments, and code enforces rules. An agent can propose sending an email; only the gate can send it.


III – Four Channels, All of Them Records

If agents never talk, how does anything move? Through four channels, and each one is a record in the API.

Sideways, between teams: state changes and events. The Assess team doesn't "tell" the Outreach team anything. It moves a lead from captured to assessed and writes the assessment. The API emits lead.assessed, and the Outreach runner wakes on it. The handoff is the data: Outreach reads the assessment record itself.

Down, from supervisor to teams: directives. These are small config records, not orders to a specific agent:

{
  "campaign": "lisbon-restaurants",
  "priority": 1,
  "rubricVersion": 3,
  "teams": { "outreach": { "paused": true, "reason": "bounce rate 3.1%" } }
}

Every runner reads the current directives at claim time and puts them into the next prompt. Pausing a team, switching a rubric or reprioritising a campaign is a single write, and every agent picks it up on its next batch. There's no broadcast and no acknowledgement to wait for.

Up, from teams to the supervisor: metrics and escalations. Metrics come for free: counts per state, errors per step, cost per team. When an agent can't decide ("this looks like a franchise, not a local business", "the prospect replied with a legal question"), it calls escalate. The escalation is a record the supervisor reads and resolves, or passes to a human.

Across time: notes on a lead. Assess notices the owner is named on the About page and the menu is a PDF, and writes that as a note. Days later, the Outreach copywriter reads it. The context travels with the work, not with the agent that happened to see it.


IV – Tools Are the Permissions

Each team is a pi agent definition: a system prompt, a versioned playbook, a model, limits, and a small set of typed tools. pi's SDK makes the important part a single option. Here's how the platform's agentic-api vibe already opens sessions:

const { session: pi } = await createAgentSession({
  model: piModel,
  modelRuntime,
  resourceLoader: loader,
  noTools: 'builtin',            // no bash, no file edits of its own
  customTools: buildTools(session),
  sessionManager: SessionManager.inMemory(workspacePath),
  cwd: workspacePath,
});

noTools: 'builtin' removes pi's own shell and file tools. What's left is exactly what we hand it. It then checks that the loaded tool set is exactly the expected one, and refuses to start otherwise.

For the prospector teams, that means permissions are just the tool list:

Team Tools
assess get_evidence, get_screenshot, add_note, submit_assessment, escalate
outreach get_lead, get_notes, draft_message, submit_for_gate, escalate
replies get_thread, classify_reply, escalate
preview get_brief, start_preview_session, publish_preview, escalate
supervisor get_metrics, get_ledger, list_escalations, set_directive, resolve_escalation, page_human

The Assess team cannot send an email. Not "is told not to", but has no tool that could. However creative the model gets, the worst it can do is submit a bad assessment, and that's what the next section is for.


V – Contracts, Retries and the Dead-Letter Queue

Every phase output has a JSON schema. submit_assessment validates on the way in:

  • Valid output moves the lead forward.
  • Invalid output gets one retry, with the validation errors fed back to pi.
  • If it fails again, the lead goes to a dead-letter queue for a human.

Bad data never flows downstream silently.

One practical finding from testing this: models wrap JSON answers in code fences. Ask for "only a JSON array" in chat, and you get ```json fences around it, so JSON.parse fails. Ask the agent to *write* out/result.json and read the file back, and in our test runs it came out clean. agentic-api now has a data workspace template for this: inputs in in/, results in out/result.json, and a system prompt that requires strict JSON in that file.

The other rule is a fresh pi session per batch. No long-running chat per team. The runner opens a session, prompts it with the playbook, directives and the batch, lets it call submit_* twenty times, and disposes it:

const { session } = await createAgentSession({
  model,
  noTools: 'builtin',
  customTools: assessTools(api, batch),
  sessionManager: SessionManager.inMemory(),
});
await session.prompt(renderPlaybook('assess', directives, batch));
session.dispose();

Long chat histories drift: a correction from three hours ago quietly changes how today's leads are scored. With a fresh session, behaviour is reproducible from three things: the playbook version, the directives and the inputs. Memory lives in records (notes, playbooks, the audit log), never in a context window.


VI – The Supervisor Is the Only Loop pi Drives

Pipeline teams are code-driven: code claims, pi judges, code completes. The supervisor is the one exception. Every fifteen minutes it's prompted with the current metrics and open escalations, and it acts through its tools:

  • set_directive to pause a team, shift a budget or switch a rubric, within bounds the API enforces;
  • resolve_escalation for the decisions it can make;
  • page_human for everything else.

It also watches circuit breakers. A team is paused automatically when:

  • its error rate goes over 20%;
  • bounces go over 2%, or complaints over 0.1%;
  • spend runs ahead of plan;
  • a platform 402 or 429 persists.

The second time the same team trips, a human gets paged.

What the supervisor can't do is as important: change pricing, change legal text, contact a lead the gate rejected, delete data, or raise its own limits. Those routes don't exist on its key.


VII – Money Is Enforced by the Platform, Not Trusted to Agents

Every team runs with its own API key on the vibe platform, and the platform, not the agent, enforces the budget.

The key system (vibe TASKS 086 and 087) gives each key:

  • a reference we choose, such as team:assess or campaign:lisbon-restaurants;
  • a dailySpendLimit and an overall spend limit;
  • an allowedVibes list, so the Assess team's key can reach web-fetcher and agentic-api but not the payment vibe;
  • a line in a unified ledger for every paid call, with the transaction id returned in an x-vibe-transaction-id header.

We store that transaction id on every pipeline step. That gives the cost per step, per lead, per team and per campaign, straight from the ledger with no estimates.

When a team runs out, the platform answers 402 and the team pauses cleanly. Revoking a key is the kill switch. The fleet runner itself also stops everything on SIGTERM.


VIII – One Lead, End to End

Here's what the design looks like for a single restaurant. Every line is a record with a time, an actor (the key reference) and a cost.

09:00  discover-3   claims area "Lisbon/Arroios, restaurants"    → lead "Tasca do Zé"        lead.discovered
09:02  enrich-1     web-fetcher extract + domains check
                    → evidence: no SSL, not mobile-friendly, Wix, info@ address                   lead.enriched
09:02  capture-2    desktop + mobile screenshots                                                   lead.captured
09:10  assess-A     batch of 20 → pi session, files in in/ → out/result.json
                    → score 78, 3 pain points each citing an evidence id, offer "mobile redesign"  lead.assessed
09:10  (code)       score ≥ 70, contact exists → qualified; company status unconfirmed → post only
09:40  preview-1    pi in agentic-api builds a preview from their real menu and photos             preview.published
10:00  outreach-2   pi writes the letter → gate: suppression ✓, channel ✓, claims↔evidence ✓,
                    identity ✓, AI disclosure ✓ → postal API                                       message.sent
day 9  (webhook)    QR code scanned → preview views logged
day 12 replies-1    pi classifies "interested" → code sends the claim and payment links            lead.replied

Note the line at 09:10. The decision that this lead gets a letter and not an email is made by code, from a rule about Portuguese marketing law: sole traders need prior consent for email, and companies don't. Agents never decide what's legal. That's another article.


IX – Why Not an Orchestrator?

We looked hard at using an existing orchestration layer. Two were already on my machine: Hermes' kanban orchestrator and Orca's orchestration. Mapped onto this design, the building blocks are almost identical:

This design Hermes kanban equivalent
The prospector API as the shared record The kanban board (SQLite)
claim with a lease and heartbeat claim_task + heartbeat_claim
Expired lease goes back to the queue reclaim_task
Handoff by changing state todo → ready → running → done, with parents
Escalations block_task, triage
Events the next team wakes on Event subscriptions
Teams Profiles

Orca covers similar ground (task DAGs, dispatch, escalation waits, decision gates), plus threaded ask-and-reply between agents, which is great for supervising a few coding agents and exactly what I wanted to avoid here.

Three differences decided it:

  1. Graph shape. These tools build a new task graph per goal: an LLM decomposes a request into a handful of cards. The prospector graph is fixed (discover → enrich → … → close) and repeated over thousands of leads a day. A card per lead per phase would be over ten thousand cards a day, each routed by a model.
  2. The orchestrator's role. There, the orchestrator routes. Here, a state change is the routing. The supervisor never touches individual leads; it only tunes directives, budgets and pauses.
  3. Where rules and money live. The compliance gate, the suppression list, per-team spend limits and the audit of every send have to live next to the lead data. A board can't refuse an email that shouldn't go out; the prospector API and the platform ledger can.

A hybrid (kanban for batches, the API for leads) would have worked too. We chose to stay on pi and build the thin runner ourselves. pi's SDK already gives us sessions, custom tools and limits, and one less moving part in a 24/7 system is worth a few hundred lines of code.


X – The Honest Accounting

This is a design, and the runner at its centre isn't built yet. Here's the line.

Built and verified:

Piece State
oak-ui 1.1 data bindings: one JSON document, sources fetched once per pass, commands with invalidation Built, all tests passing
oak-ui as a vibe runtime (app.oak.json served through the engine) Live in production and validated in a real browser
callVibeAPI for one vibe calling another Fixed and deployed
agentic-api: sessions bound to a verified owner, owner-authenticated preview publishing by API key, async prompts, turn and time limits, the data template writing out/result.json Built and tested; the new version is not activated yet
web-fetcher: mobile screenshots, contact and tech extraction, robots.txt Built and tested; the new version is not activated yet
API keys with references, spend limits, allowed vibes, a unified ledger Live

Not built yet (milestone M1):

  • The coordination tables in the prospector API: leases (claim, heartbeat, complete, expiry), the event stream, directives and escalations.
  • oak-fleet: the runner, about three hundred lines around createAgentSession, with a pool per team, crash restarts, SIGTERM as a kill switch and a health endpoint.
  • The fleet screens in the prospector's own app.oak.json, bound to the same records the agents see.
  • The outreach gate and sending providers.

The first real test is narrow on purpose: one code team (Enrich) and one pi team (Assess), running end to end on a batch of real Lisbon leads. After that we measure the only number that matters, reply rate. Everything else is plumbing.

And the lesson I'd put on the workshop wall: if two agents need to talk, one of them is missing a record. Find the record, write it down, and let them both go back to work in silence.

– Antonio

"Simplicity is the ultimate sophistication."