Agent Teams That Never Talk to Each Other: Running a 24/7 Lead Pipeline on pi
Picture a workshop where nobody speaks.
The scribe finishes a page and hangs it on a hook. The illuminator takes it down, paints the margin, and hangs it one hook further along. The binder takes it from there. The master never walks the floor. He reads the ledger, chalks tomorrow's orders on a slate by the door, and deals with whatever lands in the basket marked problems.
Nobody needs to know who else is working. Nobody waits for a reply. If a scribe falls ill mid-page, the page goes back on its hook and someone else picks it up.
That's the architecture of Oak Prospector's agent teams. It's a pipeline that finds local businesses with websites that are hurting them, gathers evidence, builds each one a preview of a better site, and handles the outreach. It's designed to run 24/7 on teams of pi coding agents, and the central rule is that the agents never talk to each other.
This article is the design. Some of it is built and running in production; the core of the runner isn't built yet. I'll be precise about which is which.
I – Why Agents Shouldn't Chat
The first version of this idea, in a previous article about OSM lead generation, was a batch script. Download a country, parse it, crawl, score, output a CSV. It worked. It also stopped at the CSV, because everything after the CSV (judging a website, writing to a real person, following up) needs judgment.
Judgment is what language models are for. The tempting next step is a "crew": an assessor agent that talks to a copywriter agent that talks to a reviewer agent, coordinated by a manager agent that reads everyone's messages.
I don't want that, for three reasons:
- Messages get lost and nobody notices. If the copywriter misses a note from the assessor, the lead just silently gets worse copy.
- Waiting is stalling. A 24/7 pipeline where agent A blocks on a reply from agent B is a pipeline that stops at 3 a.m.
- You can't audit a conversation. When a prospect asks "why did you send me this?", I want a record with a timestamp, an actor and a cost, not a transcript to reconstruct.
So the rule is: every handoff is a write to a record. The records live in one API, the prospector vibe, and that API is the only thing any agent can touch.
II – Two Loops, Only One of Them Intelligent
There are two loops in the system, and only the inner one uses a model.
The outer loop is plain code, identical for every worker:
while (running) {
const batch = await api.claim({ phase: 'assess', limit: 20 }); // lease: 15 min, heartbeat every 60 s
if (!batch.leads.length) {
await api.waitForEvent('lead.captured');
continue;
}
const result = await runPhase(batch); // code, pi, or both
await api.complete(batch.leaseId, result); // validated; moves each lead to its next state
}
A lease means nobody else gets those twenty leads for fifteen minutes. The worker heartbeats while it works. If it crashes, the lease expires and the leads go back in the queue. Every step also carries an idempotency key such as lead:123:assess:v3, so running it twice changes nothing.
The inner loop is pi: read the instructions and inputs, think, call a tool, look at the result, repeat, stop. It runs only where a decision can't be written as an if.
That split matters more than anything else in the design, because most of the pipeline doesn't need a model at all:
| Phase | Who does it | Why |
|---|---|---|
| Discover | Code | OpenStreetMap queries are deterministic |
| Enrich | Code | Fetch pages, check DNS, SSL and redirects, detect the CMS |
| Capture | Code | Desktop and mobile screenshots |
| Assess | Code builds an evidence pack → pi scores → code validates | "Is this site costing them customers?" is a judgment |
| Preview | pi builds a site in a sandbox | Design and copy are judgments |
| Outreach copy | pi writes → a deterministic gate approves or rejects | Writing is a judgment; sending is a rule |
| Replies | pi classifies → code applies the rule | "Is this a no or a not-now?" is a judgment |
| Supervisor | pi reads metrics → proposes → code applies within bounds | Deciding to pause a team is a judgment |
The cheapest thing that works wins. Language models make judgments, and code enforces rules. An agent can propose sending an email; only the gate can send it.
III – Four Channels, All of Them Records
If agents never talk, how does anything move? Through four channels, and each one is a record in the API.
Sideways, between teams: state changes and events. The Assess team doesn't "tell" the Outreach team anything. It moves a lead from captured to assessed and writes the assessment. The API emits lead.assessed, and the Outreach runner wakes on it. The handoff is the data: Outreach reads the assessment record itself.
Down, from supervisor to teams: directives. These are small config records, not orders to a specific agent:
{
"campaign": "lisbon-restaurants",
"priority": 1,
"rubricVersion": 3,
"teams": { "outreach": { "paused": true, "reason": "bounce rate 3.1%" } }
}
Every runner reads the current directives at claim time and puts them into the next prompt. Pausing a team, switching a rubric or reprioritising a campaign is a single write, and every agent picks it up on its next batch. There's no broadcast and no acknowledgement to wait for.
Up, from teams to the supervisor: metrics and escalations. Metrics come for free: counts per state, errors per step, cost per team. When an agent can't decide ("this looks like a franchise, not a local business", "the prospect replied with a legal question"), it calls escalate. The escalation is a record the supervisor reads and resolves, or passes to a human.
Across time: notes on a lead. Assess notices the owner is named on the About page and the menu is a PDF, and writes that as a note. Days later, the Outreach copywriter reads it. The context travels with the work, not with the agent that happened to see it.
IV – Tools Are the Permissions
Each team is a pi agent definition: a system prompt, a versioned playbook, a model, limits, and a small set of typed tools. pi's SDK makes the important part a single option. Here's how the platform's agentic-api vibe already opens sessions:
const { session: pi } = await createAgentSession({
model: piModel,
modelRuntime,
resourceLoader: loader,
noTools: 'builtin', // no bash, no file edits of its own
customTools: buildTools(session),
sessionManager: SessionManager.inMemory(workspacePath),
cwd: workspacePath,
});
noTools: 'builtin' removes pi's own shell and file tools. What's left is exactly what we hand it. It then checks that the loaded tool set is exactly the expected one, and refuses to start otherwise.
For the prospector teams, that means permissions are just the tool list:
| Team | Tools |
|---|---|
| assess | get_evidence, get_screenshot, add_note, submit_assessment, escalate |
| outreach | get_lead, get_notes, draft_message, submit_for_gate, escalate |
| replies | get_thread, classify_reply, escalate |
| preview | get_brief, start_preview_session, publish_preview, escalate |
| supervisor | get_metrics, get_ledger, list_escalations, set_directive, resolve_escalation, page_human |
The Assess team cannot send an email. Not "is told not to", but has no tool that could. However creative the model gets, the worst it can do is submit a bad assessment, and that's what the next section is for.
V – Contracts, Retries and the Dead-Letter Queue
Every phase output has a JSON schema. submit_assessment validates on the way in:
- Valid output moves the lead forward.
- Invalid output gets one retry, with the validation errors fed back to pi.
- If it fails again, the lead goes to a dead-letter queue for a human.
Bad data never flows downstream silently.
One practical finding from testing this: models wrap JSON answers in code fences. Ask for "only a JSON array" in chat, and you get ```json fences around it, so JSON.parse fails. Ask the agent to *write* out/result.json and read the file back, and in our test runs it came out clean. agentic-api now has a data workspace template for this: inputs in in/, results in out/result.json, and a system prompt that requires strict JSON in that file.
The other rule is a fresh pi session per batch. No long-running chat per team. The runner opens a session, prompts it with the playbook, directives and the batch, lets it call submit_* twenty times, and disposes it:
const { session } = await createAgentSession({
model,
noTools: 'builtin',
customTools: assessTools(api, batch),
sessionManager: SessionManager.inMemory(),
});
await session.prompt(renderPlaybook('assess', directives, batch));
session.dispose();
Long chat histories drift: a correction from three hours ago quietly changes how today's leads are scored. With a fresh session, behaviour is reproducible from three things: the playbook version, the directives and the inputs. Memory lives in records (notes, playbooks, the audit log), never in a context window.
VI – The Supervisor Is the Only Loop pi Drives
Pipeline teams are code-driven: code claims, pi judges, code completes. The supervisor is the one exception. Every fifteen minutes it's prompted with the current metrics and open escalations, and it acts through its tools:
set_directiveto pause a team, shift a budget or switch a rubric, within bounds the API enforces;resolve_escalationfor the decisions it can make;page_humanfor everything else.
It also watches circuit breakers. A team is paused automatically when:
- its error rate goes over 20%;
- bounces go over 2%, or complaints over 0.1%;
- spend runs ahead of plan;
- a platform 402 or 429 persists.
The second time the same team trips, a human gets paged.
What the supervisor can't do is as important: change pricing, change legal text, contact a lead the gate rejected, delete data, or raise its own limits. Those routes don't exist on its key.
VII – Money Is Enforced by the Platform, Not Trusted to Agents
Every team runs with its own API key on the vibe platform, and the platform, not the agent, enforces the budget.
The key system (vibe TASKS 086 and 087) gives each key:
- a reference we choose, such as
team:assessorcampaign:lisbon-restaurants; - a
dailySpendLimitand an overall spend limit; - an
allowedVibeslist, so the Assess team's key can reachweb-fetcherandagentic-apibut not the payment vibe; - a line in a unified ledger for every paid call, with the transaction id returned in an
x-vibe-transaction-idheader.
We store that transaction id on every pipeline step. That gives the cost per step, per lead, per team and per campaign, straight from the ledger with no estimates.
When a team runs out, the platform answers 402 and the team pauses cleanly. Revoking a key is the kill switch. The fleet runner itself also stops everything on SIGTERM.
VIII – One Lead, End to End
Here's what the design looks like for a single restaurant. Every line is a record with a time, an actor (the key reference) and a cost.
09:00 discover-3 claims area "Lisbon/Arroios, restaurants" → lead "Tasca do Zé" lead.discovered
09:02 enrich-1 web-fetcher extract + domains check
→ evidence: no SSL, not mobile-friendly, Wix, info@ address lead.enriched
09:02 capture-2 desktop + mobile screenshots lead.captured
09:10 assess-A batch of 20 → pi session, files in in/ → out/result.json
→ score 78, 3 pain points each citing an evidence id, offer "mobile redesign" lead.assessed
09:10 (code) score ≥ 70, contact exists → qualified; company status unconfirmed → post only
09:40 preview-1 pi in agentic-api builds a preview from their real menu and photos preview.published
10:00 outreach-2 pi writes the letter → gate: suppression ✓, channel ✓, claims↔evidence ✓,
identity ✓, AI disclosure ✓ → postal API message.sent
day 9 (webhook) QR code scanned → preview views logged
day 12 replies-1 pi classifies "interested" → code sends the claim and payment links lead.replied
Note the line at 09:10. The decision that this lead gets a letter and not an email is made by code, from a rule about Portuguese marketing law: sole traders need prior consent for email, and companies don't. Agents never decide what's legal. That's another article.
IX – Why Not an Orchestrator?
We looked hard at using an existing orchestration layer. Two were already on my machine: Hermes' kanban orchestrator and Orca's orchestration. Mapped onto this design, the building blocks are almost identical:
| This design | Hermes kanban equivalent |
|---|---|
| The prospector API as the shared record | The kanban board (SQLite) |
claim with a lease and heartbeat |
claim_task + heartbeat_claim |
| Expired lease goes back to the queue | reclaim_task |
| Handoff by changing state | todo → ready → running → done, with parents |
| Escalations | block_task, triage |
| Events the next team wakes on | Event subscriptions |
| Teams | Profiles |
Orca covers similar ground (task DAGs, dispatch, escalation waits, decision gates), plus threaded ask-and-reply between agents, which is great for supervising a few coding agents and exactly what I wanted to avoid here.
Three differences decided it:
- Graph shape. These tools build a new task graph per goal: an LLM decomposes a request into a handful of cards. The prospector graph is fixed (discover → enrich → … → close) and repeated over thousands of leads a day. A card per lead per phase would be over ten thousand cards a day, each routed by a model.
- The orchestrator's role. There, the orchestrator routes. Here, a state change is the routing. The supervisor never touches individual leads; it only tunes directives, budgets and pauses.
- Where rules and money live. The compliance gate, the suppression list, per-team spend limits and the audit of every send have to live next to the lead data. A board can't refuse an email that shouldn't go out; the prospector API and the platform ledger can.
A hybrid (kanban for batches, the API for leads) would have worked too. We chose to stay on pi and build the thin runner ourselves. pi's SDK already gives us sessions, custom tools and limits, and one less moving part in a 24/7 system is worth a few hundred lines of code.
X – The Honest Accounting
This is a design, and the runner at its centre isn't built yet. Here's the line.
Built and verified:
| Piece | State |
|---|---|
| oak-ui 1.1 data bindings: one JSON document, sources fetched once per pass, commands with invalidation | Built, all tests passing |
oak-ui as a vibe runtime (app.oak.json served through the engine) |
Live in production and validated in a real browser |
callVibeAPI for one vibe calling another |
Fixed and deployed |
agentic-api: sessions bound to a verified owner, owner-authenticated preview publishing by API key, async prompts, turn and time limits, the data template writing out/result.json |
Built and tested; the new version is not activated yet |
web-fetcher: mobile screenshots, contact and tech extraction, robots.txt |
Built and tested; the new version is not activated yet |
| API keys with references, spend limits, allowed vibes, a unified ledger | Live |
Not built yet (milestone M1):
- The coordination tables in the prospector API: leases (
claim,heartbeat,complete, expiry), the event stream, directives and escalations. oak-fleet: the runner, about three hundred lines aroundcreateAgentSession, with a pool per team, crash restarts,SIGTERMas a kill switch and a health endpoint.- The fleet screens in the prospector's own
app.oak.json, bound to the same records the agents see. - The outreach gate and sending providers.
The first real test is narrow on purpose: one code team (Enrich) and one pi team (Assess), running end to end on a batch of real Lisbon leads. After that we measure the only number that matters, reply rate. Everything else is plumbing.
And the lesson I'd put on the workshop wall: if two agents need to talk, one of them is missing a record. Find the record, write it down, and let them both go back to work in silence.
– Antonio