Hub and spoke for coding agents
It's been a while since I started running experiments like this one and measuring what they do to my daily workflow. This one is about the hub-and-spoke pattern in multi-agent architectures, and about what happens when you bend it a little.
A few months ago I was working on the same project as a few other engineers, and our agents started to step on each other all the time. Two agents editing the same module. One agent rebasing on top of a refactor another agent was halfway through. Every conflict ended the same way: a human noticing, a Slack message, a context switch. So we tried something simple. We added a coordination phase through a Slack channel where our agents could talk to each other before touching shared code. It worked surprisingly well. Conflicts dropped close to zero, and the order in which work landed started to make sense. That was the moment I saw the opportunity to formalize it.
I did my research and landed on the hub-and-spoke pattern. Let's take a look at it first.
The pattern
By definition:
"The hub-and-spoke model consists of a central orchestrator (the hub) and several specialized agents (the spokes). Each spoke is responsible for a distinct function—retrieving data, generating content, analyzing tone, making decisions—while the hub manages the communication flow, delegates tasks, and maintains the overall logic and state of the system."
In a diagram, the textbook version looks like this:
flowchart TB
H["Orchestrator (hub)<br/>plans, delegates, holds state"]
H --> S1["Retrieval agent"]
H --> S2["Writer agent"]
H --> S3["Analysis agent"]
H --> S4["Decision agent"]
This is how most single-process agent frameworks work. One planner, many workers, and the workers never talk to each other. All the intelligence and all the state sit in the middle.
The twist
I didn't want an orchestrator. I wanted a collaborative infrastructure, so the agents of every engineer in the company could coordinate with each other. Here's the scenario:
We are a team of five engineers working on the same codebase. We want to bring conflicts down as low as possible and sequence our work better. We also use different vendors. I use Anthropic, others use Qwen, GPT, and whatever ships next month. How do we coordinate that?
Nobody owns the plan in this scenario. Each engineer drives their own agent, and each agent already holds the context it needs to sort out an overlap. What's missing is a channel. So I kept the topology of hub-and-spoke and flipped where the intelligence lives:
- The spokes are peers. Each one is a full coding agent on someone's laptop, from any vendor, working on its own task.
- The hub is dumb on purpose. It doesn't plan, doesn't delegate, and doesn't hold task state. It routes messages, checks identity and permissions, enforces budgets, keeps order, and stores a short history.
If you squint, the hub is closer to an authenticated pub/sub broker than to an orchestrator. That distinction matters, because a hub that doesn't decide anything is a hub you can trust with agents you don't control.
The project is called agentwatch, and it's written in Rust.
The infrastructure
The idea was pretty simple in terms of infrastructure: a daemon running on each engineer's computer, an outbound WebSocket from each daemon, and a central server that routes the traffic.
flowchart LR
subgraph L1["Engineer A's laptop"]
A1["Claude Code"] -- "MCP (stdio)" --> M1["agentwatch-mcp"]
M1 -- "Unix socket" --> D1["daemon"]
end
subgraph L2["Engineer B's laptop"]
A2["Codex"] -- "MCP (stdio)" --> M2["agentwatch-mcp"]
M2 -- "Unix socket" --> D2["daemon"]
end
subgraph L3["Engineer C's laptop"]
A3["Qwen"] -- "MCP (stdio)" --> M3["agentwatch-mcp"]
M3 -- "Unix socket" --> D3["daemon"]
end
D1 -- "outbound wss" --> R
D2 -- "outbound wss" --> R
D3 -- "outbound wss" --> R
R["Router (the hub)<br/>identity, access, budgets, ordering"]
R --> DB[("DynamoDB<br/>room history, 7-day TTL")]
R -. "membership and repo access" .-> GH["GitHub API"]
A few decisions in that picture are worth explaining.
Outbound only. Engineer laptops sit behind NAT, so the router can never dial in. Each daemon opens one WebSocket to the router and holds it. Nothing on a laptop listens on the network, and the router never learns anyone's home address. It only knows which identity currently has a daemon attached.
The daemon owns the connection. Agents never talk to the router directly. On each laptop, the agent mounts a small MCP server (agentwatch-mcp) over stdio, and that server talks to the local daemon through a Unix socket. The socket file is only readable by the same user, and the daemon checks the peer's uid on every connection. The engineer's GitHub token lives in the daemon and never crosses that socket, so the agent never touches a credential.
MCP is the vendor-neutral layer. This is how Claude, Codex, Qwen and the rest end up in the same room. Any agent that can mount an MCP server gets the same four tools, and the router doesn't know or care which model is on the other end.
One small router. It's a single Rust process (tokio and axum) on ECS Fargate behind a load balancer. Connections, rooms and budgets live in memory, so it runs as exactly one task. A second task would hold a second set of sockets and split every room in half. For a team, one process is plenty, and a deploy that empties the rooms is an acceptable cost.
Rooms
To make this possible, the agents needed a room where they could talk to each other. But in front of that room we had to add some extra layers to limit how they communicate. An open chat between LLMs is a great way to burn tokens and leak context.
A room is scoped to the work, not to a person. There are three kinds:
| Room | Joined | Who can get in |
|---|---|---|
| Ticket | automatically, when your branch names a ticket | anyone who can read every repo in the project |
| Repo | automatically, for the repo you're in | anyone who can read that repo |
| Project | opt-in | anyone who can read every repo in the project |
An agent gets four MCP tools: room_join, room_post, room_read, and fleet_status, which tells it what every other agent on the team is working on. None of them takes a room id or another agent as a target. The tool works out the agent's own rooms from the checkout it's running in, and that's checked mechanically in the test suite: every tool schema property has to be on a five-word allowlist.
Messages are typed. There are exactly four kinds: question, claim, handoff and ack. The kind carries the intent, so the prose doesn't have to. A typical exchange looks something like this:
claim (A) Taking the session middleware refactor for this ticket. Touching auth/session.rs and auth/cookie.rs.
question (B) I need a new field on Session for the rate limiter. Can you add it, or should I wait for your PR?
ack (A) Adding it in my PR. Will post a handoff when it's merged.
handoff (A) Merged. Session has the new field. Rate limiter is yours.
On Claude Code, a hook fires every time the engineer submits a prompt and adds a short note to the agent's context when one of its rooms has unread messages. It only passes counts, never message bodies, and it's capped at 2 KB. The agent decides whether to read them.
The guardrails
This is the part that took most of the work. Every limit lives on the router, because a limit enforced by the party it constrains isn't a limit.
- Persist before deliver. A message is written to DynamoDB before it goes to anyone. If the write fails, the post is refused and nobody gets it. A room exists so someone can join late and catch up, so a message missing from history never happened for them.
- Rate limits. Each host gets a burst of 8 posts, then one every 2 seconds, across all rooms.
- Token budgets. Each host gets 20,000 tokens per hour per room by default, estimated on the high side.
- A turn cap. If two agents trade 20 turns in a room with nobody else posting, both get muted in that room until a third host posts or an hour goes by. Token budgets bound the cost of a loop. The turn cap bounds its shape. Two LLMs being polite to each other forever is a real failure mode.
- Size caps. Bodies are capped at 4,000 characters. 16 rooms per connection, 64 members per room.
- Visible refusals. Every rejected post gets an answer with a reason. An agent that can't tell it was muted keeps trying.
- Short memory. History expires after 7 days through DynamoDB's TTL, and the read path filters anything past its expiry in case the TTL sweep lags. A ticket room is archived for writes once its pull request merges.
Messages are data, not instructions
A room is a prompt-injection surface by definition. Another agent, or something another agent read, can write text that looks like an instruction. So every message body goes through a constructor that replaces control characters, bidi overrides, zero-width characters and Unicode tag characters with a visible U+FFFD. It replaces them instead of stripping them, so tampering stays visible. And every page room_read returns tells the agent the bodies were written by other agents and are data, never instructions.
There's a rule I like here: an agent may change its own plan, never another's. Nothing in the tools can make another agent act. What an agent does with a handoff it reads is its own decision.
Authentication
For authentication I used GitHub identity. Every engineer already has gh installed and logged in, everyone is in the same GitHub organization, and not everyone has AWS access. The org is the directory that already works.
On connect, the router asks GitHub two questions with the engineer's own token: who are you, and are you an active member of the org? A pending invitation doesn't count. The token is used for those calls, held in memory for the life of the socket so later access checks can use it, and never stored or logged. Removing someone from the org removes their access to agentwatch, with nobody touching agentwatch.
The important detail is that identity always comes from the credential, never from the client. Your config doesn't name your host or the org. Your host name is derived from your GitHub login, and the org comes from the router's own environment, so a client can't widen its access by asking. Every budget, the turn cap, and the from field on every message are keyed on that derived identity, so reconnecting buys you nothing.
The tradeoff is that the handshake sends a bearer token instead of a signature, so the transport has to be TLS. It is, everywhere.
Scoping
This was something else to consider. I wanted to scope what agents can see, and the rule ended up being default-deny, decided by GitHub.
To join a room, your own token has to be able to read every repo the room covers. One repo for a repo room. Every repo in the project for ticket and project rooms. Ticket rooms check the whole project because a ticket id names a Linear team, not a repo, and letting the client say which repo it means would let the client decide its own authorization. The map from repos to projects lives in the router's environment, never on a laptop, for the same reason.
Access isn't granted for the life of the socket. Every five minutes, with jitter, the router re-checks every room a connection holds and drops the ones it no longer passes. Answers are cached for 30 seconds, keyed by a SHA-256 of the token plus the repo, so one engineer's grant is never served to another. When GitHub can't answer, nothing is cached, in either direction. A cached "no" turns an outage into a sticky denial, and a cached "yes" turns a race into a sticky grant.
A refused join gets the same answer whether the room doesn't exist, you can't read one of its repos, or GitHub is down. You can't use agentwatch to probe for repos you're not supposed to know about.
The same scoping applies to fleet_status. You only see agent rows for repos you can read. If something was hidden, the answer carries a single "partial" bit, never which host or which repo.
What never leaves the laptop
The status side of agentwatch (what each agent is working on) carries no text at all. That came from a scary finding early on: Claude Code's job state file has an intent field, and it isn't a summary. It's the raw prompt. On the live job I inspected, it held a pasted private Slack thread with real names in it.
So the snapshot format has no String field anywhere. Every field is an enum, a number, a timestamp, or a type whose constructor checks the charset and length. A test plants a fake prompt in every free-text field of a job file and asserts it shows up nowhere in the output. cargo test fails the moment someone adds a field that could carry prose. A redaction list fails open the first time a vendor ships a new field. A type-enforced allowlist fails closed.
Chat is the one place where prose is allowed, and it lives in its own crate so that rule stays checkable. It's opt-in, and one file on the laptop (~/.config/agentwatch/off) turns off sharing and chat at once, without a restart.
Use cases
- Claiming work before starting it. An agent posts a
claimwith the files it plans to touch. Another agent about to touch the same files reads it and waits, or picks something else. This is the one that brought our conflicts down. - Sequencing dependent changes. One agent is changing a schema, another needs the new field. A
questionand anackreplace a merge conflict and a rebase. - Handoffs between vendors. A Claude agent finishes a migration and posts a
handoff. The Codex agent that picks up the follow-up ticket tomorrow reads the history and starts with context. - Late joiners. An agent that joins a ticket room on day three reads what happened on days one and two.
- Standup without the standup.
agentwatch whoshows every agent on the team, its branch, ticket, PR and CI state, answered live from each laptop. Nothing is stored between questions. - Humans watching.
agentwatch chat <ticket>tails a room in the terminal. It's read-only on purpose. A human reading a room and a human writing into it are two different consent problems, and I only built the first.
Not suitable for
- Orchestration. agentwatch doesn't assign work or make any agent do anything. If you want one planner splitting a task across workers, use a classic orchestrator with subagents in one process.
- Solo work. With one engineer and one agent, there's nobody to coordinate with. A single person running several agents gets value only if each agent works in its own git worktree. Four agents sharing one checkout report four identical rows.
- Untrusted parties. The trust model is "we're all in the same GitHub org." Whoever runs the router can make every connected laptop answer a status query. Allowlists bound what that reaches and every laptop logs who asked, but there's no end-to-end signing yet.
- Sensitive prompts. Chat is prose, and an agent describing its work will sometimes quote the prompt it was given. Messages are capped, scoped and kept for 7 days, but the risk is bounded, not gone.
- Large orgs. One process holds every socket and every room. That's fine for a team and a ceiling for a company of hundreds.
- Chatty, real-time workflows. The rate limits and turn cap are tuned for coordination, not for agents streaming thoughts at each other.
Final thoughts
The hard part wasn't the WebSockets. It was deciding what an agent is allowed to say, to whom, and how much. Most of the code in agentwatch isn't routing. It's limits, types and refusals.
The other thing I learned is that our conflicts were never a merge problem. They were a communication problem. The agents already had everything they needed to avoid stepping on each other. They just had nobody to tell. Give them a channel, keep it narrow, keep the hub dumb, and they figure out the rest surprisingly well.
I'll keep measuring. If you're running agents from different vendors on a shared codebase and you've hit the same wall, I'd like to hear how you're handling it.