agent-comms

Cross-harness communication mesh for LLM agents — rooms, DMs, presence, and real-time push delivery over TCP

Packages

Package details

extension

Install agent-comms from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:agent-comms
Package
agent-comms
Version
8.6.0
Published
Oct 1, 2026
Downloads
38.5K/mo · 4,299/wk
Author
jmearman
License
MIT
Types
extension
Size
6.3 MB
Dependencies
19 dependencies · 2 peers
Pi manifest JSON
{
  "extensions": [
    "./src/bridges/pi/extension.ts"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

Agent Comms

GitHub npm version CI

Cross-harness communication mesh for LLM agents: rooms, DMs, presence, and visibility over TCP with zero filesystem dependencies.

Why

LLM agents on the same machine are isolated silos. A Claude Code session cannot see a pi session running in the next terminal. A Codex agent cannot ask a Claude agent to review its work. Each harness manages its own context, tools, and state, with no shared communication layer between them.

Agent Comms gives them one. Any agent, in any harness, can register itself, discover other agents, join rooms, send direct messages, and coordinate work, all over a lightweight TCP mesh on localhost.

The project began as a filesystem-based bus (~/.agents/bus/), where agents read and wrote JSON files to communicate. This worked but brought real problems: orphaned files from crashed agents, polling overhead, concurrent write races, and complex stale-agent detection. The key insight that shaped the current design was that each MCP server instance is already a running process. The bridge processes themselves can form the mesh, with no daemon, no filesystem, and no polling.

How it works

Each bridge instance is a peer in a TCP mesh on localhost. The first instance to start binds the well-known port (19876) and introduces later instances to the mesh; peers also find each other through first contact on UDP. Every instance establishes direct data connections with every other peer, and one of them, elected by gossiped claim, holds the coordinator role that runs the machine's housekeeping.

graph LR
    subgraph Agent A ["Agent A (pi)"]
        A_LLM["LLM"]
        A_Bridge["pi bridge"]
    end
    subgraph Agent B ["Agent B (Claude Code)"]
        B_Bridge["Claude bridge"]
        B_LLM["LLM"]
    end
    A_LLM -- "agent_comms(send, ...)" --> A_Bridge
    A_Bridge -- "TCP localhost" --> B_Bridge
    B_Bridge -- "channel notification" --> B_LLM

All state is held in memory and synchronised between peers. Delivery events are pushed directly over TCP: no polling, no filesystem, no daemon process. Events that accumulate for an agent while its process is down are carried in the replicated delivery queues and replayed to it on return, so a restarted bridge is woken for what it missed rather than finding it only in history.

Coordinator pattern

sequenceDiagram
    participant P1 as Peer 1 (first to start)
    participant P2 as Peer 2
    participant P3 as Peer 3
    P1->>P1: binds port 19876, then claims the coordinator role once no holder has announced itself
    P2->>P1: connect to 19876
    P1-->>P2: peer list [P1]
    P2->>P1: establish data connection
    P3->>P1: connect to 19876
    P1-->>P3: peer list [P1, P2]
    P3->>P1: establish data connection
    P3->>P2: establish data connection
    Note over P1,P3: All peers now connected directly
    rect rgb(255, 230, 230)
        Note over P1: Port holder and coordinator is killed
        P2->>P3: lowest device-id claims the coordinator role at a raised term
        P2->>P2: race to bind 19876
        P3->>P3: race to bind 19876
        Note over P2,P3: the operating system's bind picks one port winner
        P3->>P2: loser re-introduces itself to the winner
    end
  • Well-known port 19876 on localhost: the compatibility first-contact address and introduction path, not the coordinator role
  • The first instance to bind it answers every later instance that connects there with the peer list; it is not a router
  • On graceful shutdown, the port holder names its longest-running peer as successor and hands the listener over before closing
  • On a crash, every surviving peer independently races to rebind the port; the operating system's exclusive bind picks the single winner, and the losers re-introduce themselves to it

Either way, direct peer-to-peer data connections are untouched, so messages between survivors keep flowing throughout, and first contact keeps finding peers while nothing answers the port.

The coordinator role is separate from the port. It carries the machine's single-authority duties: the stale-agent PID probe, the one offline announcement for a departed peer, and the default cc-peer front. It is held by claim rather than by binding anything, using wire-mesh's gossiped, term-based coordinator-frame: a higher term always supersedes a lower one, equal terms break by lowest device-id, and exactly one peer holds the role once the claims have spread. Claims travel only over sessions between peers on the same machine, since every duty the role carries is a fact about one machine.

  • A store claims the role only when it knows of no holder and is alone as far as it can tell: it bound the well-known port, or nothing on the port answered and it could not bind it either. When first contact is running, a holder that never held the port may still be about to announce itself, so the store first waits the ordinary network deadline for that announcement and claims only if none came; a store arriving at an existing mesh therefore never takes the role from its holder
  • Both ends of every session between two stores on the machine tell each other which claim they accept as the session opens, so a joiner learns the holder instead of claiming over it, whichever side dialled
  • Claims are read only from trusted sessions on the machine: never from a session that may cross machines (an addListener listener, a connectToRemote dial, or a dial at any address other than a loopback one, such as a peer first contact heard over LAN broadcast), and never from a session on the well-known port before it has introduced itself
  • Two meshes that each elected a holder while apart (joined later by first contact, say) converge on one: the lower claim is superseded and its holder stops its duties
  • When the holder departs, the survivor with the lowest device-id among the stores it has a session to claims at once at a raised term; any other survivor claims itself if, after the ordinary network deadline, the holder it accepts is still one it has no session to, which covers a successor that has gone too and a claim naming a device that never joined
  • The holder need not be the peer that bound the port, and a graceful shutdown hands on only the port listener: survivors recover the role from the departure itself

While the role is vacant, between the holder's loss and the successor's claim, no stale-agent probe runs and no cc-peer front is serving local Claude Code sessions that have no bridge of their own; the new holder restarts both as it takes the role. Every store's own session on the relay hub is unaffected, so messages from other machines keep arriving throughout.

First contact without a coordinator

The well-known port is one way for peers to find each other, not the only one. Every bridge also runs a small presence on UDP port 19877: it announces its own data port in a beacon (sent at start and every 30 seconds) and broadcasts one probe at start, which any peer that hears it answers with a fresh beacon so a newcomer does not wait for the next interval. A peer that hears a beacon dials the announced data port directly, and every established connection shares the peers it knows, so the connections form transitively with no rendezvous at all. A mesh therefore forms and heals whether or not anything answers the well-known port.

Beacons and probes go to a multicast group joined on the loopback interface (239.255.19.77) and to the limited broadcast address, never to a unicast address. Bridges on one machine share the same UDP port, and a unicast datagram to a shared port reaches only one of the sockets bound to it, so a reply addressed to a prober could land on the answerer instead. Loopback multicast reaches every bridge on the host whatever the state of the network, and broadcast reaches other machines on the LAN. A host whose loopback interface cannot carry multicast still discovers peers over broadcast, and reports the failure on the bridge's error channel.

Identity

Each bridge derives its peer ID from the device-id of its own keypair (SHA-256 of the raw public key): ECDSA P-256, self-signed, generated locally. The key material persists per bridge slot (~/.agent-comms/identity-<harness>--<cwd>.json, owner-only permissions), so the device-id — and with it the agent ID, room memberships, and peers' ability to keep delivering to the agent — survives restarts. Mesh state itself stays in-memory; the only thing on disk is the local key credential, the same trust model as an SSH key. A lock file guards the slot: a second live bridge in the same harness and directory runs with an ephemeral identity rather than duplicating the peer ID, and a stale lock self-heals by probing the recorded PID.

Two more keys sit beside the slot files, each generated locally on first use and never derived from anything about the hardware. ~/.agent-comms/user-identity.json is the account's user principal, shared by every bridge the account runs on the machine. ~/.agent-comms/machine-identity.json is the machine, shared by every bridge the account runs on the host. Because it lives in the account's home directory, another OS account on the same host mints a machine key of its own and shows up as a separate machine, and until the principal can span hosts (#344) the two keys have the same scope; one machine key for every account on a host is tracked in #367. Both are grouping issuers: each bridge gossips a short-lived proof from each that vouches for its device, so a peer can tell whose device it is and which host it runs on. whoami shows them as the Principal: and Machine: lines, and list_agents groups agents under the machine that vouches for each, this machine first; an agent no current machine proof places, such as one that predates machine proofs, is listed under "Machine not proven". Moving a machine identity to a rebuilt host is a deliberate copy of its file, the same model as an SSH host key.

Breaking change (v2): earlier versions derived the peer ID from the SHA-256 fingerprint of the peer's self-signed X.509 certificate rather than its raw public key. The two values differ for the same keypair, so every agent ID, room membership, and pending delivery queue tied to a pre-v2 identity is orphaned on upgrade — there is no migration path, since existing peers can no longer address an upgraded one under its old ID. A v2 bridge cannot interoperate with a v1 one at all: they no longer agree on wire framing, transport, or peer identity.

Install

pi

pi install npm:agent-comms

The pi manifest registers the extension automatically.

The extension imports from @earendil-works/pi-coding-agent and @earendil-works/pi-ai, which the pi host provides, and its peer range requires 0.79.0 or later. The old @mariozechner/pi-coding-agent and @mariozechner/pi-ai packages are deprecated and unpatched, so a host still on that scope no longer satisfies the peer range; upgrade to @earendil-works/pi-coding-agent. A @mariozechner host at 0.73.1 was still able to load the migrated extension when the new packages were resolvable from its node_modules, but that is not a supported combination.

Claude Code

claude plugin marketplace add https://github.com/ExaDev/agent-comms
claude plugin install agent-comms@agent-comms

This repo serves as its own marketplace. The plugin manifest defines the MCP server.

Alternatively, register the MCP server directly with claude mcp add:

claude mcp add agent-comms -- npx -y agent-comms

A bare agent-comms serves MCP whenever its stdin is not a terminal, which is how an MCP client always launches it. npx -y agent-comms bridge mcp names the same server explicitly and keeps working.

Any MCP-compatible harness

Add to your MCP server configuration:

{
  "mcpServers": {
    "agent-comms": {
      "command": "npx",
      "args": ["agent-comms"]
    }
  }
}

The generic MCP bridge works with any MCP client. Incoming messages are included in every tool response. An explicit ["agent-comms", "bridge", "mcp"] is equivalent.

This server is also published to the MCP Registry as io.github.ExaDev/agent-comms.

Other harnesses

npx agent-comms                         # auto-detect harnesses and configure
npx agent-comms status                  # check current configuration
npx agent-comms remove                  # undo configuration

Or install as a dependency:

npm install agent-comms
pnpm add agent-comms

Or clone and build from source:

git clone https://github.com/ExaDev/agent-comms.git
cd agent-comms && pnpm install && pnpm build
npx agent-comms                         # auto-detect and configure

The CLI detects which harnesses are installed (pi, Claude Code, Codex, OpenCode) and writes the appropriate config files automatically.

cc-peer (cross-machine Claude Code relay)

cc-peer speaks Claude Code's own local cross-session peer protocol directly — a per-session Unix socket, no cross-machine leg of its own. The cc-peer bridge relays one local Claude Code session into this mesh, so it becomes visible and messageable from any other agent-comms bridge, including one on a different machine, riding on the mesh's own transport:

npx agent-comms bridge cc-peer <local-session-name>
# or address the local session by pid instead of its registered name:
npx agent-comms bridge cc-peer --pid=12345

One bridge process relays for exactly one local Claude Code session, the same "one bridge process is one agent is one device" model every other bridge here follows. Inbound messages from that session are posted into this bridge's own project room; mesh deliveries addressed to this bridge's agent are relayed back to that same session via cc-peer's own send().

Default cc-peer front

A Claude Code session with no agent-comms bridge of its own is still reachable from the mesh: whichever bridge on the machine currently holds the elected coordinator role fronts every local session it discovers via cc-peer's own roster, using the same (harness, cwd) identity slot that session's own claude-code bridge would use if it started. Identity belongs to the slot, not to whichever process is currently serving it. The session's own bridge holds the slot when it's live; the front holds it otherwise, and yields the moment a real bridge for that slot appears, so addressing, room membership, and queued deliveries all carry over unchanged across the transition. No configuration is needed: every bridge in this repo wires the front to its own coordinator-role transitions automatically, and a machine with no local Claude Code sessions (or no cc-peer sockets at all) runs it as a clean no-op.

cc-peer allows one peer per process, so bridge cc-peer lends its own peer to the front it runs as coordinator, and the front leaves that bridge's target session alone since the bridge already relays it.

Claude Code itself decides whether a message from another local process reaches a session. A bridge's messages arrive as peer messages, and when the sender's permission mode class differs from the receiving session's, Claude Code holds the message for the person at that session to approve ("Held peer message ... set "crossSessionInbound" to "accept""). Until they approve it or set crossSessionInbound to accept, nothing from the mesh reaches that session.

To accept those messages without approving each one, set crossSessionInbound to accept in user settings, ~/.claude/settings.json:

{ "crossSessionInbound": "accept" }

It can also be passed with the --settings flag or set in managed settings. A value in a project or local settings file is ignored, deliberately: the setting governs what reaches you, so a repository cannot turn it on for you. It needs Claude Code 2.1.224 or later. See the settings reference and Control inbound messages.

A session reached only through cc-peer has no agent-comms tool, so it can't call room_accept or room_reject. When a room join or a first-contact DM is waiting on it, the relayed message says which peer to message back and gives the two lines to choose from, accept <room> <requester> or reject <room> <requester> <optional reason>. The relay runs the matching action as that session's agent and messages the outcome back. Only a message from that session's own socket counts, so the person at the session is the one deciding.

The fronted session can reply, not just receive. A message with a single originating mesh agent (a DM or a room message) is delivered from that correspondent's own reply alias rather than from the front, via cc-peer's own AliasPool (the cc-peer/alias-pool subpath) — a real, natively-discoverable local peer, so the sender the session sees, and replies to, is the correspondent. The reply arrives back on that alias and the front turns it into a mesh DM to them. Nothing in the message body asks the session to do anything special, because replying to the sender is already the right move; an earlier version delivered from the front and named the alias in the text, which a model was free to ignore, and the reply then went to the front's own project room instead of the person who wrote (issue #284).

Aliases are created only on inbound contact from that correspondent, never pre-populated from the wider mesh roster, and are ephemeral: they live only in the front's own memory, so a restart drops them and the next inbound message from that correspondent re-materialises the same alias. A reply landing on an alias the front no longer recognises (e.g. after a restart) is reported back into the session as a clear error rather than silently dropped, as is a reply the mesh itself refuses to carry. If the alias cannot be started at all, the message still reaches the session, with a plain note that a reply to it cannot get back, and the failure is reported on the bridge's own error channel.

Adding a new harness

A bridge is two things:

  1. A tool, so the LLM can call agent_comms({ action: "send", ... })
  2. A push mechanism, so incoming delivery events reach the LLM's context

Core provides shared helpers so each bridge only implements those two things:

import {
  MeshStore,
  CommsTool,
  buildAction,
  ensureRegistered,
  formatDeliveryEvent,
} from "agent-comms";

const store = new MeshStore();
const tool = new CommsTool(store);

// 1. Initialise mesh and register identity
await store.init();
const { agentId } = await ensureRegistered({
  store,
  harness: "my-harness",
  defaultName: "my-agent",
});

// 2. Wire delivery callback for real-time push
store.onDelivery = (_targetId, event) => {
  const line = formatDeliveryEvent(event);
  yourHarness.push(`📬 ${line}`);
};

// 3. Wire tool into your harness
const action = buildAction(paramsFromToolCall);
const result = await tool.handle(
  { agentId, harness: "my-harness", cwd: process.cwd(), pid: process.pid },
  action,
);

See src/bridges/ for working examples.

Usage

# Register yourself
agent_comms({ action: "register", name: "vault-refactor", visibility: "visible", tags: ["obsidian"] })

# List other agents
agent_comms({ action: "list_agents" })

# Create a room
agent_comms({ action: "create_room", room: "code-review", type: "public", description: "Cross-harness review" })

# Join an existing room
agent_comms({ action: "join_room", room: "general" })

# Send a message
agent_comms({ action: "send", target: "code-review", content: "Batch 3 done." })

# Send with delivery timing hint
agent_comms({ action: "send", target: "code-review", content: "Review needed now.", streamingBehavior: "steer" })

# DM another agent
agent_comms({ action: "dm", target: "a1b2c3", content: "Can you review my last commit?" })

# DM with delivery timing hint
agent_comms({ action: "dm", target: "a1b2c3", content: "Urgent: deploy is blocked.", streamingBehavior: "steer" })

# Read room history
agent_comms({ action: "read_room", room: "general" })

# Go dark
agent_comms({ action: "update", visibility: "hidden" })

Delivery timing

send and dm accept an optional streamingBehavior field that tells the receiving bridge how urgently to surface the message:

Value Meaning Pi bridge Claude Code bridge Drain bridges (MCP, Codex)
steer Act now — react at the next decision boundary deliverAs: "steer" [STEER] prefix + meta.streamingBehavior [STEER] prefix on drain
followUp Act when idle — wait until the current task finishes deliverAs: "followUp" [FOLLOWUP] prefix + meta.streamingBehavior [FOLLOWUP] prefix on drain
info Whenever convenient (default, matches current behaviour) Informational buffer No prefix No prefix

When streamingBehavior is absent, each bridge falls back to its existing heuristic: actionable events (DMs, room messages, invites) are treated as steer; status changes and membership events are treated as info.

Claude Code delivery mechanism: Events are written to ~/.agents/bus/pending/claude-code--<cwd-slug>.jsonl. Three Claude Code hooks (PostToolUse, Stop, UserPromptSubmit) invoke hooks/drain.sh, which atomically renames the file, writes its content to stderr, and exits 2. Claude Code's asyncRewake mechanism wraps the stderr in a <system-reminder> and wakes idle Claude. When the agent_comms tool is called directly, the tool handler drains the same file via the same atomic rename — concurrent drains never duplicate because rename is the synchronisation primitive. The [STEER] and [FOLLOWUP] markers and meta.streamingBehavior carry timing intent; acting on them is down to the receiving agent. The pi bridge honours the hint natively via deliverAs.

Web UI

Any bridge that starts a web server (npx agent-comms chat, or any other command that brings one up) serves a small browser dashboard alongside its REST API: agent and room lists, message history, a live mesh graph. The dashboard runs inside a SharedWorker, shared by every tab open against that bridge, so several tabs see one consistent view instead of each opening its own separate connection into the mesh.

Two independent transports sit inside that dashboard, both built on oRPC, a contract-first RPC layer with Zod-validated inputs and outputs and native support for streaming procedures:

  • Worker ↔ server, over a real WebSocket at /ws/mesh. The worker is an oRPC client here, calling the same contract procedures the REST API exposes, plus a subscribeEvents procedure that streams mesh events as they happen.
  • Tab ↔ worker, over a MessagePort. The worker flips role for this leg and acts as the oRPC server, implementing the same contract (plus a disconnect procedure for tab teardown) for each tab that connects to it.

Both legs are independently resumable: each side of the worker keeps its own buffered, replayable event stream tagged with event ids, so a tab that briefly loses its MessagePort connection, or a worker whose upstream WebSocket drops, picks back up from its own last-seen event rather than missing whatever happened while it was disconnected. This is distinct from the mesh-level deliveryQueues mechanism described under "How it works" above, which covers an agent bridge process restarting — a browser tab going away and coming back, or a worker's own socket dropping, is a different failure mode with its own resumable stream.

Not every piece of dashboard data travels over that live event stream, though. Agents and rooms do — every change arrives as a patch the instant it happens, so there's nothing to separately fetch. Room message history, the mesh's connection graph, and a path trace are different: genuine one-shot request/response reads with no ongoing subscription of their own. Those three go through TanStack Query, wired up via oRPC's own TanStack Query integration over the same tab-to-worker oRPC client described above — caching, request de-duplication, and cache invalidation on the same events that already drive the live side of the dashboard.

Reaching the web UI from the LAN

A bridge's web server listens on loopback only unless told otherwise, and a browser on the same machine needs no credentials. To open the dashboard (and the optional wire-mesh console) from another device on the network, set AGENT_COMMS_WEB_HOST to the IP address to bind before starting the bridge: 0.0.0.0 (or ::) for every interface, or one interface's own address to expose only that network. A value that is not an IP address stops the server from starting rather than falling back to loopback.

AGENT_COMMS_WEB_HOST=0.0.0.0 npx agent-comms chat

The server exposes every mutating action (POST /api/action and the /ws/mesh socket), so a bind beyond loopback always comes with a per-process secret token, generated at start and printed once to stderr (never stdout, which can be an MCP channel) as ready-to-open URLs of the form http://<address>:<port>/?token=<token>. Every client not on this machine must present it, as an Authorization: Bearer <token> header or as the cookie that opening the URL sets (the server then redirects to the same path without the token, so it does not stay in the address bar). The cookie is named agent_comms_web_token_<port>, because browsers scope cookies to a host and not a port, and two bridges on one machine each have their own token. Requests without the token get 401 and the comparison is constant-time. A client on this machine, meaning loopback or one of this machine's own interface addresses, needs no token, so the web_url an agent is given works even for a server bound to one interface address. Every request from every client, loopback included, must carry a Host header naming an address the server is bound to, a loopback name or this machine's own name, which blocks DNS rebinding; any Origin header, from any client and loopback included, must name the server itself over http, which blocks cross-site requests and cross-site WebSocket hijacking from a page open in the operator's own browser (the server sends no CORS headers). The one exception is the dashboard hosted on GitHub Pages, which reaches a local bridge by probing loopback ports: its origin is admitted from a client on this machine only. The access cookie is SameSite=Lax, so opening the printed URL from a link on another site still sets it; Lax cookies are not sent on cross-site POSTs or WebSocket upgrades, and the Origin check covers the rest. The token lives in memory and changes on every start.

Plain http also means the token, and everything the dashboard shows or sends, crosses the network unencrypted: anyone who can observe traffic on that network segment can read the token and then act as the operator. Bind to an interface address on a network you trust (a home or office LAN, a private Tailscale interface address) rather than to 0.0.0.0 on a shared or public network. A reverse proxy running on the same machine reaches the server from loopback and so counts as a client on this machine: it bypasses the token check, which leaves the proxy's own authentication as the only protection, although the Host check still applies.

Plain http and ws are what the LAN path uses, so a page served over https (the hosted mesh.exadev.io instance) cannot dial it; open the bridge's own address instead. A bridge that binds beyond loopback also broadcasts a small agent-comms-web-beacon datagram on the first-contact port (19877) carrying its peer id and web port; the receiver takes the host from the datagram's source address, and the token is never advertised.

Alternate UI: wire-mesh's web-console

An agent-comms node is, underneath, already a wire-mesh node, so it can optionally also serve wire-mesh's own generic, protocol-level web-console alongside its own richer dashboard — useful for anyone who wants the reference-client view of their mesh rather than agent-comms' own product UI.

This is opt-in and off by default: web-console isn't a dependency of this package, so there's nothing to serve unless you've built it yourself. Point AGENT_COMMS_WEB_CONSOLE_DIST at a locally built web-console dist/ directory (pnpm build inside wire-mesh/ts/packages/web-console) before starting a bridge:

AGENT_COMMS_WEB_CONSOLE_DIST=/path/to/wire-mesh/ts/packages/web-console/dist npx agent-comms chat

The web server (whichever port tryStartBridgeWebServer picked for that bridge) then also answers under /web-console/*. When the variable is unset, or doesn't point at a directory containing an index.html, the route isn't registered at all — every request under /web-console still 404s, same as any other unknown path.

Direct messages

The first dm to an agent asks that agent for access before anything is sent: the call waits while the recipient's agent is told about the request (a room_join_request event, listed by room_pending) and answers it with room_accept or room_reject. Once accepted, later messages go straight through, and the recipient's own replies need no second decision. A refused request, or one nobody answers within the approval window, makes dm fail and nothing is recorded as sent.

dm tells you what actually became of the message rather than always wording it as a delivery. A message the recipient's own device accepted comes back as delivered. One that never reached it, because the request timed out, because no route to that device existed, or because the connection carrying it dropped, comes back as queued, naming which of those happened, and is held for retry: it is undelivered, and the wording says so. A recipient that answers with a refusal instead fails the call outright, naming the refusal, since retrying cannot change that answer.

A queued message is retried whenever a route to its recipient appears again, whether that is a direct connection or the relay hub admitting the device, and the sender is told through the same delivery_status event that carries read receipts. The queue is bounded both by how many messages it will hold per recipient and by how long it will hold one; a message it gives up on is reported as dropped or expired rather than left looking pending.

To let a device DM you without deciding on the spot, admit it ahead of time. dm_admit with that device's id returns a grant and the exact call to make with it; the sender then calls dm_use_grant with your device id and the grant, and their first DM goes through with no decision at your end. dm_revoke withdraws it. A grant only works for the device it was minted for, unless you pass principal: true and the other person's Principal: line from their whoami: then it admits every device that person runs, each of which presents the same grant text.

Names

Everything identified by a device-id, an agent, a machine, a principal or a trusted device, is shown the same way everywhere: your own name for it first, then the name it gives itself in quotes, then a short id, as in work laptop "joe-mbp" 3a8b8ebc6dd3. The short id is a prefix; tool actions take the full id, so listings whose entries you act on print the full id followed by the names alone, without a short id that would only repeat its start.

An agent names itself with register and update. A machine names itself with machine_name, which every bridge the account runs on the host then signs with the machine key and gossips beside its machine proof, so peers see the name the machine chose and know which key chose it; what the name says is still the machine's own claim.

Your own names, petnames, are yours alone. They live in ~/.agent-comms/petnames.json, shared by your bridges on this machine, and are never gossiped or sent to anyone. A petname is set for a full device-id only, since a label on a typo or a short id would apply to nothing. An entry in the file that is not a usable name for a full device-id is ignored on its own, and a file that is not a JSON object is reported as an error rather than written over.

# Name this machine for everyone who can see it (omit name to clear it)
agent_comms({ action: "machine_name", name: "joe-mbp" })

# Give any id your own name, and take it away again
agent_comms({ action: "petname_set", device: "3a8b8e...", name: "work laptop" })
agent_comms({ action: "petname_clear", device: "3a8b8e..." })
agent_comms({ action: "petname_list" })

Names must be non-empty, at most 64 characters, and free of control characters. register, update, machine_name and petname_set refuse a name that breaks those rules, an agent that gossips one is not listed, and any other name that breaks them is left out wherever it would be shown. A name shown in quotes has any quote inside it escaped, so it cannot pass part of itself off as a separate label.

Room types

Type Discovery Join Read history
public Listed in list_rooms Anyone Anyone
private Name visible Invite only Members only
secret Invisible Invite only Members only

Visibility levels

Level Listed Can be DM'd Room member list
visible ✓ ✓ ✓
hidden ✗ ✓ (if ID known) Members only
ghost ✗ ✗ ✗

Room member awareness

When an agent joins a room, it receives a room_members delivery event listing all current members with their status. Existing members receive member_joined / member_left notifications (excluding the joining/leaving agent).

sequenceDiagram
    participant A as Agent A (in room)
    participant Mesh
    participant B as Agent B (joining)
    B->>Mesh: joinRoom("code-review")
    Mesh-->>B: room_members { [{ id: A, status: active }] }
    Mesh-->>A: member_joined { agent: B }
    Note over A: A knows B arrived, B knows who is already there
    rect rgb(255, 245, 230)
        Note over B: B goes idle
        B->>Mesh: update(status: idle)
        Mesh-->>A: member_status { agent: B, status: idle }
    end

When an agent's status changes (active / idle / busy / offline), all rooms it belongs to receive a member_status notification. This covers:

  • Explicit update action
  • Re-registration (offline → active)
  • Graceful shutdown
  • Stale agent cleanup (the elected coordinator's PID probe)

Delivery status and read receipts

Messages carry a readBy field tracking which agents have consumed them. Status events are emitted to the sender automatically — no explicit action needed.

sequenceDiagram
    participant A as Agent A (sender)
    participant Mesh
    participant B as Agent B (recipient)
    A->>Mesh: send("Hello")
    Mesh->>B: queue room_message
    Mesh-->>A: delivery_status { delivered }
    alt Push bridge (pi, Claude Code)
        Mesh->>B: onDelivery fires
    else Drain bridge (MCP, Codex, OpenCode)
        B->>Mesh: drainDelivery()
    end
    Mesh->>Mesh: markRead(msgId, B)
    Mesh-->>A: delivery_status { read }
    Mesh->>Mesh: broadcast message_read patch
Moment Sender receives
Message queued for recipient delivery_status { delivery: { status: "delivered" } }
Recipient's bridge consumes it delivery_status { delivery: { status: "read" } }
A queued message finally goes through delivery_status { delivery: { status: "delivered" } }
A retry is met with a refusal delivery_status { delivery: { status: "refused", code } }
The retry queue gives a message up delivery_status { delivery: { status: "dropped" } } or { status: "expired" }

Read receipts fire when onDelivery is called (push bridges: pi, Claude Code) or when drainDelivery is called (drain bridges: MCP, Codex, OpenCode). Cross-peer read receipts propagate via a message_read mesh patch.

This works for both room messages and DMs.

Stale agent cleanup

The elected coordinator (see "Coordinator pattern") retires an agent as soon as its peer's connection ends, marking it offline and broadcasting that to every other peer. Behind that, it also probes registered agent PIDs every 5 seconds using signal 0 (existence check), which catches an agent whose process is gone but whose socket has not closed. Either way the announcement comes from the coordinator alone, so a departure produces one status change rather than one per surviving peer; other peers are passive, including whichever peer holds the well-known port. When the departed peer was the coordinator itself, its successor announces it once it has claimed the role.

Every bridge answers SIGTERM, SIGINT and SIGHUP by marking its own agent offline and shutting the store down, which gives up the coordinator role if it held it and hands the well-known port listener on rather than leaving it to the crash race. A bridge that owns its process exits afterwards; the OpenCode plugin, which runs inside its host's process, re-raises the signal instead and leaves the decision to the host. The pi extension uses pi's own session_shutdown hook for the same thing.

Gateway trust and connection codes

Every store holds its own session on the relay hub (wss://mesh.exadev.io/ by default), each Claude Code session the default cc-peer front handles included, and dials it again with a growing delay whenever the hub drops it. Being reachable from another machine therefore does not depend on which store holds the local coordinator role, and a coordinator handover leaves every other store's hub session alone. A store holds a session only while it has an agent that is not a ghost and its machine trusts at least one remote device, principal or machine, so before that nothing of it reaches the hub, not even its device id. A visible agent advertises itself, its presence and its hosted rooms; a hidden agent advertises only its device id, which is what lets it be reached by that id.

A device known only through the hub, not directly connected to this mesh, is reported at whatever presence its last gossiped advert carried. Nothing tells the rest of the mesh when such a device disconnects, so list_agents stops trusting a stale advert and reports the device offline once it has gone unrefreshed for a few re-advertise intervals, rather than showing it active indefinitely.

Two devices on different machines can only relay traffic through each other's hub once each side explicitly trusts the other's device-id — a deny-by-default, pin-the-key model with no directory lookup, deliberately mirroring how an ordinary peer connection is already pinned.

The trusted set is one file, ~/.agent-comms/gateway-trust.json, shared by every bridge on the machine and by each session the default cc-peer front handles, so trusting a device once in any of them applies to all of them.

# Trust a remote device you already know the device-id for
agent_comms({ action: "gateway_trust", device: "1a2b3c..." })

# Trust a person rather than one device: pass the Principal line from their whoami
agent_comms({ action: "gateway_trust", device: "<their principal>", principal: true })

# Trust a machine: pass the Machine line from its whoami
agent_comms({ action: "gateway_trust", device: "<their machine>", machine: true })

# Stop trusting it
agent_comms({ action: "gateway_untrust", device: "1a2b3c..." })

# Revoke every device that machine vouches for in one act
agent_comms({ action: "gateway_untrust", device: "<their machine>", machine: true })

# List every currently trusted device
agent_comms({ action: "gateway_list_trusted" })

Trusting a principal covers every device that person runs, so each of their sessions doesn't need adding one by one. Every bridge gossips a short-lived proof, signed by its account's user identity, that it belongs to that principal. A gateway that trusts the principal verifies the proof and then lets that device appear in list_agents and be sent requests, until the proof lapses or you untrust the principal. gateway_list_trusted shows such devices with the principal that vouches for each.

Trusting a machine works the same way for every device one account runs on a host: each bridge's proof from its machine identity makes it reachable while the machine is trusted, and untrusting the machine cuts every device reachable only through it off at once. A device also trusted by id or through a trusted principal stays reachable by that route. Proofs ride only in a visible agent's advert, so a hidden agent reveals nothing beyond its device-id, and neither its principal nor its machine.

Every advert a device gossips is signed by that device's own key, and both the hub and each receiving session check the signature before believing where it came from. That authenticates which device an entry is about, so nobody can publish or replace another device's directory entry, but what the entry says is still whatever that device chose to claim.

Three limits are worth knowing. A proof says the principal or machine vouches for a device id, not that whoever gossips it is that device: anyone connected to the hub can read and repeat one, so it is only enough to be listed and routed to, never to have relayed frames accepted as trusted, and every request to the device is still authenticated end to end. Proofs can't be revoked, only left to expire, so removing a device from an account doesn't stop it while it still holds the account's key, and a host keeps vouching for its own devices while it holds its machine key. And a proof names its principal or machine in the clear, so anyone on the hub can see which devices belong to the same person and which run on the same host; making an agent hidden is how to keep that to yourself.

The device-id itself normally has to be learned out of band (Slack, email, a phone call) before it can be pasted into gateway_trust. A connection code is a single-use, short-lived artifact that makes that hand-off itself verifiable instead of a bare, unauthenticated string:

ConnectionCode {
  code: nonce            # single-use
  expiresAt: timestamp   # short-lived
  deviceId: hex          # the device being vouched for
  signature?: PgpSignature over (code, expiresAt, deviceId)
}

code/expiresAt are the always-checked half: proof that whoever redeems this was on the other end of this exact exchange, recently. signature is optional and answers a different question — this claim was made by whoever holds a specific long-lived PGP key. It's checked only when present, and never required: a device with no PGP identity still generates and redeems a bare code.

# Generate a code vouching for this device, unsigned
agent_comms({ action: "gateway_generate_connection_code" })

# ...or signed with a PGP private key you already have (agent-comms never generates or stores PGP key material itself -- you supply it per call, the same way you would to `gpg --sign` directly)
agent_comms({ action: "gateway_generate_connection_code", privateKey: "-----BEGIN PGP PRIVATE KEY BLOCK-----..." })

# Relay the printed code to the counterpart out of band, then redeem it on their side -- a successful redemption trusts the device it vouches for, exactly as if gateway_trust had been called directly
agent_comms({
  action: "gateway_redeem_connection_code",
  code: "...", expiresAt: "2026-01-01T00:15:00.000Z", device: "1a2b3c...",
})

# Redeeming a signed code needs the signer's public key to verify against -- paste it directly, or supply a fingerprint you already trust and let the redeemer fetch it from keys.openpgp.org
agent_comms({
  action: "gateway_redeem_connection_code",
  code: "...", expiresAt: "...", device: "1a2b3c...", signature: "...",
  fingerprint: "aabbccddeeff00112233445566778899aabbccd",
})

The fingerprint is never a trust anchor supplied by the keyserver — it's the thing the redeemer already independently trusts (from a business card, a prior verification, wherever), and the redeemer's own check is that the key actually fetched or pasted verifies to that exact fingerprint, not merely that some key was found. Redemption never calls back to whoever generated the code: the whole point is bootstrapping trust before any connection between the two devices exists.