# The Umbilical

> The real-time event stream of an agent's own activity: taps, loop protection, replay, and the durable-tap recipe

Source: https://github.com/christianbalevski/adf/blob/v0.7.4/docs/guides/umbilical.md (adf v0.7.4)

The umbilical is a real-time stream of events describing an agent's own
activity, consumable by **taps** — warm lambdas that run inside the agent's
sandbox and receive matching events with low latency.

Every meaningful runtime action — tool call, LLM turn, DB read/write, file
I/O, message delivery, trigger fire, timer fire, WS open/close, lambda
invocation — emits an event. Taps filter, transform, forward, aggregate, or
react using the full `adf.*` API.

For the full catalog of event types and payload shapes, see
[umbilical-events.md](https://agentdocumentformat.org/guides/umbilical-events).

---

## Why

Three things agents previously couldn't do without instrumenting every call
site by hand:

- **Real-time replication.** Forward every `db.write` against `local_orders`
  to a peer agent — see the canonical recipe below.
- **Cross-agent tracing.** A tap on `tool.completed` that emits
  `custom.trace.span` to the daemon bus via `/events` forwarding.
- **Self-optimization.** A tap on `turn.completed` that watches token
  usage and suggests compaction when thresholds are crossed.

The underlying primitive is broader than any one of these — streams,
replication, monitoring, and governance are all taps.

---

## Tap configuration

Add `umbilical_taps` to your agent config:

```jsonc
{
  "umbilical_taps": [
    {
      "name": "order_replicator",
      "lambda": "lib/replicator.ts:onEvent",
      "filter": {
        "event_types": ["db.write"],
        "when": "event.payload.sql.includes('local_orders')"
      }
    }
  ]
}
```

Fields:

| Field | Required | Default | Description |
|---|---|---|---|
| `name` | Yes | — | Unique identifier for this tap |
| `lambda` | Yes | — | `file:function` ref; function runs with event as first arg |
| `filter.event_types` | No | `["*"]` | Exact match or prefix (`"tool.*"`). `"*"` and bare prefixes require `allow_wildcard: true` |
| `filter.when` | No | — | Expression over `event`; same semantics as trigger `when` |
| `filter.allow_wildcard` | No | `false` | Explicit opt-in for wildcard filters |
| `exclude_own_origin` | No | `true` | Suppress dispatch when `event.source` equals `"lambda:<this tap's lambda>"` |
| `max_rate_per_sec` | No | `100` | Token-bucket cap; overruns drop and log |

Taps are **warm**. The lambda sandbox persists for the agent's lifetime and
module-level state is shared across events — critical for batching, caching,
or maintaining in-memory queues.

---

## Lambda contract

```typescript
// lib/replicator.ts
import type { UmbilicalEvent } from '../types'

// Module-level state survives every invocation.
const pending: unknown[] = []

export async function onEvent(event: UmbilicalEvent): Promise<void> {
  // Filter handled at config level; this function only sees matching events.
  pending.push(event.payload)
  // ...
}
```

Event shape:

```typescript
interface UmbilicalEvent {
  seq: number              // monotonic per-agent
  event_type: string
  timestamp: number        // epoch ms
  source: string           // agent:<turn>, lambda:<file>:<fn>, system:<subsystem>
  payload: Record<string, unknown>
}
```

---

## Loop protection

Taps that act on events generate new events. Three layered mechanisms stop
infinite loops:

1. **`exclude_own_origin` (default `true`).** When the tap's own action
   produces an event, the bus suppresses redelivery to the same tap by
   comparing `event.source` to `"lambda:<this tap's lambda>"`.
2. **`max_rate_per_sec` token bucket.** Backstop for multi-hop loops
   (tap A fires tap B fires tap A) and for filter mistakes. Overruns are
   dropped and logged.
3. **Wildcard opt-in.** `"*"` or bare-prefix filters (`"tool.*"`) require
   `allow_wildcard: true` explicitly. This is a code-review signal, not a
   functional guard.

### Wildcard taps and `lambda.*`

A tap subscribed to `"*"` or `"lambda.*"` matches its own invocation because
tap invocations produce `lambda.started/completed/failed` events.
`exclude_own_origin` catches the direct case — but if the tap does something
that causes another lambda to run, that lambda's events pass the filter and
fan back through the rate limiter.

The safe pattern is to exclude `lambda.*` explicitly in your `when` clause:

```jsonc
{
  "name": "universal_observer",
  "lambda": "lib/trace.ts:onEvent",
  "filter": {
    "event_types": ["*"],
    "allow_wildcard": true,
    "when": "!event.event_type.startsWith('lambda.')"
  }
}
```

---

## Custom events

Agent code emits events with `adf.emit_event`:

```typescript
await adf.emit_event({
  event_type: 'custom.signal.regime_change',
  payload: { regime: 'risk_off', confidence: 0.82 }
})
```

The `custom.` prefix is **required** — agents cannot spoof runtime-emitted
events. The emit helper stamps `source` from the AsyncLocalStorage context
automatically (either `agent:<turn_id>` or `lambda:<file>:<fn>` depending on
where the code is running).

---

## Replay window

The umbilical is in-memory and best-effort: an observer that disconnects, or
connects late, has no way to see what it missed. The **replay window** is an
opt-in, bounded, **in-memory** ring the runtime fills at publish time, so a
reconnecting observer can catch up over a short gap instead of guessing.

```jsonc
{
  "umbilical": {
    "log": {
      "enabled": true,
      "max_events": 2000,
      "exclude_types": ["mcp.log"]
    }
  }
}
```

| Field | Default | Description |
|---|---|---|
| `enabled` | `false` | Opt-in. No config, no window |
| `max_events` | `2000` | Ring capacity. The oldest event is evicted on overflow |
| `exclude_types` | `[]` | **Additive** to the built-in exclusions |

`turn.delta` and `binding.flow_summary` are **always** excluded — they are
per-token-batch and per-heartbeat volume, and there is no way to opt back in.
Anything in `exclude_types` is dropped on top of those.

Payloads whose JSON serialization exceeds **4 KB** are replaced by
`{ "_truncated": true, "preview": "<first ~4000 chars>" }` and flagged
`truncated: true`. That bounds how much memory a single event can pin, and it
tells the client what to do: the event happened, and the detail is available
through the normal read API — not from the replay window.

### Nothing is persisted

The window lives in the runtime process, keyed by agent id, and **dies with the
agent**. Nothing is written to the agent's database.

An earlier iteration of this feature wrote a durable, hash-chained table into the
agent's own `local_*` namespace. That was removed: `local_*` is agent space (the
namespace `db_execute` writes to), and runtime bookkeeping does not belong in it
— and, more to the point, the only consumer of this window re-snapshots when it
falls behind, so durability bought nothing it actually used.

**Durable, verifiable history is deliberately deferred**, not forgotten. The
design for it — signed epoch seals binding an event chain to the agent's logical
state, so a received `.adf` can be verified end to end — is written up in
[../design/sealed-epochs.md](https://github.com/christianbalevski/adf/blob/v0.7.4/docs/design/sealed-epochs.md). Read that before
proposing to make this window durable again.

Until then: **this is not an audit trail.** For tamper-resistant records of what
an agent did, use `adf_audit`.

### Snapshot-then-tail

The recipe for observing an agent remotely without an SSE gap to reason about:

1. **Snapshot** — read state through the existing endpoints:
   `GET /agents/:id/config`, `/loop`, `/files`, `/logs`, `/tasks`, …
2. **Probe and tail** —
   `GET /agents/:id/umbilical/events?since_seq=<n>&limit=<m>`. It answers `200`
   with `{ "events": [], "last_seq": null, "log_enabled": false }` when the
   window is off, so a client can probe cheaply.
3. Feed `last_seq` back as the next `since_seq`.

**Detecting a gap.** Every response with a non-empty window carries `oldest_seq`.
If your cursor predates the window — `since_seq < oldest_seq - 1` — you have
fallen off the back and the tail you are about to receive is **incomplete**.
Re-snapshot; do not stitch. This is the whole client contract, and it is why the
window does not need to be durable.

Because `seq` is the umbilical's own monotonic per-agent counter (persisted in
workspace meta and reserved in blocks), a cursor never collides across restarts —
but the window itself does not survive an agent restart, so a restart is simply
another gap: `oldest_seq` moves forward, the client re-snapshots. Size
`max_events` for your worst expected disconnection.

`GET /events` (SSE) remains the low-latency path. The replay window is what makes
a *short* gap in that stream recoverable.

### Not the same as the durable-tap recipe

The recipe below buffers events into a `local_*` table too, but for a different
reason: **application-level at-least-once delivery**, owned by the agent's own
lambdas, with acks and retries. That pattern stays exactly as it is. The replay
window is runtime-owned, in-memory, lossy by design (bounded, truncated,
exclusion-filtered), and for *observation*. Do not build replication on it.

---

## Canonical durable-tap recipe — DB replication

The umbilical is **best-effort**. Events can drop on throttle, be lost
during agent downtime, or be dropped when a tap handler throws. Taps that
need at-least-once delivery implement durability themselves.

The pattern: buffer events into a `local_*` table, process in a separate
flush loop, delete on ack. If the process crashes mid-flush, queued rows
persist and flush retries them on next startup.

### Tap 1 — enqueue (synchronous, narrow, cannot lose events)

```typescript
// lib/replication-queue.ts
export async function enqueue(event: UmbilicalEvent): Promise<void> {
  // Single INSERT. If this fails the event is lost — but db writes are
  // synchronous and reliable locally, so the failure mode is narrow.
  await adf.db_execute({
    sql: `INSERT INTO local_replication_queue (seq, payload_json, enqueued_at)
          VALUES (?, ?, ?)`,
    params: [event.seq, JSON.stringify(event.payload), Date.now()]
  })
}
```

Config:

```jsonc
{
  "umbilical_taps": [
    {
      "name": "replication_enqueue",
      "lambda": "lib/replication-queue.ts:enqueue",
      "filter": {
        "event_types": ["db.write"],
        "when": "event.payload.sql.includes('local_orders')"
      },
      "max_rate_per_sec": 10000
    }
  ]
}
```

`max_rate_per_sec` is raised far above the default so the enqueue path never
throttles — throttling silently loses events, which defeats durability.

### Table schema

```sql
CREATE TABLE local_replication_queue (
  seq INTEGER PRIMARY KEY,       -- event.seq, monotonic per agent
  payload_json TEXT NOT NULL,
  enqueued_at INTEGER NOT NULL,
  attempt_count INTEGER DEFAULT 0
);
```

`seq` is the primary key — if the same event somehow gets enqueued twice
(e.g. handler retry after a crash between emit and ack), the second INSERT
fails and nothing is duplicated.

### Flush loop — a second tap on `timer.fired`

```typescript
// lib/replication-flush.ts
export async function flushTick(): Promise<void> {
  const pending = await adf.db_query({
    sql: `SELECT seq, payload_json FROM local_replication_queue
          ORDER BY seq ASC LIMIT 100`
  })
  for (const row of pending) {
    try {
      await adf.ws_send({
        connection_id: PEER_CONN_ID,
        data: row.payload_json
      })
      // Ack by delete — if the process crashes before this, the row stays
      // and gets retried on next tick.
      await adf.db_execute({
        sql: `DELETE FROM local_replication_queue WHERE seq = ?`,
        params: [row.seq]
      })
    } catch (err) {
      // Leave the row. Next tick will retry. Consider incrementing
      // attempt_count and dead-lettering beyond a threshold.
      break
    }
  }
}
```

### Why this works under crashes

- The enqueue tap is synchronous: when `adf.db_execute` returns, the row is
  durable in SQLite's WAL.
- The flush tap reads-then-sends-then-deletes. A crash between send and
  delete leaves the row; next flush retries.
- On receiver side, idempotency comes from the `seq` primary key in the
  analogous `local_inbox_seq` table: INSERT OR IGNORE makes duplicates a no-op.
- Replay on restart is automatic — the table is the queue.

The **SQL semantics** behind this pattern — the queue-and-ack, crash-between-send-
and-delete, and idempotent-receive-by-`seq` guarantees — are exercised by
`tests/integration/umbilical-replication-crash.test.ts`. That test deliberately
models the queue directly and does **not** run the full umbilical dispatch or the
tap machinery, so treat it as coverage of the durability recipe's data layer, not
of tap wiring end to end.

---

## External forwarding

The daemon's `/events` SSE endpoint already serves external observers. The
umbilical feeds from the same bus, so external forwarding isn't automatic —
configure a tap that forwards explicitly via `adf.emit_event` (custom namespace)
or the HTTP tool.

```jsonc
{
  "name": "external_trace",
  "lambda": "lib/forward-external.ts:onEvent",
  "filter": { "event_types": ["tool.*"] }
}
```

Forwarding `tool.*` externally will forward tool parameters — which may
include credentials. Filter or redact in the tap.

---

## Relationship to triggers

Triggers and taps answer different questions and are not collapsing into each
other:

- **Triggers** — named, typed, single-category hooks. "When X happens, do Y."
  `on_inbox`, `on_timer`, `on_document_edit`, `on_tool_call`, etc. Each trigger
  type has a specific event shape appropriate to its category. Good for
  reactive agent logic where the mental model is "run this lambda when that
  thing happens."

- **Umbilical taps** — uniform event-stream subscribers. "Observe every event
  of these types with this shape, wherever it comes from." One event envelope
  (`UmbilicalEvent`), one subscribe API, one dispatch path. Good for
  cross-cutting observability (tracing, replication, monitoring), fleet-scale
  telemetry forwarding, and any workload that wants to treat many event
  categories uniformly.

There is overlap — an `on_inbox` trigger and a tap filtered on
`message.received` can both react to an inbound message. Prefer triggers for
reactive logic (they read more naturally and their filters are
category-specific); prefer taps for observation, forwarding, and anything
that needs to compose across event types.

The two systems will continue to evolve together but separately. Don't look
for a future unification — the uniform envelope is the tap's feature, and
cramming triggers into it would re-introduce the trigger-type schema
variance that makes taps valuable in the first place.
