On this page

The umbilical is a real-time stream of events describing an agent’s own activity, consumable by taps — warm lambdas that run inside the agent’s sandbox and receive matching events with low latency.

Every meaningful runtime action — tool call, LLM turn, DB read/write, file I/O, message delivery, trigger fire, timer fire, WS open/close, lambda invocation — emits an event. Taps filter, transform, forward, aggregate, or react using the full adf.* API.

For the full catalog of event types and payload shapes, see umbilical-events.md.


Why

Three things agents previously couldn’t do without instrumenting every call site by hand:

  • Real-time replication. Forward every db.write against local_orders to a peer agent — see the canonical recipe below.
  • Cross-agent tracing. A tap on tool.completed that emits custom.trace.span to the daemon bus via /events forwarding.
  • Self-optimization. A tap on turn.completed that watches token usage and suggests compaction when thresholds are crossed.

The underlying primitive is broader than any one of these — streams, replication, monitoring, and governance are all taps.


Tap configuration

Add umbilical_taps to your agent config:

{
  "umbilical_taps": [
    {
      "name": "order_replicator",
      "lambda": "lib/replicator.ts:onEvent",
      "filter": {
        "event_types": ["db.write"],
        "when": "event.payload.sql.includes('local_orders')"
      }
    }
  ]
}

Fields:

FieldRequiredDefaultDescription
nameYes—Unique identifier for this tap
lambdaYes—file:function ref; function runs with event as first arg
filter.event_typesNo["*"]Exact match or prefix ("tool.*"). "*" and bare prefixes require allow_wildcard: true
filter.whenNo—Expression over event; same semantics as trigger when
filter.allow_wildcardNofalseExplicit opt-in for wildcard filters
exclude_own_originNotrueSuppress dispatch when event.source equals "lambda:<this tap's lambda>"
max_rate_per_secNo100Token-bucket cap; overruns drop and log

Taps are warm. The lambda sandbox persists for the agent’s lifetime and module-level state is shared across events — critical for batching, caching, or maintaining in-memory queues.


Lambda contract

// lib/replicator.ts
import type { UmbilicalEvent } from '../types'

// Module-level state survives every invocation.
const pending: unknown[] = []

export async function onEvent(event: UmbilicalEvent): Promise<void> {
  // Filter handled at config level; this function only sees matching events.
  pending.push(event.payload)
  // ...
}

Event shape:

interface UmbilicalEvent {
  seq: number              // monotonic per-agent
  event_type: string
  timestamp: number        // epoch ms
  source: string           // agent:<turn>, lambda:<file>:<fn>, system:<subsystem>
  payload: Record<string, unknown>
}

Loop protection

Taps that act on events generate new events. Three layered mechanisms stop infinite loops:

  1. exclude_own_origin (default true). When the tap’s own action produces an event, the bus suppresses redelivery to the same tap by comparing event.source to "lambda:<this tap's lambda>".
  2. max_rate_per_sec token bucket. Backstop for multi-hop loops (tap A fires tap B fires tap A) and for filter mistakes. Overruns are dropped and logged.
  3. Wildcard opt-in. "*" or bare-prefix filters ("tool.*") require allow_wildcard: true explicitly. This is a code-review signal, not a functional guard.

Wildcard taps and lambda.*

A tap subscribed to "*" or "lambda.*" matches its own invocation because tap invocations produce lambda.started/completed/failed events. exclude_own_origin catches the direct case — but if the tap does something that causes another lambda to run, that lambda’s events pass the filter and fan back through the rate limiter.

The safe pattern is to exclude lambda.* explicitly in your when clause:

{
  "name": "universal_observer",
  "lambda": "lib/trace.ts:onEvent",
  "filter": {
    "event_types": ["*"],
    "allow_wildcard": true,
    "when": "!event.event_type.startsWith('lambda.')"
  }
}

Custom events

Agent code emits events with adf.emit_event:

await adf.emit_event({
  event_type: 'custom.signal.regime_change',
  payload: { regime: 'risk_off', confidence: 0.82 }
})

The custom. prefix is required — agents cannot spoof runtime-emitted events. The emit helper stamps source from the AsyncLocalStorage context automatically (either agent:<turn_id> or lambda:<file>:<fn> depending on where the code is running).


Replay window

The umbilical is in-memory and best-effort: an observer that disconnects, or connects late, has no way to see what it missed. The replay window is an opt-in, bounded, in-memory ring the runtime fills at publish time, so a reconnecting observer can catch up over a short gap instead of guessing.

{
  "umbilical": {
    "log": {
      "enabled": true,
      "max_events": 2000,
      "exclude_types": ["mcp.log"]
    }
  }
}
FieldDefaultDescription
enabledfalseOpt-in. No config, no window
max_events2000Ring capacity. The oldest event is evicted on overflow
exclude_types[]Additive to the built-in exclusions

turn.delta and binding.flow_summary are always excluded — they are per-token-batch and per-heartbeat volume, and there is no way to opt back in. Anything in exclude_types is dropped on top of those.

Payloads whose JSON serialization exceeds 4 KB are replaced by { "_truncated": true, "preview": "<first ~4000 chars>" } and flagged truncated: true. That bounds how much memory a single event can pin, and it tells the client what to do: the event happened, and the detail is available through the normal read API — not from the replay window.

Nothing is persisted

The window lives in the runtime process, keyed by agent id, and dies with the agent. Nothing is written to the agent’s database.

An earlier iteration of this feature wrote a durable, hash-chained table into the agent’s own local_* namespace. That was removed: local_* is agent space (the namespace db_execute writes to), and runtime bookkeeping does not belong in it — and, more to the point, the only consumer of this window re-snapshots when it falls behind, so durability bought nothing it actually used.

Durable, verifiable history is deliberately deferred, not forgotten. The design for it — signed epoch seals binding an event chain to the agent’s logical state, so a received .adf can be verified end to end — is written up in ../design/sealed-epochs.md. Read that before proposing to make this window durable again.

Until then: this is not an audit trail. For tamper-resistant records of what an agent did, use adf_audit.

Snapshot-then-tail

The recipe for observing an agent remotely without an SSE gap to reason about:

  1. Snapshot — read state through the existing endpoints: GET /agents/:id/config, /loop, /files, /logs, /tasks, …
  2. Probe and tail — GET /agents/:id/umbilical/events?since_seq=<n>&limit=<m>. It answers 200 with { "events": [], "last_seq": null, "log_enabled": false } when the window is off, so a client can probe cheaply.
  3. Feed last_seq back as the next since_seq.

Detecting a gap. Every response with a non-empty window carries oldest_seq. If your cursor predates the window — since_seq < oldest_seq - 1 — you have fallen off the back and the tail you are about to receive is incomplete. Re-snapshot; do not stitch. This is the whole client contract, and it is why the window does not need to be durable.

Because seq is the umbilical’s own monotonic per-agent counter (persisted in workspace meta and reserved in blocks), a cursor never collides across restarts — but the window itself does not survive an agent restart, so a restart is simply another gap: oldest_seq moves forward, the client re-snapshots. Size max_events for your worst expected disconnection.

GET /events (SSE) remains the low-latency path. The replay window is what makes a short gap in that stream recoverable.

Not the same as the durable-tap recipe

The recipe below buffers events into a local_* table too, but for a different reason: application-level at-least-once delivery, owned by the agent’s own lambdas, with acks and retries. That pattern stays exactly as it is. The replay window is runtime-owned, in-memory, lossy by design (bounded, truncated, exclusion-filtered), and for observation. Do not build replication on it.


Canonical durable-tap recipe — DB replication

The umbilical is best-effort. Events can drop on throttle, be lost during agent downtime, or be dropped when a tap handler throws. Taps that need at-least-once delivery implement durability themselves.

The pattern: buffer events into a local_* table, process in a separate flush loop, delete on ack. If the process crashes mid-flush, queued rows persist and flush retries them on next startup.

Tap 1 — enqueue (synchronous, narrow, cannot lose events)

// lib/replication-queue.ts
export async function enqueue(event: UmbilicalEvent): Promise<void> {
  // Single INSERT. If this fails the event is lost — but db writes are
  // synchronous and reliable locally, so the failure mode is narrow.
  await adf.db_execute({
    sql: `INSERT INTO local_replication_queue (seq, payload_json, enqueued_at)
          VALUES (?, ?, ?)`,
    params: [event.seq, JSON.stringify(event.payload), Date.now()]
  })
}

Config:

{
  "umbilical_taps": [
    {
      "name": "replication_enqueue",
      "lambda": "lib/replication-queue.ts:enqueue",
      "filter": {
        "event_types": ["db.write"],
        "when": "event.payload.sql.includes('local_orders')"
      },
      "max_rate_per_sec": 10000
    }
  ]
}

max_rate_per_sec is raised far above the default so the enqueue path never throttles — throttling silently loses events, which defeats durability.

Table schema

CREATE TABLE local_replication_queue (
  seq INTEGER PRIMARY KEY,       -- event.seq, monotonic per agent
  payload_json TEXT NOT NULL,
  enqueued_at INTEGER NOT NULL,
  attempt_count INTEGER DEFAULT 0
);

seq is the primary key — if the same event somehow gets enqueued twice (e.g. handler retry after a crash between emit and ack), the second INSERT fails and nothing is duplicated.

Flush loop — a second tap on timer.fired

// lib/replication-flush.ts
export async function flushTick(): Promise<void> {
  const pending = await adf.db_query({
    sql: `SELECT seq, payload_json FROM local_replication_queue
          ORDER BY seq ASC LIMIT 100`
  })
  for (const row of pending) {
    try {
      await adf.ws_send({
        connection_id: PEER_CONN_ID,
        data: row.payload_json
      })
      // Ack by delete — if the process crashes before this, the row stays
      // and gets retried on next tick.
      await adf.db_execute({
        sql: `DELETE FROM local_replication_queue WHERE seq = ?`,
        params: [row.seq]
      })
    } catch (err) {
      // Leave the row. Next tick will retry. Consider incrementing
      // attempt_count and dead-lettering beyond a threshold.
      break
    }
  }
}

Why this works under crashes

  • The enqueue tap is synchronous: when adf.db_execute returns, the row is durable in SQLite’s WAL.
  • The flush tap reads-then-sends-then-deletes. A crash between send and delete leaves the row; next flush retries.
  • On receiver side, idempotency comes from the seq primary key in the analogous local_inbox_seq table: INSERT OR IGNORE makes duplicates a no-op.
  • Replay on restart is automatic — the table is the queue.

The SQL semantics behind this pattern — the queue-and-ack, crash-between-send- and-delete, and idempotent-receive-by-seq guarantees — are exercised by tests/integration/umbilical-replication-crash.test.ts. That test deliberately models the queue directly and does not run the full umbilical dispatch or the tap machinery, so treat it as coverage of the durability recipe’s data layer, not of tap wiring end to end.


External forwarding

The daemon’s /events SSE endpoint already serves external observers. The umbilical feeds from the same bus, so external forwarding isn’t automatic — configure a tap that forwards explicitly via adf.emit_event (custom namespace) or the HTTP tool.

{
  "name": "external_trace",
  "lambda": "lib/forward-external.ts:onEvent",
  "filter": { "event_types": ["tool.*"] }
}

Forwarding tool.* externally will forward tool parameters — which may include credentials. Filter or redact in the tap.


Relationship to triggers

Triggers and taps answer different questions and are not collapsing into each other:

  • Triggers — named, typed, single-category hooks. “When X happens, do Y.” on_inbox, on_timer, on_document_edit, on_tool_call, etc. Each trigger type has a specific event shape appropriate to its category. Good for reactive agent logic where the mental model is “run this lambda when that thing happens.”

  • Umbilical taps — uniform event-stream subscribers. “Observe every event of these types with this shape, wherever it comes from.” One event envelope (UmbilicalEvent), one subscribe API, one dispatch path. Good for cross-cutting observability (tracing, replication, monitoring), fleet-scale telemetry forwarding, and any workload that wants to treat many event categories uniformly.

There is overlap — an on_inbox trigger and a tap filtered on message.received can both react to an inbound message. Prefer triggers for reactive logic (they read more naturally and their filters are category-specific); prefer taps for observation, forwarding, and anything that needs to compose across event types.

The two systems will continue to evolve together but separately. Don’t look for a future unification — the uniform envelope is the tap’s feature, and cramming triggers into it would re-introduce the trigger-type schema variance that makes taps valuable in the first place.