# Security Architecture

> The runtime security model: trust boundaries, defense-in-depth layers, and hardening controls

Source: https://github.com/christianbalevski/adf/blob/v0.7.4/docs/guides/security-architecture.md (adf v0.7.4)

ADF Studio runs AI agents that execute code, send messages, and interact with external services. This page documents the security model — trust boundaries, defense layers, and hardening controls.

For identity and encryption specifics, see [Security and Identity](https://agentdocumentformat.org/guides/security-and-identity). For sandbox details, see [Code Execution Environment](https://agentdocumentformat.org/guides/code-execution).

## Trust Boundaries

ADF Studio has five trust boundaries. A vulnerability at any boundary can escalate privileges.

### 1. Renderer to Main Process

The Electron renderer is treated as untrusted. Even though it's our own React app, XSS from LLM output or inbound messages could inject code into the renderer context.

**Controls:**
- `contextIsolation: true`, `sandbox: true`, `nodeIntegration: false`
- Content Security Policy (CSP) set via `session.webRequest.onHeadersReceived`:
  - `script-src 'self'` — blocks inline scripts and `eval()`
  - `worker-src 'self' blob:` — allows sandbox workers
  - `connect-src` restricted to localhost — all AI provider calls go through main process IPC, so the renderer has no reason to connect externally
  - `frame-src 'none'`, `object-src 'none'` — blocks iframes and plugins
- DOMPurify sanitization on all markdown rendered from LLM output and tool results
- `will-navigate` handler blocks renderer navigation to external pages
- `shell.openExternal` validates URL protocol (`https:`, `http:`, `mailto:` only)

**What this means:** Even if an attacker gets HTML into the chat (via LLM output, inbound messages, or tool results), inline scripts are blocked by CSP, event handlers are stripped by DOMPurify, and the renderer cannot fetch external URLs.

### 2. Sandbox to Main Process

`sys_code` and `sys_lambda` execute in a Node.js Worker Thread with a V8 VM context. This is not a security sandbox in the browser sense — `vm` does not provide hard isolation. The defense is layered:

**Controls:**
- `codeGeneration: { strings: false }` — no `eval()` or `new Function()`
- All built-in prototypes frozen inside the VM context (Object, Array, Function, String, etc.)
- `fetch`, `Request`, `Response`, `Headers` deleted from worker scope — all network goes through `adf.sys_fetch()` which routes through security middleware
- Module allowlist: only `crypto`, `buffer`, `url`, `querystring`, `path`, `util`, `string_decoder`, `punycode`, `assert`, `events`, `stream`, `zlib`
- Execution timeout: effective default is `limits.execution_timeout_ms` (60s by default), and the ceiling is `min(limits.execution_timeout_ms, 300s)`, enforced by worker termination (a bare 10s fallback applies only when no timeout is supplied at all)
- RPC bridge (`adf` proxy) validates every tool call against the agent's config before execution

**What this means:** Code cannot access the filesystem, spawn processes, or make network requests except through the `adf` proxy, which enforces tool enablement, restriction checks, and middleware.

### 3. Network Boundary

The mesh server, WebSocket connections, and channel adapters accept input from the network.

**Controls:**
- Mesh server binds to `127.0.0.1` by default — LAN exposure requires explicit `meshLan` setting
- Ed25519 message signature verification (envelope and payload)
- Configurable `allow_unsigned` (default: true for local dev, should be false for internet-facing agents)
- Allow/block lists for message senders (by DID)
- Inbox middleware pipeline — user-defined lambdas can inspect and reject messages before storage

**What this means:** Network-sourced messages go through signature verification, allow/block filtering, and middleware before reaching the agent. With `allow_unsigned: false` and proper identity setup, only verified senders can deliver messages.

### 4. External Process Boundary

MCP servers and user-installed packages run external code with the user's OS privileges.

**Controls:**
- Blocked environment variables for MCP server processes: `ELECTRON_RUN_AS_NODE`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`, `LD_LIBRARY_PATH`, `DYLD_LIBRARY_PATH`
- MCP server health checks with auto-restart (60s interval, 10s timeout)
- Connection timeout: 120s (allows for initial npx/uvx downloads)
- Package installs: native addon detection and blocking (scans for `binding.gyp`, `node-gyp` postinstall scripts, `gypfile: true`); per-package limit 50 MB, total 200 MB

**What this means:** MCP servers are operator-configured and trusted by design. The controls prevent process-level injection attacks but don't sandbox the MCP server itself. When running in a container (shared or isolated), servers are further isolated from the host. See [Compute Environments](https://agentdocumentformat.org/guides/compute) for the full compute security model, including critical implications of host access.

**Settings install trust model:** the run-location default follows *who initiated the install*. A server the user installs in Settings runs on the **host** by default — the user's explicit click is the trust decision (the same boundary every conventional MCP harness assumes), it is labeled with a persistent host badge, and the server name is auto-added to the host-approved list. A server an *agent* installs via `mcp_install` defaults to the **container** and must pass the two-tier host gate (`compute.host_access` plus the app-wide toggle) to reach the host — an autonomous install is a weaker trust event than a human one. Note that a host server remains continuously drivable by agents after install, including via prompt-injected tool outputs; container placement is the per-server hardening upgrade for that exposure.

### 5. Storage Boundary

`.adf` files are user-controlled SQLite databases. Opening an untrusted `.adf` file loads configuration, triggers, lambdas, and stored content.

**Controls:**
- Dangerous tools disabled by default — `adf_shell` and `ws_*` require explicit enablement. `sys_fetch` ships enabled but is confined by the egress guard (see [sys_fetch Egress Guard](#sys_fetch--ws_connect-egress-guard-ssrf))
- `restricted` tools get automatic HIL when called from the LLM loop, and are blocked from unauthorized code
- Trigger lambdas only fire if the trigger type and scope are configured
- Identity secrets encrypted at rest with AES-256-GCM (PBKDF2 key derivation, 100k iterations, SHA-512)
- `code_access` flag gates per-key access from code execution
- File protection levels: `read_only` (immutable), `no_delete` (writable but not deletable)

**What this means:** An untrusted `.adf` file can contain malicious config, but sensitive capabilities require explicit enablement. The risk scales with what the user enables.

## Defense-in-Depth Layers

Security relies on multiple independent layers rather than any single control:

| Layer | Protects Against | Bypass Condition |
|-------|-----------------|------------------|
| DOMPurify | XSS from LLM/message HTML | DOMPurify mutation bypass (rare, patched quickly) |
| CSP `script-src 'self'` | Inline script injection | Not bypassable without `'unsafe-inline'` |
| CSP `connect-src` localhost | XSS data exfiltration | Not bypassable from renderer |
| `will-navigate` handler | Renderer hijacking to external pages | Not bypassable (Electron event) |
| Tool enablement | Unauthorized tool use from code | Config manipulation via `sys_update_config` (gated by lock fields). Agents cannot modify `restricted`/`restricted_methods` — owner only |
| `restricted` (HIL) | Autonomous execution of dangerous tools | Cannot be bypassed from unauthorized code — returns `REQUIRES_AUTHORIZED_CODE` |
| `require_middleware_authorization` | Untrusted middleware modifying messages | Unauthorized middleware silently skipped (default on) |
| Signature verification | Spoofed network messages | `allow_unsigned: true` disables this check |
| Fetch middleware | SSRF and unauthorized outbound requests | Only effective if configured |
| SQL sanitizer | Access to system tables from `db_query`/`db_execute` | Validated against allowlist after stripping comments/literals |

## Tool Access Control

Tool access is governed by three flags on each `ToolDeclaration`:

- **`enabled`** — whether the tool exists for the agent and may be called. This is the only flag that gates execution; it applies to the LLM, lambdas, and other code
- **`visible`** — whether the tool is advertised in the LLM's tool schema. This controls only what the model is shown, **not** whether it may call the tool — an enabled tool runs even when `visible: false`
- **`restricted`** — whether the tool requires authorization (optional, defaults to `false`)
- **`locked`** — whether the agent can modify this tool's config via `sys_update_config` (optional, defaults to `false`)

The access matrix (the **Advertised** column reflects `visible`; execution is gated on `enabled`, never on visibility):

| `enabled` | `visible` | `restricted` | Advertised | LLM loop | Authorized code | Unauthorized code |
|-----------|-----------|--------------|------------|----------|-----------------|-------------------|
| `false`   | —         | `false`      | No  | Off  | Off  | Off  |
| `false`   | —         | `true`       | No  | Off  | Free | Off  |
| `true`    | `false`   | `false`      | No  | Free | Free | Free |
| `true`    | `false`   | `true`       | No  | HIL  | Free | Off  |
| `true`    | `true`    | `false`      | Yes | Free | Free | Free |
| `true`    | `true`    | `true`       | Yes | HIL  | Free | Off  |

When a tool is `enabled` and `restricted`, LLM loop calls automatically get a human-in-the-loop approval prompt (whether or not the tool is visible). Authorized code can call the tool directly without approval. Unauthorized code cannot call restricted tools at all — it receives a `REQUIRES_AUTHORIZED_CODE` error.

**`sys_lambda` authorization gate:** In addition to tool-level restriction, `sys_lambda` has argument-dependent HIL. When the LLM calls `sys_lambda` targeting an authorized file, the runtime triggers a HIL approval prompt regardless of whether `sys_lambda` itself is restricted. This ensures the user has visibility whenever the agent invokes code with elevated privileges from the conversation loop. See [Authorized Code Execution](https://agentdocumentformat.org/guides/authorized-code) for details.

**Self-modification protection:** Agents can toggle `enabled` on unlocked tools via `sys_update_config` (useful for token optimization), but **cannot modify `restricted`, `restricted_methods`, or `locked`** — these are owner-only security boundaries. Disabling a tool without locking it is a suggestion, not a boundary; to enforce a tool being off, lock it or disable `sys_update_config`.

**Guard-system config is hard-denied.** The switches that decide *what needs approval in the first place* are not agent-reachable at all through `sys_update_config` — they return a plain error with no `ProtectionDenial`, so there is no HIL prompt and no one-time override path. These guard paths are:

- `security.allow_unsigned`
- `security.require_middleware_authorization`
- `security.middleware.*`
- `security.fetch_middleware`
- a wholesale `set` of `security` (replacing the entire guard block)

The membership criterion: a switch is a guard only when changing it would weaken the mechanism that gates the agent — approval-deciding (`allow_unsigned`, `require_middleware_authorization`) or self-blinding (the middleware chains, which rewrite what the agent and its audit trail see).

Contrast this with the agent's own **capability toggles** — `code_execution.*`, tool enable/disable, `limits.*`. Those remain requestable: the agent may ask for any change a human could make, and a locked one surfaces as an HIL approval the owner can grant or deny. To *hard-enforce* a capability toggle (not just gate it behind HIL), add it to `locked_fields`. Dangerous-but-ordinary capability toggles — `security.allow_local_fetch` and the `stream_bind` gates — are locked by default in the runtime (for every agent), so enabling them is a deliberate one-time override the owner grants, not a free write and not an unreachable wall. The rule: guards that govern the approval machinery are never agent-writable; the capabilities that machinery protects are HIL-gated and requestable. An instruction like "set `security.allow_unsigned: false`" is therefore an **app UI / owner-console** action, never a `sys_update_config` call.

Additionally, these tools are **excluded from code execution** entirely:
- `say` — prevents code from monopolizing chat output
- `ask` — prevents code from bypassing human-in-the-loop

The following methods are gated by `code_execution` config flags:
- `model_invoke` — direct LLM calls
- `sys_lambda` — execute lambda functions
- `task_resolve` — approve/deny intercepted tasks
- `loop_inject` — inject context into conversation loop
- `identity_status` — read only envelope protection state, never identity values or keys
- `get_identity` / `set_identity` — read/write identity secrets
- `emit_event` — emit a `custom.*` umbilical event
- `attestation_list` / `attestation_add` / `attestation_issue` — read, store, and sign attestations

Code execution methods can also be individually restricted via `code_execution.restricted_methods`. Methods in this list can only be called from code that originates from an authorized file. This prevents agents from self-approving tasks, accessing credentials, or calling other sensitive methods from agent-written code. The list defaults to `['attestation_issue']` — signing a cert about another DID is authorized-code-only out of the box — and an explicit list replaces (does not merge with) that default.

See [Authorized Code Execution](https://agentdocumentformat.org/guides/authorized-code) for the full security model, authorization flow, governance patterns, and threat analysis.

## SQL Access Control

| Table Pattern | `db_query` | `db_execute` |
|--------------|-----------|-------------|
| `adf_loop`, `adf_inbox`, `adf_outbox`, `adf_files`, `adf_timers`, `adf_audit`, `adf_logs`, `adf_tasks` | Read-only | Blocked |
| `adf_meta`, `adf_config` | Blocked | Blocked |
| `adf_identity` | Blocked | Blocked |
| `local_*` | Read-only | Full access (INSERT, UPDATE, DELETE, CREATE, DROP) |
| PRAGMA table-valued functions | Blocked | Blocked |

SQL input is sanitized before validation: comments stripped, string literals replaced, multi-statement queries rejected.

## Shell Command Security

Shell commands go through a pre-flight AST analysis before execution:

1. The command is tokenized and parsed into an AST
2. Each resolved tool (including redirects like `>` mapping to `fs_write`) is checked against the agent's tool declarations
3. Disabled tools → rejection (exit 126)
4. Restricted tools → HIL task creation, human approval prompt (exit 130)

Commands are executed via `execFile` with array arguments (not shell strings), preventing shell injection.

## sys_fetch / ws_connect Egress Guard (SSRF)

`sys_fetch` and `ws_connect` are reachable from the LLM loop, so a prompt-injected agent could otherwise drive the local daemon, cloud metadata endpoints, or anything else on the LAN. The egress guard default-denies private/LAN destinations before the request leaves the process; **loopback is allowed by default** (local dev servers are a core agent use case), with the daemon control API carved out as a hard block. The same guard also covers `ws_connect` (ws/wss ad-hoc URLs).

Allowed by default:

- **Loopback** — `127.0.0.0/8`, `::1`, `localhost` (and `*.localhost`), including the loose `inet_aton` spellings (`127.1`, `0x7f000001`, `2130706433`, `0177.0.0.1`) — **except the daemon control API port, which is never reachable**

Blocked by default (overridable with `security.allow_local_fetch: true`):

- **RFC1918 + CGNAT** — `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`, `100.64.0.0/10` (Tailscale addresses live in this CGNAT range)
- **Unspecified** — `0.0.0.0`/`::` (reaches loopback listeners on most stacks, but is not treated as trusted loopback), plus multicast/reserved and unique-local (`fc00::/7`) IPv6 forms

Always blocked (ignores `allow_local_fetch`):

- **Link-local** — `169.254.0.0/16` (including cloud metadata `169.254.169.254`), `fe80::/10`
- **The local ADF daemon control API** on loopback (including via `0.0.0.0`/`::`) — this closes the "injected agent drives the local daemon" path unconditionally: without it, a single poisoned message could make the agent POST to `127.0.0.1` control endpoints on the operator's own machine. (The daemon also requires its bearer token and refuses cross-site browser requests; see [Daemon HTTP API → Authentication](https://agentdocumentformat.org/api).)

The check runs on **DNS-resolved addresses**, not just the literal host — a rebinding record that resolves to `192.168.0.1` is rejected — and on **every redirect hop**: a public URL that `302`s to a private address is stopped at the hop. Redirects are ALWAYS followed manually and re-checked, even under `allow_local_fetch`. Only `http`/`https` URLs are permitted for `sys_fetch` (and `ws`/`wss` for `ws_connect`).

`security.allow_local_fetch` is locked by default in the runtime (for every agent — see [Tool Access Control](#tool-access-control)): an agent's write is denied but surfaces as a protection request the owner can approve as a one-time override.

## Identity and Cryptography

| Component | Algorithm | Parameters |
|-----------|-----------|------------|
| Encryption | AES-256-GCM | 12-byte IV, 16-byte auth tag |
| Envelope key wrap | X25519 ECDH + HKDF-SHA256 | Ephemeral sender key; per-envelope domain separation |
| Share-password KDF | scrypt | N=2^17, r=8, p=1, 32-byte salt |
| Legacy password KDF | PBKDF2 | 100,000 iterations, SHA-512, 32-byte salt (read support) |
| Signing | Ed25519 | PKCS8/SPKI DER format |
| Agent identity | DID:key | `did:key:z{base58btc(0xed01 + pubkey)}` |

Identity secrets in `adf_identity` are envelope-sealed at rest: a per-envelope DEK encrypts the rows, wrapped to the owner and runtime encryption keys (and optionally a share password). Key material (`crypto:signing:*`, `crypto:envelope:*`, `crypto:kdf:*`) is never readable from agent code regardless of `code_access`. See [Security and Identity](https://agentdocumentformat.org/guides/security-and-identity) for the full identity lifecycle.

## Attacker-Controlled Inputs

These are the primary sources of untrusted data that flow through the system:

| Input | Entry Point | First Defense |
|-------|------------|---------------|
| LLM output | Agent loop | DOMPurify + CSP (renderer), tool validation (runtime) |
| Inbound ALF messages | Mesh server / adapters | Signature verification, allow/block lists, inbox middleware |
| `.adf` file contents | File open | Tool enablement defaults, restricted tools, lock fields |
| HTTP responses (`sys_fetch`) | Code execution | Fetch middleware pipeline |
| MCP server responses | MCP client | Tool result size limits (`max_tool_result_tokens`) |
| User-installed packages | `npm_install` | Native addon blocking, size limits |

**Verification meta is runtime-stamped only.** The trust flags a message carries — `identity_verified`, `payload_encrypted`, `message_verified`, `payload_verified`, `ws_remote_did` — are stamped by the ingress pipeline *after* verification. Wire-supplied values for these keys are stripped on ingress (and on egress from `msg_send` `message_meta`), so a peer cannot self-certify by setting `payload_encrypted: true` on a plaintext message. Likewise, `source: 'user'` marks the owner's own voice and is owner-only: `POST /agents/:id/trigger` rejects any event whose `data.message.source` is `user`, since that endpoint is reachable by any local process. **Residual:** the channel-adapter display name (`sender_alias`) is still spoofable for non-reserved names — only `owner`/`system`/`user` aliases are blocked. Closing that is out of scope for this pass; the verified `from` DID remains authoritative.

## Default Security Posture

Out of the box, ADF Studio ships with a conservative default configuration:

- **Autonomous mode:** off — agent requires human input to act
- **Messaging receive:** off — no inbound messages accepted
- **Mesh server:** binds to `127.0.0.1` only
- **Dangerous tools** (`adf_shell`, `ws_*`): disabled by default. `sys_fetch` is enabled but confined by the egress guard (loopback/link-local/private destinations blocked)
- **allow_unsigned:** true (appropriate for local development)
- **Identity encryption:** optional, activated by setting a password

## Operational Recommendations

### Local development
Leave defaults. No password needed. Unsigned messages are fine on localhost.

### Multi-agent local setup
Defaults still work. Consider marking tools with side effects as `restricted` if agents interact with untrusted data.

### Internet-facing agents
1. Set `security.allow_unsigned: false` **via the app UI / owner console** — it is a guard-system setting the agent cannot change through `sys_update_config` (every agent has identity keys since v24)
2. Configure allow/block lists for message senders (`messaging.allow_list` / `messaging.block_list` — agent-writable via `sys_update_config`, HIL-gated)
3. Enable `require_signature` and `require_payload_signature` (`security.require_signature` / `security.require_payload_signature` — agent-writable, HIL-gated)
4. Configure fetch middleware to restrict outbound URLs (`security.fetch_middleware.*` — owner-only guard path, like #1)
5. Review all enabled tools — disable anything not needed
6. Consider publishing the owner attestation so peers can verify ownership

### Sharing .adf files
Signing keys and credentials are envelope-sealed — a shared file cannot impersonate the agent or expose its API keys, and a recipient claims it under their own identity. To hand over credentials deliberately, use the share-password flow (see [Security and Identity](https://agentdocumentformat.org/guides/security-and-identity#sharing-an-agent)). The rest of the file is still readable, so before sharing:
- Review the loop history, `adf_files`, and configuration for sensitive content
- Consider cloning with selected tables to create a clean copy

### Reviewing untrusted .adf files
Before running an untrusted `.adf` file:
- Check enabled tools — disable `sys_fetch`, `shell`, and messaging tools
- Check triggers — system-scope lambdas can run without LLM intervention
- Check `code_execution` settings — disable `model_invoke` and `get_identity` if not needed
- Set `restricted: true` on any tools you want to monitor
