# Desktop Applications with Isolated Compute

> Run visible Linux desktop applications in isolated compute, transfer files, and validate visual results without assuming generic desktop automation

Source: https://github.com/christianbalevski/adf/blob/v0.7.4/docs/knowledge/desktop-apps.md (adf v0.7.4)

ADF can run real Linux GUI processes in an agent's managed **isolated** container, on the agent’s own Linux desktop that the principal sees in the Computer tab. This article is the end-to-end recipe for a native application: inputs in, launch, interact, validate, results out. The general driving skills (screenshots, `xdotool`, opening apps and files) are in the [Computer Use](https://agentdocumentformat.org/guides/computer-use) guide.

This article composes a task recipe; it does not replace the feature contracts in the [Compute Environments guide](https://agentdocumentformat.org/guides/compute), [Computer guide](https://agentdocumentformat.org/guides/browser), or [Documents and Files guide](https://agentdocumentformat.org/guides/documents-and-files). Follow those guides for canonical configuration, browser lifecycle, file-protection, and security procedures. An installable executable workflow belongs in a skill, not in this article.

## Capability boundary

### Source-verified capabilities

- `compute_exec` runs real shell commands in the selected compute environment. An isolated agent container uses `/workspace` as its working directory.
- When isolated compute has browser display support enabled (`compute.enabled: true` and `compute.browser` not `false`), managed agent-container processes receive `DISPLAY=:99`. The display is shared by processes in that **same isolated container**: a GUI process can render there while the user watches the container's VNC/noVNC view.
- ADF starts a display stack in that container (X server, the Openbox window manager, the tint2 panel with an application menu, PCManFM for the wallpaper and files, VNC, and a loopback-published noVNC endpoint). The Studio Computer tab embeds that noVNC page and shows the X display; it is not a second desktop session. Every window gets a close button and a panel task button, so the human can close or switch windows.
- The managed Chromium opens on demand (`adf-browser start`, an attaching browser MCP server, or the panel button) and stays closed once closed. Use `adf-browser stop` / `adf-browser resume` when a task needs it held down (for example a profile swap).
- The image ships the Mousepad text editor, LXTerminal, and PCManFM file manager, plus `xdotool`, `scrot`, `xclip`, and `xterm`. Through `compute_exec` an agent can move and click the mouse, type, send key chords, focus windows by title, screenshot the desktop to a file, and read or write the clipboard.
- `fs_transfer` moves files between the agent VFS and `isolated` or `shared` managed containers. Paths are relative to the selected endpoint and must not escape it. External containers are not file-transfer endpoints in this release.
- The maintained `@playwright/mcp` integration attaches to ADF's managed Chromium through its container-loopback CDP endpoint. This is the supported agent automation mechanism for web pages.

The `shared` target is a different, multi-agent container. Its workspace is namespaced under `/workspace/{agentId}`. The source exposes the visible `DISPLAY=:99` capability for browser-enabled **isolated** containers; do not infer that a shared-container command has a GUI display. Use isolated compute when a task needs a dedicated visible desktop process.

### Desktop control is CLI-level, not a structured tool

ADF does not expose a dedicated pointer/keyboard tool. Desktop control goes through `compute_exec` with `xdotool` and `scrot`: screenshot, look at the image, act, screenshot again (worked examples in [Computer Use](https://agentdocumentformat.org/guides/computer-use)). This is coordinate-based and works for any X application, but it is slower and less reliable than an application's own CLI/API or the Playwright MCP server for web content; prefer those where they exist. Accessibility control, audio, printing, and GPU acceleration are not verified; do not promise them.

## Prerequisites

1. The agent has `compute.enabled: true` and `compute.browser` left enabled when a visible display is required. These are agent configuration changes and may require human approval.
2. Podman and the managed compute service are available. If the selected target is unavailable, execution fails closed; ADF does not silently move the command to another container or the host.
3. The agent has access to `compute_exec` and, when files must cross the boundary, `fs_transfer`. Tool enablement and restricted-command approvals still apply.
4. Check whether the application is installed. The standard desktop includes Mousepad, LXTerminal, PCManFM and Chromium. Additional applications can be installed with `apt-get update && apt-get install -y <package>` under the applicable compute policy and approvals; do not assume PDF editors, image editors or office suites are already installed.

Do not enable host access just to obtain a GUI. Host execution is a separate, high-trust target with the user's OS privileges; see [Compute Environments](https://agentdocumentformat.org/guides/compute#security-considerations).

## Workflow

### 1. Move inputs into the container

The VFS and container filesystem are separate. Transfer the input before launching the app:

```text
fs_transfer({
  from: "vfs",
  to: "isolated",
  path: "inputs/source.pdf",
  save_as: "inputs/source.pdf"
})
```

The destination is `/workspace/inputs/source.pdf` in the isolated container. For an output, reverse the endpoints:

```text
fs_transfer({
  from: "isolated",
  to: "vfs",
  path: "outputs/result.png",
  save_as: "outputs/result.png"
})
```

`from` and `to` must differ. `path` and `save_as` are relative paths; absolute paths and `..` escapes are rejected. Directory transfers are supported. Check the transfer result before relying on the file.

### 2. Check the application, then launch it on the shared display

Use `compute_exec` with the isolated target when the tool schema exposes a `target` field. If isolated is the sole authorized environment, the schema omits that field; omit `target` and ensure the configured default is isolated. Verify the binary first rather than assuming a package is present:

```text
compute_exec({
  command: "command -v my-viewer",
  target: "isolated"
})
```

For a short-lived operation, launch an available application with an explicit display, workspace path, and retained log. Apply the same target-field rule to this call:

```text
compute_exec({
  command: "mkdir -p /workspace/logs && DISPLAY=:99 my-viewer /workspace/inputs/source.pdf >/workspace/logs/my-viewer.log 2>&1",
  target: "isolated",
  timeout_ms: 120000
})
```

`compute_exec` runs `sh -c` through a one-shot container exec. Appending `&` does not make a GUI process persistent after the tool call; the exec session takes its children with it. Launch it in its own session instead: `setsid -f my-viewer /workspace/inputs/source.pdf </dev/null >/workspace/logs/my-viewer.log 2>&1` (verified: this is how `adf-browser start` launches the managed Chromium from a one-shot exec). Use a stable app-specific output or CLI/API where possible. Keep logs and generated artifacts under `/workspace` when they need to be transferred; `/tmp` is suitable for diagnostics only.

Do not copy Chromium's container-specific `--no-sandbox` setting to every GUI application. The managed browser and root-owned smoke tests have special launch constraints; `--no-sandbox` weakens a sandbox and is not a universal recommendation. Likewise, a GPU-disabled launch may be a useful experiment-specific workaround, but the repository evidence does not establish a general GPU failure cause or a universal flag set.

### 3. Watch or interact

- A human can use the Computer tab when the managed display is available. It is a noVNC view of the container's X display, so other X clients are visible there as well as Chromium, each with its own window and taskbar entry.
- For web content, configure the maintained Playwright MCP server. It attaches to the already-running managed Chromium, preserving the same tabs, cookies, and visible session.
- For a native non-browser app, use its CLI/API, drive it with `xdotool` after inspecting a `scrot` screenshot, or ask the human to operate it in the Computer tab. A process being alive does not prove that the correct window is visible or that an operation completed.

When a browser site requests sign-in, CAPTCHA, MFA, passkey, or another security review, stop automation and ask the human to take over the visible browser. Do not bypass the challenge or ask automation to handle the human's browser credentials; the human completes the sign-in or security review in the visible browser. Follow the relevant identity or channel guide for other credential setup.

### 4. Validate the actual result

Validate at three levels:

1. **Artifact:** confirm the expected file exists and has a plausible size/type.
2. **Application:** inspect the app's output, log, or API result and check its exit status.
3. **Visible evidence:** run `mkdir -p /workspace/shots && scrot -o /workspace/shots/result.png` through `compute_exec`, transfer the image to the VFS, and inspect the actual image. Assert the expected title/content and select the intended X client or browser target when more than one full-screen window is present.

Do not claim success from a PID, window-list entry, process metadata, or a successful launch command alone. A screenshot of an old `about:blank` browser tab is not evidence that another application rendered. Keep screenshots as evidence only after inspecting them; do not leave transient PIDs or experiment-specific absolute paths in reusable instructions.

## Lifecycle and persistence

Managed isolated containers are created and started for an agent when isolated compute is enabled. On an agent stop or unload, the agent's compute lifecycle resource calls `stopIsolated`: the dedicated container is **stopped, not removed**, so its filesystem—including installed packages and files under `/workspace`—can be reused on a later start. On orderly application shutdown, runtime teardown disposes the foreground and background agents and then stops remaining managed containers. An explicit rebuild or destroy action removes that state. Application processes and display daemons do not survive a container stop; they must be launched again.

An unclean host or container failure does not guarantee cleanup or durable persistence. The browser profile is stored in the container at `/var/lib/adf/browser-profile`, not automatically inside the `.adf` file; browser persistence therefore follows container lifetime unless the profile-portability procedure is used. Treat cookies and saved passwords in that profile as sensitive.

This persistence is container-local, not a guarantee of durable backup, cross-machine portability, or a complete desktop session. Transfer important inputs and outputs to the VFS or another explicitly managed destination.

## Security and isolation limits

- **Isolated is the least-privileged GUI choice, not a security proof.** It gives an agent-dedicated managed container and workspace rather than host filesystem access, but it is still a runtime/container boundary. Review the image, packages, network, and Podman configuration for your threat model.
- Managed containers use bridge networking. A GUI process can make network requests if its program does so; visible does not mean offline.
- New managed containers are network-isolated from other agents' containers. All managed containers (isolated and shared) share Podman's default bridge—required for reliable outbound internet on the macOS Podman machine—but each one runs an in-container nftables rule set, added at startup via a `NET_ADMIN` capability, that drops new inbound connections arriving from sibling containers on the bridge. Outbound internet, container loopback, established/return traffic, and the host-loopback noVNC port-forward still work. Because each container self-protects on its own inbound path, one agent cannot reach another agent's container, and a misbehaving agent can at most re-expose *itself*, never a peer. This is peer isolation on a shared bridge, not separate networks and not protection against the host. It applies only to containers **created after** this change; a container created earlier keeps the old behavior—reachable by sibling containers—until it is rebuilt, and a rebuild erases its workspace files, installed packages, and browser profile, so it is the user's choice, not automatic.
- The shared container is not agent-isolated: agents use separate workspace directories but can see the shared container's filesystem according to its permissions, and processes can contend for resources. Use isolated for sensitive or interfering desktop work.
- Managed containers share an npm/npx cache volume. Do not treat every container artifact or cache as a private secret store.
- Host execution is not a safer fallback. It requires both the agent's `compute.host_access` and the runtime's owner-controlled host-access setting, and then runs with the user's OS privileges.
- The noVNC endpoint is published on host loopback, while browser CDP is container-loopback. Do not expose either endpoint or assume that a loopback URL is a public sharing mechanism.

## Evidence and open questions

On 2026-09-05, at baseline repository revision `1347204d2f938c6dcf14db7102ad5dcb0c3b2266`, a Linux x86_64 isolated test environment with outer `DISPLAY=:99` ran ADF Studio in development mode alongside the existing persistent Chromium session. Display inspection showed both visible windows; an explicitly selected Studio Electron target reported title `ADF Studio`, `readyState` `complete`, and non-empty rendered body text, and the resulting screenshot was inspected. Podman was not installed in that outer test environment, so this bounded observation verifies the outer Electron/display path—not managed agent-container GUI compatibility or arbitrary native applications.

That historical observation predates the current per-agent desktop and is not the limit of today’s supported computer-use workflow. The current runtime provisions native desktop applications and CLI mouse/keyboard control as described above and in [Computer Use](https://agentdocumentformat.org/guides/computer-use). Application-specific behavior still needs result verification; a delivered input event is not proof that an app completed an action. Audio and a print system are not installed; GPU acceleration remains unverified.

Documentation reviewed against runtime revision `39115d6c3221149cf6238e02774b91fc888beac2` (v0.7.2), including the desktop package list and launch configuration in `src/main/services/podman.service.ts`. This update is a source review, not a new application compatibility test.
