On this page

The headless performance harness measures runtime behavior without Studio. It spins up mock-provider agents, drives synthetic turns, records latency and resource metrics, and tears the agents down.

The harness is useful for daemon work because it exercises the canonical assembler and dispatch boundary without requiring real providers or manual UI interaction. It selects the explicit benchmark profile, which keeps synchronous teardown compatibility and disables timer polling so the measurements do not include idle timer overhead.

Entry Points

Run through npm:

npm run stress:runtime

Run the script directly:

node scripts/stress-runtime.mjs --scenario=smoke

Run the benchmark test directly:

RUN_BENCH=1 NODE_ENV=production npx vitest run tests/perf/stress-harness.test.ts --reporter=verbose

The script runs scripts/rebuild-for-node.mjs first so native SQLite bindings match Node.

Scenarios

scripts/stress-runtime.mjs supports:

ScenarioPurpose
smokeShort one-agent overhead run
overheadZero-latency mock provider to measure runtime overhead
idleMany low-traffic agents
mixedModerate load with mock tool-call probability
burstShort aggressive burst traffic
allRuns all script-supported scenarios

Example:

node scripts/stress-runtime.mjs --scenario=mixed --out=/tmp/adf-mixed.json

Output

The harness prints compact reports:

[mixed] agents=20 dur=15000ms turns=10/10 (err=0)
  turn p50/p95/p99/max: 1200/1300/1300/1300 ms
  evloop lag p50/p95/p99/max: 10/12/14/20 ms
  rss peak/avg: 180.5/170.2 MB  heap peak/avg: 64.1/60.8 MB
  timer handles peak: 24  provider calls: 10

When --out is provided, reports are also written as JSON:

{
  "reports": []
}

Metrics

The harness records:

MetricDescription
Turn countsTotal, completed, and errored turns
Turn latencyp50, p95, p99, and max successful turn duration
Event loop lagp50, p95, p99, max, and mean event loop delay
MemoryPeak and average RSS and heap used
HandlesPeak active timeout handle count
Provider callsTotal mock provider calls across all agents

How It Works

The harness uses:

  • RuntimeService with review disabled
  • createHeadlessAgent
  • MockLLMProvider
  • LoadDriver for scheduled synthetic chat dispatches
  • MetricsCollector for turn timing, memory sampling, and event loop delay

It creates multiple agents in memory through the lightweight headless compatibility API with the benchmark profile, wraps each target’s dispatch function for metrics, drives dispatch objects for the scenario duration, stops the load driver, collects metrics, and unloads all agents. The harness does not call executeTurn() directly.

When to Use It

Use the harness when changing:

  • RuntimeService
  • AgentExecutor
  • Headless agent creation
  • Trigger dispatch behavior
  • Loop persistence or session restore
  • Runtime scheduling
  • Provider call flow
  • Daemon code paths that affect many agents

For narrow API changes, normal unit tests are usually enough. For lifecycle, concurrency, memory, or latency questions, run the relevant harness scenario.

Test Gating

tests/perf/stress-harness.test.ts is gated by RUN_BENCH=1, so normal npm test does not run the benchmark scenarios. This keeps regular tests fast while preserving a repeatable benchmark path.