agentStack

These pages describe v0.19.0, the current release. That is the build the installer gives you. agentstack --version says which build you have.

Governed workflows

A workflow is one script that fans a task out to many agent runs and composes their results — the map/reduce shape, but every worker is a governed agent run instead of a bare process. One review chore becomes: run a reader over each file, synthesize the findings, then run an independent verifier to refute the weak ones. Claude Code has an ungoverned version of this today; AgentStack's version pins the orchestration code and gives every step its own reviewed authority.

For the shortest practical path, start with Run a multi-agent workflow. The complete loop is:

bash
$ agentstack workflow declare ... --preview
$ agentstack workflow declare ... --write
$ agentstack trust .
$ agentstack workflow explain <name>
$ agentstack workflow run <name>
$ agentstack workflow report <run-id>

Declare, pin, and trust a workflow, run it end to end with agentstack workflow run, render its evidence tree with agentstack workflow report, and resume an interrupted run with --resume (replay from the recorded journal — byte-identical script and args, or it refuses). Every agent step runs as a governed protected run. The interpreter boundary was independently security-reviewed on 2026-07-23, and all six findings that review raised are now closed, each with its own witness. agentstack workflow is therefore listed one hop away. It is grouped out of the everyday agentstack --help list and named under Run on agentstack more. It runs at both agentstack more workflow and agentstack workflow.

That change is discoverability, not enforcement. Not one boundary moved when the command became listed. What the review settled is the posture, and a host-tier step is still cooperative-guard only, exactly as Honest limits below describes.

Why a workflow needs governing#

One workflow command spawns N agent runs, each with tool access, filesystem reach, and token spend. The control flow is decided at runtime by script code. That is exactly the thing a security tool should not run on trust. So AgentStack treats the orchestration script the same way it treats any other executable content from a repo — as untrusted input until you review and pin it.

The security model#

Honest limits#

What it can and cannot do:

It canIt cannot
Prove which pinned script ran and what authority each step hadMake a prompt-injected step escalate — roles are a closed, pre-reviewed set and ceilings are frozen
Fence a step's network reach under --lockdownContain every tool in every posture — a host-tier step is cooperative-guard only

Step outputs are model output — untrusted data. One step's result can flow into a later step's prompt by design, so a prompt-injected step can mislead its successors; it cannot widen any grant. The built-in validation step (an independent verifier under a narrower role) is the mitigation, and the report labels each step's posture rather than implying uniform containment.

What the interpreter bounds, and what it does not#

The orchestration script runs under host-set ceilings, and two of them are partial. A partial bound must not read as total, so both residuals are stated here and in the posture block agentstack workflow report prints verbatim:

Both residuals are defects in reviewed content rather than hostile-input paths. A ceiling does not contain them; the out-of-thread watchdog does. It force-exits the process at the wall ceiling plus grace, behind an armed no-I/O exit so a blocked write cannot keep a runaway alive. Removing them is what the recorded QuickJS-in-wasmtime fallback is for.

The wall-clock budget is not an engine ceiling. max_wall_seconds is inert inside the interpreter, because the engine is clock-free; the number is surfaced to the script through budget only. Enforcement lives in two places outside it, both in the CLI:

The clock is therefore live-run state, never replayable state, because a --resume must not spuriously time out replaying a run that originally used its full wall clock. A resumed run restarts at the full effective ceiling. The effective ceiling itself is the narrowest of the machine cap, the manifest request and the script's own meta request.

Writing one#

A workflow is one JavaScript file with a small, familiar API — the same agent() / parallel() / pipeline() vocabulary as Claude Code, with one change: agent() takes a role, not a model, because the harness and model are properties of the role's toolset, not something a script may choose.

js
export const meta = {
  name: 'nightly-review',
  description: "Review the day's diff, then verify the findings",
  roles: ['reader', 'writer', 'verifier'],
}

const FINDINGS = {
  type: 'object',
  required: ['findings'],
  properties: {
    findings: {
      type: 'array',
      items: { type: 'object', required: ['file', 'summary'] },
    },
  },
}

// map: one reader per file, each returning validated JSON rather than prose
const mapped = await pipeline(
  files,
  f => agent(`List issues in ${f}.`, { role: 'reader', schema: FINDINGS }),
)
const findings = mapped.filter(Boolean).flatMap(m => m.findings)

// reduce: one synthesis per key group, so no single prompt has to hold everything
const claims = await parallel(
  partition(findings, 4, f => f.file).map(group => () =>
    agent(`Synthesize and rank:\n${JSON.stringify(group)}`, { role: 'writer' })),
)

// verify: an independent refuter under a narrower role
const checked = await parallel(
  claims.filter(Boolean).map(c => () => agent(`Try to refute: ${c}`, { role: 'verifier' })),
)
return keepUnrefuted(claims, checked)

The script runs inside a sandboxed interpreter with no filesystem, network, or environment access — the only thing it can do is request governed agent runs through agent(). Everything else is plain computation.

What a role brings: harness, model, effort#

A role is a toolset, and a toolset may declare model and effort alongside its harness:

toml
[toolsets.reader]
harness = "codex"
model = "gpt-5.5"
effort = "high"

Both are delivered to the child as launch flags on its argv — never by writing the harness's own settings file, which a governed run must never touch. Which flags (or whether any exist) is the adapter's answer, not AgentStack's: each adapter descriptor declares what its CLI calls the setting and how to select it for one headless launch.

That means a harness can be unable to carry a value, and the run says so rather than dropping it silently. Two different facts, kept apart:

Either way you get a warning line per child naming the role, the harness, the dimension and the reason, and the run proceeds on that harness's own default — an undeliverable model is a capability gap, not a manifest error. A value the adapter's own catalog rejects (an effort outside its enum) is a manifest error, and a run refuses that child before launch. agentstack workflow explain <name> reports the same facts statically, spawning nothing.

Named algorithms#

Five helpers spell out the shapes scripts kept re-deriving by hand. All five are pure compositions of parallel / pipeline / shard / partition, and not one of them calls agent() — an agent run happens only when your callback calls it, through the same bridge a hand-written script uses. So a helper can never widen a role or manufacture fan-out. The role ∈ meta.roles check and the max_agents ceiling remain the only authority path.

HelperShape
mapReduce(items, { map, reduce, partitions })map every item, drop the failures, shuffle survivors into partitions buckets, reduce each
reduceByKey(items, r, keyFn, reduceFn)group by key first, so one reducer sees all of a key's items
combine(items, per, combineFn)Hadoop's combiner — pre-summarize chunks so the reduce prompt holds 20 things, not 200
verify(claims, refute)run a refuter per claim, returning { claim, verdict } rows paired by claim
keepUnrefuted(claims, verdicts, isRefuted?)pure filter, no await, no agent

They follow the house rules the existing helpers do: never throw (a failed callback becomes null), deterministic, and total under junk arguments (clamped, not thrown). An empty bucket spends no agent, and the result array still carries one slot per bucket, so your reducer count never varies with the data.

keepUnrefuted is the reason this set exists at all: the worked example above has always called it, and until now it was never defined in the prelude — copying the documented example raised a ReferenceError.

keepUnrefuted's default predicate is a text heuristic, not a trust boundary. It greps the stringified verdict for "refuted". A refuter that phrases its finding differently ("this claim is false") reads as unrefuted, a claim whose own text contains the word reads as refuted, and a prompt-injected refuter can say whatever it likes. The schema section says the same: shape, not content. Pass your own isRefuted — ideally over a schema-validated verdict field — whenever the answer matters. A null verdict (the refuter died) is not treated as refuted, because failing closed there would silently delete claims whenever a child run failed.

Getting structured results back#

Pass a schema and the promise resolves with a parsed value instead of text, so a later stage can index it rather than parse prose. A result that does not satisfy the schema fails that step closed — the script sees null and decides. There is no automatic re-ask, because a retry would spend an agent slot your ceiling never granted.

Validation constrains shape, not content. A schema-validated result is still model output, and a prompt-injected step can return perfectly schema-valid lies. It is a parsing convenience, not a trust boundary.

Splitting and grouping#

shard(items, { per }) and partition(items, r, keyFn) are plain computation over values you already have — no agent, no tokens. partition returns exactly r buckets (empty ones included, so your reducer count never varies with the data) and always places the same key in the same bucket, which is what lets one reducer see all of a file's findings. It is not a balanced split. Skewed keys make skewed buckets, and the fix is a better key.

Keeping wide runs in memory#

A run that fans out over large outputs can ask for agent(prompt, { result: 'handle' }), resolving with { digest, bytes, preview } instead of the full text. There is also a machine ceiling on total result bytes; a run that exceeds it fails closed and tells you to use handles. Handles cost about 620 bytes each, so they are for stages returning kilobytes. On short results they are simply pointless.

Before you run it#

agentstack workflow explain <name> reports the effective ceilings, which roles launch serially, and how many agent() call sites the pinned script has — statically, spawning nothing. Sites are not calls. One site inside a loop runs once per item, so real fan-out is data-dependent. The enforced bound on total spawns is max_agents, refused per call.

Status#

Everything above is the guide. This last section is project status — how mature the capability is and what its security review found, kept separate so the two are not read as one.

The full technical contract and security rationale live in the workflows capability design doc. The manifest kind, pinning, trust review, the engine, workflow run / workflow report, negotiated ceilings, and journal-replay resume all ship.

The six review findings are closed, each with a focused witness named for it: the watchdog's no-I/O exit path, interpreter memory bounds, host-native re-entrancy, a run-total budget for native calls, cross-host resume determinism, and the crate boundary. Two of those witnesses are constructions rather than assertions — the re-entrancy witness is the reproduction itself (a getter that re-enters agent() during its own argument conversion, which spawns a second child without the guard), and the crate-boundary witness fails if any other crate takes the interpreter dependency.

Maturity, stated separately from the security gate. Closing the findings is what made the command visible; it is not a claim that workflows are well-worn. Repeated-use evidence — running real workflows on separate occasions and confirming each is easier to repeat than hand-rolled orchestration — stands at 1 of 3 occasions: the 2026-07-23 acceptance run. That is a maturity signal for you to weigh, not a gate anything is waiting on. Expect the rough edges of a young capability. The vocabulary (admission, ceilings, locked child runs) is the densest in the product, and Honest limits above is the section to read before you rely on it.

Scaling work — how the drive loop behaves at width, and what it costs — is tracked separately in the workflow scaling plan, which also records two things it did not build: automatic retry and straggler speculation (nothing in the current model can prove a role is side-effect free, and a claim the enforcement cannot back does not ship), and distributed workers (the measured bottleneck is the latency tail, not a shortage of machines).

Source of truth: docs/workflows.md — this page is generated from it.