steps docs
start here Writing pipelines resourcesexprcontrol-flowagentsattempts-timeoutworkspaceinfratemplatingmcpcomplete Reference webagents-internalsaws-workersgcp-workersconformance

Agent Steps

How an agent step in a pipeline actually runs, and the features around custom tools: required tools, call budgets/pinned args, sub-agent delegation, and reusable tasks.

The execution flow

An agent step runs a tool-calling conversation loop:

  1. Parse the agent's config: model/endpoint, system prompt, granted tools, max_turns (default 30 for a hosted agent and none at all for a CLI one, 0 for no cap; a step may override it with its own max_turns: so one long-horizon step can buy more turns without every step of the same agent paying for them). Same for timeout: and attempts:, which an agents: entry may also carry — see attempts-timeout.md.
  2. Build a system message combining the agent's persona with working-directory context (any context_paths: files are delivered as synthetic read_file tool results — see below), plus a one-line disclosure of the step's resolved wall-clock deadline when one applies — see the timeout note under step 4.
  3. Loop, up to max_turns:
    • Send the conversation + tool definitions to the model.
    • If the model requests tools, execute them (read_file, list_dir, search_files, run_shell, write_file, edit_file, web_fetch, or a custom/sub-agent tool).
    • Cap any tool output at 32,000 bytes before it goes back to the model — output over that is saved to a file under the step's working directory instead of being dropped, with a short pointer message taking its place (see compaction). Two tools carry their own bound instead, both at 100,000 bytes and neither spilling: read_file (a spilled file exists precisely so the model can pull it back), degrading to plain truncation with start_line/end_line paging, and web_fetch, which cuts the body off and says so — the page is still on the web, so a narrower URL is the way to the rest.
    • Append the tool results and continue.
  4. Exit when the model stops requesting tools, max_turns is exceeded, or loop detection kills a stuck conversation.
    • A spent turn budget ends the conversation rather than destroying it: the runner makes one final request with the tools withheld, asking the model to answer from what it already gathered, and records the answer with wrapped_up: true so a degraded answer is tellable from a confident one.
    • If that final request itself fails — a 5xx, or a token ceiling breached by it — the step reports that failure, unmarked, so it classifies as errored and fires on_error.
    • The wall-clock timeout: (see attempts-timeout.md) is handled differently, and proactively rather than only at the end: the model is told its deadline once, up front, in the system message, and — once less than a fifth of that budget remains — gets one mid-conversation nudge to wrap up. Neither costs a turn or ends the conversation; tools stay granted throughout. Only the deadline itself expiring does that, by cancelling the request outright with no further message — there is no wrap-up request for a wall-clock timeout the way there is for a spent turn budget, since the model was already warned it was coming.
  5. Print the model's final response text to the terminal, followed by its verdict and note if the step declares verdicts:.
  6. Record the step's output.

One tool can be synthesized onto a step's grant beyond what tools: lists: a required verdict tool (verdicts: on the step) — documented in control-flow.md, since it exists to serve routing.

messages: — a step holds a conversation

A step's messages: is a list, one user turn per entry. The loop above runs to completion for the first one — tool calls and all — and only then is the next sent:

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash, api_key_env: OPENROUTER_API_KEY }
  system: You review code.

jobs:
- name: review
  assert:
    execution: [reviewer]
    outcome: succeeded
  plan:
  - agent: reviewer
    messages:
      - Review the diff and say whether it is safe to ship.
      - Name the file and line your answer turns on, and what would have to be different for the opposite answer.
    assert:
      stdout: parser.go

That sequencing is the whole difference from writing both asks in one message. Asked together, the model composes its answer to the second while it is still deciding the first, and can pick the evidence that fits the conclusion it was forming. Asked in turn, the second question is put to an answer that already exists and cannot be quietly revised.

Beware the obvious use. "Are you sure?" as a second message is one of the most reliable ways to make a model abandon a correct answer, and the flip has little to do with whether the first answer was right. A second message earns its place when it demands something checkable — a file, a line, a counterfactual — rather than asking for more confidence. If the question is really "is this verdict robust?", ensemble: answers it better, because independent members have no such channel.

Built-in tools

read_file, list_dir, and search_files are granted automatically whenever tools: is absent — the zero-config default is read-only:

agents:
- name: reader
  source: { model: openrouter/qwen/qwen3.7-flash }
  # no tools: -> read_file, list_dir, search_files

jobs:
- name: inspect
  plan:
  - task: fetch
    outputs: [notes]
    run: echo 'widgets ship on tuesday' > notes/plan.txt
  - agent: reader
    inputs: [notes]
    messages:
      - "What does notes/plan.txt say?"
    assert:
      tool_calls:                 # read_file really ran, with this path...
      - name: read_file
        args: { path: notes/plan.txt }
      stdout: widgets ship on tuesday   # ...and the answer could only come from what it returned
  assert:
    execution: [fetch, reader]
    outcome: succeeded

The built-ins that mutate state or reach beyond the workspace are deliberately not in the default; each is a capability the pipeline must grant explicitly:

toolwhat it does
run_shellRun a shell command in the working directory — unconfined within the step's host or container.
write_fileWrite (or with append: true, append) a UTF-8 text file. Replaces a whole file.
edit_fileReplace an exact string in an existing file — change part of a file without re-emitting it.
web_fetchHTTP GET a URL and return its body — optionally fenced to named hosts with allow:.
ask_userAsk the person running the pipeline a question and wait for the answer.
agents:
- name: writer
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools: [read_file, write_file, edit_file, run_shell]

jobs:
- name: draft
  plan:
  - agent: writer
    outputs: [report]            # an empty report/ dir exists from turn one
    max_turns: 40
    tools: [write_file]          # this STEP narrows the agent's grant to one tool
    messages:
      - "Write your findings into report/summary.md."
    assert:
      files: [report/summary.md]   # it WROTE the file, not just claimed to
  - task: publish
    inputs: [report]
    run: cat report/summary.md
    assert:
      stdout: all clear            # ...and the artifact reached the next step
  assert:
    execution: [writer, publish]
    outcome: succeeded

A step's tools: selects from what the agent already grants — it can narrow, never widen. Naming a tool the agent does not provide is a load error, so one careless step cannot hand a model a capability the pipeline never gave it; an absent step tools: means the agent's whole grant, unchanged.

edit_file

Takes path, old_string, new_string, and optional replace_all. old_string must match exactly once unless replace_all is set; zero matches and ambiguous matches are both returned as errors phrased as next-turn instructions, since both are recoverable without burning an attempt. The file's mode is preserved. Returns replacements, first_line, and match_mode, never content.

Matching is forgiving, in three strategies tried in order of decreasing exactness: exact first; then line-trimmed (every line matches modulo leading/trailing whitespace — recovers the classic local-model miss of right block, wrong indentation); then block-anchor (for a block of 3+ lines, the first and last lines anchor and the middle is judged by per-line similarity). The matched span is always the file's own text, so a forgiving match never rewrites untouched lines to the model's spelling. match_mode in the result says which strategy landed — an inexact edit is visible, not silent.

edit_file pairs with read_file by design: read_file returns raw bytes, so text copied out of one is a byte-exact old_string. Line numbers come from search_files' content mode instead — never from read_file.

search_files

Supply pattern (a regexp matched against each line), glob (a shell pattern matched against a file's path), or both; glob alone is a filename search. path defaults to ".". Three output_modes:

Unlike every other tool, search_files never spills: its bound is arithmetic — content matches accumulate against a 28KB budget, so a saturated result lands under the 32,000-byte inline cap by construction. head_limit caps results (default 50, ceiling 200); total and truncated report the true scale, so the answer to a flooded result is a narrower pattern, not a second page. .git, node_modules, vendor, binary files, and files over 2MB are skipped. ** is supported only as a leading glob segment (**/*.go).

write_file requires the file's immediate parent directory to already exist — use run_shell (mkdir -p) first if it doesn't. Like every file tool, its path is confined to the working directory and re-validated against symlink escapes.

web_fetch

Takes one argument, url (http:// or https:// only), and returns {status, content_type, body, truncated}. The body is capped at 100KB inline and cut off past it (truncated: true) — the overflow is not spilled, because the page is still on the web and a narrower URL fetches the rest.

A bare grant reaches any http(s) URL — the same trust level as run_shell, which can already curl anywhere. The mapping form's allow: is for the agent granted a browser but not a shell: each entry matches its exact hostname and any subdomain, case-insensitively, and every hop of a redirect chain is re-checked, so a permitted host that 302s elsewhere is refused mid-flight. The fence is enforced in steps, not in prompt language — a refused fetch comes back as {"error": ...} data naming the host and the list, which matters precisely because this tool exists to read pages the pipeline does not control:

agents:
- name: auditor
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools:
  - write_file
  - builtin: web_fetch
    allow: [specification.website]   # this host and its subdomains; empty/absent = any http(s) URL

jobs:
- name: audit
  plan:
  - agent: auditor
    outputs: [notes]
    messages:
      - "Check the tracker at issues.example and record what you find in notes/status.md."
    assert:
      tool_calls:
      - name: web_fetch
        args: { url: "https://issues.example/open" }   # outside allow: — refused as tool-result data
      files: [notes/status.md]
  - task: show
    inputs: [notes]
    run: cat notes/status.md
    assert:
      stdout: refused
  assert:
    execution: [auditor, show]
    outcome: succeeded

The refusal happens before any connection is attempted, so the example above runs without network — and the model, told about the fence in the tool's own description, records the refusal instead of retrying it.

Two rules keep a written fence honest, both enforced at load:

ask_user

An agent can read files, run shells, fetch pages and delegate to a sub-agent. Without this tool the one thing it cannot do is say it does not know something and get an answer — the only ways to handle a missing fact are to guess in prose, or to fail the step and make a person restart the run with the fact baked into messages:.

approval: is the closest thing that already exists, and it is deliberately not this. An approval gates an act — publish, deploy, send — and answers exactly one question, yes or no, at a place the pipeline author chose in advance. ask_user gates a fact, at a place only the model can know it needs, with an answer that is information rather than permission.

It takes a question and an optional list of options, and it is granted like run_shell or write_file — never by default, because interrupting a person is a capability the pipeline hands over explicitly:

agents:
- name: architect          # a bigger model, asked before a person is
  source: { model: openrouter/anthropic/claude-opus-4 }
- name: writer
  source: { model: openrouter/qwen/qwen3.7-flash }
  max_questions: 5         # every step of this agent, unless the step says otherwise
  tools:
  - write_file
  - builtin: ask_user
    answered_by: architect   # escalate before parking for a person
    timeout: 5m              # how long a person is waited on
    default: patch           # what the model is told when nobody answers
    options_required: true   # refuse an answer that is not one of the offered options

jobs:
- name: release-note
  plan:
  - agent: writer
    max_questions: 2         # this step may interrupt somebody twice
    outputs: [decision]
    messages:
      - "Draft the release note. Ask which version bump this is and whether to mention the storage migration, then write each answer on its own into decision/bump.txt and decision/migration.txt."
    assert:
      tool_calls:
      - name: ask_user
      files: [decision/bump.txt, decision/migration.txt]
  - task: tag
    inputs: [decision]
    run: cat decision/migration.txt
    assert:
      stdout: "yes"          # the second answer reached the next step — as a file
  assert:
    execution: [writer, tag]
    outcome: succeeded

An answer does not escape the conversation. It is a tool result and a transcript entry: it is not readable through context: { from: ... }, and a chosen option is not a routing key. A pipeline that needs the answer downstream has the agent write it to an output, exactly as above. This is the rule the verdict design already landed — decisions travel as verdicts, everything else travels as explicit files — and verdicts: is still the right tool when a decision should steer the plan.

Who answers, in order

A question goes to the first channel that can serve it:

  1. An answer given in advance. steps run --answer 'which bump=minor' (repeatable, also on test, watch and web) answers every question whose text contains that substring, case-insensitively. It is a supported way to run unattended, not only a test seam.
  2. answered_by: <agent> — a declared agents: entry answers, with its own model, dials and tool grant. This is deliberately not the same as granting that agent as a sub-agent tool: a sub-agent is delegation the asking model chose, while a responder is an escalation a person can still intercept. The recorded row says which one answered. A responder that fails, or answers with nothing, hands the question on rather than resolving it.
  3. A person, inline. When steps run has a terminal, the question is asked right there.
  4. A person, parked. With no terminal — CI, a supervised steps web — the run parks, prints the command that answers it, and waits:
question 1: Which bump is this release?
question 1: options: major | minor | patch
question 1: waiting up to 5m0s — steps questions answer 1 <answer> -p pipeline --db .steps/pipeline.yml.db

$ steps questions -p pipeline --db .steps/pipeline.yml.db    # another shell, same directory
ID  JOB           STEP    ASKED                          QUESTION
1   release-note  writer  2026-08-25T09:14:02.000000000Z Which bump is this release?
                                                         options: major | minor | patch

$ steps questions answer 1 minor -p pipeline --db .steps/pipeline.yml.db
answered: question 1

That is a local steps run, which keeps its state beside the YAML; under steps web the line reads -p <name> alone, because the read commands default to the daemon's .steps/steps.db.

The row is written before any of those are tried, so nothing about the audit trail depends on which channel answered — or on the run still being alive when one does. Every conversation inside a run can ask: a plan step, a sub-agent (on behalf of the same run), a task's fix: agent mid-repair, and a hook. Each question is filed under the name of the agent that asked it, not the step that invoked it. The web UI's questions page reads the same rows and writes the same answer; --read-only withholds that control the way it withholds approve/reject (see web.md).

Nobody answered

When the wait expires, what happens depends on whether the grant declared a default::

tool_result ask_user:
  answered: false
  answer:   "patch"
  source:   default
  note:     "nobody answered within 5m; the declared default was used"

The model is told. An indistinguishable default would be the runtime saying a person confirmed something no person saw — the same audit lie approval: refuses when it classifies an expiry as aborted rather than rejected.

With no default:, an expiry aborts the step, matching approval: exactly. Aborted, not failed: nobody decided anything.

Asked once per run, however many steps ask

Within one run, an answer is memoized by the question's text and the options offered. One mechanism covers three otherwise-separate problems:

The memo is per run, not per pipeline: a new run is a new set of circumstances, and yesterday's answer standing in silently for today's question is the failure this whole design exists to avoid.

Bounding, and caching

max_questions: caps ask_user calls per step (default 3), with an agent-level default and the usual step-wins override. The call past the budget comes back as ordinary tool-result data naming it, never an aborted attempt. max_turns: is not a substitute: a model that asks twenty questions burns somebody's afternoon rather than a token budget, and the failure mode is social.

Caching diverges from approval: deliberately. An approval is never cached — re-asking is the point. A question is a fact, so a step that hits the step cache does not run, cannot ask, and replays its recorded answer; nobody is asked twice. There is no "never cacheable" carve-out for a step that grants ask_user, because a steps web polling every 30 seconds and re-asking the same question forever is not a gate, it is a reason nobody uses the feature. volatile: true remains the opt-out for a question that genuinely must be re-asked every run.

On a CLI-backed agent

ask_user works unchanged on @claude/... sources, and it is the one builtin a CLI never runs itself: the call is bridged back to steps, so the answer lands here — where the row, the memo and the responder ladder live — rather than in the child's own transcript. Two consequences worth knowing: the call is recorded in the trajectory under its bridged name, mcp__steps__ask_user, and timeout: on this builtin IS accepted for a CLI agent, unlike every other builtin, because the deadline really does bind.

Working directory, inputs, and dir:

An agent step's dir: sets its working directory and names the artifact it operates in (its first path component — dir: repo/cmd names repo). That artifact must be one of the step's own declared inputs:, flow-validated like any input — an agent pointed at a directory nothing fetched ("summarize the repository" with no get) fails at plan time, before the model is ever called — or one of its own outputs:, which starts empty and is what the step produces. See workspace.md.

agents:
- name: reader
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: inspect
  plan:
  - task: fetch
    outputs: [repo]
    run: |
      mkdir -p repo/cmd
      echo 'package main' > repo/cmd/main.go
  - agent: reader
    inputs: [repo]
    dir: repo/cmd                 # start here; `repo` is the artifact it names
    messages:
      - "What package does main.go declare?"
    assert:
      tool_calls:
      - name: read_file
        args: { path: main.go }   # relative to dir:, not to the step's root
      stdout: package main
  assert:
    execution: [fetch, reader]
    outcome: succeeded

Every tool path the model uses is relative to dir:, which is the point: a model handed the subdirectory it is meant to work in does not spend turns navigating to it, and cannot wander out — the file tools stay confined to the step's workspace regardless.

Custom tools, required:, and call guards

A custom tool is a tools: entry with name/description/run — a templated shell command whose parameter schema is inferred from the {{ .args.* }} references in its run:. It can be marked required: true: the step can't complete until that tool has succeeded. It may also set max_calls: (a per-conversation budget), args: (pinned values the model never sees), and timeout: (a deadline for one call — see agents-internals.md):

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools:
  - read_file
  - name: post_review
    description: Post the review. body is the review text.
    run: |
      echo posting to {{ .args.repo | shellquote }} {{ .args.body | shellquote }}
    required: true         # the step fails unless this succeeds
    max_calls: 1           # at most once per conversation
    args:
      repo: jtarchie/ci    # pinned — the model neither sees nor can override it

jobs:
- name: review
  plan:
  - agent: reviewer
    messages:
      - "Review the change and post your conclusion."
    assert:
      tool_calls:
      - name: post_review
        args: { body: looks correct }   # only model-authored args are assertable —
                                        # naming the pinned `repo` here is a load error
  assert:
    execution: [reviewer]
    outcome: succeeded

Delivering files the pipeline will read

An agent's answer is not its final message. Everything downstream — a put:, a task, another agent — reads files, and a model that summarizes its work in prose instead of writing it has produced nothing while sounding finished. assert.files: already states which files a step owes (control-flow.md), and assert.tool_calls: states the procedure it must follow. Both are checked after the step ends, and the step fails naming what is missing. With nudge: true on the assert, the model is also told while it can still act on it:

agents:
- name: responder
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools: [write_file]

jobs:
- name: answer
  plan:
  - agent: responder
    outputs: [answer]
    messages:
      - "Answer the question. Write your answer to answer/reply.md."
    assert:
      files: [answer/reply.md]
      nudge: true
  - task: deliver
    inputs: [answer]
    run: cat answer/reply.md
    assert:
      stdout: widgets.json
  assert:
    execution: [responder, deliver]
    outcome: succeeded

The model in that example answers in prose first and writes nothing. Rather than ending the step there, the conversation puts it back:

You are trying to finish, but this step declared files it must leave behind and they are not there: answer/reply.md does not exist. Your final message is not the deliverable — a later step of this pipeline reads these files, and text you write in this conversation reaches nobody. Write them now using the tools you have, then finish.

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools:
  - name: run_tests
    description: Run the test suite and report the result.
    run: echo "42 passed"

jobs:
- name: review
  plan:
  - agent: reviewer
    messages:
      - "Review the change. Run the tests before you answer."
    assert:
      tool_calls:
      - name: run_tests
      nudge: true
  assert:
    execution: [reviewer]
    outcome: succeeded

Why it is opt-in. The nudge was never the protection: a model that answers in prose already fails the step, loudly, naming the file. What the nudge adds is automatic recovery, and whether recovery is part of a step's behavior is the author's call. It matters most for steps test against a real model: with an unconditional nudge, deleting "write it to answer/reply.md" from the prompt left the fixture green, because the nudge rescued the model every time — a test that cannot fail. Setting nudge: true declares that recovery is part of the step, so the fixture tests the step including recovery; leaving it out gives a strict fixture that tests the prompt. A step without the flag that fails an unmet files: or tool_calls: says so in its error, so the choice is discoverable from the failure itself.

What a nudge may say. It may name an unmet obligation — the deliverable (files:), the procedure (tool_calls:). It may never name a wanted conclusion: nudge: true beside stdout: or verdict: is a load error, because a classifier told which verdict is expected, or a step told which substring is being looked for, satisfies the assert without the assert having tested anything.

Sub-agent delegation (agent: tools)

A tools: entry can be a sub-agent tool — { agent: <name>, description: <text> } — exposing another agents: entry to the parent model as a callable tool taking a single request string. "Delegate and get an answer back":

agents:
- name: summarizer
  source: { model: openrouter/qwen/qwen3.7-flash }
  description: Condenses a file to a paragraph.   # what a PARENT sees by default
  tools: [read_file]
- name: lead
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools:
  - read_file
  - agent: summarizer            # a sub-agent, exposed as a callable tool
    description: Summarize a file; pass the path in `request`.

jobs:
- name: digest
  plan:
  - task: fetch
    outputs: [notes]
    run: echo 'a very long account of the outage' > notes/log.txt
  - agent: lead
    inputs: [notes]
    messages:
      - "Have your summarizer condense notes/log.txt, then report."
    assert:
      tool_calls:
      - name: summarizer         # the child was reached as an ordinary tool call
      stdout: there was an outage
  assert:
    execution: [fetch, lead]     # the child records nothing of its own
    outcome: succeeded

Self-healing tasks (fix:)

fix: attaches an agent to a task step: if the run misses the step's own success criteria, the agent is invoked to repair whatever broke, and then the command runs again. Those criteria are the step's assert: when it declares one and a non-zero exit otherwise — so a task asserting code: 3 is not repaired on exit 3, and one exiting 0 with the wrong output is. A passing run never constructs the agent at all. This example is real: the task fails, the fixer creates the missing file, the re-run passes:

agents:
- name: fixer
  source: { model: openrouter/qwen/qwen3.7-flash }
  system: You fix failing checks. Make the smallest change that works.
  tools: [read_file, search_files, edit_file, write_file, run_shell]

jobs:
- name: build
  plan:
  - task: check
    run: test -f config.json
    fix: fixer                 # scalar: just the agent name
  assert:
    execution: [fixer, check]  # the fixer ran and the re-run passed; a green first
    outcome: succeeded         # try would record [check] alone

The mapping form takes per-task overrides:

    fix:
      agent: fixer
      messages:
        - Only fix compile errors; never touch a test assertion.
      dir: repo
      tools: [read_file, edit_file]   # narrow the agent's grant for this task
      attempts: 2
      timeout: 10m

Its messages: is a list of user turns, exactly as an agent step's is: the fixer runs the first to completion before the second is sent. The captured failure output is appended to the first, since it is what the repair starts from.

How the loop terminates. The agent is seeded with the verdict that failed plus the command's captured output, and given the parent task itself as a zero-arg rerun tool (its run:, never its fix:, so a rerun cannot recurse). It can edit, rerun, and see the new output. When the conversation ends, steps runs the command one final time and judges that run — by the same rule that triggered the repair. There is no repeat-until-green loop: one agent conversation, then one verdict. A still-red command fails the step normally, firing on_failure, and its error names the fixer that could not rescue it.

What a fix agent needs. It must be able to edit, and the default tool grant deliberately excludes the write tools — grant them explicitly.

Restrictions. A fix agent may not grant sub-agents, may not use MCP tool grants, and may not set image: (it runs in the parent task's image). Its conversation records no cache node or job_run of its own.

Caching. A task with a fix: makes its chain uncacheable, since whether it succeeds may depend on what a model did. The run prints note: <step> makes this chain uncacheable (fix: agent) at the point that happens.

Top-level tasks: reuse

A top-level tasks: list (mirroring resources:/agents:) lets a run:/fix: pair be defined once and reused across jobs. A job's task: step is disambiguated by whether it carries its own run::

agents:
- name: fixer
  source: { model: openrouter/qwen/qwen3.7-flash }
  tools: [read_file, write_file, run_shell]

tasks:
- name: unit
  run: echo running unit tests
  fix: fixer                 # the run:/fix: PAIR, defined once, reused by both jobs

jobs:
- name: quick
  plan:
  - task: unit               # run: absent -> resolves the tasks: entry
    assert:
      stdout: running unit tests
  assert:
    execution: [unit]
    outcome: succeeded
- name: full
  plan:
  - task: unit
    assert:
      stdout: running unit tests    # the same tasks: entry, resolved again
  - task: integration        # run: present -> inline; tasks: never consulted
    run: echo running integration tests
    assert:
      stdout: running integration tests
  assert:
    execution: [unit, integration]
    outcome: succeeded

assert:
  execution: [quick, full]   # pipeline level: the jobs `steps test` must have run

Neither execution assert names fixer, and that is the assertion: a command that satisfies its assert never constructs its fix agent. Had either step assert: gone unmet, the fix: these steps inherit would have run and then that same assert would have judged the re-run — a reused definition brings its repair behavior with it. The repair path itself is exercised in attempts-timeout.md; what is reused here is the definition.

This resolution runs identically at plan time and run time, so a task's cache hash is always computed from its resolved run: string. An undefined reference is an ordinary error at plan time. An agent step's connection/dials/tool-grant resolve the same way.

External files: run_file:, system_file:, message_files:, and file:

A task's run:, an agent's system: persona, an agent step's messages:, and a fix:'s messages: can all be loaded from a file instead of written inline — useful since a persona is often long freeform prose, and a run: a full shell program:

tasks:
- name: unit
  run_file: ci/unit.sh              # loads the task's run: from a file

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }
  system_file: prompts/reviewer.md  # loads the agent's persona from a file

jobs:
- name: build
  plan:
  - task: unit
    assert:
      stdout: unit tests pass       # ci/unit.sh's text, resolved at load
  - task: smoke
    run_file: ci/smoke.sh           # a plan step's OWN run:, from a file
    assert:
      stdout: smoke ok
  - agent: reviewer
    message_files: [prompts/review.md]  # loads the step's message from a file
    assert:
      stdout: Build looks fine
  assert:
    execution: [unit, smoke, reviewer]
    outcome: succeeded

Every *_file: path is resolved once, at load time, relative to the pipeline YAML's own directory — so everything downstream sees the resolved text and cannot tell it apart from the same value written inline. Editing an included file busts the cache exactly like editing an inline value would.

A path may use .. to escape the pipeline's own directory: the pipeline file is trusted input, and a file placed beside it by the same author is at the same trust level — a shared ../tasks/ directory is a legitimate layout. Setting both a field and its *_file: sibling is a load-time error, and so is an empty included file.

A top-level tasks:/agents: entry additionally accepts a whole-document file:, loading a complete Task/Agent definition from a separate YAML file so it can be shared across pipelines:

tasks:
- name: unit
  file: ci/unit.yml         # supplies any task field but name:
  timeout: 5m               # any field set here overrides the document's

agents:
- name: reviewer
  file: ci/reviewer.yml     # the same for a whole agent definition
  max_turns: 10             # ...and the same override rule

jobs:
- name: build
  plan:
  - task: unit
    assert:
      stdout: from the shared task   # the run: came out of ci/unit.yml
  - agent: reviewer
    messages:
      - "Review the build."
    assert:
      stdout: Nothing to flag
  assert:
    execution: [unit, reviewer]
    outcome: succeeded

The entry's own inline fields win over the loaded document's — except privileged:, which is on if either says so, since an inline false cannot be told apart from leaving it out — and the loaded document may not itself use file:/run_file: — includes are resolved one level deep only, which is what makes cycle detection unnecessary.

The run-time form: an agent step's message_files: from a fetched artifact

An agent step's message_files: additionally accepts a {artifact, path} mapping, naming a file inside an artifact a get step fetched, read at run time rather than load time:

resource_types:
- name: repos
  config:
    check: |
      printf '[{"ref": "abc123"}]'
    in: |
      mkdir -p .ci
      echo 'Review this change for correctness.' > .ci/REVIEW.md

resources:
- name: repo
  type: repos
  source: {}

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: review
  plan:
  - get: repo
    trigger: true
  - agent: reviewer
    inputs: [repo]
    message_files: [{ artifact: repo, path: .ci/REVIEW.md }]
    assert:
      stdout: The change is correct
  assert:
    execution: [repo, reviewer]
    outcome: succeeded

This is the one place a step's config can come from a fetched artifact, and it is deliberately narrow — a task's run: and a whole agent definition cannot, for two reasons:

The artifact named must be declared in the step's own inputs: (checked at load), and the read is confined with the same symlink-aware guard the file tools use. It cannot be resolved at load time — the file doesn't exist until the get runs — which costs nothing, since an agent step's chain is already unskippable.

context_paths: — files delivered as synthetic read_file results

An agent step can declare context_paths: — files whose contents are injected at conversation start as synthetic read_file tool results. The model sees the file contents as if it had called read_file itself, without consuming a turn:

agents:
- name: coder
  source: { model: openrouter/qwen/qwen3.7-flash }
  max_context_bytes: 100000    # what every step of this agent gets by default

jobs:
- name: build
  plan:
  - task: conventions
    outputs: [repo]
    run: echo 'always run go vet before committing' > repo/CONVENTIONS.md
  - agent: coder
    inputs: [repo]
    context_paths: [repo/CONVENTIONS.md]
    max_context_bytes: 400000  # ...except this step, which is handed more
    messages:
      - "State this project's convention in one line."
    assert:
      stdout: go vet           # answered from the injected file, no read_file turn spent
  assert:
    execution: [conventions, coder]
    outcome: succeeded

The point is not convenience but guarantee: conventions every invocation must follow are present from the first turn, instead of costing a read_file round trip the model might not bother with.

Paths are relative to the step's working directory and confined to its workspace, so in practice the file lives inside a declared input. They are read at run time (per attempt), which is what distinguishes them from system_file:: the persona is the pipeline author's own text, resolved once at load; context_paths: is content that arrives with a fetched artifact and can change between runs. A missing or escaping file fails the step at preparation, before a token is spent. A file that is merely too big (over max_context_bytes:, default 100KB — 0 lifts the ceiling entirely) is truncated instead, with a note pointing at read_file's paging — the author writes a path, not a size, and pr/pr.diff is a correct path that would otherwise start failing the day the pull request grew.

context_paths: is a step-level field, not agent-level — the agent definition has no notion of which inputs are available. It requires read_file in the tool grant (which it is by default). Sub-agents and fix agents do not inherit the parent step's context_paths:. max_context_bytes: is spelled on either, and the step's wins (as above) — two steps sharing one agent routinely hand it different evidence. context_window: deliberately has no step spelling for the mirror-image reason: it describes the model, and the model belongs to the agent.

In an across: matrix, each entry renders {{ .vars.<name> }} per cell, so a cell arrives already holding the code it was assigned instead of spending its first turns navigating to it:

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: review
  plan:
  - task: fetch
    outputs: [repo]
    run: |
      echo 'package api'     > repo/api.go
      echo 'package storage' > repo/storage.go
  - across:
    - var: dim
      values: [api, storage]
    agent: reviewer
    inputs: [repo]
    context_paths: ["repo/{{ .vars.dim }}.go"]
    messages:
      - "Review the {{ .vars.dim }} package."
  assert:
    execution:                 # one cell per value, each under its own coordinates
    - fetch
    - reviewer [dim=api]
    - reviewer [dim=storage]
    outcome: succeeded

One path per entry, rendered per cell. A {{ .vars.x }} naming an axis the matrix does not declare is a load error naming the entry (context_paths[0]).

Caching: the paths (not contents) enter the step's hashed content — the files live inside the workspace, so their content is already chained through the input artifacts' own hashes. A matrix cell hashes the path it rendered to, which is what makes two cells reviewing different files two different steps.

Reading another step's decision (context: { from: ... })

A verdict is the one thing every judging step produces. A classifier that simply falls through, or a shell command that wants to branch on what a model decided, needs a way to ask for it. from: is that ask, and it is declared on the reader:

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }
- name: editor
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: revise
  plan:
  - agent: reviewer
    messages:
      - "Review the change."
    verdicts: [approve, revise]      # no routing — this one just decides
    assert:
      verdict: approve               # what it decided, pinned at the source
  - agent: editor
    messages:
      - "Apply the review."
    context:
      from:
        reviewer: note               # verdict | note | full
  - task: gate
    context:
      from:
        reviewer: verdict
    run: grep -q 'approve' upstream/reviewer
    assert:
      code: 0                        # the verdict file really landed at upstream/reviewer
  assert:
    execution: [reviewer, editor, gate]
    outcome: succeeded

Model dials, and pipeline-wide defaults:

An agents: entry carries the sampling dials for its model, and defaults: supplies what every agent that names nothing gets:

defaults:
  model: openrouter/qwen/qwen3.7-flash   # any agent whose source: names no model
  delegate_budget_percent: 10            # the share a sub-agent takes of what is left
  preflight:
    timeout: 30s                         # per pre-run probe
    cache: 5m                            # a target verified this recently is trusted (max 15m)

agents:
- name: drafter
  source: { model: openrouter/qwen/qwen3.7-flash }
  temperature: 0.2            # dials, all optional — unset means the provider's own
  top_p: 0.9
  max_tokens: 2048
  reasoning_effort: low       # low | medium | high, for models that take one
  delegate_budget_percent: 25 # this agent's helpers get a bigger share
  preflight: false            # ...and this one skips the probe (a slow local model)
- name: titler                # no source: at all — defaults.model is its model

jobs:
- name: draft
  plan:
  - agent: drafter
    messages:
      - "Draft the release note."
    assert:
      stdout: Drafted
  - agent: titler
    messages:
      - "Title the release note."
    assert:
      stdout: Titled
  assert:
    execution: [drafter, titler]
    outcome: succeeded

Built-in agent profiles: @builtin/<name>

Four agents ship with steps — explorer, planner, reviewer, builder — each a persona, a tool grant, and a set of dials chosen for one role. A profile supplies everything about an agent except the one thing only you know, which is which model to call:

agents:
- name: "@builtin/reviewer"                          # quotes required -- YAML reserves a leading @
  source: { model: openrouter/qwen/qwen3.7-flash }   # the one thing the profile cannot supply

jobs:
- name: gate
  plan:
  - agent: "@builtin/reviewer"
    messages:
      - "Review the change for correctness."
    assert:
      stdout: no correctness problems
  assert:
    execution: ["@builtin/reviewer"]
    outcome: succeeded

An entry for a built-in name supplies what it sets and inherits the rest. It is not all-or-nothing: naming @builtin/reviewer to give it a model keeps the persona and tool grant that were the reason to reference it. Anything you do state wins, exactly as it does for a file: include — so tools: [read_file] on the entry above narrows the grant without touching the persona.

The dials are the point, and they come as a set. A turn count on its own is not an opinion about a role: 50 turns against a deadline that cannot fit them is a step that reliably dies on time rather than on turns, and a large context ceiling with too few turns is a step that reads a lot and cannot act on it. Each profile therefore states the dials its own turn count makes wrong, and inherits the defaults where they are already right:

profilemax_turnsmax_context_bytestimeoutthe role
explorer15default10mfind a thing and answer in one message; still going at ten minutes means stuck
planner25400,000defaultread a codebase or a change entire, then write a plan about it
reviewer30400,000defaulttrace control flow, callers and error paths across several files
builder50400,0001hrun shell, read, edit, re-run — the one role 30 minutes genuinely does not fit

max_context_bytes: is raised on the three reading-heavy roles because the 100,000-byte default is the wrong ceiling for a step whose whole job is to see something entire — measured on this repo's own review pipeline, not guessed. Where a profile states nothing, defaults: and then the built-in default apply as usual.

Changing a profile's dials re-runs work. max_turns: and max_context_bytes: fold into a step's hash the way every model dial does, so an agent referencing a built-in whose context ceiling changed is no longer the same step and will not be skipped. timeout: never hashes, so that one is free.

Budgets: budget.tokens

An agent step can loop, hold a long conversation where every turn re-sends the whole history, and retry. budget: is the ceiling on that — the AI equivalent of timeout::

agents:
- name: writer
  source: { model: openrouter/qwen/qwen3.7-flash }
  budget:
    tokens: 200000      # per invocation of this agent

jobs:
- name: publish
  budget:
    tokens: 500000      # cumulative, across every agent step in the job
  plan:
  - agent: writer
    messages:
      - "Write the release announcement."
    assert:
      stdout: Announcement written
  assert:
    execution: [writer]
    outcome: succeeded

An agent's ceiling covers the agent and everything it delegates to. A sub-agent draws on its parent's remaining allowance rather than adding to it, so budget.tokens bounds the whole delegation subtree instead of one conversation in it — otherwise a capped agent could delegate its way past its own ceiling without ever exceeding it.

defaults:
  delegate_budget_percent: 10   # the default; every agent unless it says otherwise

agents:
- name: lead
  budget:
    tokens: 400000              # bounds `lead` AND every helper it calls
  delegate_budget_percent: 25   # this one's helpers do the heavy lifting
  tools: [{ agent: researcher }]

A run that resumes continues its job budget from what earlier attempts already spent, rather than starting the allowance over — otherwise budget: would be a per-attempt ceiling wearing the name of a per-run one, and every resume would buy another full one.

Reporting happens whether or not you set one, which is the point: it is what tells you which ceilings are even sensible. Every job that ran an agent step prints what it cost — and records it, so the question survives the terminal:

$ steps runs cost -p pipeline
RUN                 TOKENS   CACHED        COST   STEPS
r-8f2a1c         4,102,338      38%    unpriced       9

$ steps runs cost -p pipeline r-8f2a1c
STEP                                TOKENS   CACHED   DURATION  FINISH
reviewer [dim=state-mutation]      412,880      61%       1m02s  stop
reviewer [dim=api]               1,204,551      22%      14m30s  length  <-- truncated

Things worth knowing:

Failover: fallback:

When an agent's model is unreachable, try a backup instead of retrying a dead connection:

agents:
- name: writer
  source:
    model: openrouter/qwen/qwen3.7-flash
  fallback:
  - source:
      endpoint: https://backup-provider.example.com/v1/
      model: equivalent-model
      api_key_env: BACKUP_KEY

jobs:
- name: publish
  plan:
  - agent: writer
    messages:
      - "Write the announcement."
    assert:
      stdout: Announcement written   # the primary answered, so no source changed
  assert:
    execution: [writer]
    outcome: succeeded

fallback: fires two ways, automatically — declaring it is what opts an agent into both, there's no separate switch for the second:

agents:
- name: writer
  source:
    model: openrouter/qwen/qwen3.7-flash
  fallback:
  - source:
      endpoint: https://backup-provider.example.com/v1/
      model: equivalent-model
      api_key_env: BACKUP_KEY

jobs:
- name: publish
  plan:
  - agent: writer
    attempts: 1   # no room to retry — the very first failure trips the cascade
    messages:
      - "Write the announcement."
    assert:
      stdout: Announcement written via the fallback   # this time the fallback actually served the run
  assert:
    execution: [writer]
    outcome: succeeded

CLI-backed agents: @claude/sonnet

An agent's source.model normally names a hosted model steps calls over HTTP. Prefix it with @ instead and steps runs a coding-agent CLI as a subprocess:

agents:
- name: reviewer
  source:
    model: "@claude/sonnet"     # quotes required -- YAML reserves a leading @
  tools: [read_file, run_shell]
  settings: project             # opt in to the repo's checked-in .claude/ scope
  reasoning_effort: high        # the one generation dial a CLI takes
  budget:
    usd: 0.50                   # CLI agents meter in dollars, not tokens

jobs:
- name: review
  plan:
  - task: fetch
    outputs: [repo]
    run: echo 'package main' > repo/main.go
  - agent: reviewer
    inputs: [repo]
    dir: repo
    messages:
      - "Review this code."

The quotes are not stylistic: a leading @ is a reserved indicator in YAML, so an unquoted value is a parse error before steps ever sees it. @claude/sonnet reads as "the claude CLI, asked for sonnet" — the part after the slash is passed through untouched.

What changes, and what doesn't

This is delegation, not a different transport. The CLI owns the conversation: its own turn loop, its own tools, its own context window. steps owns everything around it, unchanged — the workspace, the merkle hash that decides whether the step runs at all, timeout:, the recorded trajectory and response, assert:, and verdicts:/to: routing. steps also reads the CLI's transcript as it streams, so the turns it takes are published and stored exactly as a hosted agent's are — a delegated step is not a quieter one. You get the CLI's own tooling inside a pipeline that still caches, routes, and fans out.

Authentication comes from the CLI's own credential store — the subprocess inherits HOME, so a subscription login works with no api_key_env: at all. Set api_key_env: only to forward a specific key as ANTHROPIC_API_KEY.

The tool grant becomes the CLI's permissions

The CLI is passed --tools "" unconditionally — no built-in of its own, ever — and every granted tool, built-in or custom alike, reaches it over the same loopback MCP bridge that already served custom run: tools, mcp: grants, and the synthesized verdict tool. A hosted agent is a brain whose hands are steps' own tool implementations, executing wherever steps decides; a CLI agent is now the same shape — a brain that happens to live in a subprocess. The bridged tools are the same implementations a hosted agent runs: path confinement, output caps, allow: enforcement, all apply unchanged, including for a tool (read_file, run_shell, ...) that used to run as the CLI's own native Read/Bash.

Anything not granted is absent, not merely unapproved: --tools "" gives the CLI no built-in surface to fall back on, and --allowedTools names only what the bridge exports. That is deny-by-default — a capability this build of steps has never heard of is withheld because it was never granted, rather than surviving because nobody remembered to forbid it. The CLI's own configured MCP servers are excluded too (--strict-mcp-config).

Attestation: the fence is enforced, not merely asked for. --tools "" is policy inside an upstream binary steps does not pin, so each attempt checks the CLI's own stream-json init event — which lists the session's tools — against exactly the bridged grant, and kills the child the instant they disagree. This is detection, not prevention: the check runs after init is parsed, so it cannot stop a single surplus call made in the same breath as init itself, but every call after that is refused. A mismatch is an infrastructure condition (it fires on_error, not a step failure to: can route on) — retrying would just re-trigger the same fence.

A bridged call's trajectory entry is recorded de-namespaced: mcp__steps__read_file records as read_file, identical to what a hosted step's own trajectory shows for the same call, so one assert.tool_calls: [{name: read_file}] reads the same on either agent kind. A name this build does not recognize as bridged (Bash, Task, anything the CLI's own natives could in principle still report despite --tools "") is kept verbatim — a second, human-readable signal of the same fence failure the attestation check exists to catch.

A step is not your session

A CLI agent step runs with no configuration scopes by default. Your personal ~/.claude never applies — no user settings, hooks, plugins, skills — and the repo's own .claude/ scope loads only when the agent opts in with settings: project (as above). A pipeline whose behavior depends on who ran it is not a pipeline. The opt-in is hashed, so granting or revoking it invalidates the step's cache. It is also markedly cheaper: dropping user-level config cut a trivial one-step pipeline from ~76K prompt tokens to ~25K in a measured run.

--setting-sources, --strict-mcp-config and --tools "" are now the entire fence a settings: project step runs inside — three flags, all load-bearing. That matters most for settings: project together with image:: the CLI process itself is always host-side (see "Containerizing a CLI agent" below), so a fetched PR's own .claude/settings.json hooks execute on the orchestrator, not inside whatever container the step's tools run in. image: narrows what the tools can reach; it does not, and cannot, contain a settings: project agent's own config loading. Grant settings: only to a repo whose .claude/ you trust the same way you trust its run: scripts.

Verdicts are enforced at exit

A hosted agent that tries to finish without its required verdict gets forced into one more call via tool_choice. There is no such lever across a process boundary, so the rule moves to the exit: a step that declared verdicts: and finished without calling the verdict tool has failed. The failure is routable, so a failure: entry catches it. The verdict itself is captured in the parent process the moment the tool is called, over the bridge — the CLI is never trusted to report what it decided.

attempts: resumes the conversation

On the hosted path attempts: retries one HTTP request underneath a conversation that survives. A CLI agent gets the same guarantee by a different mechanism: the step names a session up front, and every retry rejoins it rather than starting the task over. The retried process is told what went wrong and to continue. Only infrastructure failures are retried — the process failed to start, exited nonzero, or died without reporting a result. A CLI that ran fine and concluded the task failed is an answer, not an outage.

What a CLI agent cannot do

These are load errors, not silent no-ops, because a setting that reads as configured while binding nothing is worse than one that is rejected:

rejectedwhy
source.endpoint:there is no request to aim anywhere
temperature:, top_p:, max_tokens:the CLI chooses its own sampling and its own output limits
source.string_tool_choice:no tool_choice on the wire to spell
compact_after_tokens:, context_window:the CLI compacts its own conversation
budget.tokens:nothing counts tokens until the subprocess exits (use budget.usd:)
sub-agent tools, in either directiona sub-agent nests inside a turn loop there is none of
a CLI agent as a task's fix: agentsame reason

required:, max_calls: and args: on a tool are accepted, same as a hosted agent's: every call now reaches the bridge (or, for required:, is checked at exit against what the bridge observed — see "Verdicts are enforced at exit" above), so there is no longer a turn loop these would promise a constraint nothing applies. timeout: on a tool is accepted regardless of whether it names a built-in or a custom/MCP tool, for the same reason: every call now reaches the bridge, so there is no longer a native path a per-call deadline would silently miss. network: none together with image: is accepted too — see below.

Containerizing a CLI agent

image: places the step's tools, never the CLI process itself — the CLI is always a host subprocess of this one, exactly as a hosted agent's conversation always runs here while its tools may run in a container. That single move is what makes containerizing a CLI agent unremarkable: credentials never approach the container (the subscription login / api_key_env: stays in this process, forwarded to the CLI's own host-side environment exactly as an uncontainerized step's would), there is no macOS-vs-Linux asymmetry to reason about, and network: none is now a coherent, useful way to sandbox a CLI agent's shell commands — cutting the tools' egress no longer touches the bridge the verdict comes back on, because the bridge was never inside the container to begin with. See infra.md.

Budgets are in dollars

A CLI agent takes budget: {usd: 0.50} rather than budget: {tokens:}. The two runners meter different things and neither converts into the other honestly — each takes the unit it can enforce, and the other spelling is a load error. A job-level budget: stays in tokens and still counts what a CLI agent spent (reported on exit, folded into the job total).

The ceiling bounds the STEP, not the attempt. Each attempt is handed what is left of usd: rather than the whole figure, so an agent with attempts: 3 and budget: {usd: 0.50} spends at most about fifty cents in total, not a dollar fifty. When the remainder reaches zero the attempts stop and the step fails naming the spend — the same thing an exhausted max_turns: does, and for the same reason: paying a provider to discover there is nothing left to spend is the outcome the ceiling exists to prevent.

About: a subprocess that CRASHES never reports what it spent — the dollar figure rides the terminal event a dead child never emits — so its share is priced from the token usage it streamed before dying, using a rate card covering the models the CLI runtimes run. A model outside that card is debited nothing, which is how every model behaved before the ceiling carried across attempts at all. The estimate is only ever used to decide how much budget is left; recorded cost stays exactly what the provider reported, so a run whose cost nobody reported still records none.

fallback: works in both directions — a CLI agent can fall back to a hosted provider, and a hosted agent to a CLI. Preflight checks a CLI target by looking for its binary on PATH — the same check whether or not the step names an image:, since the CLI is always a host subprocess.

Ensembles: asking several agents the same question

A single model has blind spots. Ask one reviewer "is this correct?" and you get one opinion with no signal about how much to trust it; ask three and require a majority, and one model's bad day stops being decisive:

agents:
- name: reviewer-a
  source: { model: openrouter/qwen/qwen3.7-flash }
- name: reviewer-b
  source: { model: openrouter/qwen/qwen3.7-flash }
- name: reviewer-c
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: gate
  plan:
  - ensemble:
      verdicts:                       # the vocabulary EVERY member votes in,
        - reject: revise              # and where the BLOCK's decision goes
        - approve: publish
      decide: majority                # or: unanimous, any, or an agent name
      member_errors: fail             # or: exclude
      agents:
      - {agent: reviewer-a, messages: ["Review the diff for correctness."]}
      - {agent: reviewer-b, messages: ["Review the diff for style."]}
      - {agent: reviewer-c, messages: ["Review the diff for security."]}
  - task: revise
    run: echo sending back
  - task: publish
    run: echo shipping
    assert:
      stdout: shipping
  assert:
    execution:                        # every member voted, then the majority's
    - reviewer-a                      # target ran — revise is absent, which is
    - reviewer-b                      # what "routed past it" looks like
    - reviewer-c
    - publish
    outcome: succeeded

⚠️ N agents cost N times one

Three reviewers cost three reviews, every run. This is the step where a job-level budget: earns its keep.

The decision rules

Two things that are never silent

The rest

What's not on this page

The mechanics underneath an agent step — malformed tool-call repair, loop detection, OpenRouter prompt caching, and conversation compaction — are in agents-internals.md. Reach for it when behavior surprises you; you don't need it to write a pipeline.