steps docs
start here Writing pipelines resourcesexprcontrol-flowagentsattempts-timeoutworkspaceinfratemplatingmcpcomplete Reference webagents-internalsaws-workersgcp-workersconformance

Control Flow

Four distinct, easy-to-conflate mechanisms for shaping how a job's plan executes, plus the self-verification (assert:) that makes fixtures out of them. All are opt-in; a pipeline that uses none of this hashes and behaves exactly as if the feature didn't exist.

On a failing step with both a hook and a route: the hook fires first (react), then the route reroutes (route).

Hooks (on_success/on_failure/on_error/on_abort/ensure)

Any plan step, or the whole job, can carry hooks. A hook is itself a full step (task/put/agent — never get), so it can run: a command, put: a resource, or invoke an agent:, and may recursively carry its own hooks.

jobs:
- name: build
  plan:
  - task: work
    run: echo building
    on_success:                # step-level hook
      task: record
      run: echo built cleanly
    ensure:
      task: cleanup            # runs whatever happened above
      run: echo tidying up
  on_success:                  # job-level hooks (inline alongside plan:)
    task: announce
    run: echo the whole job passed
  on_failure:                  # exactly one on_* fires, chosen by classification
    task: page
    run: echo a step said no
  on_error:
    task: alert
    run: echo the infrastructure said no
  on_abort:
    task: note-abort
    run: echo someone pressed ctrl-c
  ensure:                      # and this one runs whatever happened
    task: sweep
    run: echo tidying the job
  assert:
    execution: [work, record, cleanup, announce, sweep]   # step, its hooks, then
    outcome: succeeded                                    # the job's — page, alert
                                                          # and note-abort are absent

Conditional steps (when:)

A task/put/agent step (including a hook step) can carry when: — a shell command whose exit code decides whether the step runs: 0 runs it, nonzero skips it.

jobs:
- name: guarded
  plan:
  - task: scout
    outputs: [risk]
    run: echo low > risk/level.txt
  - task: deep-review                       # skipped: the guard exits 1
    inputs: [risk]
    when: grep -q high risk/level.txt       # scalar shorthand
    run: echo auditing everything
  - task: report                            # runs: the guard exits 0
    inputs: [risk]
    when: { run: test -s risk/level.txt }   # mapping form
    run: cat risk/level.txt
    assert:
      stdout: low
  - task: escalate                          # skipped, so alert/ is never written
    inputs: [risk]                          # the guard reads it, so it declares it
    outputs: [alert]
    when: grep -q high risk/level.txt
    run: echo paging > alert/page.txt
  - task: page                              # skipped too: an unwritten input reads
    inputs: [alert]                         # as absent to the guard, not as an error
    when: test -s alert/page.txt
    run: echo paged the on-call
  assert:
    execution: [scout, report]              # deep-review is absent — that IS the skip
    outcome: succeeded

Tolerated failure (try:)

A try: wraps a single inner step (task, put, agent, or another try) so a failure or infrastructure error of that step doesn't stop the plan:

jobs:
- name: ship
  plan:
  - task: build
    run: echo ok
  - try:
      task: notify           # fails — and the plan shrugs
      run: "false"           # (assert: is rejected on both halves of a try:)
  - task: after
    run: echo still running
    assert:
      stdout: still running  # the plan really did carry on past the failure
  assert:
    execution: [build, notify, after]   # notify ran and failed; the wrapper hid it
    outcome: succeeded

The wrapper is transparent: the only thing it changes is whether the plan walker stops. Everything that observes the outcome sees the truth.

The timeout case deserves its own fixture — an expired timeout: classifies as failed (attempts-timeout.md), and the shrug covers it exactly as Concourse's does:

jobs:
- name: shrug
  plan:
  - try:
      task: slow
      run: sleep 5
      timeout: 1s          # expires — a failure, per Concourse — and try: eats it
  - task: after
    run: echo carried on
    assert:
      stdout: carried on
  assert:
    execution: [slow, after]
    outcome: succeeded     # a timed-out step inside try: is not a red job

Step transitions (to:/max_visits:/verdicts:)

A step routes to another step in the same get-segment based on its outcome, including jumping backward to form a bounded loop. Two spellings, one per kind of outcome: a task/put/verdict-less agent carries to:, a binary success/failure map; an agent carries verdicts:, an ordered list that declares its vocabulary and its targets together.

A task loop, modelless and complete — bump advances state (and always succeeds, so its write is captured), check decides and routes backward until it's satisfied:

jobs:
- name: converge
  plan:
  - task: seed
    outputs: [state]
    run: touch state/visits.txt
  - task: bump
    inputs: [state]
    outputs: [state]                # read-modify-write: each visit continues the last
    run: echo visit >> state/visits.txt
  - task: check
    inputs: [state]
    run: test "$(wc -l < state/visits.txt)" -ge 3
    to: { failure: bump }           # backward — so max_visits is required
    max_visits: 5
  - task: done
    run: echo converged
    assert:
      stdout: converged
  assert:
    execution: [seed, bump, check, bump, check, bump, check, done]   # three laps
    outcome: succeeded

And verdict routing — the agent's decision picks the next step:

agents:
- name: critic
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: review
  plan:
  - task: draft
    outputs: [notes]
    run: echo 'first draft' > notes/draft.txt
  - agent: critic
    inputs: [notes]
    messages:
      - "Read notes/draft.txt. Approve it, or send it back."
    verdicts:
      - approve: publish       # route: record the verdict and jump forward
      - revise: draft          # backward — the loop max_visits: bounds
      - failure: escalate      # reserved: the step errored or decided nothing
    max_visits: 3
    assert:
      verdict: approve
  - task: escalate
    run: echo paging a human
  - task: publish
    run: echo publishing
    assert:
      stdout: publishing
  assert:
    execution: [draft, critic, publish]   # escalate is absent: the route jumped past it
    outcome: succeeded
jobs:
- name: fall-through
  plan:
  - task: flaky
    run: "false"             # no assert: here — one would decide success itself,
    to: { failure: next }    # and the failure is what the route consumes
  - task: report
    run: echo carrying on
    assert:
      stdout: carrying on
  assert:
    execution: [flaky, report]
    outcome: succeeded       # the route consumed the failure, so the job is green

The word is reserved on the value side: a step actually named next inside a to:-using segment is a load error naming the collision.

Consuming a verdict without routing on it. A verdict is also readable downstream: a later step declaring context: { from: { <step>: verdict|note|full } } is handed what that step decided — as a synthetic read_step result for an agent, as a file for a task. It needs no route between the two steps, which is what makes a bare-verdict classifier's decision usable. See agents.md.

Assert (self-verification) + steps test

assert: lets a pipeline verify its own behavior — the mechanism that turns a hooks/control-flow fixture into a runnable regression test. This whole fixture deliberately fails and stays green, because the assertions say the failure is the point:

jobs:
- name: failing-fixture
  plan:
  - task: boom
    run: "false"
    on_failure:
      task: alarm
      run: echo the failure was observed
  assert:
    execution: [boom, alarm]    # these ran, in this order — and that clears the failure
    outcome: failed             # and the plan really did fail
jobs:
- name: expected-exit
  plan:
  - task: probe
    run: |
      echo checking
      exit 3
    assert:
      stdout: checking
      code: 3
  assert:
    execution: [probe]
    outcome: succeeded       # the matching assert is what makes exit 3 a success
agents:
- name: triage
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: label
  plan:
  - agent: triage
    messages:
      - "Classify this report: the app crashes on launch."
    verdicts: [bug, feature, question]    # all bare: record the choice, route nowhere
    assert:
      verdict: bug
  assert:
    execution: [triage]
    outcome: succeeded

Naming a verdict outside the declared list, or setting it on a step with no verdicts:, is a load error: the assert could never match on any run.

jobs:
- name: wrote-something
  plan:
  - task: draft
    run: |
      mkdir -p answer
      echo "the answer is 42" > answer/reply.md
    outputs: [answer]
    assert:
      files: [answer/reply.md]
  assert:
    execution: [draft]
    outcome: succeeded

A missing file or an empty one fails the assert. Two shapes are load errors instead, because neither could ever pass: a path naming no declared output, and a bare artifact name like answer — that is the output directory, and a directory is never a non-empty file. Name a file inside it.

do: — several steps as one

A plan is already sequential, so do: is not about ordering. It is about containment: the block is a single plan step, so one hook on it observes the whole group's outcome.

jobs:
- name: ship
  plan:
  - do:
    - task: migrate
      run: echo migrating
    - task: deploy
      run: echo deploying
    - task: smoke-test
      run: echo smoke passed
    on_failure:
      task: rollback           # fires if ANY of the three failed
      run: echo rolling back
  assert:
    execution: [migrate, deploy, smoke-test]   # the children, in order; the block
    outcome: succeeded                         # itself records nothing

Without it that rollback has two spellings and both are worse: repeat the hook on all three steps and keep them in sync, or hoist it to the job, where it also fires for failures that have nothing to do with the group.

In a green run that rollback never fires, so the claims it exists for get their own deliberately-failing fixture — the deploy breaks, and the assertions pin both halves of the contract:

jobs:
- name: ship-broken
  plan:
  - do:
    - task: migrate
      run: echo migrating
    - task: deploy
      run: "false"
    - task: smoke-test
      run: echo smoke passed
    on_failure:
      task: rollback
      run: echo rolling back
  assert:
    execution: [migrate, deploy, rollback]   # smoke-test absent: the failure stopped the
    outcome: failed                          # block; rollback present: the block's hook saw it

in_parallel: — several steps at once

A plan is otherwise strictly sequential, so independent work waits on itself: three downloads run one at a time, and one slow resource check stalls everything behind it.

jobs:
- name: verify
  plan:
  - in_parallel:
      limit: 2          # max in flight; omit for unbounded
      fail_fast: true   # cancel the siblings on the first failure
      steps:
      - task: lint
        run: echo linting
      - task: test
        run: echo testing
      - task: vet
        run: echo vetting
  assert:
    execution: [lint, test, vet]   # declaration order, not completion order
    outcome: succeeded             # outcome: is what catches a SWALLOWED branch

To regression-test a parallel block, use assert.outcome — assert.execution alone structurally cannot catch a swallowed branch failure (both builds run the same branches, so the assert matches either way and then clears the difference under test). This is that fixture, the exact shape of the original bug:

jobs:
- name: verify-broken
  plan:
  - in_parallel:
      fail_fast: false   # let the siblings finish — the failure must still count
      steps:
      - task: lint
        run: echo linting
      - task: test
        run: "false"
      - task: vet
        run: echo vetting
  assert:
    execution: [lint, test, vet]   # every branch ran to completion
    outcome: failed                # and the one red branch still failed the job

race: — first success wins

Run a fast/cheap path and a slow/reliable path at the same time, keep whichever finishes successfully first, cancel the other:

jobs:
- name: summarize
  plan:
  - race:
      steps:
      - task: fast
        outputs: [summary]
        run: echo quick take > summary/text.txt
      - task: slow
        outputs: [summary]
        run: sleep 2 && echo thorough take > summary/text.txt
  - task: read
    inputs: [summary]              # the winner's summary/, whoever wrote it
    run: cat summary/text.txt
    assert:
      stdout: quick take           # `fast` won, so its artifact is the block's
  assert:
    execution: [fast, read]        # the cancelled loser records nothing
    outcome: succeeded

⚠️ This costs more, not less

Running both branches always costs both, every time, even when the fast one wins. The value is "never wait for the slow path", not "spend less". If you reached for race: to save money, you want caching or plain retries instead.

⚠️ It is unsafe for side-effecting steps

Cancelling a loser stops only future work. A branch that already filed an issue, sent a notification, or pushed a commit keeps that side effect. race: is safe for read/generate-only steps; for anything that changes the outside world, run one branch.

Cancellation is also bounded, not instantaneous: killing sh -c "sleep 5; …" kills the shell but not the sleep it forked, so a cancelled step is given a short grace period before it is abandoned.

The rest of the rules:

across: — one step, once per combination

Run the same step for every combination of some values, instead of writing out a near-identical step per cell and keeping them in sync by hand:

jobs:
- name: matrix
  plan:
  - across:
    - var: go_version
      values: ["1.25", "1.26"]
    - var: package
      values: [agent, pipeline]
    task: matrix-test
    run: echo testing {{ .vars.package }} on go {{ .vars.go_version }}
  assert:
    execution:                     # the product, last axis varying fastest
    - matrix-test [go_version=1.25 package=agent]
    - matrix-test [go_version=1.25 package=pipeline]
    - matrix-test [go_version=1.26 package=agent]
    - matrix-test [go_version=1.26 package=pipeline]
    outcome: succeeded

across: is a modifier, not a container: the step it sits on is still a task (or a put, or an agent), it just runs once per cell. {{ .vars.<name> }} substitutes into the command, the image, the prompt, the working directory, the step's own name, and each entry of an agent step's context_paths: — so a fan-out cell can be handed the file it was assigned rather than told to go find it.

The headline: per-cell caching

Concourse re-runs the entire matrix on any change. Here each cell is hashed and cached individually, so changing one value in one axis re-runs only the cells that value appears in.

That works because cells are siblings, parented on the step before the block rather than on the block itself — the block's own hash folds in every cell, so parenting cells under it would make one cell's edit change every cell's identity, which is exactly the whole-matrix re-run this exists to avoid.

Cells that are puts or agents are never skipped, for the same reasons those steps never are anywhere else: side effects and non-determinism.

The rest

A width a step decides: from_file:

An axis can take its values from a JSON array an earlier step wrote, so the matrix is as wide as that step said — "review each of these findings", where nobody knew the findings when the pipeline was authored:

jobs:
- name: fan-out
  plan:
  - task: scan
    outputs: [findings]
    run: printf '["alpha","beta"]' > findings/items.json
  - across:
    - var: item
      from_file: findings/items.json    # instead of values:
    task: investigate
    inputs: [findings]
    run: echo investigating {{ .vars.item }}
  assert:
    execution:                     # width decided by what scan wrote, not by the file
    - scan
    - investigate [item=alpha]
    - investigate [item=beta]
    outcome: succeeded

This is "the step plans, the pipeline executes": one step produces a work list, and each item becomes its own cell — independently hashed, cached and reported — instead of one agent grinding through the whole list in a conversation that outgrows its window.

Nothing carries the list but an ordinary artifact. The producer declares outputs: as it would for any other file; the axis names a path inside that artifact. The first path component is the artifact, exactly as it is for dir:, and it must be fetched or produced earlier in the plan — a load error otherwise.

Collected outputs: outputs: on the matrix

A matrix's cells are clones of one step, so they declare the same output names — which used to mean the last capture silently erased the others'. outputs: on an across: step means the block collects: each cell writes its declared output exactly as any step does, and the capture lands under the cell's own coordinates. Downstream consumes the lot as ONE artifact:

jobs:
- name: review-matrix
  plan:
  - across:
    - var: dim
      values: [api, errors]
    outputs: [findings]                  # the block collects
    task: review
    run: echo {{ .vars.dim }} looks fine > findings/report.txt
  - task: merge
    inputs: [findings]                   # findings/api/report.txt, findings/errors/report.txt
    run: test "$(ls findings | wc -l)" -eq 2 && cat findings/*/report.txt | tr '\n' ' '
    assert:
      stdout: api looks fine errors looks fine   # exactly two cells, both contents intact
  assert:
    execution:
    - review [dim=api]
    - review [dim=errors]
    - merge
    outcome: succeeded

The cell's own view is unchanged — it writes findings/report.txt and never sees the coordinates. The directory layout is one segment per axis value, declaration order, so it is derived entirely from declared things.

A ceiling that degrades: budget:

A wide matrix of agent cells is one whose total cost is easy to underestimate — especially when a step decided the width mid-run. budget: on the block caps what its cells spend together:

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: bounded-review
  plan:
  - across:
    - var: dim
      values: [api-boundaries, error-paths, concurrency]
    budget:
      tokens: 1000            # cells stop being admitted once this is spent
    agent: reviewer
    messages:
      - "Review the {{ .vars.dim }} dimension."
  - task: publish
    run: echo publishing what we got
    assert:
      stdout: publishing what we got
  assert:
    execution:                     # cell one overspent the allowance, so cells two
    - reviewer [dim=api-boundaries]   # and three were never admitted — and the plan
    - publish                         # carried on anyway, which is the whole point
    outcome: succeeded
budget: across stopped after 1 of 3 cells (spent 2,000 of 1,000 tokens)

Making it bind at any width: reserve_per_cell:

Admission can only see what finished cells reported, and a cell only finishes once something waits for its slot. So on its own the ceiling is blind to the cells it is deciding against: the first max_in_flight: cells are admitted against a total of ~0, and at a width covering every cell there is no serialization point anywhere in the block, so the budget bounds nothing.

A reservation fixes that by charging the allowance up front for work not yet reported, and pausing once it is committed:

- across:
  - var: dim
    from_file: dimensions/items.json
  max_in_flight: 6            # full width; the budget binds anyway
  budget:
    tokens: 3600000
    reserve_per_cell: 900000  # what admission charges a cell that has not reported
  agent: reviewer

Four cells are admitted against reservations alone; the fifth waits. As each finishes, its reservation is replaced by what it actually spent, and the next admission decides on that real number. Cells that come in under their reservation hand the difference back — the common case, and the reason this is a reservation and not a hard pre-allocation.

Waiting, not dropping. A refusal caused by spend is permanent, because spend only grows. A refusal caused by reservations is temporary, because in-flight cells release theirs as they finish — so the block pauses instead of truncating. A six-cell matrix costing ten tokens a cell runs all six against a 3,600 allowance; it does not stop at four having spent forty.

Reserve is a pacing knob. Reserve high and the matrix runs narrower, waiting for real numbers sooner — less overshoot, less concurrency. Reserve low and more cells start blind, which is where overshoot comes from. Neither setting loses a cell.

Concurrent cells: max_in_flight:

Serial cells are the right default for a hand-written matrix. They are the wrong default for a wide one — especially a from_file: fan-out, where the cells are N independent agents an earlier step decided on:

jobs:
- name: wide
  plan:
  - across:
    - var: dimension
      values: [api-boundaries, error-paths, concurrency, performance]
    max_in_flight: 4        # four cells at a time
    task: review
    run: echo reviewing {{ .vars.dimension }}
  assert:
    execution:              # concurrent, but merged back in declaration order
    - review [dimension=api-boundaries]
    - review [dimension=error-paths]
    - review [dimension=concurrency]
    - review [dimension=performance]
    outcome: succeeded

Sharding: parallelism:

Sometimes the axis is just "run N copies" — a test suite split into shards, where each copy needs only its slot number and the total. parallelism: N is that matrix without the hand-written axis: across: over an implicit 1-based index, plus a count var no hand-written axis could carry (an axis knows its values, not its width):

jobs:
- name: sharded
  plan:
  - task: unit
    parallelism: 3
    run: sh -c 'echo "slot {{ .vars.index }} of {{ .vars.count }}"'
  assert:
    execution:
    - unit [index=1]
    - unit [index=2]
    - unit [index=3]
    outcome: succeeded

approval: — a human in the plan

Agent pipelines eventually gate something that cannot be undone — publishing, deploying, sending. approval: is the place in a plan to stop and ask. It parks the run until someone decides, so it can't be executed by the docs suite:

agents:
- name: writer
  source: { model: openrouter/qwen/qwen3.7-flash }

jobs:
- name: publish
  plan:
  - agent: writer
    messages:
      - "Draft the release announcement into draft/summary.md."
    outputs: [draft]
  - approval:
      message: "Draft is in draft/summary.md — publish?"
      timeout: 24h
    on_abort:
      task: nag
      run: echo nobody decided within 24h    # expiry classifies as aborted
  - task: post
    inputs: [draft]
    run: cat draft/summary.md
$ steps run pipeline.yml
…
approval 1: Draft is in draft/summary.md — publish?
approval 1: waiting up to 24h0m0s — steps approvals approve 1 -p pipeline --db .steps/pipeline.yml.db  |  steps approvals reject 1 -p pipeline --db .steps/pipeline.yml.db

$ steps approvals -p pipeline --db .steps/pipeline.yml.db    # another shell, same directory
ID  JOB      REQUESTED             MESSAGE
1   publish  2026-08-05T14:02:11Z  Draft is in draft/summary.md — publish?

$ steps approvals approve 1 -p pipeline --db .steps/pipeline.yml.db
approved: approval 1

The --db is there because a local run keeps its state beside the YAML, in .steps/pipeline.yml.db, while the read commands default to a daemon's .steps/steps.db; a run under steps web prints -p <name> alone.

Three outcomes, deliberately different

classificationwhy
approvedthe plan continues
rejectedfaileda person decided; on_failure fires and a to: route can act on it
expiredabortednobody decided anything

Conflating the last two would make a silent expiry indistinguishable from a rejection, and "the deploy was rejected" is a very different thing to read in a log from "the deploy was never looked at".

⚠️ v1 scope: anyone who can run the CLI can approve

Stated deliberately rather than left to be discovered. There is no separate identity system. The recorded approver comes from $STEPS_APPROVER, $USER, or $LOGNAME: it is an audit record, not an authorization check.

The rest