steps docs
start here Writing pipelines resourcesexprcontrol-flowagentsattempts-timeoutworkspaceinfratemplatingmcpcomplete Reference webagents-internalsaws-workersgcp-workersconformance

Infrastructure Features

Two independent opt-in features for running pipelines beyond the simple one-shot host-execution case: containerized execution (image:) and cross-job downstream triggers (steps web). Container examples on this page validate but aren't executed by the docs suite (they need a docker daemon); the watch/trigger examples run as shown.

Container execution (image:)

By default every pipeline-defined command (a resource type's check/in/out, a task's run:, an agent's run_shell/custom tools) runs on the host via sh -c. Setting image: on a resource_types: entry, a top-level tasks: entry, or an agents: entry runs that entity's commands in a container from that image instead — one container per step, started on the step's first command and removed when the step ends.

resource_types:
- name: releases
  image: alpine/git             # check/in/out run in this image
  config:
    check: |
      git ls-remote --tags https://github.com/jtarchie/ci.git | tail -1 | awk '{print "[{\"ref\": \""$2"\"}]"}'
    in: echo {{ .version.ref | shellquote }} > ref

resources:
- name: tags
  type: releases
  source: {}

tasks:
- name: build
  image: golang:1.26            # this task's run: (and fix-loop re-runs) run here
  inputs: [tags]
  run: cat tags/ref && go version

agents:
- name: reviewer
  source: { model: openrouter/qwen/qwen3.7-flash, api_key_env: OPENROUTER_API_KEY }
  image: python:3.12            # this agent's run_shell/custom tools run here
  tools: [read_file, run_shell]

jobs:
- name: verify
  plan:
  - get: tags
  - task: build
    image: golang:1.25          # every one of these is a STEP-level override,
    env: [BUILD_TAG]            # applying to this step and no other use of the
    user: root                  # tasks: entry above
    network: host
    privileged: true
    container_limits:
      cpu: 512
      memory: 2147483648
  - agent: reviewer
    inputs: [tags]
    messages:
      - "Sanity-check the fetched tag."

TMPDIR when the daemon runs in a VM

On macOS the docker daemon runs inside a Linux VM (Docker Desktop, colima, Rancher), and only some host paths are shared into it — your home directory is, macOS's own $TMPDIR (/var/folders/…) is not.

steps builds each run's workspace under $TMPDIR and bind-mounts it into the container. If $TMPDIR isn't shared, that mount does not fail: docker silently creates an empty directory at the mount target. The container then runs against a workspace that has none of your inputs and writes results nowhere the host can see — typically surfacing as can't create out/result.txt: nonexistent directory, or as a step that "succeeds" and produces no outputs.

Point TMPDIR at a shared path before running:

export TMPDIR="$HOME/.steps-tmp" && mkdir -p "$TMPDIR"

Native Linux is unaffected.

CLI agents

For a CLI-backed agent (source.model: "@claude/..."), image: containerizes the step's tools — exactly as it does for a hosted agent — never the CLI process itself. The CLI is always a host subprocess of the steps process; only the run_shell/custom-tool/MCP execution a bridged call performs happens inside the container, through the same toolEnv.runner a hosted agent's tools already run through.

That single fact is why a CLI agent's containerization has none of the machinery a "containerize the whole CLI" design would need:

One thing containerizing a CLI agent's tools does not fence: web_fetch is an in-process HTTP implementation on both the hosted and the CLI path, not a shell tool, so network: none narrows run_shell and custom tools without touching it.

Remote workers (tags:)

Every step of a job runs on the machine steps runs on. tags: places one somewhere else — a GPU box, a different OS or arch, a machine that holds hardware or credentials the orchestrator does not:

jobs:
- name: train
  assert:
    execution: [prepare, train]
    outcome: succeeded
  plan:
  - task: prepare
    outputs: [data]
    run: echo seed > data/seed.txt
  - task: train
    tags: [gpu]
    inputs: [data]
    outputs: [model]
    run: |
      echo "trained from $(cat data/seed.txt)" > model/report.txt
      echo "worker: ${STEPS_WORKER:-none}"
    assert:
      stdout: "worker: gpu"
      files: [model/report.txt]

The pipeline names a capability; the invocation names the machine:

steps run --worker gpu=ssh://jt@gpu-box pipeline.yml

Keeping machines out of the pipeline file is what lets the same pipeline run on somebody else's fleet — the split Concourse draws between a step's tags: and a worker's advertised ones.

Resources on workers

A source only reachable from a worker's network — a git host inside a VPC, a registry behind a bastion — has to be checked, fetched and pushed from there. tags: on the resource places all three:

resource_types:
- name: probe
  config:
    check: printf '[{"ref":"v1","where":"%s"}]' "${STEPS_WORKER:-here}"
    in: printf '%s/%s' {{ .version.where | shellquote }} "${STEPS_WORKER:-here}" > where.txt
    out: printf '{"ref":"pushed","where":"%s"}' "${STEPS_WORKER:-here}"

resources:
- name: repo
  type: probe
  tags: [vpc]
  source: {}

jobs:
- name: mirror
  assert:
    execution: [repo, inspect, repo, compare, repo]
    outcome: succeeded
  plan:
  - get: repo
  - task: inspect
    inputs: [repo]
    run: cat repo/where.txt
    assert:
      stdout: vpc/vpc
  - get: mirror
    resource: repo
    tags: [edge]
  - task: compare
    inputs: [mirror]
    run: cat mirror/where.txt
    assert:
      stdout: vpc/edge
  - put: repo
    inputs: [repo]

The life of an aws:// worker

Three things happen on three different clocks, and knowing which is which explains most of the behaviour above. A placed step is one carrying tags:; a shim is the steps binary running in worker mode on the far end; a rung is which of the three aws:// forms you wrote.

steps run --worker aws=aws://...  pipeline.yml
│
├─ first placed step
│   │
│   ├─► ACQUIRE ── once per JOB, not per step
│   │     aws://i-0abc...          already running; nothing is acquired
│   │     aws://stopped/i-0abc...  StartInstances, wait for "running"
│   │     aws://launch/lt-0def...  CreateFleet, wait for "running"
│   │
│   ├─► BOOTSTRAP ── once per CONNECTION (SSM SendCommand)
│   │     fetch the binary from a 20-minute presigned URL, unless
│   │       <root>/steps-shim/<content-hash>/steps is already there
│   │     start  steps _shim --listen 127.0.0.1:0 --once
│   │
│   └─► CONNECT ── SSM port-forward session to that loopback port
│         inputs over, command runs on the host, outputs back
│         the shim serves ONE connection and exits
│
├─ later placed steps: reuse the machine, fresh --once shim each time
│    (a step that redials after a dropped tunnel bootstraps again too)
│
└─ job ends (succeeded, failed, or cancelled)
    │
    └─► RELEASE ── when the LAST user of the machine is done
          aws://i-0abc...          nothing; the machine is yours
          aws://stopped/i-0abc...  StopInstances, after ?idle= if set
          aws://launch/lt-0def...  TerminateInstances, after ?idle= if set

What that leaves on the instance between steps: the cached binary under <root>/steps-shim/<content-hash>/, and nothing else. No steps process runs, because --once means the shim exits after the one connection it was started for.

What runs the command: the worker's host, unless the step names an image:, in which case the worker's own docker daemon runs it in a container against the copy of the tree the shim unpacked. A stock Amazon Linux 2023 AMI has no daemon, so an aws:// worker needs one installed — through the launch template's user data, or baked into the AMI — for placed steps that name an image, and needs nothing at all for those that do not.

What the instance never holds: AWS credentials. Artifact bytes arrive over presigned URLs the orchestrator mints per transfer, so the instance profile needs AmazonSSMManagedInstanceCore and nothing else — not even read access to the bucket its own inputs came from.

A gcp:// worker lives the same life on different machinery. ACQUIRE is instances.insert from the template (or instances.start for the parked rung); there is no separate BOOTSTRAP errand, because the IAP tunnel terminates at sshd and the binary rides in over sftp exactly as ssh:// pushes it — cached on the worker by content hash the same way; CONNECT is the relay websocket carrying the SSH session; RELEASE is instances.delete (or stop). What a GCE instance never holds: GCP credentials of its own for steps' purposes — the orchestrator's ADC signs the tunnel, and artifact bytes still arrive over presigned URLs when a store is configured. One machine-shape note with money attached: name a spot provisioning model with instanceTerminationAction: DELETE in the template, or every preempted worker leaves a stopped instance whose disk keeps billing.

Artifact store (--artifact-store)

The step cache's remote half: cached step outputs mirrored to a content-addressed store on S3, so bytes evicted locally — or never present on this machine — are materialized back instead of re-earned by running the step. Opt-in, and CLI-only by design: the flag names infrastructure, and pipelines stay portable.

steps run --artifact-store s3://my-bucket/team-prefix pipeline.yml

Container network (network:)

image: isolates a command's filesystem view but not its network — a containerized run_shell an agent wrote has the same egress the host does. For a step whose commands are model-generated, that is usually the isolation you actually wanted:

agents:
- name: analyzer
  source: { model: openrouter/qwen/qwen3.7-flash, api_key_env: OPENROUTER_API_KEY }
  image: python:3.12
  network: none        # can read the workspace, can't reach anything
  tools: [read_file, run_shell]

jobs:
- name: analyze
  plan:
  - task: fetch
    outputs: [data]
    run: echo 42 > data/metrics.txt
  - agent: analyzer
    inputs: [data]
    messages:
      - "Analyze data/metrics.txt offline."

This is not a full sandbox — a command can still reach the host filesystem by absolute path, and network: host opts back out entirely.

Container privileges and limits (privileged:, container_limits:)

Both sit wherever image: does, and both require image: — a host-executed command has no cgroup to cap and no privilege to raise, so accepting either there would promise something it does not do.

resource_types:
- name: images                  # publishing an image needs a daemon of its own
  image: docker:27-dind
  privileged: true
  user: root                    # dind's daemon will not start unprivileged
  network: bridge               # ...and it has to reach the registry
  container_limits:
    cpu: 1024
    memory: 4294967296
  config:
    check: |
      printf '[{"tag": "latest"}]'
    out: |
      docker build -t app . && printf '{"tag": "latest"}'

resources:
- name: app-image
  type: images
  source: {}

tasks:
- name: integration
  image: docker:27-dind
  privileged: true              # docker-in-docker needs it
  network: bridge
  container_limits:
    cpu: 512                    # --cpu-shares
    memory: 2147483648          # --memory, in BYTES (2 GiB)
  run: ./run-integration.sh

agents:
- name: builder                 # an agent whose run_shell drives that daemon
  source: { model: openrouter/qwen/qwen3.7-flash, api_key_env: OPENROUTER_API_KEY }
  image: docker:27-dind
  privileged: true
  user: root
  env: [DOCKER_HOST]            # named, never valued — see env: below
  container_limits:
    cpu: 1024
    memory: 4294967296
  tools: [run_shell]

jobs:
- name: test
  plan:
  - task: integration
  - agent: builder
    messages:
      - "Build the image and report what failed."
  - put: app-image

Container user (user:)

On Linux, a bind mount carries host uids straight through. A container running as root — which most images do — writes root-owned files into the step's working directory, and three things break: an agent creates a file with a containerized tool and can't edit it with a host-side one, workspace capture hits permission errors, and whatever's left behind needs root to delete.

So on Linux the default is the uid:gid that started steps, not the image's user. Elsewhere the mismatch doesn't arise (Docker Desktop's VM maps ownership on bind mounts), so off Linux the default stays the image's own user.

tasks:
- name: install-deps
  image: ubuntu
  user: root          # this image installs packages at run time; it needs root
  run: apt-get update && apt-get install -y jq && echo ready

jobs:
- name: setup
  plan:
  - task: install-deps

Passing environment through (env:)

Commands run with a deliberately narrow environment: a host command sees a fixed allowlist (PATH, HOME, locale, proxy settings — not the operator's credentials, and not SSH_AUTH_SOCK, which a pipeline that needs git-over-ssh opts back in by name), and a containerized command sees only its image's own environment. That default is the trust boundary: an agent directing run_shell should not get read access to everything the operator happened to export.

env: opts specific variables back in, by name:

tasks:
- name: deploy
  env: [OPENROUTER_API_KEY]     # the name; the value stays in the operator's env
  run: |
    if [ "${OPENROUTER_API_KEY+set}" = set ]; then
      echo "credential reached the command"
    else
      echo "credential was filtered out"
    fi

jobs:
- name: release
  plan:
  - task: deploy
    assert:
      stdout: credential reached the command   # delete the env: line and this fails
  assert:
    execution: [deploy]
    outcome: succeeded

Downstream triggers (trigger: true + steps web)

By default steps is a one-shot, single-job CLI. steps web adds a long-running mode: it holds the pipelines uploaded to it by steps pipeline set, polls every resource named by any get ..., trigger: true step across every job of each one, and automatically runs whichever jobs are affected when that resource's latest version changes — including a version produced by another job's own put:

resource_types:
- name: countfile
  config:
    check: |
      printf '[{"n": "%s"}]' "$(cat {{ .source.path | shellquote }} 2>/dev/null || echo 0)"
    in: cat {{ .source.path | shellquote }} > n.txt 2>/dev/null || echo 0 > n.txt
    out: |
      next=$(( $(cat {{ .source.path | shellquote }} 2>/dev/null || echo 0) + 1 ))
      echo "$next" > {{ .source.path | shellquote }}
      printf '{"n": "%s"}' "$next"

resources:
- name: counter
  type: countfile
  source: { path: counter.txt }

jobs:
- name: publish
  plan:
  - put: counter
  assert:
    execution: [counter]
    outcome: succeeded
- name: notify
  plan:
  - get: counter
    trigger: true      # steps web runs this job when publish lands a new version
  - task: announce
    inputs: [counter]
    run: echo "counter is now $(cat counter/n.txt)"
    assert:
      stdout: counter is now     # which number depends on when the poller ran;
  assert:                        # under `steps test` the plan resolved before the put
    execution: [counter, announce]
    outcome: succeeded

assert:
  execution: [publish, notify]
steps web --interval 30s --max-concurrent 1
steps pipeline set -c pipeline.yml

Get renaming (resource:)

A get step's resource: names the resource to fetch when it should differ from the step's own name — mirroring Concourse's get.resource:

resource_types:
- name: greetings
  config:
    check: |
      printf '[{"word": "hello"}]'
    in: echo {{ .version.word | shellquote }} > word.txt

resources:
- name: repo
  type: greetings
  source: {}

jobs:
- name: aliased
  plan:
  - get: source          # the artifact (and directory, step name, to: target) is "source"
    resource: repo       # the resource whose check/in runs is "repo"
  - task: show
    inputs: [source]
    run: cat source/word.txt
    assert:
      stdout: hello
  assert:
    execution: [repo, show]   # recorded under the RESOURCE name, not the alias
    outcome: succeeded

The artifact name is the get: value; the resource fetched is resource:, defaulting to the get: value when omitted. This lets one resource appear under a task-friendly name, or twice in a plan under two names. Pair it with a task's input_mapping: (see workspace.md) to feed a reusable task's pinned input name from an aliased get.

Circuit breaker: max_consecutive_failures:

steps web runs unattended, and a job that fails on every new version will keep firing on every new version — burning model spend on a failure no automatic retry is going to fix:

jobs:
- name: nightly-summary
  max_consecutive_failures: 3
  plan:
  - task: summarize
    run: echo summarizing
    assert:
      stdout: summarizing
  assert:
    execution: [summarize]
    outcome: succeeded        # a green run also clears the breaker's count
Fri 02:00  nightly-summary failed (1/3 consecutive)
Sat 02:00  nightly-summary failed (2/3 consecutive)
Sun 02:00  nightly-summary PAUSED after 3 consecutive failures — resume with: steps jobs resume nightly-summary -p <pipeline>

How much history to keep: run_history:

steps web runs for weeks, and every build it does writes a run row, an event per step, what each agent step spent, the full text of every agent conversation, and a cached node per step. None of that used to be cleaned up. Measured on a pipeline answering Slack mentions overnight, one build cost about 23KB — so a hundred builds a day added a couple of megabytes a day, forever, and three quarters of it was cached nodes and agent transcripts.

defaults.run_history: caps it per job, keeping the newest:

defaults:
  # Runs are much bigger than versions, so the two caps are not the same number:
  # a version is a few dozen bytes and a run is tens of kilobytes.
  run_history: 20
  version_history: 50

jobs:
- name: summarize
  plan:
  - task: write
    run: echo summarized
    assert:
      stdout: summarized
  assert:
    execution: [write]
    outcome: succeeded

passed: — only run against versions that are green upstream

Without it, steps web will trigger deploy on a commit the test job already failed on, and there is no way to say otherwise. This is a correctness gap, not a convenience:

resource_types:
- name: commits
  config:
    check: |
      printf '[{"ref": "abc123"}]'
    in: echo {{ .version.ref | shellquote }} > ref

resources:
- name: repo
  type: commits
  source: {}

jobs:
- name: unit
  plan:
  - get: repo
    trigger: true
  - task: test
    inputs: [repo]
    run: echo tests pass for "$(cat repo/ref)"
    assert:
      stdout: tests pass for abc123
  assert:
    execution: [repo, test]
    outcome: succeeded

- name: deploy
  plan:
  - get: repo
    trigger: true
    passed: [unit]           # only a version unit went green on
  - task: release
    inputs: [repo]
    run: echo deploying "$(cat repo/ref)"
    assert:
      stdout: deploying abc123
  assert:
    execution: [repo, release]   # unit ran green first, so the version was released
    outcome: succeeded

assert:
  execution: [unit, deploy]      # declaration order, which is why unit's green counts
commit abc123 → unit    FAILED
commit abc123 → deploy  waiting: no version has passed [unit] yet
commit def456 → unit    ok
commit def456 → deploy  ok

max_in_flight: — how many builds of one job at once

By default a job's builds are unlimited, bounded only by steps web --max-concurrent. Cap it per job when the work is not safe to overlap but does not need full serialization:

jobs:
- name: integration
  max_in_flight: 2     # at most two builds of this job at a time
  plan:
  - task: test
    run: echo testing
    assert:
      stdout: testing
  assert:
    execution: [test]
    outcome: succeeded

serial: / serial_groups: — stop jobs racing each other

steps web --max-concurrent 4 runs jobs concurrently. For anything that deploys, publishes, or otherwise mutates the outside world, that is a hazard:

jobs:
- name: deploy-staging
  serial: true                  # never two builds of me at once
  serial_groups: [deploy-lock]  # and never at the same time as anyone else in the group
  plan:
  - task: deploy
    run: echo deploying staging
    assert:
      stdout: deploying staging
  assert:
    execution: [deploy]
    outcome: succeeded
- name: deploy-prod
  serial_groups: [deploy-lock]
  plan:
  - task: deploy
    run: echo deploying prod
    assert:
      stdout: deploying prod
  assert:
    execution: [deploy]
    outcome: succeeded

assert:
  execution: [deploy-staging, deploy-prod]
10:00:01  deploy-prod (v1) started
10:00:04  deploy-staging waiting: lock held by deploy-prod
10:03:20  deploy-prod (v1) done
10:03:20  deploy-staging started

interruptible: — what a shutdown does to a running build

steps web gets SIGTERM (a restart, a redeploy, a machine going down) while a job is mid-deploy. Whether that build is allowed to finish is the question this answers:

jobs:
- name: deploy-prod
  plan:
  - task: deploy
    run: echo deploying
    assert:
      stdout: deploying
  # default: shutdown WAITS for a running build to finish
  assert:
    execution: [deploy]
    outcome: succeeded

- name: nightly-report
  interruptible: true       # ...this one can just die
  plan:
  - task: report
    run: echo reporting
    assert:
      stdout: reporting
  assert:
    execution: [report]
    outcome: succeeded

assert:
  execution: [deploy-prod, nightly-report]

Webhook-triggered checks

steps web polls on an interval: short means fast reaction and lots of API calls, long means slow reaction. A webhook removes the tradeoff — react instantly, poll rarely as a safety net:

resource_types:
- name: commits
  config:
    check: |
      printf '[{"ref": "abc123"}]'
    in: echo {{ .version.ref | shellquote }} > ref

resources:
- name: repo
  type: commits
  source: {}
  webhook_token_env: GITHUB_WEBHOOK_TOKEN    # the variable NAME, not the token

jobs:
- name: build
  plan:
  - get: repo
    trigger: true
  - task: compile
    inputs: [repo]
    run: echo building "$(cat repo/ref)"
    assert:
      stdout: building abc123
  assert:
    execution: [repo, compile]
    outcome: succeeded
steps web --listen 0.0.0.0:8080 --read-only
curl -X POST 'http://localhost:8080/p/pipeline/check/repo?token=…'   # or: Authorization: Bearer …

The route lives under the pipeline it checks, on the same address the UI is served from. The poll loop used to open a second port of its own, so a deployment that wanted both had two addresses and two HTTP surfaces to expose; one daemon means one listener. pipeline in the path is the pipeline's name — the YAML's base name unless --name says otherwise, the same string as its /p/<name>/ page.

One listener means one exposure. Reaching this endpoint from outside the machine means binding --listen to a routable address, and that address also serves the UI — which has no authentication at all (see web.md). Anyone who can reach the port can read every run, transcript and log, and — without --read-only — trigger any job, decide any approval, and answer any question. Pair the two flags, as above: --read-only withholds every browser control while leaving this token-authenticated route working, which is exactly the shape a webhook receiver wants. Put it behind a reverse proxy if the UI needs to be reachable too.