steps docs
start here Writing pipelines resourcesexprcontrol-flowagentsattempts-timeoutworkspaceinfratemplatingmcpcomplete Reference webagents-internalsaws-workersgcp-workersconformance

The daemon

steps web                                        # starts empty
steps pipeline set -c pipeline.yml               # upload one into it

steps web is the long-running mode: it serves the browser UI at http://127.0.0.1:8088, holds whatever pipelines have been set into it, polls their trigger: true resources, and runs the jobs both of those enqueue.

It takes no pipeline arguments. A pipeline arrives by steps pipeline set and by nothing else, which is what makes three things true that were not before: vars belong to the pipeline you set rather than to the process, a pipeline's identity is a name you chose rather than whatever its file was called, and a change happens when somebody asks for it rather than when a file moves. Nothing here watches a file.

Each served pipeline is routed under /p/<name>/, where the name is the one set was given. All of them share one state database — .steps/steps.db unless --db says otherwise — and stay strangers inside it; see One database, several pipelines.

Setting a pipeline

steps pipeline set -p app -c app.yml -v repo_uri=https://github.com/acme/app

The refusal lands where you asked. That is the whole reason this is HTTP rather than a row written into a database: set needs a synchronous answer from the machine that will run the pipeline — is api_key_env: set there, is the stdio MCP binary on its PATH, does workspace.root: exist. A configuration that fails those is refused, the daemon goes on serving what it had, and the terminal that asked prints why:

$ steps pipeline set -c app.yml
http://127.0.0.1:8088 refused it: app cannot run here:
  agent "reviewer"  $OPENROUTER_API_KEY is not set (source.api_key_env)

What it does NOT check is the network. A set that passes can still meet a revoked key or a server that is down at run time; steps validate --live is the command that asks.

The rest of the family

steps pipeline list                      # what this daemon holds
steps pipeline get -p app                # the configuration it is serving
steps pipeline pause -p app              # stop polling, admitting and triggering
steps pipeline unpause -p app
steps pipeline rename -p app --to legacy # keeps the history
steps pipeline destroy -p app            # forgets it, and everything under it

Every verb takes --target (or STEPS_TARGET), defaulting to http://127.0.0.1:8088. There is no login and no saved targets, because there is nothing to log in to — see Security.

What a restart does

Nothing. The daemon loads every pipeline's current configuration from its state database at startup and serves it, so a restart picks up where the last process left off — including on a machine where the YAML never existed. A daemon that holds nothing serves an index saying how to set one.

What it shows

RouteAnswers
/With several pipelines served: what this process holds, and one run feed across all of them, newest first. With one, it redirects straight through
/p/:pipelineWhich jobs exist, how each last run went, and which jobs feed which — as a list, or as a dependency graph laid out from the passed: constraints, each node carrying its latest status
…/runsOne run history across every job of the pipeline, newest first — the cross-job view the per-job history can't give
…/jobs/:jobThis job's dependencies in both directions, its run history with a duration trend, the resource versions it has passed against, and the resolved limits each agent step runs under
…/runs/:runThe transcript: every step in plan order, what it did, and — for agent steps — what the model said and which tools it called
…/nodes/:hashWhat a merkle hash is made of, and every run that reused it: the cache's receipt
…/config/:shaThe pipeline as the runs pinned to that hash executed it — readable after the file on disk has moved on
…/approvalsPending approval: steps, and the decisions already made
…/questionsPending ask_user questions, and the answers already given
…/resourcesLatest checked version per resource, and any job the circuit breaker has paused
/docsThese docs, rendered with syntax-highlighted examples — the same pages steps docs shows in a terminal

Press / anywhere for a jump palette over pipelines, jobs, and recent runs — across every pipeline this process serves, not only the one whose page you are on. The one you are on ranks first, and a hit from anywhere else says which pipeline it belongs to.

Agent dials

A job page lists the resolved limits of each agent step in its plan: turns, context ceiling, deadline, and spend budget, after the step, the agent and the built-in default have all had their say. It exists so "why did this step stop at 30 turns" is answerable without cross-referencing three files, and it shows uncapped rather than 0 for a dial an author explicitly removed — 0 in a limit column reads as the opposite of what it means. An ensemble: that decides with a judge (decide: <agent>) lists the judge as a row of its own, because it runs — and spends — as an agent step the plan never spells out.

Turns is one word for two units and the header says so: a hosted agent's turn is one request/tool-execute round driven by steps, while a CLI agent's is whatever the child reports as num_turns — one per tool round, pooled across every messages: entry — which runs far higher for the same work. A cap that looks generous beside a hosted source can truncate a CLI one mid-task.

The budget column carries one unit or the other, never both, because the two spellings are exclusive by source kind: a hosted agent is metered in budget.tokens and a CLI agent in budget.usd (see attempts-timeout.md). Read uncapped there together with the turn column: for a CLI agent budget.usd is the only ceiling anything enforces mid-conversation, so an uncapped budget beside uncapped turns means the step is held by its deadline and nothing else. The one exception is spelled out in the same cell: an across: block's own budget: caps what its cells spend together (see control-flow.md), so such a step reads uncapped per cell · 50,000 tokens for the matrix rather than uncapped.

It covers the agents a step names. A task's fix: agent and a step's sub-agent tools: grants run under limits of their own and are not listed.

The transcript

The run page is the point of the whole thing. It renders a run the way the terminal does — steps in order, prefixed by kind, colored by outcome — with the things a scrollback cannot give you:

Live runs

A run still in flight streams to the page over server-sent events: steps appear as they start, tool calls arrive as the model makes them, and the page settles into its final state when the run ends.

The stream is built on the same rows the finished-run page reads, so a run watched live and the same run opened an hour later show the same thing, and a dropped connection costs nothing but a reconnect. It also means the UI shows runs it did not start — a steps run in another terminal against the same pipeline appears here as it happens.

Recording is the runner's job, not the UI's: every run persists its events (run_events), whether or not anything is watching. A job started from a terminal leaves the same record as one started from the browser.

Following a run you started

Triggering does not drop you back on a list to refresh. A trigger lands on a short waiting page that reports what the queue is doing and forwards itself to the live transcript the moment a worker picks the job up — a queued job has no run id until then, which is why there is a waiting room rather than a redirect.

While a run is live, the browser tab carries its status: ◐ running, ✓ passed, ✗ failed, with a matching favicon dot. The title updates the instant the run ends, so a run left in a background tab reports its outcome without being reopened.

The jobs board refreshes itself every couple of seconds, in place — it keeps your list/graph choice and scroll position rather than reloading the page — and pauses while the tab is hidden.

Triggering, approving, resuming

Five controls, each doing what a CLI verb does:

--read-only withholds all five: the controls disappear from the pages and the routes refuse. The queue is still drained, polling still runs, and steps pipeline set still works — that flag is a statement about the browser's surface, not about what the process does on its own or about how it is deployed. --listen 0.0.0.0:8088 --read-only is a build box that still has to notice new versions; read Security before you expose one.

The webhook route is the one exception, deliberately. POST /p/<slug>/check/<resource> still works under --read-only, and the job it enqueues still runs. It is not a UI control: it carries the resource's own token, which is a stronger check than the five above have, and withholding it would mean a read-only box could not be the thing GitHub notifies — which is most of why a build box is exposed at all. --read-only says a browser cannot start work here; it does not say nothing can. If that is what you want, do not give the pipeline a webhook_token_env: resource — with none, the route is a 404. See infra.md.

Aborting a run

Stopping one run leaves the daemon and every other run alone:

steps runs abort -p app 46UMHVPYRA6YHB7M   # stop a running run
steps runs abort -p app --queued build     # drop build's queued run before it starts

or ■ Abort on the run's page. It means what it means in Concourse:

It goes through the daemon rather than the database — what it stops is a context inside that process — so it needs a running steps web, takes --target like steps pipeline does, and is refused under --read-only. A run this daemon is not executing — one a steps run started against the same file, or one a crashed daemon left marked running — is refused with a message saying so.

One daemon

steps web is the whole long-running mode: it serves the UI, holds the pipelines, polls every trigger: true resource, and drains the queue all of that fills. There is no separate watcher and no one-shot — a front end that drains a queue nothing fills is a runner that looks alive and notices nothing, and two processes against one state database claim each other's work.

steps web                      # serve, and poll every 30s
steps web --interval 5m        # slower
steps web --max-concurrent 4   # up to four queued jobs at a time

There is no --once. It was the cron form of a runner: load a file, poll once, exit, never bind. A process that never binds has nothing to be set into, and a server is its own scheduler — run it under systemd as a service rather than a timer, and let --interval be the schedule.

Applying a change

Set it again. There is no watcher and no --watch:

steps pipeline set -c pipeline.yml       # edit, set, and it is serving

A save-to-apply loop is a one-line wrapper around this — watchexec -w pipeline.yml -- steps pipeline set -c pipeline.yml -n — and that is where it belongs. A watcher inside the daemon is an implicit set that fires with nobody attached, which is exactly why it needed a held-configuration banner nobody was looking at: deleting it deletes the banner, the hold state, and the question of what a daemon does with a configuration it will not accept. It refuses it, to your face.

A run in flight finishes against the configuration it started under. The swap is immediate for everything after it: the pages, the trigger poller, the webhook endpoints, and the next job the queue admits — including the serial:, serial_groups: and max_in_flight: a job is admitted under, which live in the database and are rewritten from the new configuration on every set. What the running job is executing does not change underneath it, and the run records which configuration that was — the CONFIG column steps runs prints and the revision named on the run page, where it links the configuration itself. A --resume is the exception, and deliberately: it continues a failed run under the configuration it is resumed with, which is usually the one that fixed it.

Adding or removing a trigger: true get takes effect too. The poll loop re-decides what it can check each time the configuration changes, so a resource a set adds is preflighted and then polled, and one a set removes stops being checked. A pipeline with nothing to poll is a state the loop sits in, not a reason it was never started.

A changed workspace: is adopted. The provider that materializes build directories is rebuilt when the block moves, and a run already in flight keeps the one it started with until it finishes — so a set never deletes the tree a build is working in. One this machine cannot provide is refused like any other part of the configuration.

One database, several pipelines

A daemon holds every pipeline set into it in ONE state database — .steps/steps.db unless --db names another:

steps web --db /var/lib/steps/state.db
steps pipeline set -p app -c app.yml
steps pipeline set -p infra -c infra/pipeline.yml

A bare path is a sqlite file, and so is sqlite:///var/lib/steps/state.db — the scheme is how a second driver will be chosen, the way --worker takes ssh:// and aws://, and sqlite is the only one today. A scheme no driver answers to is refused before anything is opened.

One file to back up, and one file to delete. What it is not is a merge: inside the database every row carries the pipeline it belongs to, so histories, resource versions, queues, serial groups and the merkle cache stay separate. Two pipelines each with a job named build running an identical task do not share a cache entry, and one pipeline's run_history: cap never reaps another's runs.

Reading it back needs no pipeline argument. steps runs lists what the file holds and interleaves the newest runs of all of it, which is the terminal's version of the web root:

$ steps runs --db /var/lib/steps/state.db

PIPELINE  PATH
app       /src/app/app.yml
infra     /src/infra/pipeline.yml

WHEN                 PIPELINE  JOB      STATUS     RUN
2026-08-30 09:14:02  infra     deploy   succeeded  UNVFHMCHVWY6GV6N
2026-08-30 09:12:40  app       build    failed     46UMHVPYRA6YHB7M

PATH is where the configuration was last set FROM — a file on whoever's machine ran set, recorded so a reader can tell two checkouts apart, and never opened here.

The read commands take -p <name>, not a path, because a served pipeline has no file on this machine:

steps runs -p app                      # what ran
steps runs steps -p app                # why a step did what it did
steps runs cost -p app 46UMHVPYRA6YHB7M
steps approvals -p app
steps questions -p app
steps jobs -p app

They read the database directly rather than going through the daemon, so they work against a stopped one — and they take --db when it is not the default. steps run, steps test, steps validate and steps plan still take a file path and still default to .steps/<filename>.db beside it: those are the local commands, and a local command has a file.

Run ids stay globally unique, but --resume and --replay still refuse an id belonging to a different pipeline in the same file: continuing another pipeline's run would reuse its workspace and step indexes against this pipeline's plan.

There is no migration path. A database written by a different schema is refused on open with a message saying so; the answer is to delete the file, which costs run history and cache and nothing else.

Security

There is no authentication, because there is nothing to authenticate against: this is the local runner's own front end, in the same trust domain as the shell that started it. It binds 127.0.0.1 by default.

steps pipeline set is a remote-shell endpoint. Say that plainly: a pipeline is arbitrary commands, so anyone who can reach this port can run anything they like as the user running the daemon. No token, no password, no allow-list. --read-only does not close it either — that flag has always been a statement about the browser's surface, and the set endpoint is the deployment path, not a button on a page.

That is the reason for the loopback default, and it is a stronger reason than the trigger controls ever were. Binding to a routable address hands the machine to whoever can reach the port. --listen 0.0.0.0:8088 exists for someone who has decided that is what they want; put it behind something that authenticates — an SSH tunnel, a reverse proxy, a network nobody else is on — and treat --read-only as being about the browser only.

Loopback keeps other machines off the port, not other web pages. A page open in a browser on this machine can re-point its own hostname at 127.0.0.1 after it loads — DNS rebinding — and from then on its requests reach this port carrying that name as both Host and Origin, which is exactly what a same-origin request looks like. So everything under /api/ — what steps pipeline and steps runs abort talk to — refuses a request carrying a header only a browser attaches: an Origin of any value, which every browser sends on the PUT, POST and DELETE those verbs use, or a Sec-Fetch-Site saying anything but none, which is a URL typed into the address bar. The CLI sends neither, and no page in the UI calls /api/. Host is not checked, because a reverse proxy in front of the daemon forwards the name the client used.

The browser's own controls — trigger, approve, answer, resume, abort — are POSTs refused when their Origin names another host, which stops another site aiming a form at your port. A rebinding page is not another host: it can read every page, configurations included, and press those controls, unless --read-only has withheld them. It cannot set, destroy, rename or pause a pipeline.

Flags

steps web — facts about the box, all of them:

--listen         address to serve on (default 127.0.0.1:8088)
--interval       how often to poll trigger: true resources (default 30s)
--max-concurrent maximum queued jobs running at once, per pipeline (default 1)
--pin / --force  pin a version field; ignore the cache and re-run every step
--no-preflight   skip the pre-poll health check of models and MCP servers
--read-only      serve without trigger, approval, answer, resume, or abort controls
                 (steps pipeline set is NOT withheld — see Security)
--keep-workspace leave build workspaces on disk
--answer         answer an ask_user question in advance (repeatable)
--worker         map a step tag to a machine, e.g. --worker gpu=ssh://jt@box
--db             state database: a sqlite path or sqlite:// url (default .steps/steps.db)

steps pipeline <verb> — facts about one pipeline:

--target         the daemon to talk to (default http://127.0.0.1:8088, $STEPS_TARGET)
-p / --pipeline  its name on the daemon (set defaults to the YAML's base name)
-c / --config    the YAML to upload (set only)
-v / --var, --vars-file   pipeline vars, substituted before the upload (set only)
-n               do not diff or ask; apply it (set, destroy)
--to             the new name (rename only)