steps docs
start here Writing pipelines resourcesexprcontrol-flowagentsattempts-timeoutworkspaceinfratemplatingmcpcomplete Reference webagents-internalsaws-workersgcp-workersconformance

Running a pipeline on a GCP worker

How to build, by hand, the GCP side of a gcp:// worker — and then run a pipeline that uses it.

hack/gcp-fixture.sh does all of this in one command for this repo's own tests. This page is the same thing typed out, so you can see what each resource is for and adapt it. Every command is gcloud with a project already configured. Your own account needs two roles on the project: roles/compute.instanceAdmin.v1 (create, start, stop, delete instances; write their metadata; read guest attributes) and roles/iap.tunnelResourceAccessor (open the tunnel). steps itself authenticates with Application Default Credentials — gcloud auth application-default login once, and both the compute calls and the tunnel sign with it.

What you are building, and why it is almost as small as the AWS one

A worker is a Compute Engine instance steps can reach through IAP TCP forwarding. GCP has no SSM-shaped exec channel, so the SSH contract is the transport — the tunnel terminates at the instance's own sshd, and everything ssh:// does (push a binary over sftp, run it for one session) happens over that tunnel unchanged. That decides the shape of everything below:

your laptop ──wss──▶ IAP relay (tunnel.cloudproxy.app) ──35.235.240.0/20──▶ sshd ──▶ steps _shim
     │                                                                                  │
     └───────────────── presigned URLs (optional store) ──▶ S3 ◀── artifact bytes ──────┘

Set these once so the commands below are copy-pasteable:

export PROJECT=$(gcloud config get-value project)
export ZONE=us-central1-a
export NAME=steps-worker

1. The firewall rule

gcloud compute firewall-rules create "$NAME-iap" \
  --direction=INGRESS --action=ALLOW --rules=tcp:22 \
  --source-ranges=35.235.240.0/20 --target-tags="$NAME"

The one piece of ingress this design needs: Google's IAP range, to 22, only for instances carrying the $NAME network tag. There is deliberately no 0.0.0.0/0 anywhere.

2. An instance template

An instance template is one immutable object holding a complete machine shape: image, machine type, disks, network, service account, provisioning model, metadata. You never edit one; a different shape is a different template — which is why gcp:// has no ?version=.

gcloud compute instance-templates create "$NAME" \
  --machine-type=e2-small \
  --image-family=debian-12 --image-project=debian-cloud \
  --no-service-account --no-scopes \
  --no-address \
  --tags="$NAME" \
  --provisioning-model=SPOT --instance-termination-action=DELETE \
  --metadata=enable-guest-attributes=TRUE,enable-oslogin=FALSE,startup-script='#!/bin/bash
apt-get update && apt-get install -y docker.io'

Each flag is a decision:

3. The static instance

gcloud compute instances create "$NAME" \
  --zone="$ZONE" \
  --machine-type=e2-small \
  --image-family=debian-12 --image-project=debian-cloud \
  --no-service-account --no-scopes \
  --no-address \
  --tags="$NAME" \
  --metadata=enable-guest-attributes=TRUE,enable-oslogin=FALSE

One machine you own and steps merely dials — the static worker. Deliberately not --source-instance-template="$NAME": that template is SPOT with --instance-termination-action=DELETE, so a machine built from it can be reclaimed and destroyed mid-run — and the static rung acquires nothing, so there is no re-placement to fall back on. Spot belongs on the rungs that own the machine's whole life (gcp://stopped/, gcp://launch/), where an eviction is a re-placement rather than a worker that stopped existing.

4. The worker's binary

CGO_ENABLED=0 GOOS=linux GOARCH=amd64 go build -o /tmp/steps-linux-amd64 .

Cross-compiles the binary the worker will run. CGO_ENABLED=0 is what makes it a single static file that can be pushed and executed anywhere. Match GOARCH to the machine type — amd64 for e2/n2, arm64 for t2a/c4a.

5. The pipeline

jobs:
- name: all-phases
  plan:
  - task: make
    outputs: [big]
    run: dd if=/dev/urandom of=big/blob bs=1M count=64 2>/dev/null

  - task: on-host
    tags: [gcp]
    inputs: [big]
    outputs: [r1]
    run: |
      wc -c < big/blob > r1/out
      uname -m >> r1/out

  - task: on-a-launched-machine
    tags: [burst]
    inputs: [big]
    outputs: [r2]
    run: uname -m > r2/out

  - task: publish
    inputs: [r1, r2]
    run: |
      echo "static worker:";    cat r1/out
      echo "launched machine:"; cat r2/out

The pipeline names capabilities (tags: [gcp]), never machines — the same split as everywhere else.

steps run \
  --worker "gcp=gcp://$NAME/var/tmp/steps?project=$PROJECT&zone=$ZONE&binary=/tmp/steps-linux-amd64" \
  --worker "burst=gcp://launch/$NAME?project=$PROJECT&zone=$ZONE&binary=/tmp/steps-linux-amd64" \
  pipeline.yml

The invocation names the machines. The parts of that worker URL that matter:

There is deliberately no ?capacity=: the template decides its own provisioning model, so a spot job names a spot template.

6. Tear it down

gcloud compute instances delete "$NAME" --zone="$ZONE" --quiet
gcloud compute instance-templates delete "$NAME" --quiet
gcloud compute firewall-rules delete "$NAME-iap" --quiet

Check for orphans, because a leaked instance is the expensive mistake:

gcloud compute instances list
gcloud compute disks list --filter="-users:*"

The second one matters on its own: a disk that outlives its instance keeps billing with nothing pointing at it — which is exactly what --instance-termination-action=DELETE in the template exists to prevent.

When it does not work

the IAP relay refused the connection (HTTP 403) — your account lacks iap.tunnelInstances.accessViaIAP (grant roles/iap.tunnelResourceAccessor), or the IAP API is disabled on the project.

nothing listening there, or no firewall rule allows the IAP range — the relay reached the VPC and nothing answered: the firewall rule from step 1 is missing or its target tag does not match the instance, or sshd is not running. On a machine acquired seconds ago steps waits this out (sshd is still booting); if it never resolves, it is the firewall.

ssh: unable to authenticate — the project enforces OS Login, which ignores metadata SSH keys. Set enable-oslogin=FALSE in the instance (or template) metadata, as step 2 does.

guest attributes are disabled … set enable-guest-attributes=TRUE … or pin the key with ?hostkey= — the template did not set the metadata key. Either fix the template or pin: ssh-keyscan the machine once from somewhere that can reach it, or read the key out of the serial console log, and put its SHA256:… fingerprint in the URL (URL-encode it — the base64 can contain +).

The pushed binary will not run, on Container-Optimized OS — most COS paths are mounted noexec. Name /var/lib/toolbox in the worker URL's path.

The startup script never installs docker — a --no-address instance has no route to the internet without Cloud NAT. Add NAT to the subnet's region, or bake an image with docker preinstalled.

See also