Keyless CI/CD: Workload Identity Federation for GitHub Actions
Swap service account keys in GitHub secrets for short-lived OIDC tokens, lock the trust to your own repo, and deploy to Cloud Run with no JSON key.
Somewhere in most GitHub organizations there is a secret called GCP_SA_KEY. It holds a service account JSON key that was created once, pasted into repository settings, and never rotated. It can deploy to production, it never expires, and anyone who can exfiltrate it from a compromised dependency, a leaked log line or a careless echo has the same power as your pipeline, from anywhere on the internet, for as long as the key lives. This article is for engineers who own a GitHub Actions pipeline that deploys to Google Cloud and want to delete that secret for good. I'll walk through how the token exchange actually works, set up a pool and provider with gcloud, lock the trust down so only the right repository and branch can use it, deploy to Cloud Run, and cover the errors you will almost certainly hit on the way.
Why the JSON key has to go#
A service account key is a long-lived bearer credential. Google's own guidance on service account keys is blunt about it: if you can avoid keys, avoid them. In CI the problems compound:
- No expiry. User-managed keys are valid until someone deletes them. The default is "forever".
- No binding to context. The key doesn't know it's supposed to be used from a GitHub runner on the
mainbranch. It works equally well from a laptop in another country. - Wide blast radius. CI service accounts tend to accumulate roles over time, because the person who hit a permission error at 11pm added one more role and moved on.
- Rotation is manual and rarely done. Rotating means creating a new key, updating every repo that uses it, and deleting the old one without breaking a release.
Workload Identity Federation (WIF) removes the stored secret entirely. GitHub already issues every workflow job a signed OpenID Connect token describing exactly who it is: which repository, which branch, which environment, which workflow. Google Cloud can be configured to trust that issuer, check the claims, and hand back a short-lived Google credential. Nothing sits in your repository settings that is worth stealing.
Constraints I design for#
Before writing any config, these are the requirements I hold the setup to:
- No long-lived credentials in GitHub, in the repo, or in logs.
- Trust scoped to my organization first, then narrowed per repository. A pool that trusts "any GitHub repository" is worse than a key.
- Production deploys only from a protected context: a specific branch or a GitHub environment with required reviewers.
- Least privilege per pipeline. One service account per deploy target, not one god account for the whole org.
- Auditable. I want to answer "which workflow run deployed revision X" from Cloud Logging.
- Boring to operate. No custom token brokers, no extra infrastructure.
The options#
There are three realistic ways to authenticate a GitHub workflow to Google Cloud.
| Approach | Stored secret | Credential lifetime | Scope control | Operational cost | My take |
|---|---|---|---|---|---|
| Service account JSON key in GitHub secrets | Yes, long-lived | Until deleted | None beyond the SA's roles | Manual rotation, leak response | Retire it |
| WIF with service account impersonation | None | Access token, 1 hour by default | Attribute condition + per-SA workloadIdentityUser binding | One-time setup | Default choice |
| WIF with direct resource access | None | Federated token | Attribute condition + roles granted to the principalSet directly | One-time setup, fewer moving parts | Good for simple, supported cases |
Direct resource access means granting IAM roles straight to the federated identity, skipping the service account. It works for many APIs, but not every Google Cloud product supports federated principals, and anything that needs to "act as" a runtime service account (Cloud Run deploys do) still ends up involving a service account somewhere. I default to impersonation because every tool in the ecosystem understands it and the audit trail is clearer.
How the token exchange works#
Step by step:
- The job, running with
permissions: id-token: write, asks GitHub's OIDC endpoint for a JWT. The token's issuer ishttps://token.actions.githubusercontent.comand its claims includesub,repository,repository_owner,repository_id,ref,environment,workflow_refand more. The full list is in GitHub's OIDC documentation. - The
google-github-actions/authaction sends that JWT to Google's Security Token Service (STS). - STS looks up the workload identity pool provider, verifies the signature against GitHub's published keys, applies your attribute mapping, and evaluates your attribute condition. If the condition fails, the exchange is rejected right there.
- On success STS returns a federated access token. With impersonation, the action then calls the IAM Credentials API to generate an access token for your service account, which succeeds only if the federated identity holds
roles/iam.workloadIdentityUseron that service account. - Subsequent
gcloudcalls and client libraries in the job use the short-lived token. When it expires, it's useless.
Two gates, then: the provider's attribute condition decides which GitHub tokens are accepted at all, and the IAM binding decides which of those identities may become which service account. You want both.
Implementation#
The commands below assume a project with ID my-project, a GitHub organization acme, and a repository acme/api. Replace them with yours.
Step 1: enable the APIs#
gcloud services enable \
iam.googleapis.com \
iamcredentials.googleapis.com \
sts.googleapis.com \
run.googleapis.com \
artifactregistry.googleapis.com \
--project=my-projectForgetting iamcredentials.googleapis.com is the single most common reason the first run fails.
Step 2: create the pool#
A pool is a container for external identities. I create one per trust domain (one for GitHub, a different one for, say, another CI system) rather than one per repository.
gcloud iam workload-identity-pools create github \
--project=my-project \
--location=global \
--display-name="GitHub Actions"Step 3: create the provider, with a condition#
This is the step where the security of the whole setup is decided.
gcloud iam workload-identity-pools providers create-oidc acme-github \
--project=my-project \
--location=global \
--workload-identity-pool=github \
--issuer-uri="https://token.actions.githubusercontent.com" \
--attribute-mapping="\
google.subject=assertion.sub,\
attribute.repository=assertion.repository,\
attribute.repository_owner=assertion.repository_owner,\
attribute.ref=assertion.ref" \
--attribute-condition="assertion.repository_owner == 'acme'"The attribute mapping translates GitHub claims into attributes Google IAM can reason about. google.subject is mandatory and becomes the identity shown in audit logs; the attribute.* keys are yours to define and are what you'll reference in principalSet bindings later. Map only what you'll actually use in conditions or bindings.
The attribute condition is a CEL expression evaluated on every exchange. Here is why it is not optional: GitHub's OIDC issuer is shared by every repository on github.com. The issuer URL is identical for your private monorepo and for a repository an attacker created five minutes ago. Without a condition, the provider accepts tokens from all of them, and the only thing standing between an attacker's workflow and your project is whether your IAM bindings happen to be tight. One overly broad principalSet (for example binding the whole pool with /*) and anyone on GitHub can deploy to your project. Google now pushes hard on this: its GitHub guidance requires an attribute condition when the provider trusts GitHub's issuer, and I would add one even if it didn't.
repository_owner and repository are names. If an organization is renamed or deleted, someone else can register the old name. For high-value projects, map the immutable numeric claims instead, for example attribute.repository_owner_id=assertion.repository_owner_id, and use assertion.repository_owner_id == '123456' in the condition. You can get the ID from gh api orgs/acme --jq .id.
Narrower conditions are fine too. If only two repositories should ever reach this project, say so:
gcloud iam workload-identity-pools providers update-oidc acme-github \
--project=my-project \
--location=global \
--workload-identity-pool=github \
--attribute-condition="assertion.repository_owner == 'acme' && assertion.repository in ['acme/api', 'acme/web']"Step 4: create a deployer service account#
One service account per deploy target keeps blast radius small.
gcloud iam service-accounts create api-deployer \
--project=my-project \
--display-name="Deploys the api service from GitHub"
SA="api-deployer@my-project.iam.gserviceaccount.com"
# Deploy Cloud Run services and push images
gcloud projects add-iam-policy-binding my-project \
--member="serviceAccount:${SA}" --role="roles/run.developer"
gcloud projects add-iam-policy-binding my-project \
--member="serviceAccount:${SA}" --role="roles/artifactregistry.writer"
# Allow it to attach the runtime service account to the Cloud Run service
gcloud iam service-accounts add-iam-policy-binding \
api-runtime@my-project.iam.gserviceaccount.com \
--project=my-project \
--member="serviceAccount:${SA}" --role="roles/iam.serviceAccountUser"Note the last binding is scoped to the single runtime service account, not granted project-wide. A project-level serviceAccountUser lets the deployer act as every service account in the project, which quietly defeats the point.
Step 5: bind the federated identity to the service account#
Now let tokens from one repository impersonate api-deployer. You need the project number here, not the ID:
PROJECT_NUMBER=$(gcloud projects describe my-project --format="value(projectNumber)")
gcloud iam service-accounts add-iam-policy-binding \
api-deployer@my-project.iam.gserviceaccount.com \
--project=my-project \
--role="roles/iam.workloadIdentityUser" \
--member="principalSet://iam.googleapis.com/projects/${PROJECT_NUMBER}/locations/global/workloadIdentityPools/github/attribute.repository/acme/api"The principalSet matches every federated identity whose mapped repository attribute equals acme/api. The two principal formats worth knowing:
principalSet://…/attribute.<name>/<value>matches a group of identities by attribute.principal://…/subject/<subject>matches exactly onegoogle.subject, for examplerepo:acme/api:ref:refs/heads/main.
Step 6: the workflow#
name: deploy
on:
push:
branches: [main]
permissions:
contents: read
id-token: write
jobs:
deploy:
runs-on: ubuntu-latest
environment: production
env:
REGION: europe-west1
IMAGE: europe-west1-docker.pkg.dev/my-project/apps/api:${{ github.sha }}
steps:
- uses: actions/checkout@v7
- id: auth
uses: google-github-actions/auth@v3
with:
workload_identity_provider: projects/123456789012/locations/global/workloadIdentityPools/github/providers/acme-github
service_account: api-deployer@my-project.iam.gserviceaccount.com
- uses: google-github-actions/setup-gcloud@v3
- name: Build and push
run: |
gcloud auth configure-docker europe-west1-docker.pkg.dev --quiet
docker build -t "$IMAGE" .
docker push "$IMAGE"
- uses: google-github-actions/deploy-cloudrun@v3
with:
service: api
region: ${{ env.REGION }}
image: ${{ env.IMAGE }}
flags: --service-account=api-runtime@my-project.iam.gserviceaccount.comid-token: write is what allows the job to request the OIDC token at all. Setting permissions at the workflow level also drops every permission you don't list, which is a good default in its own right. The workload_identity_provider value is the full resource name with the project number; gcloud iam workload-identity-pools providers describe acme-github --location=global --workload-identity-pool=github --format="value(name)" prints it exactly. The resource name isn't a secret, so I keep it in plain YAML or a repository variable rather than in secrets.
For third-party actions in a deploy pipeline, I pin to a full commit SHA rather than a tag. A deploy job with an identity token is exactly the job you don't want a compromised tag to run in.
Step 7: restrict by branch and environment#
Repository-level trust is a start, but most teams need "anyone can run CI, only main can deploy to production". Be aware of one quirk: when a job references a GitHub environment, the sub claim becomes repo:acme/api:environment:production instead of repo:acme/api:ref:refs/heads/main. You can use that to your advantage, because environments carry protection rules (required reviewers, allowed branches) that GitHub enforces before the job even starts.
The cleanest pattern I've found is to map the environment and bind production to it:
# Add environment to the mapping (update keeps other flags as they are)
gcloud iam workload-identity-pools providers update-oidc acme-github \
--project=my-project --location=global --workload-identity-pool=github \
--attribute-mapping="\
google.subject=assertion.sub,\
attribute.repository=assertion.repository,\
attribute.repository_owner=assertion.repository_owner,\
attribute.ref=assertion.ref,\
attribute.environment=assertion.environment"
# Only jobs running in the protected 'production' environment of acme/api
gcloud iam service-accounts add-iam-policy-binding \
api-deployer-prod@my-project.iam.gserviceaccount.com \
--project=my-project --role="roles/iam.workloadIdentityUser" \
--member="principal://iam.googleapis.com/projects/${PROJECT_NUMBER}/locations/global/workloadIdentityPools/github/subject/repo:acme/api:environment:production"Point the production workflow's service_account at api-deployer-prod, then pair that with a GitHub environment rule that only allows deployments from main, and a pull request from a fork has no route to the production service account: it can't enter the environment, so its token never carries the matching subject. For staging, a principalSet on attribute.ref/refs/heads/main or a separate staging service account bound to the repository is usually enough.
If staging and production live in separate Google Cloud projects, give each project its own pool and provider. Then a mistake in the staging bindings can't grant anything in production, and the conditions stay short enough to read in a review.
Troubleshooting#
The failure messages are terse. These are the ones I see most, and what they usually mean.
| Symptom | Likely cause | Fix |
|---|---|---|
Unable to get ACTIONS_ID_TOKEN_REQUEST_URL or the auth step says OIDC isn't available | Missing id-token: write, or the workflow runs from a fork PR where GitHub withholds the token | Add the permission at workflow or job level; don't deploy from fork PRs |
The given credential is rejected by the attribute condition | The token's claims don't satisfy the CEL condition | Compare the condition with the actual claims; casing of org names and the environment vs ref subject quirk are common culprits |
Permission 'iam.serviceAccounts.getAccessToken' denied | The workloadIdentityUser binding is missing, points at the wrong pool, or uses the project ID instead of the number | Re-check the principalSet string character by character |
invalid_target or "pool/provider does not exist" | Typo in workload_identity_provider, or the provider is disabled/deleted | Print the name with providers describe --format="value(name)" |
IAM Service Account Credentials API has not been used | API not enabled | Enable iamcredentials.googleapis.com |
| Works in one run, fails in the next minute after a change | IAM propagation | Wait a few minutes; IAM changes are eventually consistent |
Deploy step fails with iam.serviceaccounts.actAs denied | Deployer lacks serviceAccountUser on the runtime SA | Grant it on that specific service account |
To see what GitHub actually put in the token, decode it inside a throwaway workflow on a non-production branch. The auth action has a debug mode, and you can also fetch the token yourself with curl using ACTIONS_ID_TOKEN_REQUEST_URL and ACTIONS_ID_TOKEN_REQUEST_TOKEN and print only the payload claims. Never print the full token in a log, even a short-lived one.
Auditing#
Keyless doesn't mean invisible. Two log sources tell the story:
- STS token exchanges show up in Cloud Audit Logs under
sts.googleapis.com, including which provider was used and the mapped subject. - Service account impersonation is logged by
iamcredentials.googleapis.comasGenerateAccessToken. These are Data Access logs, which are off by default for most services, so turn them on for IAM Credentials in the projects you care about.
A Logs Explorer query I keep saved:
protoPayload.serviceName="iamcredentials.googleapis.com"
protoPayload.methodName="GenerateAccessToken"
protoPayload.authenticationInfo.principalSubject:"workloadIdentityPools/github"The principalSubject contains the GitHub subject, so you can tie a token back to a repository and environment, and from the Cloud Run revision's creation time back to the workflow run.
Workload Identity Federation itself has no charge, and STS and IAM Credentials calls aren't billed per request at the time of writing. The cost that does show up is logging: Data Access audit logs count toward Cloud Logging ingestion, which is billed per GiB beyond the free allotment (see Cloud Logging pricing). For a CI pipeline that's a handful of entries per deploy, so it's negligible. Enabling Data Access logs for every service in a busy project is a different story, so scope it to IAM Credentials.
Once federation works, close the old door. Delete the JSON keys you replaced, remove the GitHub secrets that held them, and enforce the iam.managed.disableServiceAccountKeyCreation organization policy constraint (the managed successor to the legacy iam.disableServiceAccountKeyCreation, which organizations created since May 2024 already enforce by default) so nobody quietly recreates one next quarter.
Trade-offs and failure modes#
- Bindings become the security boundary. A condition that trusts your whole org plus a
principalSetonrepository_ownermeans any repository in the org can impersonate that service account, including a new one created by a junior engineer as a test. Bind to repositories or subjects, not owners. - Environment subjects change the
subclaim. Addingenvironment:to a job that previously authenticated by branch changes its subject. Existingprincipal://…/subject/…ref…bindings stop matching. Plan the migration. - Reusable workflows complicate the picture. When a job calls a reusable workflow in another repository,
job_workflow_reftells you which workflow actually ran. If you centralize deploy logic, consider conditioning on it, so only your vetted deploy workflow can mint production credentials. - Self-hosted runners are a different trust story. The OIDC token proves which workflow asked for it, not that the runner machine is clean. A persistent self-hosted runner shared across repositories can leak credentials between jobs. Use ephemeral runners.
- The token lives for the whole job. A malicious step later in the job can use the credential. Keep deploy jobs short and don't run untrusted code (test suites of arbitrary PRs, for example) in the same job that authenticates.
- Break-glass is gone. With no key in a drawer, emergency deploys go through a human with their own IAM access. That's a feature, but write the runbook before you need it.
Checklist#
Checklist
- Required APIs enabled: IAM, IAM Credentials, STS
- One pool per trust domain, provider issuer set to
https://token.actions.githubusercontent.com - Attribute condition present and restricted to your org (ideally by
repository_owner_id) or explicit repositories - Attribute mapping includes only claims you use:
repository,repository_owner,ref,environmentas needed - One deployer service account per target, with
serviceAccountUserscoped to the runtime service account only roles/iam.workloadIdentityUserbound to a specific repositoryprincipalSetor subject, never the whole pool- Production bound to a protected GitHub environment subject
- Workflow sets
permissions: id-token: writeand nothing broader than needed; actions pinned to SHAs - Data Access audit logs enabled for
iamcredentials.googleapis.com - Old JSON keys deleted, GitHub secrets removed, key-creation org policy enforced
When not to do this#
If your pipeline runs somewhere that can't produce a verifiable OIDC token (an old on-premises CI server with no identity provider, for instance) federation has nothing to federate; fix the runner's identity story first, or use attached service accounts if the build already runs on Google Cloud. Likewise, if your builds already run in Cloud Build, they have a Google-native identity and you don't need GitHub-to-Google federation at all. For everything else that deploys from GitHub Actions, I no longer consider a stored key an acceptable default.
Related articles
All articles →How I structure Google Cloud IAM so people get access through groups, workloads get one keyless service account each, and every grant is small and auditable.
13 min
What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.
14 min
One validated config module per service, secrets bound at deploy time, and no fallbacks in code: a config setup that fails loudly instead of leaking quietly.
15 min