Scale-to-zero as a budget strategy

Early-stage systems rarely go broke on traffic. They bleed money while idle. How to design a GCP stack whose cost stays near zero when no one is using it.

Published

Updated

—

Reading time

13 min

Look at the first few months of invoices for most new systems and the pattern is the same. The line items that grow with usage are small, and the ones that don't change with usage take most of the budget. A managed database running all night for nobody, a load balancer forwarding rule, a NAT gateway, a pair of VPC connector VMs, three environments that each replicate all of the above. Traffic is the cheap part of the bill until you have a lot of it. This article is for engineers and founders building on Google Cloud who want the monthly bill to follow usage instead of the architecture diagram. I'll go through what really scales to zero on GCP, the fixed costs that hide next to it, how I put guardrails around spend, and how I keep every component ready to be switched to warm capacity the day it needs to be.

The constraint: idle time is most of your time#

A system with a few hundred active users is idle most of the day. Requests arrive in bursts, background jobs run for seconds, and between them nothing is happening. Fixed-capacity infrastructure charges for every one of those idle hours, and a small team often runs three copies of it (dev, staging, prod), each with the same idle floor.

Scale-to-zero reverses that. You pay per request, per vCPU-second, per document read, per GiB published. When there's no traffic, the bill drops to storage plus a small remainder. The trade-off is well known: cold starts, per-unit prices that are higher than reserved capacity at sustained load, and a design that has to tolerate instances appearing and disappearing. Early on, I think that trade-off is almost always worth taking, because idle cost is certain and load is still hypothetical.

I've applied this to a large-scale social platform I'm building, where cost has priority over latency until usage justifies the reverse. Everything below is written so it applies to any GCP workload with similar constraints.

What scales to zero on GCP, and what doesn't#

The distinction that matters is whether a service has a per-hour floor that you pay with zero traffic. Pricing details change, so treat the links as the source of truth at the time of writing.

ComponentScales to zero?Idle cost floorNotes
Cloud Run services (request-based billing, min instances 0)YesNone beyond image storageFree monthly allowance; cold starts on first request
Cloud Run functionsYesNoneSame runtime model as Cloud Run
Cloud Run jobsYesNoneBilled only while tasks run
FirestoreYesStored data onlyPer-operation billing; daily free quota
Pub/SubYesRetained message storage, if enabledThroughput-billed; free monthly allowance
Cloud StorageYesStored data onlyEgress and operations billed per use
Cloud SchedulerNearlySmall per-job monthly feeA few jobs free per billing account
Secret ManagerNearlyPer active secret versionDestroy old versions
Cloud SQLNovCPU + memory + storage, 24/7Can be stopped; storage still bills
AlloyDB, SpannerNoInstance/node capacity, 24/7Spanner has smaller granular sizes, but still a floor
MemorystoreNoInstance capacity, 24/7Billed by provisioned GiB
GKE Standard nodesNoNode VMs, 24/7Plus cluster management fee
GKE AutopilotPartlyCluster management feePods bill per request; fee covered by a free-tier credit for one cluster
External Application Load BalancerNoPer forwarding rule, hourlyBefore any data processed
Cloud NATNoHourly gateway/instance charge + per-GiBOften added only for static egress IPs
Serverless VPC Access connectorNoMinimum connector VMs, 24/7Direct VPC egress avoids connector VMs

Two rows need more comment. Firestore and Cloud Storage aren't zero while idle, because you still pay for bytes at rest, but that cost tracks data you actually have, not capacity you reserved. And "can be stopped" (Cloud SQL) isn't the same as "scales to zero": a stopped database doesn't serve requests, so it only helps in environments where someone is willing to start it by hand.

The hidden fixed costs#

The obvious always-on components are easy to spot. These are the ones that show up on the bill without anyone having decided on them.

Load balancer forwarding rules. A global external Application Load Balancer is the standard way to put a custom domain, Cloud Armor and a CDN in front of Cloud Run. Its forwarding rules are billed hourly whether or not traffic flows. At the time of writing, the first five rules in a project share a single flat hourly rate that comes to roughly the price of a small VM per month, per environment. Cloud Run's own domain mapping or Firebase Hosting in front of Cloud Run can avoid it for simple cases, with fewer features.

Cloud NAT. It usually arrives because a partner wants to allowlist a static egress IP. A NAT gateway has an hourly component and a per-GiB data-processing charge on top of normal egress. If only one integration needs a fixed IP, route only that service through the VPC (--vpc-egress=private-ranges-only for everything else), and ask whether the partner will accept mTLS or a signed request instead.

Serverless VPC Access connectors. Connectors run on a minimum number of small VMs that are always on. Direct VPC egress lets Cloud Run send traffic into a VPC without a connector and without its idle VMs. For new services I use Direct VPC egress unless a documented limitation rules it out.

Min instances. One --min-instances=1 on a service with 1 vCPU and 512 MiB is a small cost. Ten services, each with one warm instance, in three environments, is thirty instances billed around the clock. Idle min instances bill at a reduced rate under request-based billing, but not zero.

Log ingestion. Cloud Logging bills ingestion per GiB above a monthly free allotment per project. A chatty health check, request logging duplicated at the app and platform level, or debug logs left on in staging can make logs the biggest line item on a quiet system. Exclusions on the _Default sink are free and take effect immediately.

Artifact Registry and retained data. Every deploy pushes an image, and old images keep costing storage until cleanup policies delete them. The same applies to Pub/Sub topic retention, snapshot storage, and Cloud Storage buckets without lifecycle rules.

Where the idle floor lives: the right-hand boxes bill every hour, whatever traffic the left-hand boxes see.

The decision#

My default for a new system on GCP is:

  1. Compute: Cloud Run services with min-instances=0, request-based billing, a conservative max-instances, and startup CPU boost on. Async work on Cloud Run jobs or Pub/Sub push to Cloud Run.
  2. Data: Firestore for operational data unless the access patterns need relational joins or strong multi-row constraints that Firestore can't model. If they do, Cloud SQL, accepted as a known floor, in the smallest shared-core tier for non-production environments.
  3. Messaging: Pub/Sub with push subscriptions to Cloud Run, so consumers scale to zero along with producers.
  4. Edge: Firebase Hosting or Cloud Run domain mapping until I need Cloud Armor, a CDN, or multi-region routing. Then the load balancer, as a deliberate decision with a line in the cost model.
  5. Networking: Direct VPC egress when a VPC is needed at all. NAT only when a real integration demands a static IP.
  6. Guardrails: billing budgets with forecast alerts, max-instances on every service, log exclusions from day one.

Each of these can be moved to warm capacity later without re-architecting, and that's what makes the strategy safe.

Implementation#

Deploy Cloud Run for zero idle cost#

deploy.shbash
gcloud run deploy api \
  --image=europe-west1-docker.pkg.dev/my-project/app/api:2026-09-30 \
  --region=europe-west1 \
  --no-allow-unauthenticated \
  --min-instances=0 \
  --max-instances=10 \
  --cpu-throttling \
  --cpu-boost \
  --concurrency=80 \
  --cpu=1 --memory=512Mi \
  --timeout=60s

--cpu-throttling selects request-based billing: CPU is allocated only while a request is being processed, and you aren't billed between requests. --cpu-boost gives extra CPU during startup, which shortens cold starts at no idle cost. --max-instances is your first cost guardrail. A traffic spike, a retry loop, or a bot can't scale you past it, and it also protects downstream databases from a connection storm. Keep --concurrency high for I/O-bound Node or Go services. Every request that shares an instance is one fewer instance billed.

Make "warm" a variable, not a redesign#

The point of designing for zero is that it's reversible. I keep warm capacity as a single per-environment setting, so going from zero to warm is a one-line change and a plan, not a migration.

infra/cloud_run.tfhcl
variable "warm" {
  description = "Keep one instance warm and allocate CPU always (latency over cost)."
  type        = bool
  default     = false
}
 
resource "google_cloud_run_v2_service" "api" {
  name     = "api"
  location = "europe-west1"
  ingress  = "INGRESS_TRAFFIC_ALL"
 
  template {
    containers {
      image = var.api_image
      resources {
        limits            = { cpu = "1", memory = "512Mi" }
        cpu_idle          = !var.warm # true = request-based billing
        startup_cpu_boost = true
      }
    }
    scaling {
      min_instance_count = var.warm ? 1 : 0
      max_instance_count = 10
    }
    max_instance_request_concurrency = 80
  }
}

The same approach works for the other components. Keep the data access layer behind an interface so a hot path can move from Firestore to a cached read later. Put the consumer behind a Pub/Sub subscription so it can move from push to pull on a warm worker pool without the producer knowing. Keep the domain on a DNS record you control so the load balancer can be added in front without a migration. None of these costs anything today, and each one saves you a rewrite later.

Budgets and alerts#

A Cloud Billing budget sends notifications; it doesn't cap spend. Set it anyway, and include a forecast threshold so you're warned before the month closes rather than after.

infra/budget.shbash
BILLING_ACCOUNT="000000-AAAAAA-BBBBBB"
 
gcloud billing budgets create \
  --billing-account="$BILLING_ACCOUNT" \
  --display-name="staging-monthly" \
  --budget-amount=50EUR \
  --filter-projects="projects/my-project-staging" \
  --threshold-rule=percent=0.5 \
  --threshold-rule=percent=0.9 \
  --threshold-rule=percent=1.0,basis=forecasted-spend \
  --notifications-rule-pubsub-topic="projects/my-project-ops/topics/billing-alerts"

Wiring the budget to a Pub/Sub topic lets you act on it programmatically: post to chat, open an incident, or in non-production projects, lower max-instances or disable billing on the project. Disabling billing stops resources and can delete them, so I only automate it for sandbox projects. Budget data also lags actual usage by hours, so a budget is a smoke alarm, not a circuit breaker. The real hard limits are max-instances, API quota overrides on per-call APIs, and application-level rate limits.

Cut log ingestion at the sink#

infra/logging.shbash
gcloud logging sinks update _Default \
  --add-exclusion=name=drop-health-checks,filter='resource.type="cloud_run_revision" AND httpRequest.requestUrl:"/healthz"'
 
gcloud logging sinks update _Default \
  --add-exclusion=name=drop-debug-staging,filter='severity<=DEBUG'

Only add the second exclusion in non-production projects. In production, keep debug logging off in the application instead, so you can turn it back on during an incident without editing the sink.

A simple monthly cost model#

Before building anything, I write the cost model as a small spreadsheet with two columns per component: an idle floor (what it costs with zero traffic) and a unit cost (what each extra thousand users adds). The numbers below are illustrative, rounded, and built to show the shape of the model. They are not quotes. Rebuild them from the pricing pages linked above for your region and date.

Component (one environment)Always-on design (illustrative €/month)Scale-to-zero design (illustrative €/month)
API compute2 warm instances: 60Cloud Run, min 0, within free allowance: 0–5
DatabaseSmall Cloud SQL HA instance: 90Firestore, low volume: 0–5
CacheSmallest Memorystore instance: 35None (cache in-process or in Firestore): 0
EdgeLoad balancer forwarding rules: 20Domain mapping / Hosting: 0
Egress IPCloud NAT gateway: 5–10Not needed: 0
VPC accessConnector, minimum VMs: 10–15Direct VPC egress: 0
Logs, images, secrets5–155–15
Idle floor, one environment~225–245~5–25
Three environments~675–735~15–75
CostRead the shape, not the numbers

The illustrative totals aren't the point. What matters is that the left column costs the same at zero users and at a thousand, while the right column rises with usage from near zero. Multiply by three environments and a year, and the gap funds a lot of engineering time. Model it for your own stack before you commit, and revisit it when monthly active users change by an order of magnitude.

Add the usage-driven lines separately: requests per user per day, document reads per request, Pub/Sub bytes per event, log bytes per request. That second half of the model shows where scale-to-zero stops being cheaper, which is the next section.

Trade-offs and failure modes#

Cold starts are real. A Node service with a heavy dependency graph can take a second or more to serve its first request after scaling from zero. Startup CPU boost helps, and so do smaller images, lazy-loading rarely used modules, and not opening every client connection at boot. If the p99 of the first request matters (a login page, a payment callback), that service is the first candidate for warm = true.

Per-unit prices cross over. Serverless is priced for elasticity. A service busy at 60–70% of an instance around the clock is cheaper on instance-based billing, committed use discounts, or a VM than on request-based billing. Watch the utilisation of your busiest services, and switch them to warm or instance-based billing when the model says so. Only those services, not everything.

Firestore rewards careful modelling and punishes the opposite. Per-operation billing is ideal at low volume. An unbounded listener, a feed built by fan-out reads, or a page that reads 200 documents to show 20 can make it the most expensive component at scale. Scale-to-zero doesn't protect you from wasteful queries. It just gives you a bill that shows them clearly.

Connection-heavy dependencies. Scaling from zero to fifty instances in a few seconds means fifty new connection pools. A Cloud SQL instance with a modest connection limit will refuse connections before Cloud Run reaches its limit. Cap max-instances, keep pools small per instance, and use the Cloud SQL Auth Proxy or language connectors with a pool sized to max_connections / max_instances.

Runaway scale is also scale-to-zero's failure mode. The same elasticity that makes idle cheap makes a bug expensive: a retry loop between two services, or a Pub/Sub subscription without backoff redelivering against a failing endpoint, scales out until it hits max-instances. Set limits everywhere, and alert on instance count as well as on spend.

Checklist#

Checklist

  • Every Cloud Run service has min-instances=0, request-based billing, and an explicit max-instances
  • Startup CPU boost is on and images are slim
  • Warm capacity is a per-environment variable, not a code change
  • Direct VPC egress instead of connectors; NAT only for a named integration that requires a static IP
  • Load balancer added only when Cloud Armor, CDN or multi-region routing is actually needed
  • Non-production environments carry no always-on databases, or they are stopped outside working hours
  • Billing budgets per project with a forecast threshold, routed to Pub/Sub
  • Log exclusions for health checks and noisy debug output
  • Artifact Registry cleanup policies, bucket lifecycle rules, and old secret versions destroyed
  • A two-column cost model (idle floor, unit cost) reviewed at each order-of-magnitude change in usage

When not to do this#

Don't use scale-to-zero on a path with a tight latency SLO that a cold start would break, such as synchronous payment callbacks with short provider timeouts, real-time bidding, or interactive APIs where users notice a slow first request. Keep those warm and accept the floor. It's also the wrong model for workloads that hold long-lived connections (WebSocket gateways, game or chat servers with persistent sessions, streaming ingestion), because an instance holding an open connection is billed as active anyway, so the savings disappear. And once a service runs at steady, high utilisation around the clock, reserved or instance-based capacity is simply cheaper. Treat scale-to-zero as the starting point, and move each component off it when its own numbers say so.

Share
All articles →

What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.

14 min

Where the money goes when users upload photos, and a GCS upload-and-variants pipeline that lets a CDN serve small, immutable files for years.

16 min