Cloud Run cold starts: a field guide

What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.

Published

Updated

—

Reading time

14 min

Scale-to-zero is the best and worst thing about Cloud Run. You pay nothing while nobody is calling your service, and then the first person who does waits while Google finds a machine, starts your container and runs your initialization code. For an internal cron endpoint nobody notices. For a login API, a webhook with a tight timeout, or the first request of a user's session, a two-second stall is a bug report. This guide is for engineers running HTTP services on Cloud Run who want to stop guessing about cold starts: what they are made of, how to measure them properly, which fixes are cheap and which cost real money, and how to decide whether the money is worth it.

What a cold start actually is#

A cold start happens whenever Cloud Run has to route a request to a container instance that doesn't exist yet: after scale-to-zero, during a traffic spike that outgrows the current instances, after a new revision is deployed, or when the platform replaces an instance. The request waits for four phases:

The request waits for scheduling, image fetch, container start and your app's init. Only the last one is fully under your control.
  1. Scheduling. Cloud Run picks a host with capacity for your instance's CPU and memory. You mostly can't influence this, beyond not asking for unusual sizes.
  2. Image fetch. The container image is pulled from Artifact Registry. Cloud Run streams image content rather than downloading the whole thing up front, and Google states that image size doesn't affect startup time. What your process reads at startup still has to arrive, though, so a lean image is still good hygiene.
  3. Container start. The sandbox is created and your entrypoint executes. The execution environment (gen1 or gen2, more below) affects this.
  4. Application init. Your runtime boots, loads modules, reads config, perhaps opens connections, and finally listens on $PORT. Cloud Run doesn't send traffic until the port is open or your startup probe passes.

In my experience, phases 1 to 3 are usually the smaller share for a reasonably built image. Phase 4 is where the seconds hide, and it is entirely yours.

Constraints#

What I'm optimizing for, in order:

  • Cost first. Scale-to-zero is the reason many of us are on Cloud Run. I don't want to give it up by default.
  • Predictable p99 on user-facing paths. Average latency hides cold starts. The first request after idle is what users remember.
  • No platform lock-in in application code. Fixes should be good engineering anywhere, not Cloud Run tricks.
  • Measurable. Every change is verified with metrics, not vibes.

Measuring before fixing#

Don't touch the Dockerfile until you know which phase dominates.

Request logs#

Every request to a Cloud Run service produces a request log entry with httpRequest.latency. When a request triggered a new instance, Cloud Run adds a text note to that entry: "This request caused a new container instance to be started and may thus take longer and use more CPU than a typical request." That makes cold requests easy to isolate in Logs Explorer:

Logs Explorertext
resource.type="cloud_run_revision"
resource.labels.service_name="api"
logName:"run.googleapis.com%2Frequests"
textPayload:"caused a new container instance to be started"

Compare their latency with warm requests over the same period and you have the user-facing cost of a cold start.

The startup latency metric#

Cloud Run exports run.googleapis.com/container/startup_latencies, a distribution of how long new instances take to become ready. Chart the p50 and p99 per revision in Metrics Explorer. This is the number I track across deploys: if a dependency bump adds 800 ms of init, it shows up here the same day.

Startup probes#

By default Cloud Run uses a TCP startup probe: the instance is ready as soon as the port accepts connections. If your app opens the port before it can actually serve (common with frameworks that listen first and warm caches afterwards), configure an HTTP startup probe so readiness is honest:

service.yaml (excerpt)yaml
spec:
  template:
    spec:
      containers:
        - image: europe-west1-docker.pkg.dev/my-project/apps/api:latest
          startupProbe:
            httpGet:
              path: /healthz
            initialDelaySeconds: 0
            periodSeconds: 1
            timeoutSeconds: 1
            failureThreshold: 30

A short periodSeconds matters: with a 10-second period, a service ready at 1.1 s may not be marked ready until the next probe fires.

Your own timestamps#

Platform metrics tell you how long. Structured logs tell you where. I log phase markers relative to process start, as JSON so Cloud Logging parses them:

src/startup-timing.tsts
const t0 = performance.timeOrigin;
 
export function mark(phase: string): void {
  // One JSON line per phase; Cloud Logging turns it into jsonPayload fields.
  process.stdout.write(
    JSON.stringify({
      severity: "INFO",
      message: `startup:${phase}`,
      phase,
      msSinceProcessStart: Math.round(performance.now()),
      processStartedAt: new Date(t0).toISOString(),
    }) + "\n",
  );
}

Call mark("modules-loaded"), mark("config-ready") and mark("listening") at the obvious points. The first time I did this on a Node service, "modules-loaded" accounted for most of the init time, which is the next section's subject.

The options#

Here's the full menu, roughly ordered from "do it anyway" to "spend money on it".

LeverWhat it shortensCost impactEffortCaveat
Lazy client initApp initNoneLowFirst request that uses the client pays for connecting
Trim / bundle dependenciesApp init (module loading)NoneMediumBundling needs care with native modules
Smaller base image (slim, distroless)Little directly (images are streamed); fewer files to read at startNoneLowDistroless has no shell for debugging
Startup CPU boostApp initExtra CPU billed during startupTrivialOnly helps CPU-bound init
Execution environment gen1 vs gen2Container startNone directlyTrivialgen2 needed for some features
Concurrency tuningHow often cold starts happenCan lower instance countLowNeeds load testing
Min instancesRemoves cold starts for baseline trafficPaid while idleTrivialSpikes above the minimum still cold start
Instance-based billing (CPU always allocated)Background warm-up, keeps work going between requestsHigher per-instance costTrivialOnly worth it with steady traffic

The decision#

My default for a user-facing API: fix the application first (lazy init, lean dependencies, a slim image), turn on startup CPU boost, set concurrency deliberately, and only then decide on min-instances with an actual number in front of me. Most services I've looked at get most of the improvement from the free levers. Min-instances is a pricing decision, not a performance fix, and I treat it that way.

Implementation#

Lazy client initialization in Node#

The classic mistake is creating every client at module top level. A Firestore client, a Pub/Sub client, a database pool, a secrets fetch: each loads a large dependency tree and some open connections immediately, all before your server listens, and all for requests that may not need them.

The fix is a memoized async getter. The module is loaded only when first needed (dynamic import()), the client is built once, and concurrent first callers share the same promise:

src/clients/lazy.tsts
export function lazy<T>(factory: () => Promise<T>): () => Promise<T> {
  let instance: Promise<T> | undefined;
  return () => {
    if (!instance) {
      instance = factory().catch((err) => {
        instance = undefined; // allow a retry on the next call
        throw err;
      });
    }
    return instance;
  };
}
src/clients/index.tsts
import { lazy } from "./lazy.js";
import { env } from "../config/env.js";
 
export const getFirestore = lazy(async () => {
  const { Firestore } = await import("@google-cloud/firestore");
  return new Firestore({ projectId: env.GCP_PROJECT });
});
 
export const getPubSub = lazy(async () => {
  const { PubSub } = await import("@google-cloud/pubsub");
  return new PubSub({ projectId: env.GCP_PROJECT });
});
src/routes/profile.tsts
import type { FastifyInstance } from "fastify";
import { getFirestore } from "../clients/index.js";
 
export async function profileRoutes(app: FastifyInstance): Promise<void> {
  app.get<{ Params: { id: string } }>("/profiles/:id", async (req, reply) => {
    const db = await getFirestore();
    const snap = await db.collection("profiles").doc(req.params.id).get();
    if (!snap.exists) return reply.code(404).send({ error: "not_found" });
    return snap.data();
  });
}

The trade-off is explicit: the first request that touches Firestore pays for loading and connecting. That's still better than every cold start paying for every client, including the ones the health check and 90% of routes never use. If one client is needed on nearly every request, warm it right after listen() without awaiting it:

src/server.tsts
await app.listen({ host: "0.0.0.0", port: Number(env.PORT) });
mark("listening");
void getFirestore(); // warm in the background; errors surface on first real use

With request-based billing, CPU is only allocated while requests are being processed (and during startup), so background warm-up after listen() may be throttled until the first request arrives. It still helps, because it starts early, but don't count on it finishing.

Avoid heavy top-level imports#

Node resolves and compiles every module in the static import graph before your first line of application code runs. A few habits that matter:

  • Import narrowly. Pull the specific submodule rather than a whole SDK barrel when the library supports it.
  • Keep optional features out of the startup path. PDF generation, image processing, admin-only routes: load them with import() inside the handler.
  • Bundle for production. Tools like esbuild collapse thousands of files in node_modules into one or a few files, which cuts filesystem lookups at boot. Mark packages with native bindings or dynamic requires as external.
  • Measure with --cpu-prof. Running node --cpu-prof dist/server.js locally and opening the profile in Chrome DevTools shows exactly which imports dominate.

Node 22 also offers an on-disk module compile cache via module.enableCompileCache() or NODE_COMPILE_CACHE. On Cloud Run the filesystem is in-memory and per instance, so a cache written at runtime doesn't survive to the next cold start; it only helps if you generate it at build time and ship it in the image. I treat that as an experiment to measure, not a default.

A lean Dockerfile#

A multi-stage build keeps compilers and dev dependencies out of the runtime image. For the runtime stage I use a distroless Node image:

Dockerfiledockerfile
# ---- build ----
FROM node:22-slim AS build
WORKDIR /app
RUN corepack enable
COPY package.json pnpm-lock.yaml ./
RUN pnpm install --frozen-lockfile
COPY tsconfig.json ./
COPY src ./src
RUN pnpm run build \
 && pnpm prune --prod
 
# ---- runtime ----
FROM gcr.io/distroless/nodejs22-debian12:nonroot
WORKDIR /app
ENV NODE_ENV=production
COPY --from=build /app/node_modules ./node_modules
COPY --from=build /app/dist ./dist
COPY --from=build /app/package.json ./
# distroless nodejs images use node as the entrypoint
CMD ["dist/server.js"]

If you bundle with esbuild into a single file, you can often skip copying node_modules entirely except for externals, which shrinks the image further and removes most of the module resolution work at startup.

On base images:

  • Distroless has no shell or package manager, which is smaller and a meaningfully smaller attack surface. You lose docker exec debugging; I consider that acceptable in production.
  • Alpine is small, but it uses musl instead of glibc. Prebuilt native Node modules sometimes don't ship musl binaries, and you find out at runtime. I only use it when I've checked every native dependency.
  • -slim Debian is the pragmatic middle ground if you need a shell.

Platform settings#

Startup CPU boost gives the instance extra CPU while it starts, then drops back to your configured amount:

tune-service.shbash
gcloud run services update api \
  --region=europe-west1 \
  --cpu-boost \
  --execution-environment=gen1 \
  --concurrency=80 \
  --max-instances=20

CPU boost helps exactly the part of init that is CPU-bound: module compilation, JIT warm-up, TLS setup. It won't help if you're waiting on a network call to a secrets store.

Execution environment. Cloud Run offers two execution environments. Gen1 is based on gVisor and is documented as having faster cold starts. Gen2 runs on a microVM with full Linux compatibility and generally faster sustained CPU and network performance, and it's required for things like network file system mounts. If you don't set one, Cloud Run chooses for you. For a small, latency-sensitive HTTP API, gen1 is worth trying; for CPU-heavy or I/O-heavy work, gen2 usually wins after startup. Measure both with the startup latency metric rather than trusting either description.

Billing mode. With request-based billing (historically "CPU only allocated during request processing"), you pay for CPU while requests are in flight. With instance-based billing (historically "CPU always allocated", the --no-cpu-throttling flag), you pay for the instance's whole lifetime but the CPU is available between requests for background work. Instance-based billing doesn't eliminate cold starts on its own; it makes post-listen warm-up and background tasks reliable, and it can be cheaper for services with steady, high utilization. Check the billing settings documentation for the current naming and flags.

Concurrency#

Every request that can be served by an existing instance is a cold start avoided. Cloud Run's default concurrency is 80 requests per instance and the maximum is 1000. Node services doing I/O-bound work often handle far more than one request at a time comfortably, so setting concurrency to 1 (a surprisingly common copy-paste from single-threaded examples) multiplies instance count and cold starts during bursts. Going the other way, a CPU-heavy handler at concurrency 250 makes every request slow. Load-test with a realistic mix, watch per-instance CPU and memory, and pick the highest concurrency that keeps p99 acceptable.

Min instances#

Minimum instances keeps a number of instances warm and ready:

terminalbash
gcloud run services update api --region=europe-west1 --min-instances=1

This removes cold starts for traffic that fits inside the warm instances. It does nothing for a spike that exceeds them, and it won't necessarily cover the first requests to a freshly deployed revision. If deploy-time cold starts matter, deploy with --no-traffic, let the new revision's instances come up, then migrate traffic with gcloud run services update-traffic.

CostMin-instances is a standing charge

An idle min instance is billed continuously. At the time of writing, Cloud Run bills idle minimum instances at a reduced rate compared with active ones under request-based billing, but the charge runs 24/7 whether or not anyone calls you. Multiply it out before you commit: one instance, times its vCPU and memory, times roughly 730 hours a month, times the idle rate on the Cloud Run pricing page, times every service and every region you set it on. A single small instance is usually cheap. --min-instances=2 copied across thirty microservices and two environments is not. Committed use discounts can reduce this for stable baselines.

A pattern I use: min instances only on the handful of services that sit on the critical path of an interactive request (authentication, the main read API), zero everywhere else, and never in staging or preview environments unless I'm specifically measuring latency there.

When latency is worth money#

This is the question that actually matters, and it has an honest answer only with numbers. I frame it like this:

  1. How many requests hit a cold instance? Count them from the request logs above over a week. For a service with steady traffic, it may be a fraction of a percent. For a low-traffic service that idles between calls, it may be most of them.
  2. Who experiences them? A cold start on a background webhook consumer costs nothing. A cold start on a checkout or login request has a conversion cost, even if you can't measure it precisely.
  3. What does removing them cost? Monthly min-instance cost from the calculation above.
  4. What's left after free fixes? If lazy init took cold starts from three seconds to 600 ms, the case for paying drops sharply.

An illustrative example, not a benchmark: a service whose cold start is 1.5 s, hit cold on 2% of requests, all internal, clearly doesn't justify a standing charge. The same service fronting a sign-in flow where the first request of every session is often cold probably does, and one warm instance is typically the cheapest line item in that decision. The point is to write the numbers down, rather than setting min instances everywhere "for performance".

Trade-offs and failure modes#

  • Lazy init moves failures later. A misconfigured database URL used to crash the container at startup, failing the deploy. Now it fails the first real request. Keep a readiness check that validates configuration (not connectivity) at startup, and validate env vars eagerly; they're cheap.
  • Startup probes that are too strict turn slow-but-healthy starts into crash loops. Give failureThreshold × periodSeconds headroom above your p99 startup time.
  • Min instances don't protect against spikes. A burst above the warm pool still cold starts. Concurrency and fast init are the only defense there.
  • Deploys cause cold starts too. Each new revision starts fresh instances. Frequent deploys to a latency-sensitive service mean frequent cold starts unless the new revision is warm before traffic migrates.
  • Instances aren't kept warm forever. Cloud Run may keep idle instances around for a while after traffic stops, but it's not guaranteed and not documented as a fixed window. Don't design around "it stays warm for 15 minutes".
  • Distroless without a debugging plan is painful at 2am. Keep a debug variant of the image, or rely on logs and Cloud Trace.

Checklist#

Checklist

  • Cold requests isolated in request logs; warm vs cold latency compared
  • container/startup_latencies charted per revision, p50 and p99
  • HTTP startup probe configured with a short period and adequate threshold
  • Startup phase markers logged as structured JSON
  • SDK clients created lazily with a memoized async getter; no connections at module top level
  • Heavy, rarely-used modules loaded with dynamic import(); production bundle measured with --cpu-prof
  • Multi-stage Dockerfile, dev dependencies pruned, slim or distroless runtime
  • Startup CPU boost enabled; gen1 vs gen2 measured, not assumed
  • Concurrency set from a load test, not left at 1 by accident
  • Min instances only where the cold-request count and the monthly cost justify it, calculated from the pricing page

When not to do this#

If a service is called by a scheduler, a queue push subscription or another backend with generous timeouts and retries, cold starts are a non-event: leave it at zero, skip the tuning, and spend the time elsewhere. And if a workload needs consistently low single-digit-millisecond latency under continuous heavy load, you're paying for always-on capacity either way, and a platform built around scale-to-zero may not be the right tool; compare it honestly against a small always-on deployment on GKE or a managed instance group.

Share
All articles →

Early-stage systems rarely go broke on traffic. They bleed money while idle. How to design a GCP stack whose cost stays near zero when no one is using it.

13 min

Where the money goes when users upload photos, and a GCS upload-and-variants pipeline that lets a CDN serve small, immutable files for years.

16 min