DevOps & CI/CD

Configuration and secrets in multi-service Node apps

One validated config module per service, secrets bound at deploy time, and no fallbacks in code: a config setup that fails loudly instead of leaking quietly.

Published

Updated

—

Reading time

15 min

Configuration bugs rarely look like bugs. A service boots with PORT unset and listens on a default nobody chose. A JWT secret silently falls back to "dev" in a staging environment that is reachable from the internet. A third-party API key gets rotated, one of twelve services is still pinned to the old version, and the only symptom is a trickle of 401s that someone eventually blames on the vendor. None of these crash a unit test. All of them have cost real teams real incidents. This article is for engineers running more than a handful of Node services (the examples use Cloud Run and Secret Manager on Google Cloud, but the patterns carry over to any platform) who want configuration that fails loudly at boot, secrets that never pass through CI, and a way to tell whether fifteen services agree about what environment they are in.

The constraints I design for#

Before picking tools, I write down what the setup has to survive. For a multi-service Node backend these are the ones that matter to me:

  • Many services, few people. Every service must read config the same way, or nobody will remember where the special cases are.
  • Fail at boot, not at first request. A missing variable should stop the container from ever receiving traffic. Cloud Run will keep the previous revision serving if the new one never becomes ready, which is exactly the behaviour I want.
  • No secret exists in source, in images, or in CI logs. Not even a "harmless" development default.
  • Rotation must not require a code change. Ideally it does not require a redeploy of every consumer either.
  • Least privilege per secret. The image-processing service has no business reading the payment provider's webhook key.
  • Auditable. At any moment I want to answer "which services read STRIPE_KEY, and which version?" with one command.

The Twelve-Factor App guidance still holds up here: config that varies between deploys lives in the environment, strictly separated from code. What twelve-factor does not tell you is how to keep that environment honest once you have a dozen services and three environments.

Where config can live: options compared#

There are more ways to get a value into a Node process than most teams realise, and they are not equivalent.

ApproachGood forWeaknessMy verdict
Hard-coded constantsTruly static values (HTTP timeouts, page sizes)Changing them means a releaseFine for non-secret, non-environment values
Plain env vars (--set-env-vars)Non-secret per-environment values: URLs, project IDs, feature flagsVisible to anyone who can view the service; easy to driftDefault for non-secrets
.env file loaded at runtimeLocal developmentEnds up committed, copied into images, or shared on chatLocal only, never deployed
Secret Manager exposed as env varSecrets that change rarelyValue is fixed at instance startDefault for secrets, with pinned versions
Secret Manager mounted as a volumeSecrets you want to rotate without redeploying; multi-line material such as PEM keysYour code must read a file, and possibly re-read itUse when rotation without a deploy matters
Fetching via the Secret Manager API in codeDynamic or per-tenant secretsExtra latency, per-access cost, SDK dependency at bootRarely worth it for static service config

The rule I follow: non-secret values are plain environment variables, secrets are Secret Manager references bound by the platform, and application code never talks to Secret Manager directly unless there is a genuinely dynamic need.

The decision: one config module per service#

Every service gets exactly one file that is allowed to touch process.env. It reads the environment, validates it against a schema, coerces types, and exports a frozen, typed object. Everything else in the codebase imports that object. If the environment is wrong, the module throws during import, the process exits non-zero, and the platform never routes traffic to it.

This single rule does three things at once. It makes the full list of a service's configuration discoverable in one file. It turns "undefined is not a function" at 3 a.m. into a readable startup error. And it gives you a lint target: a rule that bans process.env everywhere except src/config/env.ts is trivial to write and catches most of the drift before review.

A TypeBox-based config module#

I use TypeBox for new services because the schema is also the TypeScript type, and it is already the schema library of choice for Fastify. The module below is complete and runnable with @sinclair/typebox 0.34.

src/config/env.tsts
import { Type, type Static } from "@sinclair/typebox";
import { Value } from "@sinclair/typebox/value";
 
const EnvSchema = Type.Object({
  NODE_ENV: Type.Union([
    Type.Literal("development"),
    Type.Literal("test"),
    Type.Literal("production"),
  ]),
  PORT: Type.Integer({ minimum: 1, maximum: 65535, default: 8080 }),
  LOG_LEVEL: Type.Union(
    [Type.Literal("debug"), Type.Literal("info"), Type.Literal("warn"), Type.Literal("error")],
    { default: "info" },
  ),
  GCP_PROJECT_ID: Type.String({ minLength: 1 }),
  PUBLIC_BASE_URL: Type.String({ pattern: "^https?://" }),
  // Secrets: required, no defaults, ever.
  JWT_SECRET: Type.String({ minLength: 32 }),
  DATABASE_URL: Type.String({ minLength: 1 }),
});
 
export type Env = Static<typeof EnvSchema>;
 
function loadEnv(source: NodeJS.ProcessEnv): Env {
  const raw: Record<string, unknown> = {};
  for (const key of Object.keys(EnvSchema.properties)) {
    const value = source[key];
    // Treat empty strings as missing, and strip the trailing newline
    // that sneaks into secrets created with `echo`.
    if (value !== undefined && value.trim() !== "") raw[key] = value.replace(/\r?\n$/, "");
  }
 
  const withDefaults = Value.Default(EnvSchema, raw);
  const converted = Value.Convert(EnvSchema, withDefaults);
 
  if (!Value.Check(EnvSchema, converted)) {
    const problems = [...Value.Errors(EnvSchema, converted)].map(
      (e) => `  ${e.path || "/"}: ${e.message}`,
    );
    // Only paths and messages: never print the values themselves.
    throw new Error(`Invalid environment configuration:\n${problems.join("\n")}`);
  }
  return Object.freeze(converted as Env);
}
 
export const env = loadEnv(process.env);

A few details are deliberate:

  • Only declared keys are copied. The function iterates the schema, not process.env, so a typo like DATABSE_URL shows up as "/DATABASE_URL: Expected required property" rather than being silently ignored.
  • Value.Convert handles string-to-number coercion. Everything in the environment is a string; the schema says PORT is an integer, so "8080" becomes 8080 and "eighty" fails validation.
  • Error messages carry paths, not values. Startup errors end up in your log sink, and a validation error that echoes a malformed secret has just published it.
  • The throw is fine here. This runs at module load, before the HTTP server exists, so a plain Error that kills the process is exactly the intended behaviour.

If you do not want a dependency#

For a small worker, a hand-written validator is perfectly reasonable. The important property is not the library, it is that validation happens once, at startup, in one place.

src/config/env.tsts
function required(name: string): string {
  const v = process.env[name]?.trim();
  if (!v) throw new Error(`Missing required env var ${name}`);
  return v;
}
 
function int(name: string, fallback: number): number {
  const v = process.env[name];
  if (v === undefined || v === "") return fallback;
  const n = Number(v);
  if (!Number.isInteger(n)) throw new Error(`${name} must be an integer`);
  return n;
}
 
export const env = Object.freeze({
  port: int("PORT", 8080),
  projectId: required("GCP_PROJECT_ID"),
  webhookSecret: required("WEBHOOK_SECRET"),
});

Notice what int gets a fallback and what required does not. Defaults are acceptable for tuning knobs. They are never acceptable for credentials.

No secret fallbacks, not even for development#

This line appears in a surprising number of production codebases:

tsts
const secret = process.env.JWT_SECRET ?? "dev-secret";

It is a vulnerability, not a convenience. The day the variable is missing in a deployed environment (a renamed secret, a copy-pasted deploy command, a new region), the service starts happily and signs tokens with a string that is sitting in your Git history. Anyone who has read the repository can now mint valid sessions. The same applies to database passwords, HMAC keys, and API keys: a missing credential must be a boot failure. If developers need a value locally, it goes in their untracked .env file.

Local development with .env files#

Node has had built-in .env loading since 20.6 (--env-file), and 22.9 added --env-file-if-exists, so dotenv is optional now. See the --env-file documentation:

package.jsonjson
{
  "scripts": {
    "dev": "node --env-file-if-exists=.env --watch dist/server.js"
  }
}

The conventions I hold teams to:

  • .env and .env.* are in .gitignore, with an explicit !.env.example exception.
  • .env.example is committed, lists every key the config module declares, and contains placeholders only (JWT_SECRET=replace-with-32+-random-chars).
  • A CI check compares the keys in .env.example with the keys in the schema, so the example never goes stale.
  • A secret scanner (gitleaks, GitHub secret scanning, or similar) runs on every push. Someone will eventually commit a real .env; the scanner is how you find out in minutes instead of months.
  • Deployed environments never read a .env file. The Dockerfile has a .dockerignore entry for it.

Secret Manager on Cloud Run#

On Google Cloud, secrets belong in Secret Manager and are bound to Cloud Run services by reference. The service definition says "expose version 3 of jwt-secret as JWT_SECRET"; the platform fetches the value using the runtime service account's identity when an instance starts. The value never appears in the deploy command, the image, or the CI job.

CI deploys references; only the runtime identity can read values.

Creating secrets without the trailing newline#

The most common Secret Manager bug I have seen is not a security issue at all: it is a newline. echo "value" | gcloud secrets create ... stores value\n, and then an HMAC comparison fails for no visible reason. Use printf:

terminalbash
printf '%s' "$(openssl rand -base64 48)" | \
  gcloud secrets create jwt-secret \
    --replication-policy=user-managed \
    --locations=europe-west1 \
    --data-file=-

User-managed replication in a single region is also a data-residency choice; automatic replication is simpler if you do not care where copies live.

Per-secret IAM, not project-wide#

Grant roles/secretmanager.secretAccessor on each secret to the specific runtime service account that needs it. Granting the role at project level means every service can read every secret, which quietly defeats the point of having separate identities.

terminalbash
gcloud secrets add-iam-policy-binding jwt-secret \
  --member="serviceAccount:auth-api@my-project.iam.gserviceaccount.com" \
  --role="roles/secretmanager.secretAccessor"

Give each Cloud Run service its own service account. The default compute service account has broad permissions and is shared by everything that did not ask for something better.

Env var vs volume, pinned vs latest#

Cloud Run supports two ways to expose a secret: as an environment variable or as a file in a mounted volume.

terminalbash
gcloud run deploy auth-api \
  --image=europe-west1-docker.pkg.dev/my-project/services/auth-api:1.42.0 \
  --service-account=auth-api@my-project.iam.gserviceaccount.com \
  --set-env-vars=NODE_ENV=production,GCP_PROJECT_ID=my-project \
  --set-secrets=JWT_SECRET=jwt-secret:3,DATABASE_URL=database-url:7,/secrets/signing/key.pem=signing-key:latest

The last entry is the volume form: the value appears as a file at /secrets/signing/key.pem. The trade-offs:

  • Env var secrets are resolved when an instance starts. A new version does nothing until instances are replaced. Google's documentation recommends pinning a specific version for this mode, and I agree: with latest, two instances of the same revision can start minutes apart and hold different values. Pinning makes every rotation an explicit deploy that you can roll back.
  • Volume secrets referencing latest are read at file-read time. That enables rotation without a redeploy, but only if your code re-reads the file rather than caching it forever. It also suits multi-line material such as PEM keys, which is awkward in an env var.

Two flag details catch people. --set-secrets replaces the entire set of secrets on the service; use --update-secrets to add or change one without dropping the others. And --set-env-vars splits on commas, so a value such as ALLOWED_ORIGINS=https://a.example,https://b.example is parsed as two variables, the second one invalid. Change the delimiter using gcloud's escaping syntax:

terminalbash
gcloud run services update web-api \
  --update-env-vars="^@^ALLOWED_ORIGINS=https://a.example,https://b.example@LOG_LEVEL=info"

In practice I keep most of this out of ad-hoc commands altogether: each service has a declarative service YAML (or Terraform resource) checked in, and the deploy step applies it.

CostWhat Secret Manager actually costs

At the time of writing, Secret Manager pricing charges per active secret version per location per month, plus a per-10,000 fee for access operations, with a small free tier. Two consequences. Old versions you never disable keep costing money, so destroy or disable versions after a rotation has settled. And code that calls the access API on every request turns a fixed monthly cost into one that scales with traffic, which is one more reason to let the platform resolve secrets at instance start.

Rotation#

Secret Manager can send rotation reminders to a Pub/Sub topic on a schedule, but it does not rotate anything for you. You write the handler that generates a new credential at the provider and adds a new version. The safe sequence for a pinned secret is:

  1. Add the new version (gcloud secrets versions add ... --data-file=-).
  2. If the upstream system allows two valid credentials at once, enable both.
  3. Deploy each consumer pinned to the new version, one service at a time.
  4. Watch error rates. If they rise, roll back the revision; the old version is still enabled.
  5. Revoke the old credential upstream, then disable the old version, then destroy it after a grace period.

For signing keys, overlap matters even more: the verifier should accept both old and new keys (a key ID in the token header makes this clean) until every token signed with the old key has expired.

CI never sees production secrets#

The deploy pipeline needs permission to create revisions, not to read values. With Workload Identity Federation, the CI job exchanges its OIDC token for short-lived Google credentials; no service account key file exists anywhere. The deployer identity needs roles/run.developer (or a custom role) on the service and roles/iam.serviceAccountUser on the runtime service account it deploys as. It does not need secretAccessor on anything, because the deploy only contains references like jwt-secret:3.

This matters beyond tidiness. CI systems run third-party actions, cache directories across jobs, and print environment dumps when builds fail. If production secret values are never in the job, none of those can leak them. Build-time secrets (a private npm registry token, say) are a separate, scoped credential with no production data access.

Auditing config drift across services#

Once you have ten or more services, they will drift: one still reads a secret at version 2 after everyone else moved to 3, one has LOG_LEVEL=debug left over from an incident, one is missing a variable that only matters on a rare code path. I audit with a short script rather than trusting memory:

scripts/audit-config.shbash
#!/usr/bin/env bash
set -euo pipefail
REGION=europe-west1
 
for svc in $(gcloud run services list --region="$REGION" --format='value(metadata.name)'); do
  gcloud run services describe "$svc" --region="$REGION" --format=json |
    jq -r --arg svc "$svc" '
      .spec.template.spec.containers[0].env[]?
      | [$svc, .name,
         (if .valueFrom.secretKeyRef
          then "secret:\(.valueFrom.secretKeyRef.name):\(.valueFrom.secretKeyRef.key)"
          else "plain:\(.value)" end)]
      | @tsv'
done | sort -k2,2 -k1,1 | column -t

Sorting by variable name puts every service's view of JWT_SECRET or LOG_LEVEL on adjacent lines, which makes a stale version or an odd value obvious. It prints plain values, so it is for people who can already view the services, and none of them are secrets by construction. Run it per environment, diff staging against production, and consider a scheduled job that fails when two services disagree on a secret version for longer than a rotation window.

Keeping secrets out of logs#

The last leak path is your own logger. A request logger that serialises headers will write bearer tokens into Cloud Logging, where retention is long and access is broad. Pino has redaction built in and applies it at serialisation time:

src/logger.tsts
import pino from "pino";
import { env } from "./config/env.js";
 
export const logger = pino({
  level: env.LOG_LEVEL,
  redact: {
    paths: [
      "req.headers.authorization",
      "req.headers.cookie",
      'req.headers["x-api-key"]',
      "*.password",
      "*.token",
      "*.secret",
    ],
    censor: "[REDACTED]",
  },
});

Redaction paths are explicit, so wildcards only go one level deep; a nested body.user.password needs its own path. I treat the list as a living document and add to it after every review that finds a new shape. And I never log the env object, not even at debug level.

Trade-offs and failure modes#

  • Fail-fast can block a deploy during an incident. If an emergency fix ships and a new required variable is missing, the revision will not start. That is the right outcome; the fix is to add the variable, not to relax validation.
  • Pinned versions mean rotations need deploys. I accept that cost for the auditability. Where it is genuinely unacceptable, use a volume mount with latest and make the code re-read the file.
  • Secret Manager is a boot-time dependency. If the runtime service account loses access, new instances fail to start while existing ones keep running. IAM changes to secrets deserve the same review as code.
  • Validators drift from reality. A schema that declares a variable nobody reads is noise; one that misses a variable read elsewhere is a hole. The lint rule against process.env outside the config module is what keeps the schema honest.
  • Plain env vars are not secret. Anyone with run.services.get can read them. If you would not paste it into a chat channel, it is a secret.

Checklist

  • Exactly one module per service reads process.env, enforced by a lint rule.
  • The environment is validated against a schema at startup, and the process exits on failure.
  • Validation errors report names and paths, never values.
  • No credential has a default or fallback in code.
  • .env files are git-ignored, .env.example is committed and checked against the schema in CI, and a secret scanner runs on every push.
  • Secrets are created with printf '%s', not echo.
  • Each service has its own runtime service account with secretAccessor granted per secret.
  • Env-var secrets are pinned to explicit versions; latest is used only with volume mounts and code that re-reads the file.
  • --update-secrets and --update-env-vars are used for incremental changes, with a custom delimiter for comma-containing values.
  • CI authenticates with Workload Identity Federation and holds no secret-read permission.
  • A drift audit runs on a schedule across every service and environment.
  • Loggers redact authorization headers, cookies, and credential-shaped fields.

When not to do this#

If you run one service on a PaaS with a built-in secret store and a single environment, the platform's own settings screen plus a small validator is enough; a drift audit for one service is ceremony. If your secrets are genuinely dynamic (per-tenant database credentials, short-lived tokens minted per request), platform-bound env vars are the wrong model, and you want the Secret Manager API or a broker such as Vault with caching in process. And if you are already fully on Kubernetes with an operator that syncs secrets from a cloud store, keep the validation half of this article and let the operator handle the binding.

Share
All articles →

What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.

14 min

Early-stage systems rarely go broke on traffic. They bleed money while idle. How to design a GCP stack whose cost stays near zero when no one is using it.

13 min