Protecting pay-per-call APIs: validation, idempotency, locks
Every upstream call spends money and quota. A layered defence of validation, rate limits, caching, single-flight and leases keeps both under control.
The moment your backend calls a third-party API that charges per request, every endpoint in front of it becomes a spending interface. A geocoding lookup, a stats provider, a document verification service, an SMS gateway: each call costs money, draws down a quota you share with every other user, and carries your API key's reputation with the provider. A retry loop, a double-clicked button or a curious user with a script can turn a modest bill into an incident. This article is for backend engineers who own such an integration and want a defence that is cheap to run, easy to reason about and hard to bypass. The approach is layered: each layer rejects or absorbs a class of waste before it reaches the next, and the upstream call is the last thing that happens, not the first.
Constraints#
Before choosing mechanisms, be explicit about what you are protecting against. In my experience the waste comes from five places:
- Invalid input. Requests that could never succeed (empty strings, malformed identifiers, oversized payloads) but still get forwarded and billed.
- Duplicates. The same logical request arriving several times: client retries, double submits, multiple tabs, a page that fires the same query on every render.
- Concurrency. Many different users asking for the same resource at the same moment, typically right after a cache entry expires. This is the thundering herd.
- Abuse. Enumeration of a lookup endpoint, scraping through your backend to use your paid key for free, or one account consuming the whole shared quota.
- Upstream trouble. The provider returns 429 or 5xx, and naive retries multiply both the bill and the load on a service that is already struggling.
And a few constraints from the platform: the backend runs as several stateless instances behind a load balancer (Cloud Run, Kubernetes, Lambda), so any state that must be shared across instances lives in a store like Redis or Firestore. The provider's own rate limit is global to your key, not per instance. Most providers bill successful calls; some also bill failed or throttled ones, so read your provider's terms before assuming errors are free.
The layers, compared#
| Layer | Stops | Where state lives | Cost to run | Typical response when it trips |
|---|---|---|---|---|
| Input validation + normalisation | Invalid input, cache-key fragmentation | None | Free | 400 |
| Per-user and global rate limits | Abuse, runaway clients, quota exhaustion | Redis (or in-process for a single instance) | One round trip | 429 with Retry-After |
| Response cache with TTL | Repeat reads across users | Redis / Firestore / CDN | One read | Cached body |
| Idempotency keys | Duplicate paid writes | Redis / Firestore | One read + one write | Stored original response |
| Single-flight (in-process) | Concurrent duplicates on one instance | Memory | Free | Shares the in-flight promise |
| Distributed lease lock | Concurrent duplicates across instances | Redis / Firestore | One transaction | Wait for cache, or 202 |
| Brute-force lockout | Enumeration on lookup endpoints | Redis / Firestore | One increment | 429 or 403 for the lockout window |
| Circuit breaker + backoff | Retry storms during upstream failure | Memory per instance | Free | Stale cache or 503 |
| Budget alerts + kill switch | Everything the other layers missed | Billing + a config flag | Negligible | Feature disabled |
No single row is sufficient. Rate limits do nothing about a cache stampede; a cache does nothing about enumeration; a lock does nothing about a bill that grows because traffic legitimately grew. The value is in the stack.
The decision#
My default for any pay-per-call integration is: validate and normalise, rate limit per user and globally, serve from a shared cache keyed by the normalised input, collapse concurrent misses with single-flight plus a lease, wrap the actual call in a circuit breaker with capped, jittered retries, and put a kill switch and a budget alert around the whole thing. Paid writes (send a message, run a check) additionally require an idempotency key. Lookup endpoints that reveal whether something exists get a brute-force lockout.
That sounds like a lot. In practice it is a few hundred lines of shared code, written once and wrapped around every provider client.
Implementation#
The examples are TypeScript for Node.js, using ioredis and the Firestore Admin SDK. Substitute your own stores; the patterns do not depend on them.
Step 1: validate and normalise before anything else#
Validation is the cheapest layer and the one most often skipped, because the provider "will reject it anyway". It will, and it may bill you for the rejection. Normalisation matters as much as validation: " 10 Main St " and "10 main st" should be one cache entry, not two paid calls.
import { createHash } from "node:crypto";
const MAX_LEN = 200;
export type GeocodeQuery = { address: string; country: string };
export function normaliseGeocodeQuery(raw: unknown): GeocodeQuery | null {
if (typeof raw !== "object" || raw === null) return null;
const { address, country } = raw as Record<string, unknown>;
if (typeof address !== "string" || typeof country !== "string") return null;
const a = address.normalize("NFKC").trim().replace(/\s+/g, " ").toLowerCase();
const c = country.trim().toUpperCase();
if (a.length < 3 || a.length > MAX_LEN) return null;
if (!/[\p{L}\p{N}]/u.test(a)) return null; // must contain a letter or digit
if (!/^[A-Z]{2}$/.test(c)) return null; // ISO 3166-1 alpha-2
return { address: a, country: c };
}
export function cacheKey(q: GeocodeQuery): string {
const h = createHash("sha256").update(`${q.country}|${q.address}`).digest("hex");
return `geocode:v1:${h}`;
}Version the cache key (v1). When you change normalisation rules, bump it rather than serving entries built under the old rules.
Step 2: rate limit with a token bucket#
A token bucket allows short bursts up to a capacity while enforcing a steady average rate. It is the right model for user-facing traffic, where a burst of three requests on page load is normal and three hundred in a minute is not.
For a single process, an in-memory bucket is enough:
export class TokenBucket {
private tokens: number;
private last: number;
constructor(
private readonly capacity: number,
private readonly refillPerSecond: number,
private readonly now: () => number = Date.now,
) {
this.tokens = capacity;
this.last = now();
}
tryTake(cost = 1): boolean {
const t = this.now();
const elapsed = (t - this.last) / 1000;
this.tokens = Math.min(this.capacity, this.tokens + elapsed * this.refillPerSecond);
this.last = t;
if (this.tokens < cost) return false;
this.tokens -= cost;
return true;
}
/** Seconds until `cost` tokens are available; use it for Retry-After. */
secondsUntil(cost = 1): number {
const missing = cost - this.tokens;
return missing <= 0 ? 0 : Math.ceil(missing / this.refillPerSecond);
}
}With several instances, per-instance buckets let a user multiply their allowance by the instance count, and the provider's quota is global anyway. Move the bucket into Redis and make the check atomic with a Lua script. Using the server's TIME avoids clock skew between instances:
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL!);
const SCRIPT = `
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local rate = tonumber(ARGV[2])
local cost = tonumber(ARGV[3])
local t = redis.call('TIME')
local now = tonumber(t[1]) * 1000 + math.floor(tonumber(t[2]) / 1000)
local s = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(s[1]) or capacity
local ts = tonumber(s[2]) or now
tokens = math.min(capacity, tokens + (now - ts) / 1000 * rate)
local allowed = 0
if tokens >= cost then
tokens = tokens - cost
allowed = 1
end
redis.call('HSET', key, 'tokens', tostring(tokens), 'ts', tostring(now))
redis.call('PEXPIRE', key, math.ceil(capacity / rate * 1000) + 1000)
return allowed
`;
export async function allow(
key: string,
capacity: number,
refillPerSecond: number,
cost = 1,
): Promise<boolean> {
const res = await redis.eval(SCRIPT, 1, key, capacity, refillPerSecond, cost);
return res === 1;
}Apply two buckets per request: one keyed by the authenticated user (rl:user:{uid}:geocode) and one global bucket (rl:global:geocode) sized just under the provider's documented limit. The global bucket turns a provider-side 429, which may be billed and may count against you, into a cheap local rejection. Rate limit by authenticated identity rather than IP wherever you can; IPs are shared behind carrier NAT and trivially rotated by anyone determined.
Step 3: cache responses, including the negative ones#
The response cache is where most of the money is saved, because it deduplicates across users, not just within one. Two users geocoding the same address, or a thousand users viewing the same public entity, should cost one upstream call per TTL.
Choose the TTL from the data's real rate of change and the provider's terms. Some providers restrict how long you may store their responses, so check the licence before caching for a week. Cache "not found" results too, with a shorter TTL. Without negative caching, a missing entity is the most expensive thing a user can ask for, because it never gets cached.
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL!);
type Cached<T> = { found: true; value: T } | { found: false };
const HIT_TTL_S = 7 * 24 * 3600; // positive results
const MISS_TTL_S = 6 * 3600; // negative results
export async function readCache<T>(key: string): Promise<Cached<T> | null> {
const raw = await redis.get(key);
return raw ? (JSON.parse(raw) as Cached<T>) : null;
}
export async function writeCache<T>(key: string, entry: Cached<T>): Promise<void> {
const ttl = entry.found ? HIT_TTL_S : MISS_TTL_S;
await redis.set(key, JSON.stringify(entry), "EX", ttl);
}Step 4: idempotency keys for paid writes#
Caching handles reads. For paid actions that change state, such as sending a message or submitting a verification check, the client generates a unique key per logical operation and sends it in an Idempotency-Key header, the convention described in the IETF httpapi draft (still an Internet-Draft, not an RFC, at the time of writing, but widely implemented). The server stores the outcome keyed by (user, key) and replays it for any retry.
The important detail is the in-progress state. If the first attempt is still running when the retry arrives, the retry must not start a second paid call:
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL!);
const TTL_S = 24 * 3600;
type Stored = { status: "pending" } | { status: "done"; code: number; body: unknown };
export async function withIdempotency(
uid: string,
idemKey: string,
run: () => Promise<{ code: number; body: unknown }>,
): Promise<{ code: number; body: unknown }> {
const key = `idem:${uid}:${idemKey}`;
const claimed = await redis.set(key, JSON.stringify({ status: "pending" }), "EX", TTL_S, "NX");
if (claimed !== "OK") {
const prev = JSON.parse((await redis.get(key)) ?? '{"status":"pending"}') as Stored;
if (prev.status === "done") return { code: prev.code, body: prev.body };
return { code: 409, body: { error: "request_in_progress" } };
}
try {
const result = await run();
await redis.set(key, JSON.stringify({ status: "done", ...result }), "EX", TTL_S);
return result;
} catch (err) {
await redis.del(key); // let the client retry a failure that did not complete
throw err;
}
}Deleting the key on failure is a judgement call. If the provider may have performed the action before your call errored (a timeout after the request was sent), deleting risks a double charge. For those providers, pass your idempotency key through to them if they support one, or record the ambiguous state and reconcile.
Step 5: collapse concurrent misses with single-flight#
When a popular cache entry expires, every request that arrives before it is refilled will miss. Single-flight makes concurrent callers for the same key share one promise:
export class SingleFlight<T> {
private readonly inflight = new Map<string, Promise<T>>();
do(key: string, fn: () => Promise<T>): Promise<T> {
const existing = this.inflight.get(key);
if (existing) return existing;
const p = fn().finally(() => this.inflight.delete(key));
this.inflight.set(key, p);
return p;
}
}It is ten lines, free to run, and removes most duplicates on a single instance. It does nothing across instances, which is what the next layer is for.
Step 6: a lease lock across instances#
A lease is a lock with an expiry, so a crashed holder cannot block everyone forever. The holder fetches from upstream and writes the cache; everyone else waits briefly for the cache to fill.
With Redis, acquisition is SET key token NX PX ttl, and release must check the token so you never delete someone else's lease:
import { randomUUID } from "node:crypto";
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL!);
const RELEASE = `
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('DEL', KEYS[1])
end
return 0
`;
export async function acquireLease(key: string, ttlMs: number): Promise<string | null> {
const token = randomUUID();
const ok = await redis.set(`lease:${key}`, token, "PX", ttlMs, "NX");
return ok === "OK" ? token : null;
}
export async function releaseLease(key: string, token: string): Promise<void> {
await redis.eval(RELEASE, 1, `lease:${key}`, token);
}If you have no Redis and do have Firestore, a transaction gives the same guarantee. Hash the key, since document ids cannot contain slashes:
import { createHash, randomUUID } from "node:crypto";
import { getFirestore, Timestamp } from "firebase-admin/firestore";
const db = getFirestore();
const ref = (key: string) =>
db.collection("leases").doc(createHash("sha256").update(key).digest("hex"));
export async function acquireLease(key: string, ttlMs: number): Promise<string | null> {
const token = randomUUID();
const won = await db.runTransaction(async (tx) => {
const snap = await tx.get(ref(key));
const expiresAt = snap.get("expiresAt") as Timestamp | undefined;
if (snap.exists && expiresAt && expiresAt.toMillis() > Date.now()) return false;
tx.set(ref(key), { token, expiresAt: Timestamp.fromMillis(Date.now() + ttlMs) });
return true;
});
return won ? token : null;
}
export async function releaseLease(key: string, token: string): Promise<void> {
await db.runTransaction(async (tx) => {
const snap = await tx.get(ref(key));
if (snap.exists && snap.get("token") === token) tx.delete(ref(key));
});
}Add a TTL policy on expiresAt to clean up abandoned lease documents, but do not rely on it for correctness: TTL deletion is asynchronous and can lag, which is why the code checks expiresAt itself.
A lease here is an efficiency lock, not a correctness lock. If the holder stalls past its TTL, a second caller may fetch too. For a cache fill that costs one extra call, which is acceptable. If a duplicate would corrupt data, you need fencing tokens or the provider's own idempotency support, not a lease.
Step 7: circuit breaker and disciplined retries#
When the provider is failing, retries are the most expensive thing you can do. Retry only on 429 and 5xx, cap attempts low, use exponential backoff with full jitter, and honour Retry-After. Around that, a circuit breaker stops calling entirely after repeated failures and probes periodically.
export class CircuitBreaker {
private failures = 0;
private openedAt = 0;
private probing = false;
constructor(private readonly threshold = 5, private readonly cooldownMs = 30_000) {}
async call<T>(fn: () => Promise<T>): Promise<T> {
const open = this.failures >= this.threshold;
if (open) {
const coolingDown = Date.now() - this.openedAt < this.cooldownMs;
if (coolingDown || this.probing) throw new Error("circuit_open");
this.probing = true; // half-open: allow one probe
}
try {
const result = await fn();
this.failures = 0;
return result;
} catch (err) {
this.failures += 1;
if (this.failures >= this.threshold) this.openedAt = Date.now();
throw err;
} finally {
this.probing = false;
}
}
}
export class RetryableError extends Error {
constructor(message: string, readonly retryAfterMs?: number) {
super(message);
}
}
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
export async function withRetry<T>(fn: () => Promise<T>, maxAttempts = 3): Promise<T> {
for (let attempt = 1; ; attempt++) {
try {
return await fn();
} catch (err) {
if (!(err instanceof RetryableError) || attempt >= maxAttempts) throw err;
const backoff = Math.random() * Math.min(8_000, 250 * 2 ** attempt); // full jitter
await sleep(Math.max(backoff, err.retryAfterMs ?? 0));
}
}
}Your provider client throws RetryableError for 429 and 5xx (parsing Retry-After, which may be seconds or an HTTP date) and a plain error for everything else. A 400 will not succeed on retry, so retrying it only spends money. When the breaker is open, serve a stale cache entry if you have one; stale data is usually better than an error page.
Step 8: compose the layers#
With the pieces in place, the handler reads top to bottom in cost order:
import { normaliseGeocodeQuery, cacheKey } from "./input";
import { readCache, writeCache } from "./cache";
import { allow } from "../limits/redis-bucket";
import { SingleFlight } from "../concurrency/single-flight";
import { acquireLease, releaseLease } from "../concurrency/redis-lease";
import { CircuitBreaker, withRetry } from "../upstream/breaker";
import { geocodeUpstream, type GeoResult } from "./provider";
import { isEnabled } from "../config/kill-switch";
const flights = new SingleFlight<GeoResult | null>();
const breaker = new CircuitBreaker(5, 30_000);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
export class HttpError extends Error {
constructor(readonly status: number, message: string) {
super(message);
}
}
export async function geocode(uid: string, raw: unknown): Promise<GeoResult | null> {
if (!(await isEnabled("geocode"))) throw new HttpError(503, "feature_disabled");
const q = normaliseGeocodeQuery(raw);
if (!q) throw new HttpError(400, "invalid_query");
if (!(await allow(`rl:user:${uid}:geocode`, 10, 0.2))) throw new HttpError(429, "slow_down");
if (!(await allow("rl:global:geocode", 40, 40))) throw new HttpError(429, "busy");
const key = cacheKey(q);
const hit = await readCache<GeoResult>(key);
if (hit) return hit.found ? hit.value : null;
return flights.do(key, async () => {
const token = await acquireLease(key, 10_000);
if (!token) {
// Another instance is fetching. Wait briefly for it to fill the cache.
for (let i = 0; i < 10; i++) {
await sleep(300);
const filled = await readCache<GeoResult>(key);
if (filled) return filled.found ? filled.value : null;
}
throw new HttpError(503, "try_again");
}
try {
const again = await readCache<GeoResult>(key); // filled while we acquired?
if (again) return again.found ? again.value : null;
const result = await breaker.call(() => withRetry(() => geocodeUpstream(q), 2));
await writeCache(key, result ? { found: true, value: result } : { found: false });
return result;
} finally {
await releaseLease(key, token);
}
});
}The rate limit numbers here are illustrative: ten requests of burst per user refilling at one every five seconds, and a global bucket sized for a provider limit of about fifty per second. Derive yours from the provider's documented quota and your real traffic.
Step 9: brute-force lockouts on lookup endpoints#
Some paid endpoints answer "does this exist?": look up a company by registration number, a parcel by tracking code, an account by a public identifier. Those are enumeration targets, and every guess costs you a call. Rate limits slow enumeration down; a lockout on failures stops it:
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL!);
export async function isLockedOut(uid: string): Promise<number> {
return redis.pttl(`lock:lookup:${uid}`); // > 0 means locked, ms remaining
}
export async function recordLookupMiss(uid: string): Promise<void> {
const k = `miss:lookup:${uid}`;
const misses = await redis.incr(k);
if (misses === 1) await redis.expire(k, 3600); // count misses per rolling hour
if (misses >= 5) {
const lockSeconds = Math.min(24 * 3600, 60 * 2 ** (misses - 5));
await redis.set(`lock:lookup:${uid}`, "1", "EX", lockSeconds);
}
}Check the lockout before the cache, so a locked-out user cannot probe even cached entries. Combine it with negative caching (Step 3), authentication, and a bot check such as reCAPTCHA Enterprise or Firebase App Check on the client. Return the same response shape and similar latency for "not found" and "locked", so the endpoint does not leak which one happened.
Step 10: budgets and a kill switch#
Every layer above has a bug you have not found yet. The last line of defence is financial. Most providers let you set a daily quota or spend cap in their own console; set it just above your expected peak. On Google Cloud, Cloud Billing budgets send alerts at thresholds but do not stop spending on their own. You can route budget notifications to Pub/Sub and act on them programmatically, for example by flipping your kill switch.
The kill switch itself is a config document or key read with a short in-memory cache, so checking it costs almost nothing per request:
import { getFirestore } from "firebase-admin/firestore";
const db = getFirestore();
const cache = new Map<string, { value: boolean; at: number }>();
const TTL_MS = 30_000;
export async function isEnabled(feature: string): Promise<boolean> {
const c = cache.get(feature);
if (c && Date.now() - c.at < TTL_MS) return c.value;
const snap = await db.doc(`config/features`).get();
const value = snap.get(feature) !== false; // absent means enabled
cache.set(feature, { value, at: Date.now() });
return value;
}An illustrative example: a provider charging $5 per 1,000 calls, and a page that triggers one call per view. At 200,000 uncached views a day that is $1,000 a day. A shared cache with a 90% hit rate brings it to $100, and single-flight shaves the stampede spikes on top. Check your provider's pricing page at the time of writing, multiply by your real traffic, and decide your cache TTL and spend cap from the result, not after the first invoice.
Trade-offs and failure modes#
- Every layer adds a dependency. If Redis is down, do you fail open (allow the call, risk spend) or closed (reject, lose the feature)? For paid calls I fail closed on rate limits and leases and fail open on the cache read. Decide deliberately and test it.
- Stale data. Long TTLs save money and serve outdated answers. Offer a user-triggered refresh that is itself rate limited, rather than shortening the TTL for everyone.
- Lease TTL tuning. Too short and a slow upstream call lets a second fetcher in; too long and a crashed holder blocks the key. Set it to a comfortable multiple of the provider's p99 latency.
- Waiters add latency. Callers that lose the lease poll the cache. That is slower than calling upstream directly, which is the cost-for-latency trade you are choosing on purpose.
- Per-instance state is partial. Single-flight and circuit breakers live in memory, so each instance learns separately. That is fine; the shared layers catch what they miss.
- Idempotency ambiguity. A timeout after sending a paid write leaves you unsure whether it happened. Use the provider's idempotency support where it exists and reconcile otherwise.
- Lockouts can hit legitimate users. Keep the threshold generous, the first lockout short, and give support a way to clear it.
Checklist
- Inputs are validated and normalised before any upstream call, and invalid input returns 400 locally
- Cache keys are derived from normalised input and versioned
- Per-user and global token buckets exist, the global one sized under the provider's quota
- Responses are cached with a TTL that respects the provider's terms, and "not found" is cached too
- Paid writes require an idempotency key with an explicit in-progress state
- Concurrent misses are collapsed with single-flight in-process and a lease lock across instances
- Retries happen only on 429 and 5xx, are capped, jittered and honour
Retry-After - A circuit breaker stops calls during an outage and stale cache is served meanwhile
- Lookup endpoints have a failure lockout and do not leak "not found" versus "locked"
- A spend cap is set in the provider console and a billing budget alert is configured
- A kill switch can disable the feature in under a minute without a deploy
- Behaviour when Redis or Firestore is unavailable is decided and tested
When not to do this#
If the upstream call is free, unmetered and fast, most of this is overhead: a simple timeout and a modest cache are enough. If traffic is low and internal (an admin tool used by three people), a provider-side spend cap and a kill switch cover the real risk without the rest. And if you find yourself building an elaborate lock choreography to protect one expensive call, step back and ask whether the data can be fetched on a schedule instead, in a single batch job, and served from your own store. Moving the paid call off the request path entirely is often the strongest protection of all.
Related articles
All articles →Pub/Sub will deliver some messages twice, and that is by design. Here is how to build consumers whose side effects happen once anyway, with working code.
17 min
A layered CLAUDE.md, tight permissions, hooks, skills and scoped MCP servers: the setup that keeps Claude Code safe and cheap across dozens of repositories.
16 min
What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.
14 min