Refresh-on-read instead of a cron job

A scheduled sync pays to refresh every record, including the ones nobody reads. Refreshing stale data when someone looks makes cost follow attention.

Published

Updated

—

Reading time

18 min

Most systems that mirror data from somewhere else start with the same job. A scheduler fires at 03:00, a worker pages through every record, calls the upstream API for each one and writes back whatever it gets. It is easy to write, easy to explain, and it keeps everything fresh. It also has a cost curve that I've learned to distrust: it grows with the number of records you hold, not with the number of records anyone looks at. On day one those are the same number. A year later you are refreshing hundreds of thousands of rows every night so that a few thousand of them can be viewed the next day. This article is for engineers who cache third-party data (profile data pulled from an external service, exchange rates, repository stats, shipment status) and want the bill to track reads instead of inventory. I'll cover the stale-while-revalidate pattern applied at the application level, how I dedupe concurrent refreshes, why I prefer Cloud Tasks over Pub/Sub for the background refresh, and the cases where a cron job is still the right answer.

The constraint: most records are cold#

Look at the read distribution of almost any user-facing dataset and you get a long tail. A small set of records is read many times a day. A larger set is read a few times a month. Most records are read rarely or never: accounts that signed up and left, entities that were linked once and forgotten, items that only matter when their owner comes back.

A nightly sweep ignores that distribution. It pays the full price for every record, every night:

  • Upstream calls. One request per record per run, often against an API that charges per call or gives you a fixed daily quota.
  • Database writes. One write per record per run, even when nothing changed, unless you diff first, which costs a read.
  • Compute. A job whose run time grows linearly with the table, until it no longer fits in its window and you start sharding it.
  • Failure blast radius. When the upstream has a bad night, the sweep fails for everyone at once, and the retry doubles the load against a service that is already struggling.

The freshness you buy with that is also uneven. A record refreshed at 03:00 is 20 hours old when its owner opens it at 23:00, and a record nobody opens is fresh for no one.

I ran into this on a large-scale social platform I'm building, where a lot of the data shown on a page comes from paid third-party APIs. The cost constraint was simple: upstream spend has to grow with attention, and a record nobody views should cost nothing beyond its storage.

The options#

Before choosing, I laid out the approaches I'd seen work. The table compares them on the dimension that matters most here, what the cost scales with.

ApproachCost scales withFreshness when readRead latencyMoving parts
Nightly full sweep (cron)Total recordsUp to 24 h oldFast (cached)Scheduler, batch job
Sweep of recently active recordsActive recordsUp to one run interval oldFastScheduler, batch job, activity index
Synchronous refresh on readReads of stale recordsAlways freshSlow on stale reads, fails with upstreamNone extra
Stale-while-revalidate on readDistinct stale records readOne TTL, plus one refreshFast (cached)Lease, queue, worker
Upstream push (webhooks)Changes upstreamNear real timeFastWebhook endpoint, signature checks

Webhooks are the best answer when the upstream offers them, and most of the APIs I integrate with don't. Synchronous refresh on read is simple but ties your page latency and availability to someone else's API. The "recently active" sweep is a reasonable middle ground, but it needs an index of what's active, and it still refreshes records whose readers won't come back before the next run.

The decision: stale-while-revalidate, in the application#

HTTP already has a name for the behaviour I wanted. RFC 5861 defines the stale-while-revalidate Cache-Control extension: a cache may keep serving a response after it goes stale, for a bounded time, while it revalidates in the background. MDN's description puts the benefit well: revalidation hides its latency from the client, so the response appears to have been fresh all along. The same RFC defines stale-if-error, which lets a cache serve stale content when the origin fails. I use both ideas, but one layer down, in the service that owns the cached data, because the origin here is a third-party API and the "cache" is a Firestore document.

The rules I settled on:

  1. Every read returns the stored snapshot immediately, together with its fetchedAt time, so the UI can say "updated 3 hours ago".
  2. If the snapshot is older than its TTL, the read triggers one background refresh. The reader doesn't wait for it.
  3. TTL depends on who is reading. The owner of a record sees a shorter TTL than a visitor. Owners notice staleness in their own data. Visitors mostly don't, and there are many more of them.
  4. Manual refresh exists, with a cooldown. Users can ask for fresh data, but not more than once per window.
  5. At most one refresh per record is in flight. A popular record read a thousand times in the minute after it goes stale triggers one upstream call, not a thousand.
  6. A failing upstream backs off per record. Errors push the next allowed attempt further out, and readers keep getting the last good snapshot.
  7. The refresh runs on a rate-limited queue, sized to the upstream's quota, not on the request thread.

With these rules, the number of upstream calls in any period is bounded by the number of distinct records read in that period, divided by how many TTL windows fit into it. Records nobody reads cost nothing.

The reader never waits for the upstream. Only stale, unleased reads start the dashed path, and the queue sets the pace for the third-party API.

Implementation#

Each cached record is one Firestore document keyed by the upstream identity, for example snapshots/{source}:{externalId}. It holds the payload plus a few bookkeeping fields: fetchedAt, leaseUntil, failures, nextAttemptAt and lastManualAt. Keeping the bookkeeping on the same document means the hot path, a fresh read, costs exactly one document read.

The staleness check and the lease#

The read path does a plain read first. Only if the snapshot is stale, not already leased and not in backoff does it pay for a transaction to claim the lease. That ordering matters for cost: most reads of a hot record hit a fresh snapshot or an existing lease, and both of those are decided from the one read you already did.

src/refresh/read-with-refresh.tsts
import { createHash } from "node:crypto";
import { DocumentReference, Firestore, Timestamp } from "@google-cloud/firestore";
import { CloudTasksClient } from "@google-cloud/tasks";
import { env } from "../config/env";
import { logger } from "../logger";
 
const db = new Firestore();
const tasks = new CloudTasksClient();
 
export type Audience = "owner" | "visitor";
 
const HOUR = 60 * 60 * 1000;
const TTL_MS: Record<Audience, number> = { owner: 24 * HOUR, visitor: 72 * HOUR };
const LEASE_MS = 5 * 60 * 1000;
const GRPC_ALREADY_EXISTS = 6;
 
export interface Snapshot<T> {
  data: T | null;
  fetchedAt: Timestamp | null;
  leaseUntil?: Timestamp;
  nextAttemptAt?: Timestamp;
  failures?: number;
  lastManualAt?: Timestamp;
}
 
const isStale = (s: Snapshot<unknown> | undefined, ttlMs: number, now: number) =>
  !s?.fetchedAt || now - s.fetchedAt.toMillis() > ttlMs;
 
export const isBlocked = (s: Snapshot<unknown> | undefined, now: number) =>
  (s?.leaseUntil?.toMillis() ?? 0) > now || (s?.nextAttemptAt?.toMillis() ?? 0) > now;
 
export async function readWithRefresh<T>(key: string, audience: Audience) {
  const ref = db.collection("snapshots").doc(key);
  const s = (await ref.get()).data() as Snapshot<T> | undefined;
  const result = { data: s?.data ?? null, fetchedAt: s?.fetchedAt?.toDate() ?? null, refreshing: false };
 
  const ttlMs = TTL_MS[audience];
  if (!isStale(s, ttlMs, Date.now()) || isBlocked(s, Date.now())) return result;
 
  try {
    const generation = await claimLease(ref, ttlMs);
    if (generation !== null) {
      await enqueueRefresh(key, generation);
      result.refreshing = true;
    }
  } catch (err) {
    // A failed refresh must never fail the read. The lease expires on its own.
    logger.warn({ err, key }, "refresh enqueue failed");
  }
  return result;
}
 
export async function claimLease(ref: DocumentReference, ttlMs: number): Promise<number | null> {
  return db.runTransaction(async (tx) => {
    const now = Date.now();
    const s = (await tx.get(ref)).data() as Snapshot<unknown> | undefined;
    if (!isStale(s, ttlMs, now) || isBlocked(s, now)) return null; // someone else won
    tx.set(ref, { leaseUntil: Timestamp.fromMillis(now + LEASE_MS) }, { merge: true });
    return s?.fetchedAt?.toMillis() ?? 0; // the generation this refresh replaces
  });
}
 
export async function enqueueRefresh(key: string, generation: number) {
  const parent = tasks.queuePath(env.GCP_PROJECT, env.TASKS_LOCATION, env.REFRESH_QUEUE);
  const taskId = createHash("sha256").update(`${key}:${generation}`).digest("hex");
  try {
    await tasks.createTask({
      parent,
      task: {
        name: `${parent}/tasks/${taskId}`,
        httpRequest: {
          httpMethod: "POST",
          url: env.REFRESH_WORKER_URL,
          headers: { "Content-Type": "application/json" },
          body: Buffer.from(JSON.stringify({ key, generation })).toString("base64"),
          oidcToken: { serviceAccountEmail: env.TASKS_INVOKER_SA, audience: env.REFRESH_WORKER_URL },
        },
      },
    });
  } catch (err) {
    if ((err as { code?: number }).code !== GRPC_ALREADY_EXISTS) throw err;
  }
}

Three details in this code carry most of the weight.

The transaction re-checks everything. Two instances can both see a stale, unleased snapshot on their plain read. Both then enter runTransaction, and Firestore reruns a transaction when a concurrent write touches the documents it read, so only one of them commits the lease. The loser re-reads, finds the lease, and returns null. Reads have to come before writes inside a Firestore transaction, which this function respects.

The task name is a second dedupe layer. Cloud Tasks rejects a create call with ALREADY_EXISTS when the ID matches an existing task or one deleted or executed recently. The documentation says a deleted task's ID can take up to 24 hours to be released. Hashing the key together with the generation (the fetchedAt being replaced) means "refresh this record from this version" can only be enqueued once, even if a lease expires early or a retry replays the enqueue. The hash also follows the documented advice: named creates are slower because of the duplicate lookup, and sequential IDs, such as ones built from timestamps, increase latency and error rates. A SHA-256 hex string is well within the 500-character limit and uses only allowed characters.

The read awaits the enqueue. It's tempting to fire the enqueue and return. On Cloud Run with request-based billing, CPU is only allocated during request processing, so work left running after the response is sent may stall until the next request arrives on that instance. I pay the few tens of milliseconds on the rare read that starts a refresh, rather than switching the whole service to instance-based billing. If that latency matters, put the stale check behind the response in a service that already runs with instance-based billing, and accept that cost explicitly.

The worker and per-record backoff#

The worker is a plain HTTP handler on Cloud Run that Cloud Tasks calls with an OIDC token. It checks whether someone has already refreshed past the generation it was asked to replace, calls the upstream once, and writes the outcome.

src/refresh/worker.tsts
import { FieldValue, Firestore, Timestamp } from "@google-cloud/firestore";
import type { Snapshot } from "./read-with-refresh";
import { fetchUpstream } from "./upstream";
 
const db = new Firestore();
const HOUR = 60 * 60 * 1000;
const BASE_BACKOFF_MS = 0.25 * HOUR;
const MAX_BACKOFF_MS = 24 * HOUR;
 
export async function handleRefresh(key: string, generation: number): Promise<void> {
  const ref = db.collection("snapshots").doc(key);
  const s = (await ref.get()).data() as Snapshot<unknown> | undefined;
  if ((s?.fetchedAt?.toMillis() ?? 0) > generation) return; // already refreshed
 
  try {
    const data = await fetchUpstream(key, { timeoutMs: 10_000 });
    await ref.set(
      {
        data,
        fetchedAt: FieldValue.serverTimestamp(),
        failures: 0,
        leaseUntil: FieldValue.delete(),
        nextAttemptAt: FieldValue.delete(),
      },
      { merge: true },
    );
  } catch (err) {
    const failures = (s?.failures ?? 0) + 1;
    const cap = Math.min(MAX_BACKOFF_MS, BASE_BACKOFF_MS * 2 ** (failures - 1));
    const delay = cap / 2 + Math.random() * (cap / 2); // jitter spreads retries
    await ref.set(
      { failures, nextAttemptAt: Timestamp.fromMillis(Date.now() + delay), leaseUntil: FieldValue.delete() },
      { merge: true },
    );
    // Returning normally acks the task. The next read after nextAttemptAt retries.
  }
}

The choice that makes this cheap is in the catch. An upstream failure is recorded and the task is acknowledged. I don't let the queue retry upstream errors, because a retry is only worth paying for if someone will read the result. If the record is read again after nextAttemptAt, the read starts a new attempt. If it isn't, the failure costs nothing more. The queue's own retry policy still covers infrastructure failures: a worker that crashes, times out or fails to write to Firestore throws before it acknowledges, and Cloud Tasks redelivers.

Treat permanent errors differently from transient ones. A 404 from the upstream means the record no longer exists there, and I store that as a terminal state rather than backing off forever. A 429 or 5xx is transient, and the exponential backoff with jitter handles it. If many records fail at once, that's an upstream outage, not a per-record problem, and a small circuit breaker that pauses enqueues for the whole source is cheaper than letting every record discover it separately.

The queue: rate limits sized to the upstream#

infra/refresh-queue.shbash
gcloud tasks queues create upstream-refresh \
  --location=europe-west1 \
  --max-dispatches-per-second=5 \
  --max-concurrent-dispatches=10 \
  --max-attempts=3 \
  --min-backoff=30s \
  --max-backoff=300s

--max-dispatches-per-second and --max-concurrent-dispatches are the queue's rate controls. I set them from the upstream's quota, not from what my worker can handle, because the upstream is the constraint I pay for. A burst of reads after a newsletter goes out then turns into a queue that drains at a steady rate instead of a spike of 429s. The retry flags apply only to the infrastructure failures described above. One system limit to keep in mind: a single queue dispatches at most 500 tasks per second, so very high refresh volumes need several queues.

Manual refresh with a cooldown#

Manual refresh reuses the same lease and the same enqueue. It skips the TTL check but not the lease, the backoff or its own cooldown.

src/refresh/manual-refresh.tsts
import { Firestore, Timestamp } from "@google-cloud/firestore";
import { enqueueRefresh, isBlocked, type Snapshot } from "./read-with-refresh";
 
const db = new Firestore();
const COOLDOWN_MS = 24 * 60 * 60 * 1000;
const LEASE_MS = 5 * 60 * 1000;
 
export async function manualRefresh(key: string): Promise<{ queued: boolean; retryAfterMs?: number }> {
  const ref = db.collection("snapshots").doc(key);
  const outcome = await db.runTransaction(async (tx) => {
    const now = Date.now();
    const s = (await tx.get(ref)).data() as Snapshot<unknown> | undefined;
    const sinceManual = now - (s?.lastManualAt?.toMillis() ?? 0);
    if (sinceManual < COOLDOWN_MS) return { retryAfterMs: COOLDOWN_MS - sinceManual };
    if (isBlocked(s, now)) return { retryAfterMs: LEASE_MS };
    tx.set(ref, { lastManualAt: Timestamp.fromMillis(now), leaseUntil: Timestamp.fromMillis(now + LEASE_MS) }, { merge: true });
    return { generation: s?.fetchedAt?.toMillis() ?? 0 };
  });
 
  if ("retryAfterMs" in outcome) return { queued: false, retryAfterMs: outcome.retryAfterMs };
  await enqueueRefresh(key, outcome.generation);
  return { queued: true };
}

Return the remaining cooldown to the client so the button can show when it becomes available again. A refresh button that silently does nothing produces support tickets. One that says "available again in 6 hours" doesn't. If the upstream is paid per call, the cooldown is also your per-user spending limit, and it belongs in the same layered defence I describe in protecting pay-per-call APIs.

Cloud Tasks or Pub/Sub for the background refresh#

Both work, and I've run the refresh on both. The comparison comes down to four properties, all from Google's own Cloud Tasks versus Pub/Sub guide.

PropertyCloud TasksPub/Sub
Deduplication at creationYes, by task nameNo
Dispatch rate controlPer queue: rate and concurrencyNot built in; flow control only on pull subscribers
Scheduled deliveryYes, scheduleTime up to 30 days aheadNo
Fan-out to several consumersNo, one explicit targetYes, one topic, many subscriptions
Maximum payload1 MB10 MB

For refresh-on-read, the first three rows decide it. Creation-time dedupe backs up the lease. Rate control protects the upstream. And scheduled delivery lets me defer a refresh, for example to the moment a backoff expires, without a polling loop.

Pub/Sub is the better choice when the refresh result needs to reach several consumers (a search indexer, an analytics sink, a notification service), or when you already publish a "record viewed" event for other reasons. In that case, the dedupe has to live in the consumer. Ordering keys serialise delivery per record, but they don't dedupe, and a redelivery replays every later message for that key, including acknowledged ones. The generation check at the top of the worker is what makes the handler safe to run twice, the same approach I use in idempotent Pub/Sub consumers. Pub/Sub's subscription retry policy gives you exponential backoff between 0 and 600 seconds, which suits redelivery after a crash but not per-record backoff measured in hours. That still belongs on the document.

What it costs, cron versus on-read#

The shape of the comparison matters more than the numbers, so here is an illustrative model rather than a quote. Take 1,000,000 cached records. Assume that on a typical day 20,000 owners open their own records and 40,000 distinct records are opened by visitors, with the owner TTL at one day and the visitor TTL at three days.

Monthly volume (illustrative)Nightly cronRefresh-on-read
Upstream calls30,000,000about 0.6M (owners) + at most 1.2M (visitors)
Firestore writes for refreshed data30,000,000same as upstream calls
Extra lease transactions and tasks0about one each per refresh
Firestore reads on the read pathunchangedunchanged, one per view

Even the worst case for visitors, where every day brings a completely new set of 40,000 records, stays below 2 million refreshes a month, against 30 million for the sweep. In practice the visitor set overlaps heavily from day to day, so the real figure is lower. Firestore bills per document read and write, and Cloud Tasks bills per billable operation above a monthly free allowance. Check both pages for current prices in your region, and add the upstream's own per-call price, which is usually the largest line.

CostCost follows attention, and so does the bookkeeping

Refresh-on-read isn't free per refresh. Each one costs a lease transaction, a task creation and dispatch, and a worker invocation on top of the upstream call and the write the cron job would have paid anyway. That overhead is roughly a constant factor per refresh. The saving comes from doing far fewer refreshes. If most of your records are read every day, the two designs converge and the cron job, which batches better, may win. Model your own read distribution before switching.

The model also shows the compute side. A cron job that sweeps a million records needs a long-running worker or a sharded Cloud Run job every night, whatever the traffic. The on-read worker scales to zero between refreshes, which fits the approach I describe in scale-to-zero as a budget strategy.

Trade-offs and failure modes#

The first read of a new record has nothing to serve. Stale-while-revalidate needs something stale. For a record that has never been fetched, you choose between fetching synchronously (the reader waits, and a timeout needs a fallback) and returning an empty state with refreshing: true and letting the client poll or subscribe to the document. I fetch synchronously when a user links a new record, because they're waiting for it anyway, and use the async path everywhere else.

Readers see old data, sometimes for more than one TTL. The reader who triggers the refresh gets the stale snapshot. Only the next reader sees the new one. With a three-day visitor TTL, a visitor can see data that's three days old plus however long the queue takes. Show fetchedAt in the UI. Stale data shown with its age is honest. Stale data presented as current is a bug.

Bots and crawlers trigger refreshes. If every read can start a paid upstream call, a crawler walking your public pages is a cost amplifier. I only allow authenticated reads, or reads from verified human sessions, to start a refresh. Anonymous and crawler reads get the snapshot, never a refresh.

Leases can be lost. If the worker dies after the lease is claimed, the record is stuck until leaseUntil passes. Keep the lease a little longer than the worker's timeout plus queue latency, and never longer than you're willing to wait for recovery. Five minutes works for upstream calls with a ten-second timeout.

Hot records concentrate writes. A record read thousands of times a minute is still refreshed once per TTL, but each stale read that loses the lease race costs a transaction read. The plain-read check before the transaction keeps that small. If a single document becomes very hot, serve it from an in-process cache for a few seconds in front of Firestore.

Records nobody reads go stale, by design. Anything computed over all records (search indexes, rankings, aggregate statistics) sees data of mixed age. That's the most important limitation, and it's the subject of the last section.

Checklist#

Checklist

  • Every cached record stores fetchedAt, and the UI shows it
  • TTLs are chosen per audience, with owners on a shorter TTL than visitors
  • Fresh reads cost one document read; the lease transaction runs only on stale, unleased reads
  • The lease is claimed in a transaction that re-checks staleness, lease and backoff
  • Task IDs are hashed from the record key and the generation being replaced
  • The refresh queue's dispatch rate and concurrency are set from the upstream quota
  • Upstream failures set a per-record nextAttemptAt with exponential backoff and jitter, and are acknowledged
  • Permanent upstream errors are stored as a terminal state, not retried
  • Manual refresh has a cooldown and returns the time remaining
  • Anonymous and crawler reads never start a refresh
  • A per-source circuit breaker pauses enqueues during an upstream outage

When not to do this#

Keep the cron job when freshness is required whether or not anyone reads the data. Scheduled reports and digests have to be built from current data at a fixed time, and an email summary can't wait for a page view. Compliance deadlines work the same way: retention deletions, credential and token expiry checks, and reconciliations that an auditor expects to run on a schedule. Anything that has to detect changes, such as alerts on a price threshold or a status change, needs polling, because nobody is reading the record when it changes. Search indexes and cross-record aggregates also need every record reasonably fresh, which on-read refresh deliberately doesn't give you. And if your dataset is small, or most records are read every day, a nightly batch is simpler, batches upstream calls more efficiently, and costs about the same. I often run both: refresh-on-read for what people look at, and a small, slow sweep for the handful of fields that must be fresh for everyone.

Share
All articles →

Early-stage systems rarely go broke on traffic. They bleed money while idle. How to design a GCP stack whose cost stays near zero when no one is using it.

13 min

Where the money goes when users upload photos, and a GCS upload-and-variants pipeline that lets a CDN serve small, immutable files for years.

16 min

What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.

14 min