Dead-letter queues: a small team's safety net

A dead-letter queue turns a stuck message into a ticket instead of an outage. How I set one up on Pub/Sub, alert on it, and replay safely after a fix.

Published

Updated

—

Reading time

15 min

A message that can never succeed is the most expensive thing a queue can hold. It retries forever, burns compute on every attempt, fills the error logs with the same stack trace, and, if the subscription uses ordering keys, it blocks every message behind it. On a large team, someone on call eventually notices the retry graph and deals with it. On a team of two, nobody is watching that graph at 3 a.m., and the first sign is a user asking why their order confirmation never arrived. A dead-letter queue (DLQ) is how I make that failure mode boring: after a bounded number of attempts, the broker moves the message somewhere safe, an alert fires, and a human looks at it during working hours. This article is for engineers running consumers on Google Cloud Pub/Sub, though the pattern applies to SQS, RabbitMQ, or any queue with redelivery. I'll cover what poison messages look like, how to configure dead-lettering and its IAM, how to alert on it, how to inspect what landed there, and how to replay it without making things worse.

This builds on idempotent Pub/Sub consumers, which covers how a consumer classifies failures and makes side effects happen once. Here the focus is on what happens to the messages that fall through.

The constraints I design for#

These are the requirements I write down before configuring anything:

  • No message is lost silently. Acking a message you failed to process is data loss. Dropping it after N retries without keeping it is also data loss.
  • No message retries forever. A permanently failing message must stop consuming compute and log volume within minutes, not days.
  • A human finds out, within hours. A DLQ nobody checks is a slower way of losing data.
  • Recovery is a command, not a project. After the bug is fixed, putting the stuck messages back through the pipeline should take one reviewed command.
  • It costs nothing while it's empty. Most of the time the DLQ holds zero messages, and it should cost close to zero.

What poison messages look like#

Before building the safety net, it's worth being precise about what it catches. In my experience, messages end up dead-lettered for one of five reasons:

  • Malformed payloads. A producer shipped a schema change before the consumer, a field is null that was never supposed to be, or someone published a test message by hand from the console.
  • References to data that doesn't exist. The event points to a record that was deleted, or to one that hasn't been written yet because the producer published before committing.
  • Bugs that only some inputs trigger. A Unicode edge case, an integer overflow, a timezone boundary. The consumer is fine for 99.99% of messages and crashes on this one every time.
  • Downstream rejections that won't change. A third-party API returns 400 or 422 for this input. Retrying a validation error doesn't help.
  • Long outages. A dependency is down for longer than the retry budget. These messages aren't poison at all; they're healthy messages that ran out of attempts. This case matters for replay, because they usually succeed as soon as you send them again.

The first four should ideally be caught by the consumer itself and quarantined on the first attempt. The DLQ is the backstop for the ones your classification missed, and for the fifth case, which no classification can catch.

Options#

ApproachStops infinite retriesKeeps the messageNeeds consumer codeCost when idle
No DLQ, retry until retention expiresNoUntil retention (7 days by default), then goneNoRetry compute plus storage
App-level quarantine (write to a database, then ack)Yes, for errors you classifyYesYes, in every consumerOne collection or table
Broker DLQ, one shared topic for all subscriptionsYes, for every failureYes, if the DLQ has a subscriptionNoOne topic plus one subscription
Broker DLQ, one topic per consumer subscriptionYes, for every failureYes, if the DLQ has a subscriptionNoIdle topics and subscriptions cost nothing

App-level quarantine and a broker DLQ aren't alternatives. Quarantine handles the errors you anticipated, fast and with context. The DLQ handles everything else, including crashes, timeouts and bugs in the quarantine code itself.

The decision#

Every push or pull subscription that triggers a side effect gets its own dead-letter topic, and every dead-letter topic gets exactly one pull subscription that exists only to hold messages for inspection. I alert on that subscription's backlog, and I replay through a dedicated replay topic per consumer, never through the original topic. Per-consumer DLQs cost nothing extra when idle, and they make the alert tell you which consumer is broken without anyone having to read attributes first.

One dead-letter topic, one inspection subscription and one replay topic per consumer. Only the broken consumer sees replayed messages.

Implementation#

The examples use an orders topic with an orders-fulfilment push subscription delivering to a Cloud Run service.

Step 1: topics and subscriptions#

infra/dlq.shbash
PROJECT_ID="my-project"
CONSUMER="orders-fulfilment"
URL="$(gcloud run services describe fulfilment --region=europe-west1 --format='value(status.url)')"
PUSH_SA="pubsub-push@${PROJECT_ID}.iam.gserviceaccount.com"
 
# 1. The dead-letter topic and its inspection subscription.
gcloud pubsub topics create "${CONSUMER}-dlq"
gcloud pubsub subscriptions create "${CONSUMER}-dlq-inspect" \
  --topic="${CONSUMER}-dlq" \
  --message-retention-duration=31d \
  --expiration-period=never
 
# 2. The source subscription, with a retry policy and dead-lettering.
gcloud pubsub subscriptions create "${CONSUMER}" \
  --topic=orders \
  --push-endpoint="${URL}/pubsub/orders" \
  --push-auth-service-account="${PUSH_SA}" \
  --ack-deadline=60 \
  --min-retry-delay=10s \
  --max-retry-delay=600s \
  --dead-letter-topic="${CONSUMER}-dlq" \
  --max-delivery-attempts=10
 
# 3. The replay path: same endpoint, its own topic, same DLQ.
gcloud pubsub topics create "${CONSUMER}-replay"
gcloud pubsub subscriptions create "${CONSUMER}-replay" \
  --topic="${CONSUMER}-replay" \
  --push-endpoint="${URL}/pubsub/orders" \
  --push-auth-service-account="${PUSH_SA}" \
  --ack-deadline=60 \
  --min-retry-delay=10s \
  --max-retry-delay=600s \
  --dead-letter-topic="${CONSUMER}-dlq" \
  --max-delivery-attempts=10 \
  --expiration-period=never

Three flags here matter more than they look.

--max-delivery-attempts accepts 5 to 100 and defaults to 5. I use 10 with a 10-second to 10-minute backoff, which gives a transient outage somewhere in the order of half an hour of retries before the message is parked (illustrative: the exact spacing between attempts is up to Pub/Sub). The documentation notes that the count is approximate because forwarding is best-effort, so don't build logic that expects exactly N attempts.

--expiration-period=never on the inspection subscription is the one people forget. Subscriptions expire after 31 days of inactivity by default, and activity means pulls or successful pushes. A healthy DLQ subscription is never pulled, so after a month of good behaviour it deletes itself, and the next dead-lettered message is published to a topic with no subscriptions. The documentation is blunt about that case: messages published to a topic with no subscriptions are lost.

--message-retention-duration=31d is the maximum for a subscription. It's the upper bound on how long you have to notice and act. Your alert should fire long before it matters.

Step 2: the IAM grants for forwarding#

Dead-lettering is performed by the Pub/Sub service agent, not by your consumer's identity. It needs to publish to the dead-letter topic and to acknowledge the message on the source subscription:

infra/dlq-iam.shbash
PROJECT_NUMBER="$(gcloud projects describe "$PROJECT_ID" --format='value(projectNumber)')"
PUBSUB_AGENT="serviceAccount:service-${PROJECT_NUMBER}@gcp-sa-pubsub.iam.gserviceaccount.com"
 
gcloud pubsub topics add-iam-policy-binding "${CONSUMER}-dlq" \
  --member="$PUBSUB_AGENT" --role=roles/pubsub.publisher
 
for SUB in "${CONSUMER}" "${CONSUMER}-replay"; do
  gcloud pubsub subscriptions add-iam-policy-binding "$SUB" \
    --member="$PUBSUB_AGENT" --role=roles/pubsub.subscriber
done

I grant these on the individual topic and subscriptions, not at project level. The service agent only needs to touch the resources involved in dead-lettering, and the bindings live next to the resources they protect (see least privilege in practice for why I avoid project-wide grants). If either binding is missing, Pub/Sub can't forward the message, so it keeps being redelivered to the failing consumer instead. Your retries continue, your DLQ stays empty, and the alert you rely on never fires. I check both bindings in CI with gcloud pubsub topics get-iam-policy rather than trusting that someone remembered.

Step 3: alerts that page the right person#

Pub/Sub exports the metrics I need in Cloud Monitoring:

  • pubsub.googleapis.com/subscription/num_undelivered_messages: the backlog of a subscription. On the inspection subscription, anything above zero means a message was dead-lettered and nobody has dealt with it.
  • pubsub.googleapis.com/subscription/oldest_unacked_message_age: age in seconds of the oldest message in that backlog. This is the one that protects against retention expiry.
  • pubsub.googleapis.com/subscription/dead_letter_message_count: a delta count of messages forwarded from the source subscription. Useful for a dashboard showing which consumer produces dead letters and how fast.

Both gauges are sampled every 60 seconds and can take up to a couple of minutes to appear, so don't expect a page within seconds. I use two conditions on the inspection subscription:

monitoring/dlq-alert.yamlyaml
displayName: "DLQ: orders-fulfilment"
combiner: OR
severity: WARNING
documentation:
  mimeType: text/markdown
  content: |
    Messages were dead-lettered from `orders-fulfilment`.
    Runbook: inspect `orders-fulfilment-dlq-inspect`, fix, replay via `orders-fulfilment-replay`.
conditions:
  - displayName: "DLQ backlog above zero"
    conditionThreshold:
      filter: >-
        resource.type = "pubsub_subscription" AND
        resource.labels.subscription_id = "orders-fulfilment-dlq-inspect" AND
        metric.type = "pubsub.googleapis.com/subscription/num_undelivered_messages"
      comparison: COMPARISON_GT
      thresholdValue: 0
      duration: 300s
      aggregations:
        - alignmentPeriod: 300s
          perSeriesAligner: ALIGN_MAX
  - displayName: "Oldest dead letter older than 7 days"
    conditionThreshold:
      filter: >-
        resource.type = "pubsub_subscription" AND
        resource.labels.subscription_id = "orders-fulfilment-dlq-inspect" AND
        metric.type = "pubsub.googleapis.com/subscription/oldest_unacked_message_age"
      comparison: COMPARISON_GT
      thresholdValue: 604800
      duration: 300s
      aggregations:
        - alignmentPeriod: 300s
          perSeriesAligner: ALIGN_MAX
alertStrategy:
  autoClose: 86400s
terminalbash
gcloud monitoring policies create \
  --policy-from-file=monitoring/dlq-alert.yaml \
  --notification-channels="projects/${PROJECT_ID}/notificationChannels/CHANNEL_ID"

The first condition tells you something broke. The second is a reminder that you still haven't dealt with it, and it fires with three weeks of retention to spare. I route the first to a chat channel and the second to email, because a week-old dead letter is a process failure, not an incident.

I also keep a standing alert on the source subscription's oldest_unacked_message_age. That catches the opposite failure: the consumer is down and messages are piling up before they reach the DLQ.

CostWhat the safety net costs

At the time of writing, Pub/Sub pricing charges throughput at $40 per TiB after a free 10 GiB per month, with a 1 KB minimum per request, and charges storage at $0.27 per GiB-month for unacknowledged messages retained more than a day after publishing. An empty dead-letter topic and subscription cost nothing. Illustratively, 10,000 dead-lettered messages of 2 KB each come to about 20 MB: the forwarding throughput and a month of storage together round to a fraction of a cent. The real cost of a DLQ is the retries before it (ten delivery attempts means ten handler invocations) and the logs each failed attempt writes. Google has also announced per-policy charges for Cloud Monitoring alerting starting no sooner than September 2027, so check the observability pricing page if you plan hundreds of alert policies.

Step 4: inspecting what landed#

When the alert fires, the first job is to look without changing anything. Pulling without --auto-ack leases the messages for the ack deadline and then returns them to the backlog:

Inspect dead lettersbash
gcloud pubsub subscriptions pull orders-fulfilment-dlq-inspect \
  --limit=10 --format=json \
  | jq '.[] | {
      id: .message.messageId,
      attempts: .message.attributes.CloudPubSubDeadLetterSourceDeliveryCount,
      source: .message.attributes.CloudPubSubDeadLetterSourceSubscription,
      published: .message.attributes.CloudPubSubDeadLetterSourceTopicPublishTime,
      body: (.message.data | @base64d)
    }'

When Pub/Sub forwards a message, it wraps the original data and attributes in a new message with a new message ID, and adds attributes that identify where it came from: CloudPubSubDeadLetterSourceDeliveryCount, CloudPubSubDeadLetterSourceSubscription, CloudPubSubDeadLetterSourceSubscriptionProject and CloudPubSubDeadLetterSourceTopicPublishTime. That's enough to join the message to the consumer's logs. If the consumer logs the event ID from the payload on every attempt, a single log query shows every failed attempt with its error.

gcloud pull isn't guaranteed to return everything in the backlog, even with a higher --limit, so for anything beyond a quick look I use the replay tool in dry-run mode, which pages through the subscription and prints a summary.

Step 5: a replay tool that can't make things worse#

Replay is where a DLQ turns into a second incident. I've seen three ways it goes wrong: republishing to the original topic (which delivers the message to every subscription on that topic, not just the broken one), dropping the original attributes (so the consumer routes it differently), and replaying in a loop while the bug is still deployed. The tool below is built to avoid all three.

tools/replay-dlq.tsts
import { PubSub, v1 } from "@google-cloud/pubsub";
 
const MAX_REPLAYS = 3;
 
function originalAttributes(attrs: Record<string, string>) {
  // Keep the producer's attributes; drop what dead-lettering added.
  return Object.fromEntries(
    Object.entries(attrs).filter(([k]) => !k.startsWith("CloudPubSubDeadLetter")),
  );
}
 
async function main() {
  const [project, consumer, mode = "dry-run", limitArg = "100"] = process.argv.slice(2);
  const limit = Number(limitArg);
 
  const sub = new v1.SubscriberClient();
  const subscription = sub.subscriptionPath(project, `${consumer}-dlq-inspect`);
  const replayTopic = new PubSub({ projectId: project }).topic(`${consumer}-replay`);
 
  let handled = 0;
  let overLimit = 0;
 
  while (handled < limit) {
    const [res] = await sub.pull({ subscription, maxMessages: Math.min(50, limit - handled) });
    const received = res.receivedMessages ?? [];
    if (received.length === 0) break;
 
    for (const { ackId, message } of received) {
      if (!ackId || !message) continue;
      const attrs = message.attributes ?? {};
      const replays = Number(attrs["x-replay-count"] ?? "0");
 
      if (mode !== "apply" || replays >= MAX_REPLAYS) {
        if (replays >= MAX_REPLAYS) overLimit++;
        console.log(JSON.stringify({ id: message.messageId, replays, attrs }));
        // Leave it in the DLQ and make it visible again straight away.
        await sub.modifyAckDeadline({ subscription, ackIds: [ackId], ackDeadlineSeconds: 0 });
      } else {
        await replayTopic.publishMessage({
          data: Buffer.from(message.data as Uint8Array),
          attributes: { ...originalAttributes(attrs), "x-replay-count": String(replays + 1) },
        });
        // Ack only after the publish succeeded: a crash in between means a duplicate, never a loss.
        await sub.acknowledge({ subscription, ackIds: [ackId] });
      }
      handled++;
    }
  }
 
  console.error(`handled=${handled} mode=${mode} over-replay-limit=${overLimit}`);
  await sub.close();
}
 
main().catch((err) => {
  console.error(err);
  process.exit(1);
});

Run it with npx tsx tools/replay-dlq.ts my-project orders-fulfilment for a dry run, and add apply 500 to replay up to 500 messages. A few design choices matter here:

  • Replay targets one consumer. The -replay topic has one subscription pointing at the same endpoint, so the other consumers on orders never see these messages. The handler doesn't care which subscription delivered the message.
  • Original attributes survive. The payload bytes and the producer's attributes are passed through unchanged. Only the dead-letter bookkeeping is stripped, and one counter is added.
  • Replays are bounded. A message that has already been replayed three times stays in the DLQ. If three fixes haven't worked, a human needs to read it.
  • Publish, then ack. If the tool dies mid-batch, the worst outcome is a message that is both replayed and still in the DLQ. That's a duplicate, and the consumer is idempotent, so it's harmless. The reverse order would risk losing it.
  • Consumers must be idempotent. This is non-negotiable. Some of these messages partly succeeded before failing, and a replay runs the handler from the top.

Step 6: the triage runbook#

When the alert fires, I follow the same steps every time, and they're written in the alert's documentation field so whoever gets paged doesn't have to remember them:

  1. Look, don't touch. Run the replay tool in dry-run mode. Count the messages and group them by error.
  2. Find the failed attempts in the logs. Search the consumer's logs for the event IDs, and read the error on the last attempt.
  3. Classify. Outage (the dependency is back, so replay now), bug (fix, deploy, then replay), or bad data (the message is wrong and will never be right).
  4. For a bug, fix and deploy first. Replaying into the same code just sends the messages back to the DLQ after ten more attempts.
  5. Replay a small batch. Run the tool with apply 5, watch the consumer logs, then replay the rest.
  6. For bad data, record and discard. Save the payloads to a bucket or ticket with the reason, then ack them. Discarding is a decision someone makes and writes down, not the default.
  7. Close the loop. If the consumer should have quarantined these messages on the first attempt, add that classification so the next batch never reaches the DLQ.

Trade-offs and failure modes#

Too few attempts turns outages into dead letters. With the minimum of five attempts and a short backoff, a ten-minute dependency outage can dead-letter a large share of traffic. That's recoverable with the replay tool, but it's noisy. Size the attempts and backoff so a routine outage is absorbed by retries.

Too many attempts hides poison. At 100 attempts with a 10-minute maximum backoff, a poison message can retry for over half a day before anyone hears about it, writing an error log every time. If the consumer already quarantines known-bad input, a modest attempt count is enough.

Ordering keys. With message ordering, a failing message blocks later messages on its key until it's dead-lettered. After that, ordering continues past the gap, so the consumer processes message 5 before message 4 is replayed. If order matters for correctness, carry a version in the payload and reject stale writes, as described in the idempotency article.

The DLQ as a work queue. A DLQ that always holds a few hundred messages trains everyone to ignore the alert. If a class of failure is expected, handle it in code: quarantine, skip, or retry with its own schedule. The DLQ should be empty most days.

Silent misconfiguration. Missing IAM bindings, an expired inspection subscription or a typo in the dead-letter topic name all look the same from the outside: an empty DLQ and a quiet alert. I compare dead_letter_message_count against the inspection subscription's backlog on a dashboard, and I test the path in staging by publishing a deliberately malformed message and waiting for the alert.

Replaying personal data. Dead-lettered payloads can contain personal data, and retention in the DLQ is up to 31 days. Restrict who can pull from the inspection subscription (a roles/pubsub.subscriber grant on that subscription only), and make sure the replay tool's dry-run output doesn't end up in a shared chat channel.

Checklist#

Checklist

  • Every subscription that triggers a side effect has its own dead-letter topic
  • Every dead-letter topic has exactly one inspection pull subscription
  • Inspection subscriptions use --expiration-period=never and 31-day retention
  • The Pub/Sub service agent has publisher on the DLQ topic and subscriber on each source subscription
  • A retry policy with exponential backoff is set, and max delivery attempts is sized to outlast a routine outage
  • Alerts fire on inspection backlog above zero and on oldest dead letter older than a week
  • A separate alert watches the source subscription's oldest unacked message age
  • Consumers log the event ID on every attempt so dead letters can be joined to their errors
  • Replays go through a per-consumer replay topic, keep original attributes and are capped per message
  • The replay tool publishes before it acks, and consumers are idempotent
  • The triage runbook lives in the alert's documentation
  • The dead-letter path is tested in staging with a deliberately malformed message

When not to do this#

If a message is worthless once it's late (a presence heartbeat, a cache invalidation that the next write will supersede, a metrics sample), dead-lettering it just creates work. Let it expire, or ack it on failure and count the failures in a metric. The same goes for pipelines where the source of truth can regenerate every message on demand: a nightly rebuild job doesn't need a DLQ when rerunning the job is the recovery. Save the safety net for messages that represent something a user did or is waiting for.

Share
All articles →