Serving user images without burning money

Where the money goes when users upload photos, and a GCS upload-and-variants pipeline that lets a CDN serve small, immutable files for years.

Published

Updated

—

Reading time

16 min

The first version of user images in almost every product is the same: the client posts a file to the API, the API writes it to a bucket, and the frontend renders the bucket URL in an <img> tag. It works on day one and gets expensive quietly. A phone photo is 2 to 5 MB. A feed that shows twenty of them as 300-pixel thumbnails downloads tens of megabytes to draw a screen, every scroll, for every user, and you pay egress on every byte. Meanwhile the API is proxying uploads it never needed to touch, and the photos still carry the GPS coordinates of the user's kitchen. This article is for engineers building anything with user-generated images (avatars, listings, posts, attachments) on Google Cloud, although the shape of the solution works on any provider. It covers where the money actually goes, the pipeline I use, and the code to build it.

The constraints#

These are the requirements I start from when I design image handling, including for a large-scale social platform I'm building:

  • Egress is the bill. Storage is cheap; bytes leaving the cloud are not. Every design decision is judged by how many bytes it puts on the wire.
  • The API server should never carry image bytes. Proxying uploads burns CPU, memory, and request time on the most expensive tier of the system.
  • Untrusted input. Uploaded files are hostile until proven otherwise: wrong MIME types, decompression bombs, polyglot files, and metadata that leaks location.
  • Deletion must be real. Under GDPR, when a user deletes a photo or their account, the image has to stop being served, including from caches, within a bounded time.
  • Predictable cost at 10x traffic. Nothing in the serving path should cost more per request as traffic grows, and ideally the marginal request is served entirely from cache.

The cost anatomy of a user image#

Before comparing architectures, it helps to name every meter that runs. On Google Cloud, at the time of writing (check Cloud Storage pricing and Cloud CDN pricing for current rates in your region):

  • Storage, billed per GB-month and by storage class. Standard is the default; Nearline, Coldline and Archive are cheaper to hold but charge retrieval fees and have minimum storage durations (30, 90 and 365 days).
  • Operations. Class A operations (writes, lists) cost roughly ten times more than Class B operations (reads). Generating six variants per upload is six Class A writes; serving directly from the bucket is one Class B read per view.
  • Network egress from Cloud Storage to the internet, billed per GB and tiered by monthly volume. For image-heavy products this is usually the largest line by far.
  • CDN charges, if you put Cloud CDN in front: cache egress per GB (cheaper than origin egress at higher tiers), cache fill per GB when the CDN fetches from the bucket, a per-request lookup fee, and the load balancer's forwarding rule charged per hour.
  • Transformation compute, the CPU time to decode, resize and re-encode. AVIF encoding in particular is CPU-heavy.

Two observations drive everything below. The biggest lever is the number of bytes per view, which is decided by the variant you serve, not by the CDN. And the second biggest is how often the origin is touched, which is decided by your object naming and caching headers.

Options compared#

There are two separate decisions: when to transform (at upload, or on request) and what serves the bytes.

OptionTransform modelCost shapeOperational weightGood fit
Public GCS bucket, direct URLsAt upload (your code)Origin egress on every view; no fixed costMinimalLow traffic, internal tools
Cloud CDN with a backend bucketAt upload (your code)Cache egress + lookups + fill; small fixed LB costLoad balancer, certificate, DNSMost products on GCP at meaningful traffic
Firebase HostingNot for user content: serves deployed files, dynamic content needs a function or Cloud Run rewritePer-GB transfer, plus compute if you rewrite to a serviceLow for static sites, awkward for user uploadsYour app shell, not your users' photos
On-the-fly resizer (Cloud Run) behind Cloud CDNPer unique URL, on first requestCompute on every cache miss; cost spikes with cache churnYou own a hot, CPU-bound serviceMany layouts that need arbitrary sizes
Third-party image CDN (Cloudflare Images, imgix, Cloudinary and similar)On request, managedPer image stored, delivered or transformed, depending on vendorLowest; vendor lock-in on URLsSmall teams that want to buy the whole problem

On-the-fly resizing looks flexible, but it has a nasty property: every distinct width and format combination is a cache miss that costs compute, and if the width comes from the URL, anyone can generate unlimited misses by iterating ?w=301, ?w=302, and so on. If you go that route, allow-list the sizes, at which point you have reinvented fixed variants with worse latency on first view.

The decision#

I generate a small, fixed set of variants at upload time, store them in a public bucket under immutable names, and serve them through Cloud CDN with a one-year Cache-Control. Originals never become public: the client uploads directly to a private staging bucket with a V4 signed URL, an event-triggered function validates and processes the file, and the staging object is deleted.

The API issues permissions, not bytes; the CDN serves immutable variants.

The variant set is deliberately small: three widths (320, 768 and 1536 pixels) in two formats (AVIF and WebP). That covers thumbnails, in-feed images and full-screen views on high-density displays. Six objects per image is a trade: more storage and more Class A writes at upload, in exchange for never sending a 3 MB file to render a 300-pixel box.

Implementation#

Step 1: issue a V4 signed upload URL#

The API authenticates the user, decides what they are allowed to upload, and returns a short-lived URL that permits exactly one PUT of one object with one content type. The V4 signing process binds the method, object name, expiry and any signed headers into the signature.

src/routes/uploads.tsts
import { randomUUID } from "node:crypto";
import { Storage } from "@google-cloud/storage";
 
const storage = new Storage();
const stagingBucket = storage.bucket(process.env.STAGING_BUCKET!); // validate in your config module
 
const MAX_BYTES = 10 * 1024 * 1024;
const ALLOWED_TYPES = new Set(["image/jpeg", "image/png", "image/webp", "image/avif"]);
 
export async function createUploadUrl(userId: string, contentType: string, declaredSize: number) {
  if (!ALLOWED_TYPES.has(contentType)) {
    throw new Error("unsupported_type"); // map to a 400 in your framework
  }
  if (!Number.isInteger(declaredSize) || declaredSize <= 0 || declaredSize > MAX_BYTES) {
    throw new Error("too_large");
  }
 
  const uploadId = randomUUID();
  const objectName = `incoming/${userId}/${uploadId}`;
 
  const [url] = await stagingBucket.file(objectName).getSignedUrl({
    version: "v4",
    action: "write",
    expires: Date.now() + 10 * 60 * 1000, // 10 minutes
    contentType,
    extensionHeaders: {
      "x-goog-content-length-range": `0,${MAX_BYTES}`,
    },
  });
 
  return {
    uploadId,
    url,
    headers: {
      "Content-Type": contentType,
      "x-goog-content-length-range": `0,${MAX_BYTES}`,
    },
  };
}

Things worth knowing about this code:

  • On Cloud Run there is no key file to sign with. The client library falls back to the IAM signBlob API, so the service's runtime service account needs roles/iam.serviceAccountTokenCreator on itself. The first deploy without that grant fails with a permission error on signing, not on upload.
  • The client must send exactly the signed headers. A browser that sends a different Content-Type gets a 403 SignatureDoesNotMatch. Return the headers alongside the URL so the frontend cannot get them wrong.
  • The declared size is advisory. The signed x-goog-content-length-range header asks Cloud Storage to reject bodies outside the range. If you need a guarantee enforced purely by policy, a V4 signed POST policy with a content-length-range condition is the belt-and-braces option. Either way, the processing function re-checks the real size.
  • The staging bucket is private, has uniform bucket-level access and public access prevention enforced, and needs a CORS configuration allowing PUT from your web origin.

The object name contains the user ID and a random upload ID. It is never shown to anyone and never served, so it does not need to be pretty.

Step 2: validate and generate variants on upload#

An Eventarc trigger on google.cloud.storage.object.v1.finalized for the staging bucket invokes a Cloud Run function. Using a separate staging bucket matters: if the function wrote its output into the bucket that triggers it, every write would fire another event.

The function uses sharp, which wraps libvips and is fast and memory-efficient for this job.

functions/make-variants/index.tsts
import { createHash } from "node:crypto";
import { cloudEvent } from "@google-cloud/functions-framework";
import { Storage } from "@google-cloud/storage";
import sharp from "sharp";
 
type StorageObjectData = { bucket: string; name: string; size?: string };
 
const storage = new Storage();
const PUBLIC_BUCKET = process.env.PUBLIC_BUCKET!;
const MASTER_BUCKET = process.env.MASTER_BUCKET!;
const MAX_BYTES = 10 * 1024 * 1024;
const WIDTHS = [320, 768, 1536] as const;
const FORMATS = ["avif", "webp"] as const;
const ACCEPTED = new Set(["jpeg", "png", "webp", "avif"]);
const IMMUTABLE = "public, max-age=31536000, immutable";
 
cloudEvent<StorageObjectData>("makeVariants", async (event) => {
  const { bucket, name, size } = event.data!;
  if (!name.startsWith("incoming/")) return;
 
  const source = storage.bucket(bucket).file(name);
  if (Number(size ?? 0) > MAX_BYTES) {
    await source.delete({ ignoreNotFound: true });
    return;
  }
 
  let input: Buffer;
  try {
    [input] = await source.download();
  } catch (err: any) {
    if (err?.code === 404) return; // redelivered event, already processed
    throw err;
  }
 
  // Reject anything that is not really an image we accept, and cap decoded
  // pixel count so a tiny file cannot expand into gigabytes of memory.
  const opts = { limitInputPixels: 50_000_000, failOn: "error" as const };
  const meta = await sharp(input, opts).metadata().catch(() => null);
  if (!meta?.format || !ACCEPTED.has(meta.format)) {
    await source.delete({ ignoreNotFound: true });
    return;
  }
 
  const [, userId, uploadId] = name.split("/");
  // Deterministic per upload: a retried event produces the same names.
  const imageId = createHash("sha256").update(uploadId).update(input).digest("base64url").slice(0, 22);
 
  // autoOrient() applies EXIF orientation. sharp drops EXIF
  // (including GPS) from output unless you ask it to keep metadata.
  const base = sharp(input, opts).autoOrient();
 
  const jobs = WIDTHS.flatMap((width) =>
    FORMATS.map(async (format) => {
      const resized = base.clone().resize({ width, withoutEnlargement: true });
      const encoded =
        format === "avif" ? resized.avif({ quality: 50, effort: 4 }) : resized.webp({ quality: 75 });
      const buffer = await encoded.toBuffer();
      await saveOnce(PUBLIC_BUCKET, `v/${imageId}/${width}.${format}`, buffer, `image/${format}`, IMMUTABLE);
    }),
  );
 
  const master = await base.clone().resize({ width: 2560, withoutEnlargement: true }).webp({ quality: 90 }).toBuffer();
  jobs.push(saveOnce(MASTER_BUCKET, `m/${userId}/${imageId}.webp`, master, "image/webp", "private, no-store"));
 
  await Promise.all(jobs);
  // Record { uploadId, userId, imageId, width: meta.width, height: meta.height } in your database here,
  // using uploadId as the document key so retries are idempotent.
  await source.delete({ ignoreNotFound: true });
});
 
async function saveOnce(bucket: string, name: string, data: Buffer, contentType: string, cacheControl: string) {
  try {
    await storage.bucket(bucket).file(name).save(data, {
      resumable: false,
      contentType,
      metadata: { cacheControl },
      preconditionOpts: { ifGenerationMatch: 0 }, // never overwrite an existing object
    });
  } catch (err: any) {
    if (err?.code !== 412) throw err; // 412: already written by an earlier attempt
  }
}

The details that earn their keep:

  • metadata() is the real type check. The declared content type came from the client; the decoded format comes from the file's bytes. Anything sharp cannot decode, or decodes as a format you did not accept, is deleted.
  • limitInputPixels guards against decompression bombs: a small PNG that claims to be 50,000 by 50,000 pixels. sharp has a default limit; I set a lower one that matches what a phone camera produces.
  • Metadata is stripped by default. sharp does not copy EXIF, ICC or XMP to the output unless you call keepMetadata() or withMetadata(). Location data is gone before anything becomes public.
  • Idempotency. Events can be delivered more than once. Deterministic names, ifGenerationMatch: 0, and a database write keyed by the upload ID make a retry harmless.
  • Memory. A 1536-pixel AVIF encode of a 12-megapixel photo needs a few hundred megabytes. Give the function 1 GiB, set concurrency low (one or two requests per instance), and measure before tuning.

Note that HEIC photos from iPhones are not in the accepted list: the prebuilt sharp binaries do not decode HEVC-compressed HEIC, for patent reasons. Most browsers convert to JPEG on upload through a file input, but native apps should convert before uploading.

Step 3: immutable names and long cache headers#

The public URL of a variant is https://img.example.com/v/{imageId}/768.avif. The imageId is derived from the upload and its content, and an object at that name is never overwritten. That is what makes this header safe:

texttext
Cache-Control: public, max-age=31536000, immutable

Browsers and the CDN can keep the file for a year without revalidating, because a new image, or a re-encoded variant, always gets a new name. If you later change the encoder settings, you write new objects under a new prefix (v2/) and update references; you never mutate in place. This is the single decision that turns the origin bucket from something read on every view into something read once per edge location.

I do not deduplicate identical files across users. Pure content addressing would make two users who upload the same meme share one object, and then deleting it for one user would require reference counting. Mixing the upload ID into the hash keeps names unique per upload and deletion trivially safe.

Step 4: put Cloud CDN in front of the bucket#

A backend bucket with Cloud CDN behind a global external Application Load Balancer:

terminalbash
# The variants bucket must be readable by the CDN; only derived variants live here.
gcloud storage buckets add-iam-policy-binding gs://example-img-variants \
  --member=allUsers --role=roles/storage.objectViewer
 
gcloud compute backend-buckets create img-backend \
  --gcs-bucket-name=example-img-variants \
  --enable-cdn \
  --cache-mode=CACHE_ALL_STATIC \
  --default-ttl=86400 \
  --max-ttl=2592000
 
gcloud compute url-maps create img-map --default-backend-bucket=img-backend
 
gcloud compute ssl-certificates create img-cert --domains=img.example.com --global
gcloud compute target-https-proxies create img-proxy --url-map=img-map --ssl-certificates=img-cert
gcloud compute addresses create img-ip --global
gcloud compute forwarding-rules create img-https \
  --global --load-balancing-scheme=EXTERNAL_MANAGED \
  --address=img-ip --target-https-proxy=img-proxy --ports=443

The --max-ttl of 30 days is deliberate, and it is about deletion rather than performance: in CACHE_ALL_STATIC mode it caps how long the CDN honours the one-year max-age, which bounds how long a deleted image can linger at the edge even if an invalidation is missed.

Step 5: responsive markup#

Store the original dimensions with the image record, and let the browser pick the smallest file that fits:

htmlhtml
<picture>
  <source
    type="image/avif"
    srcset="https://img.example.com/v/Qm3x/320.avif 320w,
            https://img.example.com/v/Qm3x/768.avif 768w,
            https://img.example.com/v/Qm3x/1536.avif 1536w"
    sizes="(min-width: 768px) 600px, 100vw" />
  <source
    type="image/webp"
    srcset="https://img.example.com/v/Qm3x/320.webp 320w,
            https://img.example.com/v/Qm3x/768.webp 768w,
            https://img.example.com/v/Qm3x/1536.webp 1536w"
    sizes="(min-width: 768px) 600px, 100vw" />
  <img src="https://img.example.com/v/Qm3x/768.webp" width="1200" height="800"
       alt="Description written by the uploader" loading="lazy" decoding="async" />
</picture>

The width and height attributes give the aspect ratio so the layout does not shift while images load. sizes is the part people get wrong: it must describe the rendered width, or the browser will pick the 1536-pixel file for a thumbnail. The MDN guide to responsive images is the clearest reference.

Step 6: deletion and GDPR#

When a user deletes an image, the order matters:

  1. Remove the database record, so no page references the image any more.
  2. Delete the variants (gcloud storage rm "gs://example-img-variants/v/Qm3x/**" or the equivalent library call) and the master.
  3. Invalidate the CDN path: gcloud compute url-maps invalidate-cdn-cache img-map --path="/v/Qm3x/*" --async.

Invalidations are rate-limited, so for a full account deletion with thousands of images, I invalidate a per-user path where the URL structure allows it, or accept the --max-ttl bound and document it. Also check soft delete: buckets retain deleted objects for seven days by default, which is billed as storage and keeps the data recoverable. That is useful on the master bucket, but for the variants bucket I disable it or shorten it, and either way it belongs in the retention section of the privacy policy. Browser caches on other people's devices are out of your reach; what you control is that your URLs stop resolving.

Step 7: lifecycle rules#

Object Lifecycle Management cleans up what the pipeline leaves behind. On the staging bucket, anything still under incoming/ after a day is an abandoned or rejected upload:

staging-lifecycle.jsonjson
{
  "rule": [
    { "action": { "type": "Delete" }, "condition": { "age": 1, "matchesPrefix": ["incoming/"] } }
  ]
}

On the master bucket, masters are only read when you regenerate variants, so they can move to Coldline after 30 days. Coldline has a 90-day minimum storage duration and a per-GB retrieval fee, so this only pays off if regenerations are rare batch jobs. Apply with gcloud storage buckets update gs://example-img-staging --lifecycle-file=staging-lifecycle.json. Keep the variants bucket on Standard: it is read constantly by cache fills, and retrieval fees there would cancel the savings.

An illustrative monthly cost model#

CostIllustrative, not a quote

These numbers are illustrative, with rates rounded from the public price lists at the time of writing for a European region and a single volume tier. Use the Google Cloud pricing calculator with your own traffic before making decisions.

Assume 50 million image views a month, mostly thumbnails and feed images.

ScenarioBytes deliveredDominant chargesRough monthly order
Originals served straight from a bucket (avg 2.5 MB)~125 TBOrigin egressFive figures in USD
Variants from the bucket, no CDN (avg 60 KB)~3 TBOrigin egress, Class B readsLow hundreds
Variants through Cloud CDN, 95% hit ratio~3 TBCache egress, lookups, LB fixed costLow hundreds, lower latency

The shape is the lesson. Serving the right-sized variant is a roughly fortyfold reduction in bytes; that dwarfs everything else. At this volume, adding Cloud CDN is close to cost-neutral: cache egress is cheaper than origin egress per GB, but you add lookup fees and an hourly forwarding-rule charge. You buy the CDN for latency and for how it scales into higher volume tiers, not for a dramatic saving at small scale. Storage for a million images, with six variants plus a master, is a few tens of dollars a month. Transformation compute depends on your image mix; measure the CPU seconds of your function on real uploads rather than trusting anyone's benchmark.

Trade-offs and failure modes#

  • Fixed variants mean redesigns cost a backfill. A new layout that needs a 1024-pixel width means regenerating from masters. That is a batch job, and it is why I keep masters at all.
  • Asynchronous processing creates a gap. For a second or two after upload, the variants do not exist. The client needs a pending state (show the local preview from the file input) and a way to learn the image is ready: poll the upload ID, or subscribe to a database document.
  • AVIF encoding is slow. At high effort values it can dominate function time. effort: 4 is a reasonable middle; measure before raising it.
  • A public bucket is public. Anything written there is readable by anyone with the URL. Never let the function write masters or originals to it, and consider an organisation policy that restricts allUsers bindings to buckets with a specific label.
  • Signed URL leakage. A signed URL is a bearer credential for its lifetime. Keep expiries short and scope each URL to one object name.
  • Missed invalidations. If step 3 of the deletion flow fails silently, the image lives at the edge until --max-ttl expires. Log and retry invalidations like any other side effect.

Checklist

  • Clients upload directly to a private staging bucket with a V4 signed URL scoped to one object, one content type and a short expiry.
  • The runtime service account can sign blobs, and the staging bucket has CORS for PUT.
  • A finalize-triggered function re-checks size, decodes with a pixel limit, and deletes anything that is not an accepted image.
  • Output is stripped of EXIF and written as a small fixed set of widths in AVIF and WebP.
  • Variant names are immutable and unique per upload, and carry public, max-age=31536000, immutable.
  • Writes use ifGenerationMatch: 0 and database writes are keyed by upload ID, so retries are harmless.
  • Cloud CDN fronts a variants-only bucket, with --max-ttl chosen with deletion in mind.
  • Markup uses srcset, an accurate sizes, and explicit width and height.
  • Deletion removes the record, the objects and the CDN cache entry, and soft-delete retention is a conscious choice.
  • Lifecycle rules clear abandoned uploads and age masters to a colder class.

When not to do this#

If your product shows a few thousand images a month, a private bucket plus short-lived signed read URLs, or a single public bucket with pre-sized uploads, is enough; a load balancer, certificate and function pipeline are operational weight you do not need yet. If your images are genuinely private (medical documents, ID scans), none of the caching advice applies: serve them through signed URLs or an authenticated proxy with Cache-Control: private, no-store. And if you have more layouts than engineers, a managed image CDN that transforms on request will cost more per image but less in attention, which at that stage is the scarcer resource.

Share
All articles →

Early-stage systems rarely go broke on traffic. They bleed money while idle. How to design a GCP stack whose cost stays near zero when no one is using it.

13 min

What a Cloud Run cold start is made of, how to measure each phase, and which fixes, and which min-instances bill, actually shorten it for your service.

14 min