Building & Shipping

The Discount Was Real. The Cache Was Stale.

Brett Ridenour Brett Ridenour · Published September 2026

I made a promo code on staging. I went to attach it to a checkout flow. The picker was empty. No error. No spinner. Just an empty list, sitting there, pretending my discount didn’t exist.

The first instinct is always the same. Did the write fail? Did the API return the wrong shape? Is there some tenant filter I forgot? I opened the network tab and the server was fine — a POST /promo-codes came back 201, the row was in Postgres, the promo existed. It just wasn’t visible from the screen that needed to see it.

The bug turned out to be one of the least glamorous kinds you can ship. The screen was showing a five-minute-old cache, and the mutation that created the promo had no idea that particular cache was its problem.

What the logs actually showed

The clue was in the shape of the timeline, not any single request. Here is what the staging API log looked like when I lined the events up.

  1. 16:07:34
    Opened the checkout flow editor
    GET /checkout-flows/offer-options → 200, empty array. No promos existed yet.
  2. 16:08:48
    Created a promo code
    POST /promo-codes → 201. The promo is real. It is in the database.
  3. 16:10:33
    Went back and saved the flow
    offer-options was never requested again. The editor served the empty array from 16:07 straight from memory.

Read that middle line once. The mutation succeeded. The row exists. And then the editor, on remount, just… didn’t ask. It had asked two and a half minutes earlier, and as far as it was concerned, that answer was still fresh.

That is TanStack Query doing exactly what I told it to do.

Why it happened

Two lines of configuration, in two different files, that had never been in the same room together.

The first line, in the query client setup, said this:

// apps/web/client/src/lib/queryClient.ts:33
staleTime: 1000 * 60 * 5, // 5 minutes
refetchOnWindowFocus: false,

Any list fetched in the last five minutes is served from cache. Even when a freshly remounted screen desperately needs the newer version. Even when I, the user, just created the thing it is looking for.

The second line, in the mutation hook, said this:

// apps/web/client/src/hooks/use-promo-codes.ts:79
queryClient.invalidateQueries({ queryKey: ['/api/promo-codes'] });

“When somebody creates a promo, invalidate the promo-codes list.” Which sounds correct. It is correct, for the screen that shows the promo-codes list.

The picker on the checkout flow editor, on the other hand, was keyed like this:

// apps/web/client/src/components/checkout-flows/offer-editor.tsx:30
useQuery({ queryKey: ['checkout-offer-options', locationId], ... })

Different string. Different query. TanStack has no idea those two things are the same underlying data. From its point of view, one client cares about “the list at /api/promo-codes” and the other cares about “checkout offer options for a location.” Both queries hit the promo table on the server. Only one of them ever gets told when the promo table changed.

The mutation is broadcasting on channel A. The screen is listening on channel B. They never meet.

The two query keys, side by side

The category of bug this actually is

This is not a checkout bug. This is a naming convention bug that only reveals itself when two engineers (or the same engineer, six months apart) pick different key strings for the same underlying data.

Every mature TanStack app hits this exactly once and then never again, because the fix is structural, not local. You stop letting anyone hand-roll a queryKey array in a component, and you make a factory.

lib/query-keys.ts
− Before — string soup
// Everywhere in the codebase, ad hoc
useQuery({ queryKey: ['checkout-offer-options', locationId] })
useQuery({ queryKey: ['/api/promo-codes'] })
useQuery({ queryKey: ['promos', 'active', locationId] })

// And when you invalidate, you guess which strings to hit
queryClient.invalidateQueries({ queryKey: ['/api/promo-codes'] })
+ After — one prefix, one invalidate
// One file: lib/query-keys.ts
export const queryKeys = {
promos: {
  all: (loc: string) => ['promos', loc] as const,
  list: (loc: string, filters?: PromoFilters) =>
    ['promos', loc, 'list', filters] as const,
  offerOptions: (loc: string) =>
    ['promos', loc, 'offer-options'] as const,
},
} as const

// One invalidation catches everything under the prefix
queryClient.invalidateQueries({ queryKey: queryKeys.promos.all(loc) })

The load-bearing detail is that TanStack matches query keys by prefix. If every promo-related query in the entire app starts with ['promos', locationId, ...], then a single invalidateQueries({ queryKey: ['promos', locationId] }) invalidates all of them at once. The list screen refetches. The picker refetches. Every filter combination of the list refetches. You cannot forget one, because you did not have to remember them individually in the first place.

It is the same idea as scoping CSS with a component name, or namespacing an env variable. Once you commit to the prefix, the whole class of “did I invalidate the right key” errors disappears.

What I actually shipped as the fix

Two things, in order:

  1. The immediate patch. Set staleTime: 0 on the offer-options query, and turn refetchOnMount: 'always' back on for that specific hook. Ugly, but it makes the bug impossible today without waiting for a refactor. The public checkout can keep its longer caching. Only the admin editors need to lean on freshness.

  2. The structural fix, as a GitHub issue. Add lib/query-keys.ts. Migrate every hand-rolled queryKey array to a factory entry. Add a lint rule that flags new queryKey: [...] literals outside the factory. Add one unit test that renders the offer editor, fires the create-promo mutation, and asserts the picker refetches — the guard test I should have had six months ago.

The reason it is not one PR is that “rename every query key” is a boring, high-blast-radius change that deserves its own review, and the immediate patch stops the bleeding today. Structural fixes ship better when they are not on fire.

The takeaway

If your app has more than about five useQuery calls, the query keys have already started disagreeing with each other and you just haven’t noticed yet. The disagreement is silent. It survives your tests. It waits patiently for a user to do the exact pair of actions that reveals it — and then it looks, from the outside, like your server is broken.

The database was never lying. The API was never lying. My cache was telling the truth about a moment three minutes ago, which is, in software, the worst kind of lie there is.