I ran a diff between main and production yesterday morning and did the math. Two hundred and ninety commits. A fast-forward — no branches to reconcile, no rebase drama — but 290 commits is not a hotfix. It is a season.
Inside those commits was a redesigned checkout, a rewritten financial ledger with a one-time data conversion, a durable notification queue, four security fixes, ten new database migrations, and a filter dropdown I had not thought about in a month. Any one of those is a Tuesday for a big team. Together, on a codebase that a live operator was currently taking bookings on, they were the thing I was going to lose sleep over.
This is the frame I ended up using to push it. Not the code — the shape of the plan. Because when a solo SaaS founder promotes 290 commits at once, the interesting decisions are almost never in the diff.
The audit that made the plan possible
Before I wrote a single runbook step, I made myself write a boring document nobody was going to read but me. One table per area of the release, with three columns: what changes, who sees it on day one, and can I turn this specific thing off after it ships.
That third column is the whole game.
Once I had the table, the risk profile was obvious. Four of the seven areas were dormant — code that shipped but stayed off until an operator flipped a toggle. Those got a green check and I stopped worrying about them.
The other three were the ones that actually change on day one. And each of them needed its own escape hatch — separate from the git rollback, because the git rollback is the whole 290.
Per-tenant feature flags are the escape hatch
The scariest customer-visible change in the release was a full checkout redesign. New landing filters, a new payment step, a rewritten availability engine, a “most popular” badge, a rating that was fake and finally got yanked. Real UI. Real color changes. The operator’s actual customers were about to see it.
I did not want to bet the release on it.
So the checkout redesign got a per-tenant boolean:
// apps/api/src/lib/location-flags.ts
export async function checkoutV2Enabled(locationId: string): Promise<boolean> {
const settings = await db
.from('location_settings')
.select('features_enabled')
.eq('location_id', locationId)
.single();
return settings?.features_enabled?.checkout_v2 === true;
}
Two things about this that I want to say out loud, because I got them wrong on an earlier release:
-
It is read on every request. There is no cache. It hits Postgres. That is on purpose. If I need to switch checkout back to the legacy pages at 3am because a shopper reports a broken card form, I want the flip to take effect on the next page load, not after a five-minute TTL expires or a deploy runs. The database read is a rounding error next to the payment API call that comes right after it.
-
It scopes to a single tenant. One operator can be on the new checkout while every other location stays on the old one. That is what lets me stage the redesign against a real live customer instead of a synthetic staging account, without dragging every other operator along.
The escape-hatch shape ends up being a one-line SQL update:
UPDATE location_settings
SET features_enabled = features_enabled || '{"checkout_v2": false}'::jsonb
WHERE location_id = 'loc_xxx';
I can run that from a psql prompt in the time it takes to type it. No deploy. No cache invalidation. The next request the operator’s checkout serves is the legacy page. That is the escape hatch.

The lesson from earlier releases is that the scariest thing in the diff should be the thing you can turn off the fastest, and “push the old sha” is not fast enough when it drags 289 other commits back with it.
Sequencing changes are risk changes
The other insight, which took me embarrassingly long to internalize, is that the order of the release steps is itself a risk decision. Not a checklist. A design.
Here is the order I settled on for this push:
- Step 1Apply database migrations against production PostgresTen migrations, all additive. They run before the app code that reads the new columns. If any migration fails, nothing else has happened yet — I can walk away.
- Step 2Push main to production, wait for the deployThis is the moment the new app code becomes live. If it crashes, the migrations from step 1 are harmless — the old code doesn't read them.
- Step 3Run the one-time ledger conversion (30–45 minutes)The financial system swap. During this window, money-moving actions (refunds, cancellations, card charges) return a 409 while balance reads still work. Read-only for the operator.
- Step 4Verify the conversion with a fresh dry runThe same script, in read-only mode, against the converted data. If it wants to change anything, something is wrong. I stop and investigate before flipping the checkout.
- Step 5Flip checkout_v2 = true for the one live tenantThis is when the redesign becomes visible to the operator's customers. Every prior step has been invisible to them.
- Step 6Real booking with my own card, then refund itThe end-to-end smoke test. If the whole pipeline — quote → payment → webhook → ledger → refund — comes back clean, the release is done.
The step that matters most in that list is step 5. Not because it is the hardest — the ledger conversion in step 3 is by far the scariest — but because it is the point at which risk becomes visible to the customer. Everything before step 5 is a risk I own privately. Everything after step 5 is a risk the operator’s shoppers own with me.
By putting the customer-visible flip last, I get to run the whole financial migration in a state where the operator’s shoppers still see the known-good checkout pages. If the ledger conversion goes sideways, they never notice — they just get a 409 on a refund for half an hour, which is a call I can handle. The scary money change happens under the covers first. The scary UI change comes only after the ledger passes a fresh dry run.
If I had done these in the opposite order — flip checkout first, convert the ledger later — a bug in the new checkout page and a bug in the ledger conversion would look like the same incident, and I would be triaging two systems at once while a shopper is trying to pay.
Ship the invisible risk before the visible one.
— A rule I keep having to relearn
What this actually costs
The whole release fits in a single 90-minute window if nothing goes wrong. There are five reversible checkpoints, and the one irreversible one — the ledger conversion — is bracketed by a full Postgres dump on one side and a verification dry run on the other. If a bug slips through, either the per-tenant boolean or a physical database restore gets me back to a known state.
Two hundred and ninety commits sounds terrifying until you break them into “what changes for the operator on day one” and “what stays dormant.” Four of the seven areas are dormant. Two of the three live ones are covered by rollback commands I can type in one line. The one that requires a database restore is the one I take a real backup for.
The point is not that this release is safe. It is that I made the decisions about which parts get to be scary, and where the escape hatches live, on a whiteboard before I opened a terminal. When solo founders talk about “shipping to prod,” this is the part that gets skipped in the story — the sequencing, the boolean, the boring table with the rollback column. It is also the part that decides whether the ship is uneventful or the ship is the blog post you write at 4am.
I would rather have the boring release. I have already had the other kind.