Somewhere around hour thirty of a checkout rewrite that was not converting the way it was supposed to, I asked the question out loud — in a Claude session, with no human in the room — of exactly how hard it would be to switch the production account back to the old checkout. Not forever. Just for the afternoon. Just so the next payment attempt would land.
I had been in the new code for a week. I knew it better than the old one at that point. I had the postmortem half-written in my head. I had opinions about every line of the diff. The last thing I wanted to do was give up the ground I had taken and go back to the version I had just spent two months writing a replacement for.
And still. I wrote git log origin/main..HEAD and started counting commits.
Because the question “should I roll back?” is not a question about whose code is better. It is a question about which codebase I want to spend tomorrow debugging from. That is the whole frame. Once I started asking it that way, the decision got weirdly easy.
The usual frame is wrong
The way most people talk about rollback is as a verdict. “We had to roll back” is said the way you would say “we had to call the fire department.” Something failed. Someone embarrassed themselves. The git history gets a scar. The retrospective has a section titled What Went Wrong. The engineer who shipped the thing feels like they are being asked to walk themselves out of the building.
Rollback is not a verdict on the code. It is a decision about which codebase you want to debug from tomorrow.
— A thing I kept telling myself at 2am
That framing is why so many smart people fix forward past the point where it is helping anyone. You push one more small patch. Then another. Then a third. Each one makes sense in isolation. The feature works a little better than it did two hours ago. Meanwhile the people who are supposed to be using the product keep bouncing, and your “fixing” is actually a slow leak you are papering over with increasingly specific duct tape.
The rollback-as-verdict frame makes you optimize for not being the person who rolled back. It does not make you optimize for the shoppers who were trying to pay you an hour ago.
The real question
The real question is simpler than the verdict version. It is this.
Of the two codebases I have available to me — the one I just shipped, and the one it replaced — which one is a better place to stand while I figure out what is actually happening?
Not which one is better. Which one is a better base camp.
The old codebase has one enormous property that the new one does not: production has been staring at it for months, and the ways it fails are the ways it fails. The paths that work are known. The weird customer who always triggers the same bug on the same browser is already in the backlog. You know its edges. The map is drawn.
The new codebase has the opposite property. Every bug could be the bug. Every weird customer is a potential new failure mode. Every “hm, that looks off” has a long tail of “or maybe it is this other thing I just changed.” The map is being drawn in real time, by whoever is paying you.
Rolling back is not “the new code was bad.” Rolling back is “the new code is not the right place to run production from today.” The new branch does not get deleted. It gets moved back to staging, where it belongs when you cannot yet see every way it fails.
The checklist I actually use
I run through four questions. In this order.
# 1. Is the current production codebase hurting people, right now?
# (not "could it" — "is it")
#
# 2. Do I know, precisely, which commit introduced the hurt?
# If no → rollback buys me time to find out.
# If yes → fix forward is on the table.
#
# 3. If I fix forward, how long until a shopper on prod sees the fix?
# (deploy time + cache time + propagation + any human in the loop)
# If > 1 hour and #1 is yes → rollback.
#
# 4. Is the thing I am fixing load-bearing in a way the old code was not?
# i.e. will rolling back cause a *different* set of problems?
# If yes → fix forward is probably correct, and the pain of fixing is the price.

The thing most people get wrong is step 3. They know the current prod is hurting people. They know what the fix is. They assume the fix will land in ten minutes. In practice the fix takes forty minutes because the test suite needs a nudge, and then another twenty because the deploy failed on the first try, and then another thirty of cache-eviction waiting. Meanwhile two hundred and seventeen shoppers have tried to check out and left.
If step 1 is yes and step 3 is more than an hour, the right move is almost always rollback. You get to breathe. The paying shoppers get to pay. You fix the new code in a staging environment where the only pain is yours.
Fix forward versus rollback, in a table
| Feature | Fix forward | Rollback |
|---|---|---|
| Right when the bug is isolated and the fix is small | ✓ | − |
| Right when prod is hurting and the fix is more than an hour out | − | ✓ |
| Right when the new code has load-bearing features the old code lacks | ✓ | − |
| Right when nobody can reproduce the bug locally | − | ✓ |
| Right when you keep seeing new unrelated bugs on the new branch | − | ✓ |
| Right when the old branch has a known, worse bug than the new one | ✓ | − |
| Feels better emotionally | ✓ | − |
| Is better for the shopper trying to pay you | sometimes | often |
The last row is the only one that matters.
The emotional-cost row is real and worth saying out loud. Fix forward feels like winning. Rollback feels like losing. It is one reason smart teams let shoppers suffer longer than they would ever let themselves suffer. The right instinct is to notice that gap and then override it.
The move I actually made
For the specific thing I was dealing with this week — a payment surface that was no longer completing bookings — I did not end up pulling the rollback trigger. The fourth question caught me. The old code did not have the pricing logic the new code did, and the pricing logic was the actual reason I had rewritten the thing. Rolling back would have broken a different, larger set of things.
So I did the next-best move. I shortened the loop. I shrunk the diff I was deploying from “a feature branch” to “the smallest possible change that could restore the one broken surface.” I added instrumentation that made the next failure legible from the first replay. I set a hard deadline, by the hour, for how long I would stay in fix-forward mode before I pulled the rollback trigger anyway.
That last part is the trick. You do not have to decide right now. You can decide to decide in three hours. And when the three hours hit, you check the checklist again, with fresh eyes, and the answer is almost always obvious.
The takeaway
Rollback is not a verdict. It is base-camp selection. The question is never “whose code was better.” The question is “where do I want to be standing tomorrow morning, with a cup of coffee, when the next weird customer bug comes in.”
Sometimes that is the new branch. Sometimes it is the old one. The engineer who can tell the difference ships more product than the engineer who takes it personally.