Second thoughts
Loud decisions are rarely the heavy ones
Stakes don't make a decision hard. Irreversibility and silence do. Amazon's one-way door framework names that failure mode, then buries it in a footnote.
On 26 September 2025, a technician moved a lithium-ion battery in a government data center in Daejeon. Forty minutes later it exploded, and 858 terabytes of South Korean government work files burned with no external backup.
The fire was an accident. The decision was not.
Years earlier, someone decided that civil servants would store work documents in the cloud, and made it mandatory. That sentence is the one-way door. For the roughly 125,000 officials who used the system, it removed the local copy that would have made a fire survivable. On the day it was made, it read like a storage policy.
We rank decisions by how loud they are. Loudness is not weight.
The loudest decision in the room is usually reversible
Deliberation goes to the framework debate, the linter config, the naming argument. All contested, all undoable in an afternoon. Meanwhile the schema gets designed by whoever opened the editor first.
A Hacker News commenter, May 2026, stated the triage rule: "Make sure the decisions you're spending the most time on are the most important."
The irreversible calls get skipped because they are rare. In February 2026 Anthropic classified 998,481 sampled tool calls from its public API and found "only 0.8% of actions appear to be irreversible (such as sending an email to a customer)." Different population, same shape. The dangerous slice is too thin to justify a standing process, so it inherits whatever process the other 99% gets.
In July 2026, a 651-point r/devops post described four months spent tearing apart a working monolith into 14 services after the CTO returned from a conference. The sharpest reply, at 64 points: "i think the biggest red flag isn't even microservices, it's that the operational cost doesn't seem to have been part of the decision."
Not scored low. Absent.
Reversal is priced in years
Airbnb's infrastructure team published its escape from ingest-priced observability vendors in 2026: "Our plans for this migration, begun five years ago, evolved significantly as we proceeded." The haul was 300 million timeseries, 3,100 dashboards, and more than 300,000 alerts across 1,000 services. Five years to walk back one procurement decision.
GitLab issue #581028, opened 17 November 2025, begins: "The migration of organizations from the Legacy Cell to the Protocell is a one-way door as currently planned as of 17 Nov 2025." Four rollback options were costed. None was automated.
- Airbnb's exit from one vendor
- 5 years
- 300 million timeseries, 3,100 dashboards and more than 300,000 alerts across 1,000 services. The door was never locked — it was just this expensive.
- Agent actions judged irreversible
- 0.8%
- Of 998,481 sampled tool calls, February 2026. Machines rather than people: no comparable count exists for human decisions.
- Daejeon G-Drive, 26 September 2025
- 858 TB
- Destroyed with no external backup. Roughly 125,000 officials were required to keep work documents there instead of locally.
Doors also close at different speeds. A March 2026 paper measured the drift across 357 agent traces: "in communication tasks, agents that reach even a mild risk state have an 85% chance of violating safety within five steps, while in technical tasks the probability stays below 5% from any state." Same mild risk, same five steps, a seventeen-fold gap in how fast it gets away from you.
Jeff Foster, July 2026: "You aren't holding an option with a fixed exercise price. You are holding an option whose strike only ever goes up."
Nobody pages you when the strike price moves.
The half nobody instruments
Irreversibility is one axis. The other is whether any signal ever comes back.
A commenter in March 2026 described Voyager 1's thruster revival:
"They sent a command that would either revive thrusters dead since 2004 or cause a catastrophic explosion, then waited 46 hours for the round trip with zero ability to intervene. That's a production deployment with no rollback, no monitoring dashboard, and a 23-hour latency on your logs."
In September 2025 the Forecasting Research Institute graded its panel of 89 superforecasters and 80 domain experts against 38 subquestions that had since resolved. Individual forecasters' performance "was not statistically distinguishable from two simple algorithms: a 'no-change' forecast and trend extrapolation."
A June 2026 arXiv paper shows that when choices are made simultaneously, a scoring test can identify the better-informed agent. For choices made sequentially, "no such scoring test exists." Not "we don't measure it." Cannot.
A June 2026 thread resurfaced Repenning and Sterman on why nobody gets credit for problems that never happened. The best comment, from two years of Y2K remediation: "In the end, 'nothing happened,' so all that time and money was wasted, according to nearly every company I worked with. Even had one demand a full refund. I agreed as long as I could revert all the work that I had done. They agreed, and the next day after that their entire system collapsed."
Sometimes the loop exists and still doesn't close. A November 2025 r/sre post: connection pool exhaustion, fixed, documented, then the identical incident five months later, because the ticket sat in the backlog. "so we fixed the same connection pool issue twice. documented it twice. got paged twice at 2am. very efficient." The top comment, at 151 points: "These two postmortems do not stem from the same failure. The first failure was Redis connection pool. The second failure was bug prioritization."
A postmortem nobody acts on is documentation, not feedback.
The best argument against all of this
In PNAS, accepted April 2026, Sunde, Zegners and Strittmatter analyzed move-by-move data from in-person professional chess against an engine benchmark and found "a robust negative association between the time spent on a decision and its quality," holding computational complexity, distinctiveness of alternatives, and time pressure constant. That is an association, not a lever you can pull.
Jason Wadsworth, June 2026: "Here's the uncomfortable truth: most of your one-way doors swing both ways. You just haven't pushed on them." His mechanism: we retroactively declare doors one-way to protect ourselves from admitting error, and "Cost is not the same as a welded-shut door."
The strike price can also fall. In May 2026 Bun rewrote 535,496 lines of Zig into Rust in eleven days and merged it to canary with "0 tests skipped or deleted." Roughly $165,000 of API spend, about 64 model instances running in parallel. Choosing an implementation language was the textbook one-way door. It just got a price tag, and the price was eleven days.
Wadsworth is right about the doors and wrong about what holds them shut. Welded steel stops nobody. A bill nobody will approve does. Airbnb could have left its vendors in year one. The door was open the entire time, onto five years of work. A door you can technically walk back through and never will is a one-way door with better public relations.
Bezos introduced one-way and two-way doors in Amazon's 2015 shareholder letter, and his worry ran opposite to this article's: "As organizations get larger, there seems to be a tendency to use the heavy-weight Type 1 decision-making process on most decisions, including many Type 2 decisions." Fair, and real.
But he attached a footnote to that sentence:
The opposite situation is less interesting and there is undoubtedly some survivorship bias. Any companies that habitually use the light-weight Type 2 decision-making process to make Type 1 decisions go extinct before they get large.
The most-quoted decision framework in technology names the exact failure mode described here, then sets it aside in a footnote, on the grounds that its victims die before anyone can study them. "Less interesting" is doing heroic work in that sentence. Every observation that large companies over-deliberate is drawn from a sample of companies that survived their under-deliberation.
That is also why the chess result doesn't rescue you. It attacks deliberation time; it says nothing about deliberation targeting. Stripe maintains bidirectional replication during migrations specifically to keep rollback available. Figma's PGKeeper flips traffic back automatically when error rates cross a threshold: "In Q4 2025 alone, PGKeeper prevented more than 20 incidents that would otherwise have caused user-visible outages." Neither team deliberated the door open. They paid to keep it open.
Two questions, before the meeting starts
The triage takes ninety seconds, and it is not "how big is this."
What does reversal cost in six months, not today? The strike price moves. If the answer is a sprint, stop deliberating and go. If it is Airbnb's five years, that is the decision deserving everything you have.
What signal would tell me I was wrong, and when does it arrive? Write the falsifiable version down before committing: the number, the date, the threshold. Not because your forecast is any good; the Forecasting Research Institute suggests it probably isn't. A written threshold is the only thing that turns a silent outcome into something you can be caught being wrong about.
Cross the two and the budget allocates itself.
- Reversible, fast feedback. Ship it. The world grades you cheaply.
- Reversible, no feedback. Cap the spend and set a review date, or it runs forever.
- Irreversible, fast feedback. Spend money to convert it: bidirectional replication, an auto-revert, a retention window.
- Irreversible, no feedback. The only quadrant that has earned a real process. The one getting a hallway conversation.
If you run a weighted matrix, this is where the weights belong. Score reversal cost as its own criterion, then check the flip point: how far does reversal cost have to move before the winner changes?
Rob Berry, May 2026: "The last reversible decision is the most important decision."
Not the biggest one. The last one before the door shuts — which, in Daejeon, was a storage policy.
