you can't improve judgment you don't log
intuition is just a model you refuse to eval.
builders instrument everything. every request logged, every agent traced, every funnel step measured, every button a/b tested. we would never ship a model without evals. we would never deploy a system whose failures we couldn't replay.
then we make the ten decisions that actually determine the company, who to hire, what to build, when to pivot, how to price, from memory, with zero instrumentation, and we call it intuition.
judgment is the founder's core asset. it is also the only system a founder runs with no logs, no evals, and no replay. a decision journal is the cheapest instrumentation that exists, and almost nobody keeps one.
your memory is a corrupted dataset
the reason you need logs is that the built-in alternative, remembering, actively lies to you. not occasionally. systematically.
hindsight bias rewrites your forecasts after the fact: once you know how it went, you genuinely remember having seen it coming, so every outcome quietly becomes evidence that your judgment works. outcome bias grades decisions by results instead of process: the reckless bet that got lucky files itself as brilliance, the sound bet that got unlucky files itself as a lesson to be more timid. and the sample is tiny and emotionally weighted, so the loudest memories, not the most representative ones, become your training data.
put it in builder terms: you are fine-tuning yourself, continuously, on fabricated labels. no wonder twenty years of experience so often turns out to be one year of experience with nineteen years of confirmation.
the fix is a written record your later self can't tamper with.
what to log
the format matters less than the honesty. mine is a few minutes per meaningful decision:
- the decision and the real options. including the ones you're embarrassed to be considering.
- what i expect to happen, with a number and a date. "this hire works out" is ungradeable. "in six months this hire owns the pipeline and i've stopped reviewing their output, 70%" is a forecast that can be scored.
- what would change my mind. written now, before defending the decision becomes part of my identity.
- my state. tired, scared, euphoric, rushed. it feels irrelevant when you write it. it becomes the most predictive field in the whole log.
then the part that makes it a loop instead of a diary: every quarter, grade the closed forecasts against reality.
what the review actually shows you
everyone who does this discovers the same two things, and they're never the two they expected.
first, your overconfidence has a direction. nobody is just "overconfident"; each person is miscalibrated in a specific, personal way. some always overestimate how fast they'll ship. some always overestimate other people's follow-through. some are perfectly calibrated on product and wildly off on people. this is your error signature, and no book, no framework, no advisor can give it to you, because it isn't general knowledge. it's a fact about you, and it only exists in your own scored record.
second, your worst decisions cluster. by state, tired and rushed, or by type, hiring, pricing, anything involving conflict. once you can see the cluster, the fix is often embarrassingly mechanical: don't decide that category of thing after a red-eye. get a second opinion on exactly one class of call. the log converts vague self-improvement into a patch list.
there's real evidence behind this loop. tetlock's forecasting work found that the people who beat everyone else weren't smarter or better informed; they made granular probability estimates, got scored, and updated. calibration turned out to be trainable. that's the strongest empirical case we have that judgment is a skill with a feedback loop, not a gift with a ceiling.
the objections
"founder decisions are too rare and too long-horizon to score." for endgames, yes. so score checkpoints. a pivot takes two years to fully grade, but "in 90 days we'll see signal x" grades fast, and big decisions decompose into many such forecasts. besides, even fifteen major decisions a year is seventy-five over five years: a real dataset about the only decision-maker you can't replace, and your competitors have zero rows on themselves.
"i move fast. journaling is overhead." this one has it backwards. half the value arrives before any review, in the writing itself: forcing options, forecasts, and disconfirmers into sentences is the fastest de-biaser known. it takes ten minutes, and it routinely catches the decision that would have cost a quarter. that's not overhead. that's the highest-margin ten minutes in the company.
and the honest hedge: a log can become theater. if you journal for an audience, even an audience of your future self's ego, you'll write defensible entries instead of true ones. the log only works exactly as private and exactly as blunt as you can bear to make it.
the eval you've been refusing
intuition is real. it's pattern recognition trained on your reps, and in domains where you genuinely have reps, it deserves weight. but untested intuition and tested intuition feel identical from the inside. the only way to know which one you're holding is to score it.
you already believe this. it's why you log your systems, trace your agents, eval your models. you just exempted the most important model in the building.