Every plan rests on things you believe but can't yet prove. List them, say how sure you are, and ConditionFirst works out where the risk actually sits and what it's worth paying to find out.
Aim for five to nine. Fewer and you're hiding something; more and you've usually split one belief across several lines, which double-counts the risk.
Add a few things that have to be true, and the result will appear here.
Three things: find out what you don't know, decide in advance when you'd quit, and check individual decisions as they come up.
Write these now, while you're calm. The moment you most need a stopping rule is the moment you'll most want to argue with it.
What the numbers are, what they aren't, and how to tell whether yours mean anything yet.
Not the probability your plan succeeds. Nobody can compute that. What you get is the probability implied by your own stated beliefs, propagated correctly. If the beliefs are wrong the answer is wrong — the calibration log below is how you find out whether they have been.
Not a precise figure. Every belief is held as a range, not a point, and the range is widened automatically depending on how you said you know it. Say something is 95% likely on gut feel and it will be carried as 59–99%, because that is what a gut feeling is worth.
Not a ranking it can't support. If two approaches swap places depending on where inside your own uncertainty the numbers land, the result page says they can't be told apart instead of declaring a winner.
Not a substitute for going and finding out. The most useful output here is the list of things worth checking, not the percentage at the top.
| Ranges | Your confidence becomes a range whose width depends on your evidence: hard data ±5 points, a comparable case ±12, gut feel ±20. Each of 4,000 runs samples a value from inside that range. |
| Linked beliefs | Beliefs you mark as linked move together, using a one-factor Gaussian copula. This preserves the exact confidence you entered while making linked items rise and fall as a group — so linked beliefs aren't double-counted as separate risks. |
| Kills vs hurts | A belief marked kills the plan zeroes that run. One marked hurts multiplies it by however much of the goal survives. This matters enormously: treating everything as fatal is what makes naive models spit out 4%. |
| The ceiling | Fréchet–Hoeffding bound: the lowest confidence among your deal-breakers. No pattern of luck or linkage can beat it. It's the one figure here that depends on no modelling choice at all. |
| Worth finding out | (1 − confidence) × what's at stake − cost to check. The money you'd expect to save by discovering a belief is wrong before you spend on it. Negative means checking costs more than the mistake. |
| Repeatability | The same inputs always give the same numbers — the randomness is seeded. A tool whose answer moves when you refresh isn't one. |
This is what turns your confidence figures from opinions into measurements. Each time one of your beliefs resolves, log it here with the confidence you gave it at the time. After ten, you get a real score.
The method has one unavoidable weakness: you are rating your own beliefs, and nobody is well calibrated about a plan they are invested in. The fix is not a better algorithm, it is a second person.
Use Send this plan for a blind second opinion in the menu. It produces a file with your confidence figures and every money amount stripped out, so your rater starts from the beliefs alone rather than anchoring on your numbers. They rate, save, and send it back; Merge a returned second opinion brings their figures in beside yours and the Result tab shows where you disagree.
Read the gaps, not the average. Two people landing twenty points apart on the same sentence almost always means the belief is written loosely enough that they are rating different things — which is worth finding out before you spend money on either version of it.
Tagging a belief with its kind shows you how often that class of thing actually works out. This is deliberately the first question on each card, because starting from the base rate and adjusting for what is specific to you is the single most reliable correction for overconfidence — and doing it the other way round, starting from your story and rationalising the rate afterwards, is how everything ends up at ninety percent.
The figures are conservative reference points drawn from published failure-rate literature, not precision measurements. They exist to make you justify a departure, not to overrule you. If you are twenty points above the base rate, that may be entirely correct — you should just be able to say why out loud.
This works best as a structured conversation, not a spreadsheet you live in — a couple of hours with the people who actually hold the beliefs, producing one page you can hand to whoever needs to approve the spend. The list of what's worth checking is the deliverable. The percentage is just how you got there.
Above roughly twelve beliefs the arithmetic starts collapsing on its own, regardless of how good the plan is. If that happens the fix is to merge or link lines, not to distrust the tool.