Reading this site
How sure we are
Three things. The four words that sit beside every claim here, how long you would have to hold something before your own account could tell you it worked, and what it means when this site says nothing at all.
The vocabulary
The four words
Every claim here is labelled settled, probably, might, or we can't tell. Four rungs, not five and not seven, because rungs a reader cannot tell apart are not rungs. Each one carries the same underlying question: what is the chance this choice leaves you better off than the named alternative, over thirty years, after costs and taxes.
Nothing here is ever green for good news, and that is the more important half of the reason. A coloured badge answers two questions with one mark: how sure we are, and whether what we found is welcome. Those are separate axes. A settled claim can be settled bad news, and something we cannot tell you about might be the best idea on the site. Collapse the two and the reader loses the one thing this scale exists to give them.
Settled
Confidence: Settled
No hedge, flat declarative. It follows from a contract, a tax rule, or arithmetic, and the only uncertainty left is whether you actually do it. Moving your money into cheaper share classes and the right accounts saves about 1.09% a year. No range, because there is not one worth printing.
If a claim has a probability attached, it is not settled. A 99% chance and a contractual certainty are different objects, and merging them is how a confidence scale stops carrying information.
Probably
Confidence: Probably
Roughly a 70% to 90% chance, and the sentence always carries the size and the range. The effect is larger than the smallest thing the test could have seen, the range stays on one side of zero, and it survives being marked down for how many ideas we tried before this one. One material gap remains, and the sentence beside it names that gap.
The lean toward cheap and profitable companies lives here: +0.79 percentage points a year against a cheap 65/35 index fund, at 1.0% of drift. The gap: 86% of it is measured on data starting in November 1990, and nobody has ever run it on years the idea had not already seen.
Might
Confidence: Might
Roughly 50% to 75%. The evidence points one way and keeps pointing that way, and at least one thing is wrong with the design. The window is short, or the effect is smaller than the smallest difference the test could see, or nobody has run it on years the idea had not already seen, or it rests on an assumption a reasonable person would argue with. A might always carries that smallest difference and a plan for settling it. Otherwise it is a guess wearing a lab coat.
We can't tell
Confidence: Can't tell
The range covers real help and real harm. We print the range and refuse to print a probability, because inventing one is the failure this whole scheme exists to prevent. This is the most common verdict on the site, and it must never be read as no.
One rule cuts across all four. Every probability here is a best case, because the arithmetic behind it assumes we know exactly how large the advantage is. We don't. The real chance is lower than the printed one, and we cannot tell you by how much.
Your own account
How long before you would know
Two things decide whether your own account could ever settle a question. How much better you expect to do, and how far the thing wanders from what you are comparing it against on the way. The wandering does most of the damage, because it counts twice over. A portfolio that bounces around twice as much takes four times as long to prove itself.
Take one fixed advantage and run it past two levels of wandering, and the wait comes out at 24 days at 0.10% of drift, 105 years at 4.00%. Same edge, same arithmetic, and one of those is a lunchtime while the other is longer than a life.
That is the contrast the site is built around. The fee, tax and account-placement savings are worth about 1.09% a year against the portfolio you would otherwise hold, and they hardly wander at all, so you would know inside about 3.5 months. The whole recommended portfolio, against a cheap index fund holding the same mix, is ahead by 92 basis points and drifts by 400 basis points a year, so you would know in 31 years.
Three and a half months and thirty-one years, one rule applied twice. None of which argues for doing nothing. It argues for doing the certain things first, and for justifying everything else before you buy it, because your own results are never going to be the evidence.
Silence
The three kinds of nothing
When this site reports nothing, it means one of three things, and only the third is a result.
We never looked
Plenty of ideas here carry no verdict because nobody has tested them. That is a hole in the work rather than a judgement on the idea, and a page that has one says so.
We looked and could not see
Returns bounce around. Any difference you measure between two portfolios sits inside that bouncing, and before running a test you can work out how large a difference would have to be before the data could pull it back out. Under that size, a real effect and no effect produce answers that look the same.
The moving-average timing rule is the cleanest example here. It beat a properly matched control by +0.74 percentage points a year against a control taking the same average market risk. The smallest gap that design could have detected was 3.03. The measurement is real and the arithmetic is right, and the answer is still that we cannot see. Publishing the 0.74 as a finding would report the width of our own ignorance.
This is why so many pages here end in a shrug, and it is why we can say that a century of monthly US data cannot settle a question, which sounds like an excuse and is a measurement.
We looked and found something real pointing the other way
Only this one is a result, and it is worth more than the other two kinds put together. Bitcoin bought as ballast moves 1.53x up, 1.62x down, so it falls harder than it rises, which is the opposite of the shape a diversifier is bought for. That is not a failure to find an effect. It is an effect, found, pointing the wrong way.
No evidence of an effect and evidence of no effect are different sentences, and we are careful about which one we write. Where a result leans, we say which way. Where an effect of a given size would have been visible, we say what that size was.
Corrections
Where we have been wrong
Thirty-three errors have been found in this project's own published numbers. Eleven of them moved a recommendation rather than a decimal, and each is written up at corrections with what we believed, what was wrong, what caught it, and what changed. Start with the trend-following return that was quoted here for most of the project's life, that nobody had ever measured, and that was deleted rather than replaced with a better guess.