Corrections
Corrections
Thirty-three errors have been found in this project's own published numbers. Eleven are below, chosen because each one moved a recommendation rather than a decimal.
If you read one, read the trend-following return. It was quoted confidently here for most of the project's life, nobody had ever measured it, and when we worked out what it actually was we deleted it rather than printing a better guess.
Policy
What happens when we find one
Five rules, written down before they were needed.
The number is corrected everywhere it appears, and an entry lands here with its date.
A frozen specification is never edited, even when the specification was the thing that was wrong. Its defect is recorded on the page that reads it, because a specification you can revise after seeing its result is worth nothing.
A withdrawn number is deleted rather than replaced with a better guess. A blank looks like an oversight and a plausible number looks like work, so this is harder than it sounds.
Where a correction has a mechanical fix, the fix is code that refuses. Three entries below ended with a function raising an error on the exact mistake that produced it.
Typos and clumsy sentences are fixed in silence. An entry appears here when an error moved a number, a verdict or a recommendation. The other twenty-two live beside the figures they changed.
The log
The log
Newest first.
Two different portfolios, both called the recommendation
Corrected 2026-08-24
What we believed. That the seven-fund portfolio published here was the seven-fund portfolio that had been tested, and that a 30% trend weight had been measured as better than a 25% one.
What was wrong. Two things, and the second is the one that matters. The frozen specification calls a 25% portfolio the recommendation, because that is what we recommended on the day it was frozen; we then moved five points out of a plain US index fund and into the stacked trend fund and relabelled nothing, so a reader could find two different seven-fund portfolios each presented as the portfolio. The weighted fee went from 33.4 basis points to 38 basis points on the way.
And the result that moved the weight, 0.50 points, plausibly 0.23 to 0.77, was never a comparison between those two. It scored the seven funds at 25% against an earlier eight-fund draft at 30%, which differs in four of its holdings as well. The figure record said "same seven funds". It was not the same seven funds, and nobody has ever run these seven at 30%.
What caught it. Reading the frozen file's own list of what it scored, rather than the sentence the result was being quoted in.
What changed. The 30% portfolio is named as the recommendation everywhere, the 25% one as what was tested, and every figure drawn from that run says which of the two it belongs to. The result is still published, still the only whole-portfolio comparison here that clears its own resolution, and now carries what it compared.
What has not changed is the number, because this one is still open. Strip the trend and the borrowing out of both and the two sets of holdings are worth 0.80 and 0.79 percentage points a year against the same control over the same months, which is close enough that the trend weight is the plausible remaining explanation. That is an argument rather than a paired measurement. The matched pair is written down as an open question and has not been run.
"How long before you would know" was two questions under one label
Corrected 2026-08-23
What we believed. That three and a half months for the fee and tax savings could sit beside 59 years for the whole construction and be read as one comparison. It was the site's signature line.
What was wrong. They answer different questions. One is how long before your own account could separate an effect from what you are comparing it against. The other is how long before a study could separate that same effect from zero.
What caught it. Writing every horizon on the site into one table, and finding two that could not both be following the same rule.
What changed. The headline contrast is now three and a half months against 31 years, both answering the first question. The construction's own-account horizon is 12.2 years, and the 59 years answering the second keeps its own label. Two records carrying 59 years under an own-account heading were withdrawn.
One name holding two different facts
Corrected 2026-08-23
What we believed. That the name we use for how far the portfolio wanders held one quantity.
What was wrong. It held two. Six percent is the drift from a control that borrows as much as the portfolio does. Four hundred basis points is the drift from a cheap index fund holding the same split. Any sentence reaching for the shared name got whichever had been written most recently.
What caught it. A checker that refuses to let one name carry two facts, written for a different reason and finding this on its first run.
What changed. Split in two, with the comparison written into each record. The checker runs in the build, so the collision cannot come back.
A sell rule finer than the instrument that reads it
Corrected 2026-08-23
What we believed. That a round number made a rule. If the amount of trend the fund actually delivers falls below 0.70, sell it.
What was wrong. The first reading came back at 0.681, on a range that spans the bar twice over in both directions. It neither fires the rule nor clears it.
What caught it. Taking the reading, and finding the answer inside its own uncertainty.
What changed. Nothing about the bar. It stays unfired and recorded as unresolved, because moving a bar after seeing the number it was written for is worse than setting it badly. It should have come from the exposure the fund needs at the weight we hold just to break even, 0.19 to 0.27. We look again around September 2027.
A fund's largest line read in place of its whole filing
Corrected 2026-08-23
What we believed. That MATE's stock exposure was the S&P 500 fund at 49.8% of net assets in its February 2026 holdings report. Read that line and stop, and MATE looks worse than selling stocks to buy trend.
What was wrong. The same filing shows a bet on the S&P 500 through a futures contract at another 61.8%. The stock leg is 111.6% of net assets, and the arithmetic flips from bad to slightly better than break-even.
What caught it. Reading the filing to the end rather than to its biggest line.
What changed. A stacked fund's stock leg is now every instrument delivering stocks, added together, and one of those is always a futures contract. CTAP offered the same trap two quarters later: 70.41% in a fund, 32.23% in futures underneath it, and a stock leg of 102.64%.
A trend-following return we had never measured
Corrected 2026-08-22
What we believed. That trend following would return 1.80 percentage points a year from here. The figure sized a position, and was then used to argue against holding that position at all.
What was wrong. Nobody had measured it. It was the 2012 to 2025 stretch's own compounded average with a 1.50% fee taken off, compared against a break-even computed on a different basis, with the same fee charged a second time by the model that consumed it. Three errors inside one number: a variance term worth 0.77 points, a fee counted twice, and a single figure used with no range when its own range ran from −2.67 to +11.00 and so contained the 10.98 it was being used to overturn.
What caught it. Running the arithmetic backwards. Add the fee back, add the variance term back, and 1.80 returns 4.07, which is that same file's own average on the basis the break-even used. A forecast that reconstructs a past average to the second decimal is not a forecast.
What changed. The sign of the recommendation. On the published basis the position subtracted 0.35 percentage points a year against the same portfolio without it; corrected, it adds 0.34. The recommended weight moved to 25%.
The number itself is gone, and nothing replaced it. Defensible estimates run 2.84 to 7.18 percentage points a year depending on which construction you believe, the break-even is 2.96, and picking one out of that spread would be choosing the answer rather than measuring it.
A borrowed portfolio compared against an index that borrowed nothing
Corrected 2026-08-22
What we believed. That the construction beat the index, so the construction was doing something.
What was wrong. The construction holds $1.32 of assets for every dollar put into it and the index holds $1.00. You beat that index in good years and trail it in bad ones whatever else you do. Take the credit on the way up and you take the blame on the way down.
What caught it. Building a control that borrows exactly as much as the portfolio does, and running everything again against it.
What changed. At matched borrowing the worst loss was −49.3% against −60.8%. The unflattering consequence is printed too. The one whole-portfolio comparison that clears the smallest gap its own test could see, the recommended portfolio against the investor's original at −0.50 percentage points a year, is mostly measuring 5.4 points of borrowing rather than anything about how the portfolio is built.
A visibility bar on one basis, printed beside a range on another
Corrected 2026-08-22
What we believed. That the strongest measurement on the site cleared the smallest difference its test could see by 68%.
What was wrong. The range had been worked out with error bars that allow for markets clustering, calm with calm and panic with panic. The bar printed beside it treated every month as a fresh coin flip. Two bases, one sentence.
What caught it. Reading the run's own output file rather than the page. Both numbers were in it, and only the narrower one had been published.
What changed. On matched error bars it is 0.72 percentage points a year rather than 0.47. The result still clears, by 10% rather than 68%, and the data a study would need to separate the effect from zero goes to 30 years. One arm of the tournament stops clearing.
Adding up savings measured against different comparisons
Corrected 2026-08-22
What we believed. That putting each fund in the account that taxes it least was worth +38 to +55 basis points a year. It was the largest new number in the recommendation.
What was wrong. Four things at once. The range had been truncated, dropping a reading worth 18 basis points. A rebalancing hurdle avoided had been booked as a saving worth 14, on the same page that says a hurdle avoided is not one. Choosing which shares to sell and never selling at all had both been booked, and each excludes the other. And 94.7% of the top figure came from one holding's undistributed foreign accrual.
What caught it. An independent re-derivation, which reproduced every priority, the credit cap and all twelve cells exactly. The calculation was right. What was booked from it was not.
What changed. The line reads +2 to +7 basis points a year basis points a year against a control an investor could hold, with the conditional +19 to +49 beside it and added to nothing. The aggregating function now raises an error on a cross-benchmark sum.
Assuming every dividend was the good kind
Corrected 2026-08-22
What we believed. That fund dividends qualify for the long-term tax rate. It is the standard simplifying assumption and nobody had checked it.
What was wrong. It had reversed the account-placement ranking for emerging-market stocks, which is exactly the decision the section existed to make.
What caught it. Reading the sponsors' own filed fractions, which took an afternoon. VEA files 66.27%. VWO files 34.63%. IEMG 34.82%, AVES 44.48%, IDMO 25%, DFIV 100%.
What changed. Every placement conclusion that predated the filed fractions had to be worked again rather than restated, which is slow and unpleasant and the only honest option.
Two measurements of the same thing, multiplied together
Corrected 2026-08-17
What we believed. That a fund's measured exposure to cheap stocks, 0.410, should be discounted by 0.520, the share of the textbook version an ordinary investor can actually get. That product, 0.213, was used in five places.
What was wrong. The two numbers measure the same exposure, so multiplying them discounts it twice. Work through the algebra and the second turns out to be the first plus a small remainder. The product also sits below the exposure delivered by every value fund we have audited, which should have been the tell.
What caught it. Deriving the identity between the two quantities instead of taking two published numbers as independent.
What changed. That line had been understated by roughly a factor of two. The chain is three terms now: the weight, times the difference in exposure between the fund and what it replaces, times the reward, minus the cost. The research code raises an error rather than accepting a discount at all, and the 0.40 sitting in the edge budget was deleted rather than replaced.
The pattern
What the eleven have in common
Eight were caught by the project arguing with itself: an adversarial review, a re-derivation of arithmetic somebody had already done once, a control built to have a known answer, a frozen file read beside the sentence quoting it. Two came from reading a fund filing all the way through instead of stopping at its largest line. One came from a build check written for something else.
Most of them were arithmetic that was correct when somebody did it. What broke came afterwards, when the number left the page that made it true and travelled without its window, its comparison or its error bars. That is the failure this site's content model is built against, and it is why every number here drags its source behind it.
One of the eleven is still open, and none of that means the current numbers are right.