Overfit Check

Overfitting is not a strategy looking strange. It is a strategy fitted to noise, and the only real test is whether it keeps working on days nobody could tune it against. This site records when every symphony's logic was last edited, so that test can actually be run.

A symphony URL like https://app.composer.trade/symphony/4aI4kVT5cEc0XJpTLei3/details, a bare ID, or the symphony's raw JSON. The record check runs entirely in your browser, against data downloaded with this page. Reading the logic needs the symphony itself, so pasting a URL or ID asks Composer for it, exactly as the converter does. Paste the JSON instead and nothing leaves your browser at all. Nothing is stored or logged either way.

- of symphonies still deliver their backtested annual return once the author stops editing

Two ways of asking, and they disagree

The same question measured twice, because each measurement has a flaw the other does not. Read them together rather than picking one.

A bigger backtest predicts a better future, up to a point

Strategies sorted by the return of their fitted era alone, restricted to those fitted over five years or more so a short noisy window cannot drive it. Up to roughly +240% a year the relationship is what you would hope: a better backtest really does go on to deliver more. Past that it inverts.

Turnover is what predicts a backtest not surviving

Every stored field was tested against what actually happened next, holding backtested return roughly constant so the comparison is not just regression to the mean. Turnover won, and nothing else came close. Nobody reaches for turnover when asked about overfitting.

What this cannot see

  • This is out of sample with respect to editing, not a clean holdout. It catches the author who kept adjusting until the curve looked right, which is how a Composer symphony usually becomes overfit. It does not catch one fitted in a single pass and never touched again.
  • Turnover is an association, not a cause. High turnover could mean curve fitting, or it could just mean trading costs and slippage eating the return. Both stories predict the table above. The backtests here are already net of modelled slippage and fees, so this is not simply a backtest ignoring costs, but real-world slippage can still run ahead of the model, and this data cannot separate the two.
  • Return concentration is not used here, deliberately. The share of return coming from a strategy's best few days is a true and useful thing to know, and it is on every curated strategy page. It does not predict out-of-sample failure: its correlation is near zero and its sign is unstable. Over 80% of all symphonies get more than their entire return from their best 5% of days, so as a flag it fires on four strategies in five and tells you nothing.
  • The score has no weights, and that is what makes it checkable. It is one measured ratio, 100 x (1 - delivered / fitted), not a blend of metrics. A composite would invite the exact optimisation this page exists to detect and its weights could never be falsified; a single ratio has nothing to tune, and both of its inputs are printed beside it so you can recover it yourself.
  • Never read the score without par. How much a backtest gives up is itself strongly tied to how big it was (rank correlation -0.56), because anything extreme reverts whether or not it was overfit. A score of 95 is damning against a par of 60 and unremarkable against a par of 97. The score itself contains no population data at all; par is the only place the population enters, and it is kept deliberately outside the number.
  • The score largely agrees with simply beating the market. Ranked against annualized excess return over SPY the two correlate at +0.93, with about 80% of each tail shared. It is not a fundamentally new signal. What it adds is that it answers the question actually asked, whether a strategy kept its own promise, rather than a proxy for it, and it separates them exactly where the promise was largest.
  • This page rates overfitting using measures chosen by the same people who chose which measures to keep. That is a version of the problem it reports. Saying so is cheap, so it is said.