CTA Trend Following

CTA Trend Following — Robust Regions, Not Best Cells

Twenty-six markets researched as they stood on 2015-12-31, using Kaufman's concentration method: find where the good results cluster, fund three or four separated parameters out of that region, and run the whole thing as one book.

As of 2015-12-31 — the last ten years are withheld. Every number on this tab is computed on bars up to and including 2015-12-31, and nothing after that date is read by any calculation here. The cut happens on the raw price frame before any indicator touches it, so no parameter on this tab was chosen with knowledge of what came next. The withheld data still exists and has not been touched — it is the out-of-sample test for whatever comes out of this work, and it is only worth having once the strategies are fixed.

Why this tab exists

The market cards on /markets each report fifty base cells, and two summary readings have been taken from them. Both are bad inference, and the second one produced a conclusion worth retracting.

  • The best cell. The single luckiest point on a fifty-point sweep. The site has called this out as a selection effect since 2026-07-27.
  • The median cell. The 2026-07-31 write-up ranked markets on median Sortino against buy-and-hold and concluded that physical commodities do not trend — 2 of 16 beat holding. The median asks “what happens if I pick a parameter at random?”, which nobody does. A surface can have a genuinely good, genuinely broad region and still show a poor median, because the median is dragged down by every parameter no researcher would ever have selected.

Kaufman's method instead. Smooth each parameter with its tested neighbours, find the contiguous plateau that holds at or above 70% of the smoothed peak, and fund three or four maximally separated periods from inside it. A plateau survives out of sample far more often than a spike, because a spike usually means one parameter happened to catch one move. Spreading the entries is the other half: four adjacent periods are one system tested four times, whereas SMA 33 / 51 / 90 are three different systems that happen to agree the market trends.

Where the method finds no broad region, the market is declined rather than traded at whatever its least-bad parameter was. That decision is part of the strategy, so the book below is shown both ways.

Markets funded
14 of 25
11 declined, 1 with no in-sample data at all
Funded markets beating hold
11 of 14
19 of 25 across everything
Selection metric
Sortino
Sharpe punishes the upside trend following exists to capture
Design date
2015-12-31
Nothing later was read by any calculation

The Book — 14 Markets, Equal Weight, Rebalanced Daily

Only the markets whose surface showed a broad plateau. This is the book the method would actually have run: a CTA that researches a market and finds no robust parameter region does not allocate to it.

Sortino
1.46
hold 0.43 (+1.04)
Calmar
0.52
hold 0.05 (+0.47)
Sharpe
0.98
hold 0.30 (+0.69)
CAGR
6.6%
hold 2.7% (+3.8%)
Max drawdown
-12.6%
hold -52.4% (+39.8%)

Mean pairwise correlation between the per-market ensembles is 0.067. That number is what makes a multi-market book worth running: a low figure means the 14 markets are genuinely separate bets, so the book's drawdown is far smaller than any single market's. Both legs are measured against a buy-and-hold of their own markets, never a different basket. Everything stops on 2015-12-31.

Every market — the plateau, the funded entries and what they earned

Sorted with the funded markets first, then by Δ sortino against buy-and-hold. “Plateau” is the contiguous stretch of SMA periods that all worked; the width beside it is how many of the 25 tested periods that covers. Histories differ, so drawdowns are not comparable across rows.

MarketSysPlateauWidthFunded entriesER₂₀CorrCAGRMax DDSortinoB&HΔvs medianvs best
Brazilian Real*B66–12013/2557, 70, 85, 1200.2280.708.1%-17.2%0.967-0.216+1.183+0.442-0.111
US Dollar IndexB30–456/2530, 42, 70, 900.2650.574.3%-12.6%0.968-0.077+1.044+0.450+0.029
High Grade CopperB57–859/2563, 69, 75, 850.2290.8116.1%-35.9%1.0130.291+0.722+0.353-0.050
PalladiumA30–456/2533, 39, 450.2510.8313.4%-46.6%0.8030.172+0.631+0.385-0.247
Crude OilA39–515/2542, 48, 110, 1200.2290.767.6%-32.5%0.4880.036+0.453+0.301-0.118
SoybeansA39–515/2542, 45, 48, 1000.2350.837.7%-34.7%0.6050.186+0.418+0.373+0.040
British Pound*A100–1205/25105, 115, 1200.2300.931.8%-14.3%0.4270.011+0.416+0.330-0.200
CottonA70–12011/2575, 85, 95, 1100.2430.896.1%-45.3%0.4360.063+0.373+0.192-0.097
Orange Juice*A85–1106/2585, 95, 1050.2430.927.3%-47.5%0.4690.180+0.289+0.354-0.325
Sugar No. 11A39–6610/2545, 54, 60, 950.2390.829.0%-50.5%0.5590.295+0.264+0.278-0.263
Live CattleA30–5710/2530, 36, 42, 540.2420.843.9%-17.9%0.4710.317+0.154+0.402-0.052
GoldA66–12013/2580, 90, 105, 1150.2400.896.6%-31.9%0.6580.706-0.048+0.403-0.074
30-Year T-BondA66–805/2570, 75, 80, 900.2260.921.2%-16.9%0.2110.390-0.179+0.243-0.100
2-Year T-NoteB66–805/2569, 70, 75, 800.2260.860.3%-6.1%0.2600.510-0.250+0.071-0.317
Bitcoinshort historyB30–394/2533, 36, 390.2490.8781.2%-26.1%2.607-0.092+2.700+2.232-0.267
Lithium & Battery ETF*narrowB42–421/2542, 60, 63, 660.2650.727.0%-21.6%0.598-0.420+1.018+0.850-0.012
Mexican Peso*narrowB57–571/2542, 45, 570.2280.731.9%-18.3%0.321-0.447+0.768+0.487-0.241
Japanese Yen*narrowA54–572/2551, 54, 570.2330.931.0%-14.5%0.201-0.048+0.249+0.193-0.030
CoffeenarrowA115–1202/2585, 115, 1200.2250.873.7%-35.5%0.2400.022+0.217+0.345-0.047
Lean HogsnarrowA45–451/2545, 48, 600.2980.803.4%-43.9%0.2400.053+0.187+0.491-0.216
Natural GasnarrowA45–451/2542, 45, 480.2250.931.5%-66.1%0.057-0.126+0.184+0.275-0.026
SilvernarrowA95–1104/2595, 100, 105, 1100.2510.937.9%-53.3%0.4560.292+0.165+0.401-0.035
CocoanarrowA90–1003/2590, 95, 1000.2350.924.2%-48.8%0.2610.410-0.148+0.461-0.033
10-Year T-NotenarrowA70–752/2566, 70, 750.2290.920.5%-10.5%0.1580.321-0.163+0.317-0.099
WheatdeclinedAnone90, 100, 110, 1200.2250.88-0.6%-67.4%-0.0410.198-0.238+0.243-0.064

The three ways a market fails to get funded. Declined means the smoothed surface never rose above zero — there was no region to select from in either direction. Narrow means there was a good region but it covered fewer than 20% of the tested periods, so the entries are necessarily close together and will move as one. Short history means the window holds fewer than 750 bars, which is not enough for a 120-day SMA to have been tested more than a handful of times regardless of how the surface looks. All three still get a page; none of them enter the funded book.

Not on this tab at all: Aluminium (no usable bars at or before 2015-12-31). These have no usable bars before 2015-12-31, so there is nothing to study — they remain on /markets with their full history.

Reading the last three columns. Δ is the ensemble against holding the market — the only column that says whether the work was worth doing. vs median is the ensemble against the middle of all fifty cells, and it is the direct measurement of the client's point: a positive number there means selecting a robust region beats the random pick the median stands for. vs best is the ensemble against the luckiest single cell. That one is almost always negative in sample and is supposed to be — the ensemble gives up a little of the in-sample peak in exchange for not depending on one parameter.

* Markets with a note you should read first.

  • Lithium & Battery ETF: This is NOT the lithium price. CME lithium carbonate futures are not on our feed, so the card tracks the Global X Lithium & Battery Tech ETF (LIT) — an equity basket of lithium and battery companies. It moves with equity beta and battery-sector sentiment as well as lithium itself, and is the closest available proxy rather than the commodity.
  • Brazilian Real: CME Brazilian real futures, quoted in USD per real, so a long position is long the real against the dollar. About a third of bars have high == low, so the intraday stop and target sections lean close-only.
  • Orange Juice: Orange juice also has its own seasonality card elsewhere on the site; this one is the twelve-section trend study on the same futures series, not a duplicate of it.
  • British Pound: CME British pound futures, quoted in USD per pound, so a long position is long sterling against the dollar — not the GBP/USD spot convention where the pair is quoted the other way round.
  • Mexican Peso: CME Mexican peso futures, quoted in USD per peso, so a long position is long the peso against the dollar. Yahoo reports settlement-only bars for much of the early history — 70% of bars had no intraday range in 2001-2008, falling to 39% since 2017 — so the intraday stop and target sections lean close-only on this card. The closes themselves move normally (2.5% unchanged days, ATR never zero).
  • Japanese Yen: CME Japanese yen futures, quoted in USD per yen, so a long position is long the yen against the dollar — not the USD/JPY spot convention, where long means the opposite.

Does the efficiency ratio predict which trend speed a market wants?

The reason for funding several speeds per market rather than one is that the efficiency ratio suggested different markets prefer different SMA lengths. That is a testable claim, so here it is tested rather than repeated: each funded market's ER over this window against the centre of its plateau. Spearman rank correlation across 14 markets is -0.275. The hypothesis predicts a negative sign — straighter price paths should favour shorter SMAs — and the data is close to flat, so on this evidence ER does not pick the speed for you. That is an argument for funding several speeds rather than trusting one.

What this tab is not. Every number here is in-sample. A plateau is a better bet than a spike, but it is still chosen by looking at the same data it is scored on, so these figures are an upper bound on what the method would have delivered live. The honest test is the withheld decade, and it stays withheld until the strategies are fixed — running it now would turn the only clean out-of-sample data available into just another parameter. Costs and slippage are not modelled anywhere on this tab; the ensembles trade three or four times as often as a single system, so a realistic cost assumption will hurt them more than it hurts buy-and-hold.