Skip to content
Research
17 min read

One sweep, six ways to pick a config

Ninety-six configurations, six explicit rules for picking one, and a holdout none of them were chosen on. The training table ranked them almost exactly backwards, the plateau heuristic finished last, and the best result came from a parameter I removed to prove it was needed.

Universe
Top 1,500 US names by liquidity, point-in-time
Test period
2020–2026

A parameter sweep does not hand you a strategy. It hands you a table with one row per configuration, and then you have to pick a row. That second step gets far less scrutiny than the first, which is odd, because the sweep is arithmetic and the pick is where all the judgement lives.

So I ran it as an experiment. One sweep of 96 configurations, six explicit rules for choosing a row from it, and every resulting configuration taken to a holdout period that none of them were selected on. The rules are code, not taste: each one resolves to exactly one row, so nobody gets to nudge the answer afterwards.

The short version is that the training table told me almost nothing about which configuration would do well afterwards, and where it did say something, it was pointing the wrong way.

The strategy

Everything below runs the same system. Only its parameters change.

Cross-sectional momentum is about as old and as plain as a systematic equity strategy gets. Rank a universe by how much each name has gone up recently, buy the ones at the top, and keep re-ranking. The bet is that recent relative strength persists for a while, which Jegadeesh and Titman documented in 1993 and which has been argued over ever since.

The mechanics here, in the order they happen on a rebalance day:

  1. Take the eligible universe. The 1,500 most liquid US names, reconstructed as of the decision date rather than as of today. Dead tickers are still in the data. A company that was liquid in 2021 and delisted in 2023 is present until the day it stops trading, which is the only way the backtest can lose money on one.
  2. Score each name by its trailing return over a lookback window, optionally ignoring the most recent month. That gap is the skip, and the point of it is short-term reversal: names that have just spiked tend to give some of it back, so the last month of return is the least reliable part of the signal.
  3. Take the top N by that score and hold them, either equal weight or inverse-volatility weighted, so a name that moves twice as much gets half the money.
  4. Sell laggards with hysteresis. A holding is not sold the moment it drops out of the top N. It is sold when it falls out of the top 2N. Without that buffer, a name oscillating around the boundary gets bought and sold every rebalance and pays the spread each time.
  5. Optionally stand aside in a downtrend, blocking new entries while the market index sits below its 200-day average.

That leaves six knobs, and the sweep turns each one:

  • lookback 63, 126 or 252 trading days
  • skip 0 or 21 days
  • rebalance every 5 or 21 days
  • crash filter off, or block new entries below the 200-day average
  • weighting equal or inverse volatility
  • book size 10 or 20 names

Three times two, five times over, is 96 configurations.

The setup

  • Universe: 1500 US names, point-in-time, delisted retained
  • Train: 2020-01-01 to 2024-01-01
  • Holdout: 2024-01-01 to 2026-07-31
  • Costs: 0.25% per fill
  • Capital: $25,000
  • Benchmark: an equal-weight index of the same universe

One detail matters more than the others. Universe selection, meaning which 1,500 names are eligible at all, is frozen at the training boundary in both windows. Ranking the liquidity cap across the whole period instead would pick the names that turned out to be liquid by 2026 and let a 2020 backtest trade them, which is survivorship bias sneaking back in through the liquidity filter. It was worth roughly 8 percentage points of training CAGR per 1.5 years of leakage when I had that bug, which is more than most of the effects being measured here.

The table you are picking from

Before the rules, look at the spread. Here is the training MAR, meaning CAGR divided by maximum drawdown, for all 96 configurations. The letters are where the six rules land.

medianABDEF0.201.40

96 configs, MAR on the training window

The best configuration in the table scores five times the worst. That range is the size of the decision being made, and it is entirely a decision about parameters, not about whether momentum works.

Six rules

  • A. Plateau. Score every configuration by the median MAR of itself and its immediate neighbours, one grid step away on exactly one axis, and take the highest. This is the standard advice: prefer a broad flat region over a lone spike.
  • B. Best training CAGR. Take the top row by return. The rule every methodology section warns against.
  • C. Best training Sharpe. Take the top row by Sharpe.
  • D. Best training MAR. Take the top row by return over drawdown. This is the rule my live book is picked with.
  • E. Textbook 12-1. Do not pick from the sweep at all. Use the specification the literature has used since 1993: 252-day lookback, skip the most recent month, no crash filter, equal weight, monthly, 20 names.
  • F. Best MAR, no skip. Take D's configuration and remove the skip. Not a selection rule, an ablation, to price one parameter in isolation.

Here is what each rule actually selected.

RuleLookbackSkipRebalanceCrash filterWeightingNames
APlateau (neighbourhood median)63dNone21dOffEqual10
BBest train CAGR(same as C)126dNone21dOffEqual10
CBest train Sharpe(same as B)126dNone21dOffEqual10
DBest train MAR126d21d21dOffInverse vol20
ETextbook 12-1, no filter252d21d21dOffEqual20
FBest train MAR, no skip126dNone21dOffInverse vol20

Two of these collide. On this sweep the highest-CAGR row and the highest-Sharpe row are the same configuration, so B and C are one pick wearing two hats. That is itself worth knowing: the choice between ranking by return and ranking by Sharpe, which sounds consequential, changed nothing here.

Read down the columns and the picks are not scattered across the grid. Every rule chose monthly rebalancing over weekly and every rule left the crash filter off, so two of the six axes were effectively decided before any selection rule got involved. The disagreement is concentrated in the lookback, the skip, the weighting and the book size.

What happened

TrainHoldout
CAGRMax DDSharpeMARTradesCAGRMax DDSharpeMARTrades
APlateau (neighbourhood median)63.4%-54.1%1.081.1749221.2%-41.1%0.590.52298
BBest train CAGR73.4%-57.7%1.191.2731450.2%-40.5%1.041.24189
CBest train Sharpe73.4%-57.7%1.191.2731450.2%-40.5%1.041.24189
DBest train MAR57.0%-41.0%1.171.3956659.9%-36.2%1.371.65367
ETextbook 12-1, no filter27.6%-44.7%0.730.6239061.6%-39.6%1.251.56224
FBest train MAR, no skip45.6%-45.6%1.041.0056866.0%-37.2%1.361.77374

For scale, three things that involve no selection at all, over the same holdout:

  • Equal-weight universe: 11.9% CAGR, 22.2% drawdown, MAR 0.54
  • QQQ, bought and held: 25.5% CAGR, 22.9% drawdown, MAR 1.11, Sharpe 1.18
  • SPY, bought and held: 20.2% CAGR, 19% drawdown, MAR 1.06, Sharpe 1.25

Index returns are close-to-close and pay no costs, the same treatment the equal-weight benchmark gets, so they are a little flattered relative to the strategies. Read them as a floor rather than a like-for-like line.

Train → Holdout

0.000.501.001.502.00TrainHoldoutA0.52Plateau (neighbourhood median)B1.24Best train CAGRC1.24Best train SharpeD1.65Best train MARE1.56Textbook 12-1, no filterF1.77Best train MAR, no skip
Fig. 1Every configuration's training reading joined to its holdout reading. A is drawn in red. Switch the metric: the ordering is not stable across them, and it is not stable across windows either.

The index is a harder benchmark than the universe

The equal-weight universe is the fair internal benchmark, because it is the same names the strategy picks from. It is also a low bar: it returned 11.9% over the holdout. Every pick except A cleared it comfortably.

QQQ is the bar a reader actually cares about, and it is much higher. Held from the first day of the holdout it returned 25.5% a year at a 22.9% drawdown, for a MAR of 1.11 and a Sharpe of 1.18.

That reframes two rows. B, the highest-training pick, returned twice QQQ's CAGR but at nearly twice the drawdown, finishing at MAR 1.24 against QQQ's 1.11 and with a lower Sharpe, 1.04 against 1.18. And A did not merely lose to the universe, it lost badly to an index fund.

The honest summary is that only D, E and F earned their complexity here. Every one of these strategies takes a drawdown in the high thirties to low forties. QQQ took 22.9%. Anything in this table has to be worth sitting through roughly twice the pain of holding the Nasdaq.

What costs do to the ranking

Every figure so far assumes 0.25% per fill. Turnover is not constant across these picks, so that assumption is doing work: over the holdout B trades 189 times and F trades 374, nearly double. A cost assumption that favours low turnover would favour B.

So I ran the holdout again at four cost levels.

Cost per fillABDEF
0.00%0.581.341.731.591.85
0.10%0.571.271.691.661.83
0.25%0.521.241.651.561.77
0.50%0.431.111.581.631.71

MAR

The headline survives. F is first at every cost level and A is last at every cost level, so the two claims the article leans on are not artefacts of one cost assumption. Doubling costs from 0.25% to 0.50% takes about 0.06 off F's MAR, which is small next to the gap it is winning by.

The middle of the table does not survive. D and E swap places between 0.25% and 0.50%, which is what you would expect from two configurations separated by less than 0.1 of MAR. And E is not even monotonic in cost: its holdout CAGR runs 61.6%, 65.4%, 61.6%, 66.2% as costs rise, which is impossible as a cost effect and is instead the book taking a different path. Whole-share sizing and a changed fill price shift which names are held, and over a twenty-name book that compounds.

I find that more informative than the cost sensitivity itself. Raising the cost from 0.25% to 0.50% per fill increased E's holdout CAGR, from 61.6% to 66.2%. Charging a strategy more cannot make it earn more, so that 4.6 percentage points is pure path noise, and it is larger than the gap between D and E in the main table. Differences of that size are not signal, and the middle three rows should be read as a tie rather than a ranking.

The training table was worse than useless

Rank the five distinct configurations by training CAGR, then rank them by holdout CAGR, and compare the two orderings. The rank correlation is -0.8. Not weak. Negative, and strongly so. On this sweep, over this holdout, the configurations that looked best in training were systematically the ones that did worst afterwards.

On MAR and on Sharpe the correlation is -0.1, which is another way of saying nothing at all.

I want to be careful about how much weight that number can carry. It is five configurations over one holdout in one market, so it is an illustration rather than a measurement, and a single number computed on five points is fragile by construction. But the direction is clear enough in the table above that it does not depend on the statistic: the two best training rows, B and A, finished fourth and fifth out of five.

The plateau rule lost, badly

A is the rule I would have defended before running this. Prefer the flat region, avoid the spike, do not take the maximum of anything. It selected a short-lookback, no-skip, ten-name book that returned 63.4% in training and then 21.2% in the holdout, for a MAR of 0.52.

That is the worst result in the table, and it is the only pick that failed to beat the equal-weight benchmark on MAR. All that care, and it lost to holding everything.

I do not think this kills the plateau idea, and I want to say why rather than quietly dropping it. My neighbourhood is crude: six axes, most of them with only two values, so "one step away" often means flipping a switch rather than nudging a dial. A plateau in a space that coarse is not much of a plateau. What the result does kill is the idea that the heuristic is free protection. It is a model of the parameter space, and a bad model of it costs you.

Ranking by MAR picked a different, better book

D is the only sweep-derived rule that chose a genuinely different kind of configuration: a longer lookback, the skip switched on, inverse-volatility weighting, twenty names instead of ten. It gave up a lot in training, 57% against B's 73.4%, and it was the best sweep-derived pick out of sample at MAR 1.65 against B's 1.24.

The mechanism is visible in the drawdowns rather than the returns. B ran a 57.7% training drawdown; D ran 41%. Ranking by MAR does not reward finding the configuration that made the most money. It rewards finding the one that made money without a hole in the middle of it, and that property held up out of sample when raw return did not.

This is also the configuration the live system runs, which I should disclose rather than present as a neutral finding. I picked it before running this experiment, so the experiment is a check on a decision already made, not the reason for it.

The config nobody would have picked

E is the plain textbook specification. No sweep, no tuning, the parameters a paper from 1993 hands you. In training it placed last of the five on MAR at 0.62, with a CAGR of 27.6% against a field where the top row made 73.4%. Nobody running this sweep would have shipped it.

Out of sample it returned 61.6%, the second-highest in the table, at MAR 1.56.

I cannot tell you from one holdout whether that is because the textbook specification encodes something real that four years of US data happened to disagree with, or because it got lucky. What I can say is that the sweep ranked it last and the holdout ranked it near the top, and any story about parameter selection has to survive that.

The skip did the opposite of what it does in training

F exists to price one parameter. It is D with the skip removed and nothing else touched. In training that is expensive: CAGR falls from 57% to 45.6%, MAR from 1.39 to 1. Exactly the argument for keeping the skip.

Out of sample it reverses. Removing the skip improved every measure: 59.9% to 66% of CAGR, MAR 1.65 to 1.77, and it was the best result in the whole table.

So the single parameter I could justify from theory, on a documented effect, with training evidence behind it, was working against me for the last two and a half years. I have no clean explanation. Short-term reversal is a real effect and it is also a crowded one, and the honest position is that one holdout cannot distinguish "the skip stopped paying" from "the skip did not pay in this particular window."

Plateau (neighbourhood median)1.66×Best train CAGR2.84×Best train MAR3.38×Textbook 12-1, no filter3.51×Best train MAR, no skip3.72×Equal-weight universe1.37×
1×2×3×4×5×202420252026
Fig. 2Growth of one unit over the holdout only, every configuration starting flat on the first day of the holdout. The dashed line is the equal-weight universe. C is omitted because it is the same configuration as B.

The curves say something the table cannot. E, the textbook configuration that the sweep ranked last, is in front for most of the window and is only caught near the end. D and F pass it in the final months. Read at almost any earlier point, this chart would have crowned a different winner, and the gap between the top three for most of 2025 is small enough that the final ordering rests on a handful of weeks.

A is the exception. It stays with the pack through the first year and only falls away from early 2025, which is worth noting because a shorter holdout would have made it look merely unremarkable rather than the one pick that lost to buying the whole universe.

What this does not show

Five distinct configurations, one holdout window, one market, one strategy family. This is one experiment, and the honest reading is that it constrains my choices rather than establishing a rule.

Three specific reasons for caution:

  1. The holdout is a bull market. From 2024-01-01 the equal-weight universe returned 11.9% a year with a 22.2% drawdown, against 14.9% and 41.4% in training. Momentum in a rising market flatters everything, and five of six configurations improved out of sample. A test where everything gets better is a weak test of what makes things better.
  2. The grid is coarse. Most axes have two values. The plateau rule in particular deserves a re-run on a finer grid before I trust the result against it.
  3. The selection window is one regime. Training runs 2020-01-01 to 2024-01-01, which contains one crash and one enormous recovery. Configurations were ranked on their behaviour through that, and the holdout contains nothing like it.

Checking the harness before believing the harness

Everything above depends on the backtest being honest about time, so I tested that rather than assuming it. Two checks.

The look-ahead probe

If any decision in the holdout used data from after the moment it was made, the result is worthless no matter how careful the selection rules are. The test: run one configuration from the start of the holdout to a midpoint with prices loaded only to that midpoint, then run the same configuration to the end of the holdout with prices loaded to today, and compare the two equity curves over the days they share.

A run that cannot see the future and a run that can must agree exactly on the overlap. They do: 374 shared days, zero differing values. The strategy class enforces this by construction, since it hands each decision only the candles up to the previous close and fills at the next open, but the guarantee is worth measuring rather than trusting.

The one that was not clean

The second check asked whether the 1,500 selected names depend on how far prices were loaded. They are supposed to be frozen at the training boundary. They are not, quite. Loading to the boundary, to a midpoint, and to today each select 1,500 names, and each set differs from the others by exactly one ticker at the liquidity cutoff.

That is small, and it is not nothing. Running the training window under both loads gives 57.04% CAGR against 57.70%, a difference of 0.66 percentage points, because the swapped name is actually traded. The trading calendar is identical in both, 1,006 days, so the cause is the universe rather than the dates. I had assumed the calendar. Measuring it said otherwise.

The mechanism is that the candidate pool is assembled over the loaded window before the liquidity ranking is applied to data up to the boundary. A slightly larger pool moves the cutoff, and the marginal name changes. So a name's eligibility in 2021 depends faintly on data from 2026, which is precisely the class of leak the boundary is there to prevent.

It does not move anything in this article: 0.66 percentage points is far smaller than the gaps the conclusions rest on, the training numbers come from the boundary-limited load and reproduce the sweep exactly, and the look-ahead probe above passes. But the train and holdout passes here are running universes that differ by one name out of 1,500, and it is better to say so than to let someone find it later.

What I changed

Two things, and neither is "found the right rule."

First, I stopped reading a sweep's ordering as information about the future. It is a description of what happened in the training window. On this evidence the ranking is not merely noisy, it may be inverted, and the sensible use of a sweep is to rule out configurations that fail everywhere rather than to crown the one that succeeded most.

Second, the pick now has to clear the benchmark on the holdout before it ships, which A would have failed. That is a lower bar than picking correctly, and it is one I can actually check.

The uncomfortable part is that the best result in this table came from an ablation I ran to confirm a parameter, and the second best came from ignoring the sweep entirely. I have not changed the live configuration on the strength of that, because doing so would be selecting on the holdout, which is the whole error this article is about. It goes on the list to test properly, on a window neither of us has looked at yet.

Each rule in this article is stated tightly enough to argue with: a sort key, a tie-break, and the grid it sorts. If you think the plateau rule was defined badly, and I think there is a case that it was, the definition is in the list above and the disagreement is about that definition rather than about what I meant.

Written by

Koray GocmenFounder

Builds and runs the systematic strategies behind this research.

Full source

The code, the data and the parameter files behind this article.

All research