StratVerra

How traders accidentally overfit an options strategy

Xavi ·

The first backtest is mediocre. You move the entry five minutes later and it improves.

Then you change the short strike from 20 delta to 18, widen the wings, lower the profit target, remove Tuesdays, and add a trend filter. The equity curve smooths out. The drawdown shrinks. You keep going.

At some point, you may stop researching a strategy and start teaching it the answers to one historical exam.

That is overfitting.

Every adjustment spends information

A trader runs a test on five years of data, changes a rule, and runs it again on the same five years. The second result is not an independent confirmation. Information from the period guided the change.

Repeat this often enough and the data becomes part of the strategy design, even if you never write code.

The number of completed trades does not fix this. A backtest with 2,000 trades can still be weak evidence when it was selected from thousands of attempted variations.

Keep a research log. Record the original idea, every material change, and why you made it. The list may be uncomfortable. That is the point.

Exact winners deserve doubt

Suppose a 10:07 entry performs far better than 10:05 or 10:10.

There might be a market reason for that exact minute. More likely, a handful of historical trades landed differently around the cutoff.

Stable ideas tend to work across a neighborhood. Test nearby times, deltas, widths, and exits. You are not trying to find the top performer. Look for a plateau where reasonable settings produce broadly similar behavior.

A narrow spike is a warning that the setting may fit noise.

Give each rule a job

A volatility filter might keep a premium-selling strategy out of low-credit conditions. A delayed entry might avoid the widest markets after the open. Those explanations do not prove the rules work, but they give you something to test.

“It improved the backtest” is not enough on its own.

When two versions behave similarly, prefer the simpler one. Extra rules create more ways to fit historical accidents, more edge cases to debug, and more reasons for live behavior to surprise you.

Complexity is sometimes justified. Make each condition earn its place.

Protect a part of the history from yourself

Divide the available dates chronologically.

Use the earlier period to build the rules. Keep the later period hidden until you have frozen a version. That later section is the holdout, or out-of-sample, test.

The discipline is simple and annoying: if the holdout disappoints, do not keep changing the rules and checking the same period. Once you use it to make a decision, it has joined the development data.

You can revise the strategy, but you will need fresh evidence. Depending on the available history, that may mean a later untouched period, a rolling walk-forward process, or patient observation on current data.

Test different conditions, not just different dates

A five-year sample can contain one dominant market mood.

Look at high- and low-volatility periods, sustained trends, sudden gaps, and quieter sessions. If one type of environment produced nearly all the gains, say so plainly.

A regime filter can be legitimate, but it must use information available at the time. Labeling a period “high volatility” after seeing the entire month creates accidental foresight.

Make costs part of the robustness test

Overfit strategies often have thin historical margins. Slightly worse fills expose them.

After freezing the rules, increase slippage and commissions. If the strategy falls apart after a small cost change, it may be exploiting the fill model rather than a repeatable market behavior.

Do not tune the strategy again for each cost setting. You are testing the same rules under different assumptions.

Know when to stop optimizing

Set the research budget before you begin. Decide which parameters you will vary, the sensible range for each, and what evidence would make you abandon the idea.

Without a stopping rule, every disappointing test becomes an invitation to add one more filter.

A failed idea is not wasted work. You learned that the original claim did not survive the test you gave it. That is exactly what research is supposed to do.

StratVerra saves strategy versions and reproducible backtests, which helps separate one rule set from the next. Use that history as an audit trail, not a gallery of only the winning runs.

The final article covers what comes after historical testing and why paper profit is not the only thing worth watching.

Backtests are simulations. They do not predict future results, and simulated fills can differ from live execution. This article is educational and is not investment advice.