StratVerra

Stop asking whether the backtest made money

Xavi ·

A backtest finishes with a profit. Good. Now the work starts.

The ending balance is the most visible number and often the least informative one by itself. You need to know how the strategy got there, what it had to endure, and how much of the result depended on a few trades.

Win rate can tell the wrong story

Picture ten trades. Nine make $100 each. One loses $1,200.

The win rate is 90%. The strategy lost $300 before costs.

Now flip the shape. Six trades lose $100 and four make $250. That strategy wins only 40% of the time but finishes $400 ahead before costs.

Options structures often produce lopsided payoffs, so win rate needs context. Compare the average win with the average loss. Then look at the biggest observations rather than relying on averages alone.

A high win rate can feel comfortable right up to the trade that erases months of small gains.

Expectancy describes the average trade

A simple form is:

expectancy = (win rate × average win) - (loss rate × average loss)

Positive historical expectancy means the average tested trade made money under the model. It does not mean the next trade will win or that the average will persist.

Sample size matters. A promising average across 18 trades is fragile. Hundreds of trades provide more observations, though a large sample can still be overfit or concentrated in one market regime.

Check expectancy after modeled slippage and commissions. Gross expectancy is not what reaches the account.

Drawdown turns a chart into an experience

Maximum drawdown measures the largest drop from a prior equity peak. Also look at how long the strategy stayed below that peak.

Two tests can earn the same amount. One climbs unevenly but recovers quickly. The other spends a year underwater and makes nearly everything in the final month. Those strategies place different demands on capital and patience.

Historical drawdown is not a worst-case guarantee. The future can produce something larger. Treat the tested drawdown as an observed event, not a boundary.

Find out who paid for the result

Break the backtest into years, quarters, volatility conditions, or another grouping that makes sense for the strategy.

You are looking for concentration. Did one crisis create most of the profit? Did the strategy stop working after a market structure change? Were all the losses clustered around expiration days or major announcements?

Uneven performance is normal. Dependence on one accidental period is different.

Sort trades from largest winner to largest loser. Then remove the best few trades and recalculate the broad conclusion. This is not because those winners are illegitimate. It tells you whether the result has any depth beyond them.

Do the same with losses. Read the trade records and logs. A giant loss may reveal intended strategy risk, a contract-selection surprise, an execution assumption, or a rule that did not behave as expected.

Profit factor helps, but it can flatter a small sample

Profit factor divides gross profit by gross loss.

A value above one means the tested winners outweighed the tested losers before any omitted costs. It is useful beside win rate because it reflects magnitude.

An unusually high value deserves suspicion, not applause. A short date range, one huge winner, or optimistic fills can produce a number that will not survive another sample.

No threshold can certify a strategy. Read the number in the context of trade count, drawdown, cost assumptions, and how many variations you tested before choosing this one.

Stability matters more than the champion run

Rerun the strategy with nearby settings. If a 15-delta entry looks good, inspect 13, 14, 16, and 17. If the entry is 10:00, try reasonable neighboring times.

You are not looking for a new winner. You are checking whether the result sits in a broad area of similar behavior or on one sharp peak.

Then make execution less friendly. A result that survives modest changes in parameters and costs is more interesting than the best line in a leaderboard.

A short review routine

After every serious run:

  1. Confirm the selected contracts and fills match the rules.
  2. Read the largest winners and losses.
  3. Compare gross and net results after costs.
  4. Examine drawdown depth and duration.
  5. Split results across time.
  6. Test nearby parameters without searching for a new optimum.
  7. Save the exact version and run settings.

StratVerra produces reproducible backtests and uses the same strategy logic when you move into supported execution modes. That makes it easier to compare what the historical model said with what the strategy later does in paper or live trading.

Before that step, there is one more research trap to handle: overfitting.

Backtests are simulations. They do not predict future results, and simulated fills can differ from live execution. This article is educational and is not investment advice.