A strategy can fail because the idea is weak, the code is wrong, the data is misleading, the simulation is permissive, the parameters were overfit, or the market changed. A useful research process attempts to separate those possibilities instead of collapsing them into one performance number.
1. Define the claim before measuring it
Write the mechanism in plain language. What market behaviour is expected? Why should it exist? Which observations drive a decision? Under what conditions should the strategy remain inactive? Which result would make the explanation less plausible?
Then translate that explanation into an unambiguous specification: inputs, timeframes, entry and exit rules, position sizing, state transitions, handling of repeated signals, protective controls, and exceptional cases.
Another engineer should be able to read the specification and identify when two implementations disagree.
2. Audit the evidence source
Document where the data came from, which fields exist, how timestamps and sessions are interpreted, and what transformations were applied. Look for gaps, duplicates, changes in symbol specification, unrealistic spread treatment, and periods whose quality differs from the rest.
The appropriate granularity depends on the strategy. Tick-sensitive execution logic needs a different dataset and simulator than a strategy that acts once at a bar boundary.
3. Verify the implementation independently of performance
Before optimising, construct scenarios whose expected outcome is known. Inspect event logs, state changes, orders, protective levels, and position transitions. Test boundary conditions: session changes, zero or missing values, duplicate events, rejected operations, reconnects, and the first event after initialisation.
A profitable curve cannot tell you whether the system followed the intended rules. Implementation verification needs its own evidence.
4. Establish a documented baseline
Choose a reference configuration and record the code version, dataset, instrument specification, test period, capital assumptions, costs, and platform settings. The baseline is not the “best” result; it is a stable point from which differences can be explained.
Review the distribution beneath aggregate metrics: trade count, holding time, long versus short behaviour, time and regime concentration, cost burden, drawdown path, extreme trades, and dependence on a small set of observations.
5. Search deliberately for fragility
Change parameters around the selected values rather than celebrating a single optimum. Test neighbouring periods and different market conditions. Apply plausible increases in costs or stricter execution assumptions. Remove unusually influential observations. Separate components to understand which module drives the result.
The goal is not to make every variant profitable. It is to learn whether the conclusion depends on a narrow and unexplained combination of choices.
6. Protect genuinely unseen evidence
Where the research design supports it, reserve a period or sample that is not used to invent rules or select parameters. Repeatedly checking a holdout while changing the strategy turns it into training data in practice, even if the file is still labelled “out of sample”.
Forward observation or paper execution can add operational evidence, but it also has limits. A short observation window may contain too few independent events, and simulated fills may remain different from live execution.
7. End with a research decision, not a marketing verdict
A useful review states what was tested, what failed, what changed, what remains unknown, and what the next action is. The decision may be to proceed, redesign, gather better data, narrow the claim, or stop.
Stopping is a valid research outcome. So is deciding that evidence is not yet strong enough for a public performance statement or product release.
Hypothesis, specification, code version, data version, configuration, assumptions, result summary, failure modes, limitations, and next decision.