So what is backtesting in trading, in plain terms? You take a fixed set of rules, run them over price history, and record what would have happened.
The idea sounds simple, and the trap sits in the word would. A backtest describes one version of the past, not a preview of your account.

What Is Backtesting in Trading?
Three pieces make up every test. History goes in, a rule reads it, and a result comes out the other side.
The Definition in One Line
Backtesting measures how a fixed rule would have behaved on data it did not shape. Nothing more than that, and nothing less.
Fixed matters most. A rule you adjust while watching the outcome stops being tested and starts being drawn.
The Three Pieces
History supplies the bars, complete with the gaps and quirks of whoever recorded it. Your rule supplies the decisions, stated precisely enough for a machine.
The result supplies a trade list and a curve. All three pieces deserve suspicion, and most traders only inspect the third.
What It Is Not
A backtest is not a forecast, an approval or a promise. It is a record of one experiment on one sample.
Nor does it measure your discipline. The simulated trader never hesitates, never oversleeps and never doubles up after a loss.
Nor does it prove a rule beats the market. It shows one outcome, on one record, under settings you chose yourself.
Why Traders Bother at All
Given those limits, testing still earns its place. Three benefits justify the hours.
Elimination Is Cheap
Most trading ideas fail. Finding that out on history costs a few hours, while finding out live costs the account.
So think of the tester as a filter rather than a search engine. Its job is throwing things away.
That framing also protects your time. An idea killed on Tuesday cannot occupy the next six months.
It Forces the Rules to Exist
You cannot test a hunch. Writing the entry, the exit and the size precisely enough to run is valuable even if the test never happens.
Half of what people call a strategy dissolves at this stage. That collapse is useful information.
It Produces Numbers You Can Argue With
A test yields trade count, average win, average loss and the deepest drawdown. Those figures beat any impression formed by scrolling a chart.
Feed them into our expectancy calculator to see what a single trade is worth on average. One number then replaces a lot of arm waving.
Where a Backtest Sits in the Process
Testing history is the first step of four, not the whole job. Each later step removes an illusion the earlier one allowed.

In Sample First
Run the rule once on the first stretch of data. Resist the urge to tune, since tuning turns evidence into decoration.
Then Data It Never Met
Apply the identical settings to a held-back period. This single step filters out most of what circulates online.
Then Forward, on Live Prices
Run the rule on prices arriving in real time. Our note on forward testing covers how long that phase should last.
Then Small Live Size
Real orders bring real fills, real spread and real emotion. Nothing before this point tests any of the three.
What a Backtest Actually Proves
This section is the reason the page exists. Get it right and every report reads differently afterwards.
Disproof Is Strong
A rule that lost money across years of varied conditions has failed a fair examination. That verdict holds, because those conditions genuinely occurred.
You can therefore discard the idea with confidence. Elimination is the one thing history does reliably.
Proof Is Not Available
A rule that made money has passed one examination on one sample. No promise exists that tomorrow resembles that sample.
So the honest verdict on a good result reads not yet disproved. That phrasing keeps position sizes sensible.
The Sample Problem
Any history is one path the market happened to take. Reshuffle the order of a few shocks and the curve changes shape.
Treat the tested period as a single draw. One draw supports a weak conclusion, however many bars it contains.
Where the Test Ended
The most honest picture of backtesting comes from joining two curves. The panel below does exactly that.

The Join in the Curve
Everything left of the marked point comes from simulation. Everything right of it comes from orders that actually filled.
That join appears in most real records. The line rarely continues at the same angle.
Why the Live Part Usually Sits Lower
Four causes stack up. Costs arrive in full, fills scatter, the market changes character, and the trader intervenes.
None of that requires dishonesty from anyone. It simply reflects what a simulation can and cannot hold.
What the Join Should Change
Expect a step down and plan around it. A rule that only just clears its costs on history will usually fall short live.
Build the margin in advance instead. Ideas that survive a deliberately pessimistic cost model tend to survive the join as well.
Then keep both records side by side afterwards. A live log read against its own test is the most useful document a trader keeps.
How Faithful Is the Simulation?
Quality varies enormously between runs. Four settings decide how close the model sits to reality.
The Modelling Mode
The MT4 Strategy Tester offers every tick, control points and open prices only. Every tick models movement inside each bar from M1 history, control points approximates it far more roughly, and open prices only checks the bar open alone.
Match the mode to when your rule acts. Our procedural guide on how to backtest a strategy walks through the settings screen.
Modelling Quality
The report prints a percentage for modelling quality. Around ninety percent is the practical ceiling on MT4’s own M1 history in every tick mode, and it reads not applicable for open prices only.
Higher figures in adverts come from imported real tick data. Our page on modelling quality explains what the number does and does not tell you.
The Spread Assumption
The tester applies a fixed spread unless you change it. Live spread widens on releases and around the daily rollover, so cost is understated exactly when it bites.
Mismatched Chart Errors
That line in the report counts holes in the history. A large figure means your data is broken, so it says nothing about the strategy.
The Data Behind the Test
Every result inherits the quality of its history. Three questions settle whether that history deserves trust.
Whose Record Is It?
Platform history arrives from one broker’s server. Quotes, holiday handling and the exact close of each bar all vary between firms.
Two brokers can hand you two different answers on a single rule. Neither answer is the market itself.
How Deep Does It Go?
Short files flatter almost any idea. Cover a trending stretch, a quiet range and at least one shock, or the sample proves very little.
How Fine Is It?
Bar data hides the path inside each bar. Tick data shows that path, and somebody has to import it before the tester can use it.
Fine data matters most for rules that act within a bar. For a weekly rule the difference barely registers.
Bias That Creeps In Without Anyone Lying
Most flattering results come from ordinary mistakes. Four biases account for the bulk of them.
Lookahead
A rule that reads the close of a bar and then enters at the open of that same bar knows the future. Reports built that way look wonderful and cannot happen.
Optimisation Bias
Try enough settings and one combination tops the list, even on data with no pattern in it. The winner usually describes noise.
Survivorship
You test the symbols and the ideas still in front of you. Whatever vanished never enters the sample at all.
Hindsight in Manual Replay
Bar replay makes peeking easy. One glance forward converts a test into a memory exercise, and the log stops meaning anything.
Each bias looks obvious in somebody else’s work. Assume your own run contains at least one.
Reading a Report Without Fooling Yourself
Reports invite the wrong reading order. Start at the bottom and work up.

Open the Trade List First
Behaviour lives in the trade list: how often the rule fires, how long it holds, how the losses cluster. The curve merely summarises all that.
Then the Drawdown
Find the deepest fall from a peak, then find how long the account spent underwater. Depth ends accounts, and duration ends patience.
Then the Trade Count
Thirty trades produce a confident-looking number and no information. Several hundred across different conditions start to mean something.
Net Profit Last
The headline total depends on position size, which you chose. It tells you less than any other figure on the page.
Double the size and the total doubles, while the edge stays exactly where it was. So read profit as a consequence of sizing, not as a measure of quality.
Note What Is Absent
Count what the report never shows: the modelled spread, the date range, the settings and the trade list. Absence of those items is itself a finding.
The Three Things Missing From Every Backtest
Some gaps cannot be closed by better data. Three of them deserve naming.

Slippage Variation
A test applies a rule about fills. Live fills scatter around your price, and they scatter widest when the market moves fastest.
Requotes
On some execution models the platform returns a new price instead of filling. Our note on requotes covers when that happens and why.
Your Own Hesitation
The simulated trader takes every signal. You will skip one, usually the one that follows three losses, and often the best one of the month.
Keep a written log to measure that gap. Our guide to using a trade journal shows how to record skipped signals alongside taken ones.
Download the complete indicator database
Put these concepts on your charts. One email unlocks the full library of 1,380+ indicators with compiled MT4 and MT5 files, plus my TradingView scripts. No paywall, no spam, unsubscribe any time.
Get free access to my indicator database
One email unlocks 1,380+ free MT4, MT5 and TradingView indicators — the complete library. No single-tool download; you get the whole database.
A Small Worked Example
Concrete numbers make the limits obvious. Take one simple rule and follow it through.
The Rule
Buy when a fast moving average crosses above a slow one. Exit after twenty bars, with a stop at twice the average range.
The First Run
Two hundred trades appear across five years, with a positive average result before costs. Equity climbs, and the temptation to stop reading arrives right here.
Plenty of people stop at this screenshot. Two more checks decide whether the idea means anything.
The Cost Check
Subtract a realistic spread from every one of those trades. Two hundred trades multiply a small per-trade cost into a figure that reshapes the curve.
The Unseen Period
Run the frozen settings on two years held back from the start. If behaviour holds, the idea survives; if it collapses, one afternoon saved you a year.
The Verdict
Either outcome repays the time spent. Only one of them lets you write not yet disproved in your notes.
Manual Testing and Automated Testing
Two routes exist, and they suit different methods. Neither one is more serious than the other.
Automated Runs
A program executes thousands of trades in minutes and removes human error from the arithmetic. It also removes the learning that comes from watching each decision.
Automation needs the rule expressed as code. Our explainer on expert advisors covers what that involves.
Manual Bar Replay
Step through history one bar at a time with the future hidden. Log every decision before the next bar prints.
Discretionary methods can only be tested this way. The cost is time, and the benefit is genuine familiarity with the pattern.
The Honest Comparison
Automated tests scale, while manual tests teach. Serious work on a discretionary rule usually needs a few hundred replayed occurrences.
Two Reports, One Strategy
A single rule can produce two very different reports. Nothing dishonest need happen in between.
Change the Date Range
Shift the start by three months and a curve often changes character. A record beginning after a bad stretch describes the calendar more than the rule.
Change the Modelling Mode
Open prices only can turn a losing intrabar rule into a winner on paper. The mode decided that answer, not the market.
Change the Spread
Add one pip to the modelled spread and thin edges vanish. Frequency decides how hard that single change bites.
So ask for the inputs alongside any curve. Without them, two reports cannot be compared at all.
What a Test Says About You
Nothing whatever, and that gap deserves attention. The simulated trader is an ideal version of you.
The Ideal Trader in the Report
They take every signal, never oversleep, never widen a stop and never double after a loss. You will do at least one of those things.
Measuring the Gap
Trade the rule live at minimum size, then compare the two logs trade by trade. That difference is your execution gap, and it can be measured.
Shrinking It
Automation removes hesitation and introduces different problems instead. A written checklist closes part of the gap without any code at all.
Vocabulary Worth Knowing
Reports and forums use a small set of terms constantly. Nine of them cover almost everything you will meet.
| Term | What it means |
|---|---|
| In sample | The stretch of history a rule was developed on |
| Out of sample | A held-back stretch the settings never saw during development |
| Walk forward | Optimising on one window, testing on the next, then stepping along |
| Modelling quality | How well the tester reconstructed movement inside each bar |
| Profit factor | Gross profit divided by gross loss across the trade list |
| Expectancy | The average result of one trade once wins and losses combine |
| Maximum drawdown | The deepest fall from an equity peak during the test |
| Monte Carlo | Reshuffling trade order many times to see the range of outcomes |
| Mismatched chart errors | A count of holes in the history rather than a strategy result |
Learn those nine and most reports become readable. The remaining jargon is usually decoration.
Common Misreadings
Five interpretations turn a useful exercise into wishful thinking. Each one has a simple correction.
Treating the Curve as a Forecast
The shape describes the past under your assumptions. Nothing about it constrains next year.
Testing Until It Passes
Repeated passes over one sample find its accidents. Our page on curve fitting shows the signature of a fitted result.
Ignoring the Cost Model
Check the spread setting and whether commission appeared at all. A rule that only survives at zero cost was never viable.
Judging a Rule on Twenty Trades
Short samples swing wildly for reasons that have nothing to do with edge. Variance dominates early and settles slowly.
Comparing Two Tests Run Differently
Two reports on different symbols, ranges, modes or spreads cannot be ranked against each other. Match the inputs first, or draw no conclusion.
Backtesting for Discretionary Traders
Testing is often dismissed as an algorithmic activity. That is a mistake, and a costly one.
Rules Exist Whether You Write Them or Not
Every discretionary trader follows patterns. Writing them down converts an instinct into something you can examine.
Test the Setup, Not the Whole Method
Pick one repeatable setup and test that alone. Layering three ideas together makes the result impossible to attribute.
Accept a Rougher Answer
Manual results carry more noise and fewer occurrences. They still beat a strong feeling, which is the honest comparison to make.
Our overview of forex trading strategies covers setups worth testing. Start with one, not with five.
What to Do With a Good Result
A promising report deserves a specific response. Three moves make sense.
Try to Break It
Raise the modelled spread, shift the start date, and change one setting slightly in each direction. Fragile results collapse under any of those.
Give It Unseen Data
Run the frozen settings on a period you held back. Accept whatever that run says, once.
Then Go Small
Trade minimum size and compare the live log against the test. Two records that disagree teach more than any curve.
FAQ
What is backtesting in trading, in one sentence?
It is the practice of running a fixed set of rules over historical prices to record how those rules would have behaved. The result is a trade list and an equity curve for one sample of the past. That output describes history under your assumptions rather than predicting anything.
Is a profitable backtest enough to start trading a rule?
No, because a good result only means the rule has not been disproved on that sample. Confidence should come from a held-back period, then forward testing on live prices, then a small live record. Each stage removes an illusion the previous one allowed.
How is backtesting different from forward testing?
A backtest runs over data that already exists, so it finishes in minutes and can cover years. Forward testing runs in real time on prices nobody has seen, which makes it slow and far more honest about execution. Most serious processes use both, in that order.
Why can a backtest never be fully accurate?
Three things sit permanently outside it: slippage that varies from trade to trade, requotes on some execution models, and your own hesitation. A widening spread belongs on that list too, since the tester holds spread fixed by default. Better data narrows the gap and never closes it.
Do I need programming skills to backtest?
Not for manual work, where bar replay and a written log cover a discretionary method well. Automated testing does need the rule expressed as code, either written yourself or built from an existing program. Start manually if the rule involves judgement, since coding a vague idea simply produces confident nonsense.
How often should a tested rule be retested?
Re-run the same frozen settings once a year, or sooner after an obvious change in conditions such as a lasting shift in volatility. Resist re-optimising at each review, since that quietly refits the rule to recent noise. Compare the fresh period against the original run rather than replacing it.
How many trades make a backtest meaningful?
Twenty tells you almost nothing, one hundred starts to hint, and several hundred across varied conditions supports a cautious view. Sample size matters more than the number of years covered, because variance dominates small samples completely. Even a large sample describes one path the market happened to take. Results are not guaranteed; past performance is not indicative of future results.
External references
- For background on this concept, see Sampling Bias on Wikipedia.
- For broader market context, see Quantitative Analysis in the BabyPips Forexpedia.
