TradingView Out-of-Sample Testing for NQ and MNQ

A strategy can look better every time you change a setting and still become less useful. If the same historical trades keep deciding which rules you retain, you are testing your ability to fit that history as well as the trading idea.
Out-of-sample testing separates the data used to develop a strategy from data reserved for evaluating a frozen version. For an NQ or MNQ strategy in TradingView, the key is not a special performance threshold. It is a disciplined boundary: decide the rules before looking at the test result, then record what happens without quietly rewriting the experiment.
What counts as out-of-sample data?
In-sample data is the historical period used to develop or select rules and inputs. Out-of-sample data is a separate period that did not influence those decisions. TradingView's strategy documentation describes splitting data this way and testing the selected strategy on the out-of-sample period without fine-tuning it.
A date range is not automatically out of sample because it is later on the calendar. If you have already studied its trades while choosing your stop, entry filter or trading hours, those observations have influenced the model. Calling the same range a test set does not undo that exposure.
This distinction also applies to discretionary exclusions. Removing a losing news day after seeing its result is another rule choice. If an event filter belongs in the strategy, define the event categories, timing and information available at the decision time before evaluation.
Freeze more than the entry settings
A reproducible baseline needs the entire experiment, not just a screenshot of a profitable report. Save a dated record containing:
Strategy name, script version and all input values.
Exact symbol, contract month or continuous-series identifier, timeframe, chart type, session and timezone.
Entry, exit, stop, target, quantity and position-management rules.
Commission, slippage, order-fill assumptions and calculation settings.
Development dates, test dates and the method used to restrict trades to each window.
Known data gaps, rollover treatment, news exclusions and unresolved assumptions.
The evaluation questions and the conditions that would make the result unsuitable for the intended use.
Do not change position sizing between two runs and attribute the difference entirely to the entry logic. Likewise, NQ and MNQ are different contracts with different dollar exposure. Record which one the test actually uses rather than treating their profit totals as interchangeable.
Keep costs in the baseline instead of adding them only after a favorable result. The NQ and MNQ commission-and-slippage guide explains how to test those assumptions separately.
A chronological test plan you can copy
The following dates are an invented research example, not an AORDS backtest or a recommendation that these periods are sufficient.
Development window: January 1 through December 31, 2024. Develop the rules and compare the limited alternatives you decided to investigate.
Freeze point: Save version 1, its settings and evaluation plan before examining the next window's results.
Historical test window: January 1 through June 30, 2025. Evaluate version 1 unchanged and retain the complete report and trade list.
Decision: Record the outcome, limitations and any reason to reject the version. Do not hide this record if the result is disappointing.
This plan is a genuine holdout only if the 2025 observations did not inform development. If you have already used all available history to choose the strategy, acknowledge that limitation. A future paper-observation period can provide new observations, but its simulated execution must remain clearly labeled.
Choose the time windows before seeing which split produces the best result. A single favorable six-month segment cannot establish behavior in every market condition. There is no universal split ratio or trade count that turns a strategy into a proven one.
Run the comparison without changing the experiment
Separate the scoring window from the calculation history
Indicators may need earlier bars to establish their state. Restricting new entries to the test window is not necessarily the same as starting every calculation on its first day. Document the warm-up and reset behavior, and decide how to handle a position crossing a boundary.
Use the supported date-range controls or the strategy's documented date inputs, then verify the actual reported period and trade list. Available controls depend on the script and account. Do not assume that moving the visible chart has changed the strategy's evaluation window. The backtest start-date guide explains why initialization can change results.
Keep the symbol and execution assumptions consistent
Use the same declared data conventions and simulation assumptions across the development and test reports. If a contract rollover or data-availability change makes that impossible, record the break rather than presenting the two runs as directly comparable. See the continuous versus dated MNQ contract guide.
A historical strategy report still uses TradingView's broker emulator. Holding out data does not fix lookahead bias, synthetic-price fills, unrealistic costs or execution mistakes. Review those issues before interpreting a clean-looking holdout result. TradingView documents the emulator's assumptions.
Compare a set of questions, not just net profit
How many eligible trades occurred, and were they concentrated in a few sessions?
What were the net average trade and profit factor under the stated costs?
How large were drawdowns, and which drawdown definition was used?
Did one unusually large winner dominate the result?
Did losses cluster in conditions the development window scarcely contained?
Were any trades missing, invalid or affected by unresolved simulation assumptions?
A positive total can coexist with an unacceptable drawdown or too little evidence. A negative total is a reason to investigate, not permission to change the settings until that same test turns positive. Use the expectancy and profit-factor guide to interpret the metrics rather than selecting whichever one looks strongest.
What happens when you change a setting after the test?
You have started a new development step. Keep version 1's result, describe why version 2 changed, and treat the period that motivated the change as information used in development.
For example, suppose version 1 performs poorly in the holdout. You inspect those losing trades, shorten the permitted entry window and rerun the same dates. The improved result may suggest a hypothesis worth investigating, but it is no longer an untouched test of version 2. This is an application of the test-set leakage principle: information from evaluation has entered model selection.
Repeating this process across many stops, filters and dates can produce an attractive winner by selection alone. Research by Bailey, Borwein, López de Prado and Zhu explains why ordinary holdout methods can be unreliable in investment backtests when multiple configurations are tried. Their paper, The Probability of Backtest Overfitting, supports treating a simple split as one safeguard rather than a certificate of reliability.
Keep an experiment log: version, hypothesis, settings changed, windows examined, result, decision and remaining questions. Record abandoned variants too. The number of attempts is part of the context behind the selected result.
Where walk-forward testing helps, and where it does not
A walk-forward process repeats development followed by a later test window. At each step, settings are selected using only information available before that step's test period. It can show how a predefined update process behaves across several chronological segments. QuantConnect's walk-forward documentation describes periodically selecting parameters using a trailing historical window.
The update schedule, selection rule and window lengths are themselves design choices. If you keep changing them to improve the combined historical outcome, you are optimizing the walk-forward process too. Preserve a final untouched evaluation period where feasible, or clearly state that no such period remains.
Do not describe a rolling historical simulation as live trading. A rule-based research process and a working live execution system are separate things.
Frequently asked questions
Is changing a risk setting still optimization?
Yes, if the historical outcome influences which value you keep. Stops, targets, quantity rules and trading-hour filters can all affect the selected result, even when the entry signal never changes.
Can a ready-built strategy be tested out of sample?
Yes, but ask what data informed its development. A period you personally have not seen may still have been used by its developer. New observations collected after a documented freeze point are easier to distinguish from previously used history, though they still cannot guarantee future performance.
Does a passing test mean I should trade live?
No. It answers only the questions defined for that experiment under its assumptions. Capital risk, operational readiness and execution behavior require separate assessment. The backtest-versus-live guide provides a separate workflow for those differences.
The useful outcome is an honest record: what was selected, what was held back, what changed and what remains uncertain. Freeze first, evaluate second, and keep the unsuccessful experiments visible.
Educational information only. Futures involve substantial risk. Historical and simulated results do not guarantee future results.
Get AORDS through Whop and follow the access instructions sent by email. Cancel anytime.
© 2026 AORDS. Trading involves risk. Past performance does not guarantee future results.