How to Evaluate a Trading Bot: The Due Diligence Checklist

A practical framework for judging automated strategy products, signal services, and commercial trading bots before you subscribe.

Key Takeaways

  • Most trading-bot marketing fails the first test: you cannot verify the track record, the assumptions, or the execution conditions. That matters because backtests are easy to overfit and hard to reproduce [1][2].
  • Forward-tested results are better than pure backtests, but they still need timestamps, trade logs, fees, slippage, and a clear definition of the universe being traded [3][4].
  • Survivorship bias, hidden costs, and vague “AI” claims can make a mediocre strategy look excellent on paper. A due-diligence checklist should focus on evidence, not branding [5][6].
  • AIBROKER’s cryptographic audit trail and point-in-time universe construction are designed to make strategy verification more auditable; see our methodology page for how those controls work [7].

Retail investors are being sold more automated trading products than ever: bots, signal rooms, model portfolios, copy-trading feeds, and “AI” strategy subscriptions. The packaging changes, but the due-diligence problem is the same. You are being asked to trust a performance claim that may be based on cherry-picked history, unrealistic fills, or a universe that quietly excludes the losers.

The right question is not “Does it look profitable?” It is “Can I verify what was traded, when it was traded, what it cost, and whether the result could have been achieved in the real world?” The SEC has repeatedly warned investors to be skeptical of auto-trading and trading-software claims that promise easy profits or imply that past results guarantee future outcomes [3][4].

Why this matters: if you cannot reconstruct the strategy from source data, you are not evaluating a trading bot — you are evaluating a sales page.

1) Start with the claim, not the product

Before you look at screenshots, Discord testimonials, or a glossy equity curve, write down the exact claim being made. Is the product selling a backtest, a forward test, a live signal stream, or a fully automated execution system? Those are not interchangeable. A backtest is a research artifact. A forward test is a live or paper-trading observation. A live track record is the only one that proves the strategy survived real market conditions, and even then you still need to know whether the account size, fees, and execution venue match your own [1][2].

Bailey and López de Prado’s classic paper on backtest overfitting is still the right warning label here: the more strategy variants you test, the more likely you are to find something that looks good by luck rather than skill [1]. That is especially true in marketplaces where vendors can quietly discard failed versions and promote the one surviving curve. If the seller cannot explain the research process, the result is not robust evidence; it is a marketing artifact.

Table 1. Claim type vs. evidence quality
Claim typeWhat it can proveWhat it cannot proveDue-diligence requirement
BacktestHistorical behavior under stated assumptionsLive execution quality, slippage, regime shiftsFull rules, universe, costs, and out-of-sample testing [1][2]
Forward testObserved performance in a live or paper periodLong-run robustness, capacity, drawdowns in other regimesTimestamped trade logs, broker statements, and duration [3]
Live track recordActual executed resultsFuture returnsIndependent verification, fee disclosure, and account comparability [4]

Common mistake: investors often compare a bot’s backtest to their own live account expectations. That is apples-to-oranges. A backtest usually assumes perfect knowledge of the universe and cleaner fills than a retail account can get in practice.

2) The 10-question due-diligence checklist

Use this as a gate, not a formality. If a vendor cannot answer these questions clearly, you should assume the answer is unfavorable.

Table 2. Trading-bot due-diligence checklist
#QuestionWhat a credible answer looks likeRed flag
1Is the track record backtested, forward-tested, or live?Clearly labeled with dates and methodology“Verified performance” with no definition
2Who verified the results?Independent auditor, broker statement, or reproducible logsOnly screenshots or testimonials
3What universe was traded?Point-in-time list with inclusion/exclusion rulesCurrent universe used to describe past results
4What fees were included?Subscription, commissions, spreads, financing, taxes where relevant“Net returns” with no cost breakdown
5What slippage assumption was used?Explicit basis points or execution modelZero slippage in a fast strategy
6How often is the strategy re-optimized?Fixed schedule with out-of-sample validationFrequent parameter tweaking after losses
7How many variants were tested?Disclosed research process and selection criteriaOnly the best-performing version shown [1]
8What happens in bad regimes?Drawdown history and stress scenariosOnly bull-market examples
9Can you export the trade log?CSV or broker-linked historyNo raw data access
10What would make the strategy stop working?Clear failure conditions and monitoring rules“It works until it doesn’t”

That checklist is intentionally boring. Boring is good. The best defense against a glossy sales pitch is a set of questions that force the vendor to reveal the plumbing.

3) Backtested vs. forward-tested vs. live: the verification ladder

Investors get tripped up because these labels sound more scientific than they are. A backtest can be useful, but only if the rules are fixed before the test and the data are point-in-time. A forward test is more credible because it uses live market conditions, but it may still be too short to include a real bear market or a volatility shock. A live track record is strongest, yet even that can be misleading if the account is tiny, the strategy is capacity-constrained, or the vendor cherry-picks the start date [2][4].

Here is the practical tradeoff: the more “real” the evidence, the less control the vendor has over the environment. That is why you should prefer evidence that is both live and independently auditable, not merely impressive-looking.

Table 3. Verification ladder for automated strategies
LevelStrengthWeaknessBest use
BacktestFast, cheap, repeatableOverfitting risk, unrealistic fillsIdea screening
Forward testUses live conditionsShort sample, regime dependenceEarly validation
Live track recordActual executionMay be tiny, cherry-picked, or non-comparablePrimary evidence

Practical takeaway: if a vendor leads with a backtest and buries the live record, treat the backtest as a hypothesis, not a result.

4) Hidden fees, slippage, and the math that marketing leaves out

Many bots are sold on gross returns. That is the wrong number to focus on. A strategy that turns over positions frequently can look excellent before costs and mediocre after them. The SEC has warned investors that automated trading systems can generate losses if commissions, spreads, and execution quality are not understood up front [3][4].

Slippage is especially important for short-horizon strategies. If a bot trades liquid large-cap stocks, a few basis points may be manageable. If it trades small caps, thin crypto pairs, or fast intraday signals, the gap between the quoted price and the executed price can dominate the edge. A vendor who assumes zero slippage is not being conservative; they are changing the question.

Below is a simple worked example. It is illustrative only, not actual performance data.

Table 4. Illustrative cost drag on a hypothetical bot
AssumptionValue
Gross annual return18.0%
Subscription fee2.0%
Commissions and spreads1.5%
Slippage2.0%
Net annual return12.5%

Footnote: Illustrative example only. Assumes a one-year period, a diversified U.S. equity universe, monthly rebalancing, no taxes, and a zero risk-free rate. Costs are hypothetical and meant to show how gross returns can shrink after fees and execution frictions. Not actual AIBROKER performance or any audited track record.

That 5.5 percentage-point drag is not exotic. It is the kind of gap that appears when a strategy is sold on headline returns but not on implementation details. If the vendor cannot tell you how they modeled spreads, financing, and turnover, you do not know what you are buying.

5) Survivorship bias: the quiet killer in strategy marketplaces

Survivorship bias is one of the most common ways strategy marketplaces mislead users. The platform shows you the bots that are still alive, still marketed, or still subscribed — not the ones that failed and disappeared. That makes the average result look better than the true population. We cover the mechanics of this problem in more depth in our survivorship bias guide, but the core idea is simple: if the graveyard is invisible, the living look unusually strong.

This is not a niche academic issue. It affects ranking pages, “top performers” lists, and copy-trading leaderboards. A strategy that was launched after a favorable market regime, or one that was re-optimized after a drawdown, can look like a star until the cycle turns. Bailey and López de Prado’s work is relevant here because overfitting and selection bias often travel together: the more candidates you test, the more likely the marketplace is to surface a lucky survivor [1].

Why this matters: a marketplace leaderboard is not a research sample. It is a filtered display of survivors, and the filtering rules matter as much as the returns.

6) What investors get wrong about “AI” and automation

The biggest mistake is assuming automation equals objectivity. A bot can be systematic and still be badly designed. It can be fast and still be fragile. It can be “AI-powered” and still be little more than a rules engine wrapped in a buzzword.

Another common error is confusing complexity with robustness. More indicators, more parameters, and more filters do not automatically improve a strategy. In fact, they often increase the chance of overfitting. If you want a useful companion to this section, read our backtest checklist and our systematic vs. discretionary explainer. Both help separate process from performance.

The honest assessment is that most retail investors should prefer simple, transparent rules over opaque black boxes. A strategy that is easy to explain is not automatically good, but a strategy that cannot be explained is usually hard to trust.

7) How AIBROKER approaches verification: audit trail and point-in-time universes

When AIBROKER references strategy research or backtest outputs, we link to our methodology page because the process matters as much as the result. Two controls are especially important: a cryptographic audit trail and point-in-time universe construction. The first is designed to preserve a tamper-evident record of research inputs, outputs, and revisions. The second ensures that the strategy is evaluated only on the securities that were actually available at each historical point, rather than on today’s cleaned-up list of survivors.

That matters because many bot vendors implicitly use hindsight. They test on a current universe, then describe the result as if the same set of assets had always been available. It is a subtle but serious error. If you want to understand the broader market-mechanics context, our how stock prices are set article and bid-ask spread guide are useful complements.

Table 5. Verification controls: vendor claim vs. auditable process
ProblemTypical vendor approachAIBROKER-style control
Track record tamperingStatic screenshots or PDFsCryptographic audit trail with logged revisions
Universe driftCurrent list applied to past dataPoint-in-time universe construction
Hidden re-optimizationUndisclosed parameter changesVersioned research artifacts and change history
Reproducibility“Trust us”Documented methodology and exportable inputs

8) A decision tree for subscribers

Use this quick decision tree before paying for any bot or signal service.

Decision tree: should you subscribe?
StepQuestionIf yesIf no
1Is the performance claim clearly labeled as backtest, forward test, or live?ContinueStop
2Can you inspect fees, slippage, and turnover assumptions?ContinueStop
3Is the universe point-in-time and reproducible?ContinueStop
4Can the trade log be independently verified?ContinueStop
5Does the strategy still make sense after costs?Maybe subscribeDo not subscribe

This is intentionally strict. The burden of proof should sit with the seller, not the buyer.

9) A worked example: reading a bot sales page like an analyst

Suppose a vendor advertises a bot with the following claims: 92% win rate, 38% annual return, “AI-driven,” and “verified.” A retail investor might stop there. An analyst should ask: verified by whom? Over what period? On what universe? With what costs? What was the maximum drawdown? How many trades? Was the result concentrated in one regime? Did the strategy survive a different volatility environment?

Now compare that to a more credible disclosure: the vendor publishes a dated live track record, a downloadable trade log, a fee schedule, a slippage assumption, and a description of the asset universe. That still does not guarantee success, but it gives you something to interrogate. If the strategy also has a documented research process and a clear failure mode, you are finally evaluating a product rather than a promise.

For readers who want a broader framework for judging risk-adjusted outcomes, our Sharpe vs. Calmar article and risk measurement guide are good next steps.

10) The final checklist before you pay

Here is the short version. If you remember nothing else, remember this: a trading bot should be judged like a research claim, not a consumer gadget. Ask for the raw evidence. Ask for the assumptions. Ask for the failure cases. Ask whether the result survives costs and a different market regime. And if the seller cannot answer those questions, walk away.

Practical takeaway: the best subscription decision is often the one you do not make. In automated trading, the absence of verifiable evidence is itself evidence.

Markets reward patience more reliably than novelty. A bot that cannot survive scrutiny is not a shortcut to alpha; it is a shortcut to disappointment.

Sources & Further Reading

  1. Bailey, D. H., & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. SSRN. Source
  2. López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.
  3. U.S. Securities and Exchange Commission. Investor Bulletin: Automated Trading Systems and Trading Software. Source
  4. U.S. Securities and Exchange Commission. Investor Alert: Be Cautious of Claims About Automated Trading Systems. Source
  5. U.S. Securities and Exchange Commission. Investor Bulletin: Copy Trading and Social Trading. Source
  6. Fama, E. F., & French, K. R. (1993). Common risk factors in the returns on stocks and bonds. Journal of Financial Economics, 33(1), 3–56. Source