Record hypotheses, execution, risk, and net outcomes so the journal can distinguish repeatable evidence from luck—without optimizing stops or size on a small sample.
Entry and exit prices show P&L but not cause. Before an order, record a stable setup version, horizon, thesis, evidence, invalidation condition, planned entry and order type, stop logic, size, maximum planned loss, and the market information available at that timestamp. Afterward, reconcile fills, partial fills, fees, spread or slippage, financing or borrow, corporate actions, maximum adverse and favorable excursion, exit reason, and rule adherence. Classify a trade as planned or impulsive before its outcome is known; a compliant loss and an accidental win belong in different populations. Preserve timezone, currency, data version, and an immutable link to the order record so hindsight cannot rewrite the setup. A journal does not prove edge in one trade. It creates a point-in-time dataset for testing whether eligibility, risk, and execution produce a repeatable net distribution. What actually matters is whether another reviewer can reproduce the classification without reading the P&L first. [1][2]
Table 1. Point-in-time trade record| Field | When fixed | Validation | Question |
|---|
| Thesis and setup | Before order | Version and timestamp | Was it eligible? |
| Stop and size | Before order | Risk formula | Was loss planned? |
| Fills and costs | After execution | Broker reconcile | Was it executable? |
| MAE and MFE | After exit | Price convention | What path occurred? |
| Exit and adherence | After exit | Fixed taxonomy | Was policy followed? |
Classification
Label adherence before seeing P&L or the outcome contaminates the evidence.
Expectancy, win rate, payoff ratio, MAE, MFE, and slippage answer different questions, but none should be read alone. Compute expectancy net of commission, spread, slippage, exchange fees, borrow, financing, and any tax assumption relevant to the decision. Express results in dollars and in R, the return divided by planned initial risk, so varying position sizes do not dominate. Worked example: with a 45% win rate, $220 average net winner, and $140 average net loser, expectancy is 0.45 × 220 − 0.55 × 140 = $22. If the average loser rises to $180, it becomes 0.45 × 220 − 0.55 × 180 = $0—not −$0.50. Even $22 is an estimate, not proof of edge. Report trade count, median, dispersion, downside tail, drawdown, turnover, and a confidence or bootstrap interval that respects clustering by date, asset, and regime. State whether slippage is measured against arrival price, midpoint, decision price, or another benchmark. A few positive dollars can be sampling noise or uncapturable after capacity. Publish the denominator for every metric and keep canceled, rejected, partially filled, and zero-position eligible signals in separate documented populations. Excluding inconvenient attempts can inflate fill quality or setup performance even when every retained arithmetic result is correct. [1]
Table 2. Metrics with controls| Metric | Definition | Required companion | Failure |
|---|
| Expectancy | Mean net result | Interval and clustering | Noise called edge |
| Win and payoff | Frequency and magnitude | Read together | High win-rate trap |
| MAE and MFE | Path excursions | Executable prices | Look-ahead stop |
| Slippage | Fill minus benchmark | Named benchmark | Incomparable fills |
| Turnover | Capital traded | Cost and capacity | Activity hidden |
Arithmetic
At a $180 average loss, the stated 45%/$220 example has exactly zero expectancy.
The minimum template contains trade ID; timestamp and timezone; instrument or contract; strategy and setup version; regime defined before entry; thesis and invalidation; planned entry, order, stop, and target logic; actual fills and quantities; position size and planned risk in currency and R; all costs; MAE and MFE from an executable-price convention; exit reason from a fixed taxonomy; adherence flag; and links to market and broker records. Store partial orders rather than averaging away execution detail. Validate required fields, data types, sign conventions, currency conversion, splits, dividends, and reconciliation to the account statement. Emotional notes can supplement the record but cannot substitute for facts. Use a data dictionary and reject unknown setup labels instead of allowing spelling variants to fragment the sample. A lightweight form is better than an ornate sheet only when it remains complete and accurate. Before adding a field, name the decision it will change, the quality check it needs, and who maintains it. The checklist is a control system, not a diary layout. Treat the journal as sensitive financial data. Store no brokerage password, API secret, full account number, tax identifier, or unnecessary personal note. Use least-privilege access, encryption appropriate to the storage system, versioned backups, and a tested restore. Separate analytical trade IDs from broker account identifiers, log exports, and define retention and deletion rules. If screenshots or order confirmations are attached, redact credentials and unrelated balances. A corrupted or selectively missing journal can bias analysis as surely as a bad formula, so report missing records and reconciliation exceptions rather than silently dropping them.
Table 3. Validated lightweight template| Column | Example | Control |
|---|
| Timestamp | 2026-06-18 10:14 ET | Timezone |
| Setup | breakout-v3 | Fixed list |
| Plan | entry / stop / target logic | Before order |
| Fill | Price and quantity | Statement |
| Risk | Currency and R | Formula |
| Exit | Rule / stop / discretionary | Taxonomy |
Data quality
A field without validation adds confidence faster than information.
There is no universal weekly-execution, monthly-strategy, quarterly-structure cadence. Set cadence by trade frequency, holding period, risk of waiting, and number of observations required for the intended decision. Execution incidents—missed stops, broken feeds, rejected orders, abnormal slippage, or a breached loss limit—may require immediate review. Aggregating fills can occur daily or weekly. Changing a setup, stop, or sizing method requires a defined sample, out-of-sample evidence, approval, and rollback; calendar month-end alone supplies none of those. Separate continuous risk monitoring from research governance. A dashboard should show net expectancy with uncertainty, payoff and win rate, cost components, turnover, adherence, composition by setup and regime, and missing-data rates. Define regimes before evaluation so a losing period cannot be relabeled out of the sample. More frequent review can increase discretion and multiple testing; less frequent review can leave execution failures uncontrolled. Choose the cheapest cadence that detects the relevant failure within its risk tolerance, and version every rule change. [2]
Table 4. Review governance| Process | Cadence driver | Output | Cannot do |
|---|
| Risk monitoring | Loss urgency | Alert or intervention | Rewrite edge |
| Execution review | Fill count and incident | Cost diagnosis | Change signal |
| Strategy research | Effective sample | Test proposal | Auto-deploy |
| Structural approval | Material evidence | Version and rollback | Erase history |
Cadence
Risk monitoring and strategy modification are different processes.
MAE describes the worst observed adverse excursion and MFE the best favorable excursion under a stated price convention. They can reveal that winning trades often need room or that exits capture little of observed favorable movement, but they do not identify an optimal stop or target. Intraday highs and lows may not be executable; bar data can hide ordering; gaps jump past stops; stopped trades disappear from a winners-only sample; and changing a stop changes the win rate, payoff, holding time, exposure, and subsequent path. Do not infer that an 0.8% stop is correct because 80% of historical winners stayed above it. Segment by setup without slicing until a flattering result appears, predefine the candidate rule, include every eligible trade, and test with walk-forward or held-out data plus realistic gaps and costs. Position size should reflect a plausible loss distribution, portfolio correlation, liquidity, and risk budget—not average MAE alone. Ten, fifty, or one hundred trades does not guarantee power; effective sample size depends on dispersion and dependence. The real risk is turning a diagnostic into an automatic order parameter.
MAE/MFE
Observed excursions generate hypotheses; they do not choose stops.
Revenge trading can appear as larger size, shorter delay, and more trades after a loss; premature exits as a persistent gap between executable MFE and realized return; stop drift as actual loss beyond the approved boundary. Define each label and its comparison group before inspecting results. A run of losses can occur by chance, and a large MFE does not prove a wider target was fillable. Compare planned and impulsive trades net of costs, controlling for setup, liquidity, time of day, asset, and predefined regime. Correct for repeated searches or reserve a confirmation sample when many slices are examined. Case study: Trader A completes 60 trades with a 38% win rate, winners 2.8 times losers, and slippage below 0.1%; Trader B wins 64%, but winners are 0.7 times losers and slippage is 0.3%. A can have higher expectancy, yet the conclusion still needs loss magnitude, cost definition, clustering, and uncertainty. Evidence on individual investors associates heavier trading with weaker results, which makes turnover a required field rather than proof that every active strategy fails. [10][11]
Multiple tests
A flattering slice needs a held-out confirmation sample.
Translate one finding into one falsifiable hypothesis. Example: “Using a passive limit price for setup X will reduce arrival-price slippage without lowering the eligible fill rate below the preregistered tolerance.” Specify population, benchmark, net metric, minimum information requirement, test window, risk limit, approval owner, and rollback before observing the new results. Test in replay, paper trading, or limited capital according to risk; do not change entry, stop, size, and order type together. Percentages such as cutting size 25%, allowing slippage equal to 10% of average win, or setting an 0.8% stop are illustrations until the strategy's evidence validates them. Compare before and after with intervals and market composition, not just total P&L. A change that affects real capital needs independent governance and must not be deployed automatically from the journal. The journal proposes research; it does not mutate the strategy. Record the decision whether the test passes, fails, or remains inconclusive, retain the old version for rollback, and monitor post-change drift. FINRA's day-trading material reinforces that frequent trading can carry substantial risk; records must include losses and costs, not winning percentages alone. [2][12]
Governance
One preregistered finding should produce one controlled change.
Rollback
The journal proposes research and preserves the prior strategy version.
Related analysis
trading journalperformance reviewexpectancyexecutionrisk control
Sources & Further Reading
- 1. Van Tharp, Trade Your Way to Financial Freedom (expectancy framework and position sizing concepts). Source
- 2. U.S. Securities and Exchange Commission, Investor Bulletin: Day Trading: Your Dollars at Risk. Source
- 9. CFI, Expectancy Formula.
- Barber, B. M., & Odean, T. (2000). Trading is hazardous to your wealth. The Journal of Finance, 55(2), 773–806. Source
- Odean, T. (1999). Do investors trade too much? American Economic Review, 89(5), 1279–1298. Source
- FINRA. (2026). Day trading. Source