Sharpe, Sortino, or Calmar: Match the Metric to the Failure Mode

Compute all three from the same point-in-time net return series, publish every convention and uncertainty, and reject comparisons that depend on an unreproducible proxy or flattering denominator.

Key takeaways
  • Sharpe divides average excess return by total return volatility. Sortino divides return above a stated minimum acceptable return by a defined downside deviation. Calmar relates an annualized return measure to the absolute maximum drawdown. [1][2][3] None measures risk in full. Sharpe treats positive and negative variation symmetrically and can miss an unobserved crash. Sortino selects shortfalls relative to a target under a denominator convention that must be disclosed. Calmar uses the worst peak-to-trough path observed in one sample. A high ratio can coexist with gaps, negative skew, illiquidity, low capacity, leverage, stale marks, or severe estimation error. Before choosing a metric, define the decision, investor objective, horizon, liabilities, liquidity need, implementation, and failure mode. The key point is that a ratio compresses a return history; it does not certify the strategy, forecast the next path, or replace due diligence. Use the three measures as different views of an audited return series, never as interchangeable scores whose largest value automatically wins.
  • Build every comparison from total returns in the same economic currency, frequency, point-in-time window, and fee and tax convention. Align the risk-free series or target with each observation date and publication lag. State whether Sharpe uses arithmetic mean excess returns and the sample standard deviation of those excess returns, how missing observations are handled, and how the ratio is annualized. Square-root-of-time scaling relies on assumptions that autocorrelation, smoothing, overlapping positions, and changing exposure can violate. For Sortino, state the minimum acceptable return per period and whether semideviation averages squared shortfalls over all periods or only shortfall periods; those definitions produce different denominators. For Calmar, state the return numerator—commonly CAGR or another annualized measure—the exact peak/trough and cash-flow treatment, sampling frequency, and window. Do not mix CAGR with an incompatible monthly-volatility convention or compare vendor ratios with unknown methods. Worked example: 12% divided by a 24% drawdown equals 0.50 arithmetically, but that calculation alone says nothing about sampling uncertainty, recovery, costs, or suitability.
  • The prior article quotes approximate 2010–2024 annualized return, Sharpe, Sortino, and Calmar values for SPY, a 60/40 portfolio, and an unspecified trend-following CTA proxy. It does not provide the complete return files and versions, CTA ticker or index, stock and bond proxies, portfolio weights and rebalancing dates, Treasury-bill series and availability lag, distribution treatment, currency, fees, slippage, taxes, missing-data rules, metric code, or output artifact. Those figures cannot be presented as reproducible AIBROKER analysis or used to claim that the CTA ranks best on Calmar. A valid comparison freezes data vintages, identifies investable proxies, includes delisted or changed products when relevant, implements 60/40 and the CTA exactly, executes after information availability, subtracts realistic costs, publishes code and configuration, and reconciles every metric to a downloadable output. Preserve the underlying wealth curves and drawdown episodes. Until that evidence exists, only the qualitative statement is defensible: rankings can change when the denominator and return shape change. A polished table of approximations is not a backtest.
  • Sharpe is most fragile when the sample is non-normal, dependent, smoothed, selected, or too short to include the damaging tail. Option selling, carry, short-volatility, credit, and other negatively skewed strategies can accumulate small gains before one rare loss changes the mean and standard deviation. Trend or convex strategies can have irregular positive payoffs that total volatility treats as risk. Smoothing or stale private marks reduces measured volatility without reducing economic risk; overlapping trades and serial correlation distort naive annualization. Searching many universes, parameters, or start dates and reporting the winner biases the ratio upward. A Deflated Sharpe Ratio attempts to address selection and non-normality under stated assumptions; it is not a generic haircut and does not repair look-ahead data. [4] Report skew, kurtosis, quantiles, expected shortfall, largest losses, convexity or option exposure, autocorrelation, number of trials, and an uncertainty interval. Stress scenarios outside the short history remain necessary. The real risk is interpreting a smooth sample as proof that the left tail does not exist.

Compute all three from the same point-in-time net return series, publish every convention and uncertainty, and reject comparisons that depend on an unreproducible proxy or flattering denominator.

Three ratios compress three different and incomplete risks

Sharpe divides average excess return by total return volatility. Sortino divides return above a stated minimum acceptable return by a defined downside deviation. Calmar relates an annualized return measure to the absolute maximum drawdown. [1][2][3] None measures risk in full. Sharpe treats positive and negative variation symmetrically and can miss an unobserved crash. Sortino selects shortfalls relative to a target under a denominator convention that must be disclosed. Calmar uses the worst peak-to-trough path observed in one sample. A high ratio can coexist with gaps, negative skew, illiquidity, low capacity, leverage, stale marks, or severe estimation error. Before choosing a metric, define the decision, investor objective, horizon, liabilities, liquidity need, implementation, and failure mode. The key point is that a ratio compresses a return history; it does not certify the strategy, forecast the next path, or replace due diligence. Use the three measures as different views of an audited return series, never as interchangeable scores whose largest value automatically wins.

Table 1. Metric scope
RatioNumeratorDenominatorDoes not show
SharpeMean excessVolatilityTail/path
SortinoReturn over targetDownside deviationWhole tail/upside
CalmarAnnualized returnMax drawdownFrequency/duration
PanelNet returnSeveral measuresFuture
Summary, not truth

Every ratio discards material dimensions of risk.

Comparable formulas require identical frequency, target, and return conventions

Build every comparison from total returns in the same economic currency, frequency, point-in-time window, and fee and tax convention. Align the risk-free series or target with each observation date and publication lag. State whether Sharpe uses arithmetic mean excess returns and the sample standard deviation of those excess returns, how missing observations are handled, and how the ratio is annualized. Square-root-of-time scaling relies on assumptions that autocorrelation, smoothing, overlapping positions, and changing exposure can violate. For Sortino, state the minimum acceptable return per period and whether semideviation averages squared shortfalls over all periods or only shortfall periods; those definitions produce different denominators. For Calmar, state the return numerator—commonly CAGR or another annualized measure—the exact peak/trough and cash-flow treatment, sampling frequency, and window. Do not mix CAGR with an incompatible monthly-volatility convention or compare vendor ratios with unknown methods. Worked example: 12% divided by a 24% drawdown equals 0.50 arithmetically, but that calculation alone says nothing about sampling uncertainty, recovery, costs, or suitability.

Table 2. Conventions
ChoiceDeclareErrorTest
FrequencyDaily/monthlyMixingRecalculate
Risk-freeSeries and lagFuture constantAlign
TargetPer-period MARPost-selectionSensitivity
DrawdownWindow/flowsHidden definitionReproduce
Convention

Frequency, target, risk-free series, and annualization travel with the number.

The quoted 2010–2024 SPY, 60/40, and CTA results are not reproducible

The prior article quotes approximate 2010–2024 annualized return, Sharpe, Sortino, and Calmar values for SPY, a 60/40 portfolio, and an unspecified trend-following CTA proxy. It does not provide the complete return files and versions, CTA ticker or index, stock and bond proxies, portfolio weights and rebalancing dates, Treasury-bill series and availability lag, distribution treatment, currency, fees, slippage, taxes, missing-data rules, metric code, or output artifact. Those figures cannot be presented as reproducible AIBROKER analysis or used to claim that the CTA ranks best on Calmar. A valid comparison freezes data vintages, identifies investable proxies, includes delisted or changed products when relevant, implements 60/40 and the CTA exactly, executes after information availability, subtracts realistic costs, publishes code and configuration, and reconciles every metric to a downloadable output. Preserve the underlying wealth curves and drawdown episodes. Until that evidence exists, only the qualitative statement is defensible: rankings can change when the denominator and return shape change. A polished table of approximations is not a backtest.

Table 3. Unsupported proxies
ProxyQuoted claimMissing evidenceStatus
SPY~13% and ratiosSeries/codeUnaudited
60/40~7% and ratiosWeights/rebalanceUnaudited
CTA~8.5% and ratiosIndex/tickerUnaudited
ComparisonRankingCosts/conventionsDo not publish
Failed gate

The 2010–2024 proxy table lacks enough data and code for reproduction.

Sharpe is fragile to tails, dependence, smoothing, and selection

Sharpe is most fragile when the sample is non-normal, dependent, smoothed, selected, or too short to include the damaging tail. Option selling, carry, short-volatility, credit, and other negatively skewed strategies can accumulate small gains before one rare loss changes the mean and standard deviation. Trend or convex strategies can have irregular positive payoffs that total volatility treats as risk. Smoothing or stale private marks reduces measured volatility without reducing economic risk; overlapping trades and serial correlation distort naive annualization. Searching many universes, parameters, or start dates and reporting the winner biases the ratio upward. A Deflated Sharpe Ratio attempts to address selection and non-normality under stated assumptions; it is not a generic haircut and does not repair look-ahead data. [4] Report skew, kurtosis, quantiles, expected shortfall, largest losses, convexity or option exposure, autocorrelation, number of trials, and an uncertainty interval. Stress scenarios outside the short history remain necessary. The real risk is interpreting a smooth sample as proof that the left tail does not exist.

Table 4. Sharpe failures
Return shapeRiskSharpe canCompanion
Short volRare crashOverstateTail stress
TrendIrregular payoffMis-summarizeSkew/Calmar
SmoothingAutocorrelationInflateAdjusted inference
Data miningSelectionInflateTrials/DSR
Tails

A high Sharpe does not protect against a rare loss absent from the sample.

Sortino changes when the target or semideviation convention changes

Sortino asks a useful but conditional question: how much return was earned relative to a chosen minimum acceptable return per unit of defined downside deviation? The target can be zero, a dated cash rate, inflation, a benchmark, or an investor hurdle; each answers a different question. Select it before seeing which choice flatters the strategy, align it to each period, and preserve historical vintages. Two analysts can also differ because one divides squared shortfalls by all observations while another divides only by shortfall observations. Publish the formula and count. When few returns fall below the target, the denominator can be tiny and the ratio enormous or undefined; that is low information, not automatic excellence. Upside volatility is excluded from semideviation but can accompany leverage, crowded trades, or future reversal and is not inherently harmless. Report the number, magnitude, clustering, and duration of shortfalls, plus sensitivity to predeclared targets. Compare the same target across strategies only when it represents the same objective. Sortino adds an investor-relevant threshold; it does not remove tail, liquidity, or estimation risk.

Table 5. Sortino target
TargetQuestionUseRisk
0%Negative period?NominalIgnores inflation
CashBelow cash?AlternativeWrong series
BenchmarkUnderperform?ActiveTracking
HurdleMeet objective?LiabilityFew shortfalls

Calmar uses one historical extreme and omits much of the path

Maximum drawdown is the largest decline from a prior wealth peak to a later trough within the sample. It changes with start and end dates, return frequency, total-return construction, external cash flows, leverage, valuation frequency, and window. Calmar therefore rests on one extreme estimate with high sampling uncertainty: the next drawdown can be larger, while a favorable short history can make the ratio look exceptional. Calmar does not show how many drawdowns occurred, time under water, recovery speed, clustered medium losses, gaps, liquidity, or losses that did not become the maximum. A sharp 20% decline and slow recovery can feel different from repeated 10% losses or an eighteen-month 15% bleed even when one summary ranks them similarly. Report the wealth curve, drawdown curve, depth distribution, durations, recovery periods, worst days or weeks, and adverse scenarios. An Ulcer Index or other path statistic can complement the panel when relevant. Do not infer that higher Calmar is easier for every investor to hold without connecting the path to liabilities, liquidity, behavior, and capital constraints.

Table 6. Calmar omissions
ProfileCalmar seesCalmar omitsReport
Sharp fallDepthGap/recoveryWorst days
Repeated lossesOnly worstFrequencyDistribution
Slow bleedDepthTimeUnder water
Short sampleObserved extremeFuture tailUncertainty

Choose metrics by failure mode and publish the complete risk panel

For reasonably liquid and approximately symmetric return streams, Sharpe is an initial summary. When a minimum return or liability matters, Sortino adds shortfall information. When peak-to-trough capital loss constrains the investor, Calmar highlights one path extreme. In every case, calculate all three on the identical net series and add annualized return, volatility, skew, tails, expected shortfall, maximum drawdown, drawdown duration, recovery, turnover, capacity, leverage, liquidity, and exposure by regime. Compare with an implementable benchmark and cash alternative. Adjust inference for autocorrelation and multiple research attempts; preserve every tested variant. Do not select the metric, target, window, or annualization after seeing which recommendation results. Decision tree: identify the loss that threatens the objective, select the statistic that exposes it, then require companion statistics for what that measure omits. If a plausible convention reverses the recommendation, report fragility and reduce confidence or size rather than declaring one convention correct.

Trust data, uncertainty, and adverse scenarios before any ratio

The correct order is to validate data, point-in-time availability, benchmark, and execution; examine the full distribution and path; quantify uncertainty; stress the omitted risks; and only then use ratios as compressed language. Use block bootstrap or another dependence-aware method when justified, predefined subperiods, and genuinely held-out evidence. Resampling cannot invent crises or structural changes absent from the data. A strategy should not be rejected because one ratio is low or accepted because all are high. Ask which loss threatens the objective and whether the position remains financed, liquid, and diversifying in that adverse case. Checklist: version the return file and code; reconcile fees and cash flows; declare risk-free rate, target, denominator, annualization, window, and drawdown method; count research trials; publish curves and intervals; stress tails and liquidity; and record decision authority. Practical takeaway: the panel supports comparison, while objectives, evidence, uncertainty, and governance determine whether capital is committed.

Related analysis

Sharpe RatioSortino RatioCalmar RatioRisk-Adjusted Return

Sources & Further Reading

  1. Sharpe, W. F. (1966). Mutual Fund Performance. Journal of Business, 39(1), 119–138.
  2. Sortino, F. A., & Price, L. N. (1994). Performance Measurement in a Downside Risk Framework. The Journal of Investing, 3(3), 59–64. Source
  3. Young, T. W. (1991). Calmar Ratio: A Measure of Performance. California Managed Accounts Report.
  4. López de Prado, M. (2018). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management, 44(5), 1–15. Source
  5. CFA Institute. Sharpe Ratio and its limitations in performance evaluation.
  6. S&P Dow Jones Indices. SPDR S&P 500 ETF Trust (SPY) facts and index methodology references. Source
  7. Federal Reserve Economic Data (FRED). 3-Month Treasury Bill Secondary Market Rate.
  8. AQR Capital Management. Managed futures / trend-following research and index references. Source