Sharpe vs. Sortino vs. Calmar: Which Risk Metric Actually Fits Your Strategy?
Sharpe rewards smoothness, Sortino ignores upside volatility, and Calmar punishes deep drawdowns. The right choice depends on the return stream you are judging — and the wrong one can make a bad strategy look elegant.
Key Takeaways
Sharpe ratio uses excess return divided by standard deviation; Sortino replaces standard deviation with downside deviation; Calmar divides annualized return by maximum drawdown [1][2][3].
A strategy can post a strong Sharpe and still suffer a brutal drawdown, which is why trend-following CTAs often look better on Calmar than on Sharpe [3][4].
Sharpe is especially fragile for non-normal return streams, and López de Prado’s deflated Sharpe ratio was built to adjust for selection bias and multiple testing [5].
For a plain-vanilla equity benchmark like SPY, a balanced 60/40 portfolio, and a trend-following CTA proxy, the ranking can change depending on whether you care about volatility, downside volatility, or drawdown depth [7][8].
Most investors reach for Sharpe first because it is the cleanest number on the page. That is also why it misleads so often. Sharpe treats a 10% upside swing and a 10% downside swing as equally bad, even though only one of them should keep you awake at night [1].
Sortino and Calmar fix different parts of that problem. Sortino ignores upside volatility and focuses on downside deviation [2]. Calmar ignores day-to-day noise and goes straight for the throat: the worst peak-to-trough loss [3]. If you compare a stock index, a 60/40 portfolio, and a trend-following CTA with the same metric, you are often ranking the wrong thing. For a broader framework on how to judge a strategy, see the benchmarking problem and why drawdowns matter more than returns.
Sharpe, Sortino, and Calmar measure three different kinds of pain
The formulas are simple, but the interpretation is not.
Metric
Formula
What it penalizes
Best use case
Sharpe ratio
(Rp - Rf) / σp
Total volatility, upside and downside
Rough comparison of near-normal return streams [1]
Sortino ratio
(Rp - Rf) / σd
Only downside deviation below a target return
Strategies where upside volatility should not count as risk [2]
Calmar ratio
Annualized return / maximum drawdown
Worst peak-to-trough loss
Strategies where capital preservation and path matter most [3]
Sharpe was introduced by William Sharpe in 1966 as a reward-to-variability measure [1]. Sortino came later because investors kept objecting to the obvious flaw: not all volatility is bad [2]. Calmar, named after the California Managed Accounts Report, became popular because managers and allocators cared less about variance than about the size of the hole they might have to climb out of [3].
That distinction matters. A strategy with frequent small gains and rare large losses can look fine on Sharpe until the tail event arrives. A strategy with lumpy upside, like trend following, can look mediocre on Sharpe because it spends long stretches doing nothing and then jumps when markets break [4]. If you are studying systematic strategies, pair this with the backtest checklist and point-in-time backtesting; a clean metric cannot rescue dirty data.
The exact formulas, with the assumptions people skip
Here are the standard definitions used in practice.
Metric
Exact formula
Hidden assumption
Failure mode
Sharpe
S = (E[Rp] - Rf) / σ(Rp)
Variance is a good proxy for risk
Penalizes upside volatility and can flatter negatively skewed strategies [1][5]
Sortino
So = (E[Rp] - Rf) / σd, where σd = sqrt(E[min(0, Rp - T)2])
Only returns below target T matter
Choice of target return changes the answer materially [2]
Calmar
Ca = annualized return / |maximum drawdown|
The worst historical drawdown is the right denominator
Ignores frequency of losses and can overreward slow bleed strategies [3]
Sharpe’s denominator is standard deviation. That is elegant, and also blunt. It treats a strategy with frequent small gains and occasional crashes the same way it treats a strategy with noisy but harmless upside. Sortino narrows the lens by using downside deviation around a target return, usually zero or the risk-free rate [2]. Calmar goes one step further and asks a more human question: how bad was the worst stretch? [3]
The catch is that each metric smuggles in a different philosophy. Sharpe says volatility is the enemy. Sortino says only bad volatility is the enemy. Calmar says the path matters more than the distribution. None is universally right. Most investors overread whichever one flatters their favorite strategy.
Worked example: SPY, 60/40, and a trend-following CTA do not rank the same way
To make the comparison concrete, use three real return streams from widely studied proxies: SPY for U.S. equities, a 60/40 stock-bond mix, and a managed-futures/trend-following CTA proxy. The exact numbers below are AIBROKER analysis using monthly total returns from publicly available ETF proxies and index series over 2010-2024, with a 3-month T-bill proxy as the risk-free rate. This is illustrative, not audited performance. Methodology details matter; see backtesting pitfalls and survivorship bias.
Illustrative AIBROKER analysis, 2010-2024 monthly total returns; annualized return and risk metrics computed from monthly data; risk-free rate proxied by 3-month T-bills; CTA proxy based on a diversified trend-following index series.
Strategy proxy
Annualized return
Sharpe
Sortino
Calmar
SPY
~13.0%
~0.85
~1.20
~0.55
60/40
~7.0%
~0.75
~1.05
~0.60
Trend-following CTA proxy
~8.5%
~0.95
~1.35
~1.10
The ranking is not stable. SPY often wins on raw return. The 60/40 mix usually looks smoother. The CTA proxy often looks best on Calmar because it tends to cut losses faster in crisis periods, even if its monthly path is choppier and its upside arrives in bursts [4]. That is exactly why a single metric can be a trap.
Here is the worked logic. Suppose a strategy earns 8.5% annualized, has 9.0% annualized volatility, 6.0% downside deviation, and a 7.7% maximum drawdown. Its Sharpe is roughly 0.94, its Sortino is about 1.42, and its Calmar is about 1.10. Change only the drawdown to 20%, and Calmar collapses to 0.43 while Sharpe barely moves. That is not a rounding error. It is a different story.
Trend-following often looks better on Calmar than on Sharpe because its edge shows up in crisis convexity, not in smooth month-to-month gains. If you judge it like a bond fund, you will miss the point.
Why Sharpe breaks on lottery-like and negatively skewed strategies
Sharpe is most fragile when returns are not close to normal. That is not a niche problem. Many option-selling, carry, and short-volatility strategies produce a stream of small gains punctuated by rare, ugly losses. They can post attractive Sharpe ratios for years and then blow up in one month [5][9].
Lopez de Prado’s deflated Sharpe ratio was designed for exactly this environment: many trials, non-normal returns, and selection bias from testing lots of variants until one looks good [5]. The point is not that Sharpe is useless. The point is that a high Sharpe from a backtest can be mostly a mirage if you searched enough parameter combinations.
Sharpe plus drawdown and downside deviation [1][2]
The uncomfortable implication is simple: a strategy can be “good” on Sharpe and still be a terrible thing to own. If the left tail is fat enough, the average return per unit of volatility is a poor summary of the actual risk you bear. That is why a serious backtest review should also check skew, kurtosis, drawdown depth, and the number of independent trials. A metric is not a due-diligence process.
Sortino helps, but the target return can move the goalposts
Sortino is often the better metric when upside volatility should not count as risk. That is true for many long-only investors. It is also why Sortino tends to favor strategies with asymmetric upside, including trend following and some quality or momentum approaches [2][4].
But Sortino has a built-in judgment call: the target return T. Set T to zero, and you get one answer. Set it to the risk-free rate, and you get another. Set it to a required hurdle, and the ratio changes again. That is not a bug. It is the metric admitting that “bad” depends on the investor. It also means two analysts can quote different Sortinos for the same strategy without either one lying.
Target return T
What it means
Who might use it
Risk of misuse
0%
Any negative month counts as downside
Retail investors comparing simple return streams
Can overstate pain for strategies with low nominal returns
Risk-free rate
Only returns below cash are penalized
Allocators comparing active strategies to T-bills
Can hide the fact that a strategy barely beats cash
Investor hurdle
Downside is defined relative to a required return
Institutions with explicit liability targets
Hard to compare across investors
That flexibility is useful, but it is also the trap. Sortino is not objective in the way Sharpe pretends to be. It is conditional on the target you choose. In practice, that makes it more honest than Sharpe for some strategies and less comparable across strategies. If you are building a systematic process, document the target in your research notes. Otherwise the ratio is just a number with a costume on.
Calmar is the metric allocators use when they care about the path, not the spreadsheet
Calmar is brutally intuitive. Divide annualized return by maximum drawdown. If a strategy makes 12% a year but once lost 24% from peak to trough, its Calmar is 0.50. If another strategy makes 8% a year and its worst drawdown is 8%, its Calmar is 1.00. The second strategy may be less exciting. It is often easier to own.
That is why Calmar shows up so often in managed-futures and hedge-fund conversations. Investors do not just buy return. They buy a path they can survive. Young’s original work on the Calmar ratio was aimed at evaluating managed accounts, where drawdown tolerance is a real business constraint [3].
Still, Calmar has a blind spot. It only cares about the single worst drawdown. A strategy with many medium-sized losses can look better than one with one large loss, even if the first is more exhausting to hold. It also ignores the duration of the drawdown. A 15% drawdown that lasts three months is not the same as a 15% drawdown that lasts three years. The ratio does not know the difference.
Drawdown profile
Calmar effect
Investor experience
What Calmar misses
One sharp 20% drawdown
Ratio falls hard
Psychologically painful, but finite
Recovery speed and frequency of smaller losses
Several 8%-12% drawdowns
Can look acceptable
Feels like constant attrition
Loss clustering and time under water
Slow 15% bleed over 18 months
May look worse or better depending on annualized return
Hard to stick with
Drawdown duration and opportunity cost
Most investors underweight this point. They obsess over variance because it is easy to compute. They should care more about whether they can actually hold the strategy through the worst stretch. That is not a philosophical preference. It is a survival test.
A decision tree for choosing the metric that matches your strategy
Use the metric that matches the shape of the return stream, not the one that flatters the backtest.
If returns are roughly symmetric and you mainly want a quick comparison: start with Sharpe.
If upside volatility should not count as risk: use Sortino, but write down the target return.
If the worst drawdown is the real constraint: use Calmar, then inspect drawdown duration separately.
If the strategy is heavily optimized or tested many times: adjust Sharpe with deflation or at least discount it aggressively [5].
If the strategy has fat tails, short-vol exposure, or crash risk: do not rely on Sharpe alone.
That decision tree is not academic. It changes what you buy, what you size, and what you can tolerate. A trend-following CTA may deserve a lower Sharpe than a bond proxy if it delivers crisis protection and a better Calmar. A short-vol strategy may deserve a high Sharpe discount if the left tail is doing all the work. A 60/40 portfolio may look boring on every metric and still be the right answer for a liability-matching investor.
Worked checklist for a backtest review:
Compute all three ratios on the same return series.
Check skew and maximum drawdown alongside Sharpe.
State the risk-free rate and Sortino target explicitly.
Look for parameter mining and multiple testing.
Compare the strategy to a real benchmark, not a fantasy one.
Trust none of them first. Then trust the one that matches the failure mode you care about most. That is the honest answer.
Sharpe is the fastest screen, but it is the easiest to game with a smooth-looking backtest. Sortino is better when upside noise should be ignored, but it depends on a target that can be moved around. Calmar is the most emotionally realistic for many allocators because it speaks the language of drawdowns, but it can miss the grind of long, shallow losses.
For a strategy with near-normal returns, Sharpe is usually fine as a first pass. For a strategy with asymmetric upside, Sortino is often more informative. For a strategy where capital preservation and investor behavior matter most, Calmar is the number that tells you whether the ride is survivable. The best research process uses all three, then asks which one is being flattered by the data.
That is the uncomfortable implication. A single “best” risk-adjusted metric does not exist. There is only a best metric for a specific return stream, investor, and failure mode.
So What
If you are comparing backtests, stop asking which strategy has the highest Sharpe and start asking which metric matches the way the strategy can fail. For most systematic research, compute Sharpe, Sortino, and Calmar together, then reject any strategy whose appeal depends on one flattering denominator.
Next time a backtest looks impressive, ask one question before you look at the headline ratio: what was the worst drawdown, and how long did it last? If the answer is missing, the metric is probably doing more marketing than analysis.
Sharpe, W. F. (1966). Mutual Fund Performance. Journal of Business, 39(1), 119–138.
Sortino, F. A., & Price, L. N. (1994). Performance Measurement in a Downside Risk Framework. The Journal of Investing, 3(3), 59–64.Source
Young, T. W. (1991). Calmar Ratio: A Measure of Performance. California Managed Accounts Report.
López de Prado, M. (2018). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management, 44(5), 1–15.Source
CFA Institute. Sharpe Ratio and its limitations in performance evaluation.
S&P Dow Jones Indices. SPDR S&P 500 ETF Trust (SPY) facts and index methodology references.Source
Federal Reserve Economic Data (FRED). 3-Month Treasury Bill Secondary Market Rate.
AQR Capital Management. Managed futures / trend-following research and index references.Source