←Back to Blog

The Bond Market's Bad Week Was an Episode, Not a Day: Why VaR Can't See It

The 10-year Treasury yield reached 5.23% this week, its highest since 2007, through a connected run of bad days rather than one shock. Value-at-Risk and Expected Shortfall cannot tell that run apart from the same losses scattered across a quarter. A new working paper proposes Tail Episode Risk (TER) and Conditional TER (CTER) to measure the worst connected run of extreme losses. Tested on 26 years of data across six markets, it finds that Treasury episodes are the most predictable, and that a richer forecasting model still did not beat a simple EWMA volatility baseline.

By Niraj Neupane·Quantitative Researcher · Contributing Author·
Quantitative ResearchTreasury MarketBond MarketInterest RatesValue at RiskExpected ShortfallTail RiskRisk ManagementExtreme Value TheoryModel ValidationFederal ReserveMarket RiskSSRNSignalswall street
Share:

Disclaimer: This article is for educational purposes only. It summarises working-paper research that has not been peer reviewed. Trading and investments are subject to volatility, market risks, and macroeconomic shifts that could lead to partial or total loss of capital. Nothing here is investment advice or a recommendation to use any model in a live risk or trading system. Please consult your own risk, compliance and investment professionals before acting on any of it.


Financial Gurkha | Research Note | September 26, 2026

By Niraj Neupane, CA (ICAI), Quantitative Researcher. This article applies the author's own working paper, Conditional Tail Episode Risk, to this week's Treasury selloff. The paper is available in full on SSRN.


A trading floor at dusk with a screen showing the U.S. 10-year Treasury yield climbing through 5%

On Wednesday, September 23, the U.S. Treasury market had its worst day in a year and a half. The 10-year yield rose more than 13 basis points to about 5.10%, its biggest one-day move in nearly 18 months and its highest level since July 2007. By Friday it had reached 5.23%.

Most coverage focused on that single day, and in most risk systems that single day is what gets measured. A 99% one-day Value-at-Risk number is recalculated, a limit is checked, and the day is marked as a breach or not.

That framing misses how this week actually unfolded. The damage came from a sequence: a strong growth report, a weak auction, rising rate-hike odds, another soft auction, and higher yields again. Each step fed the next. In risk terms it was an episode, a connected run of extreme losses. Our standard risk measures are built to size individual days, and they cannot tell an episode apart from the same bad days scattered across a quarter.

SSRN Working Paper, September 2026 · Abstract 7502499

Conditional Tail Episode Risk: A Path-Dependent Framework for Extreme-Loss Episodes Beyond Value-at-Risk and Expected Shortfall

Niraj Neupane

Proposes Tail Episode Risk (TER), the upper-tail quantile of the worst connected run of extreme losses over a finite horizon, and Conditional TER (CTER), which splits episode risk into the probability of entering an episode and its conditional severity. Tested by simulation, constructed stress scenarios, and 2000 to 2026 data on the S&P 500, Nasdaq, long-duration Treasuries, EUR/USD, USD/JPY and Bitcoin.

Download the paperFree · Opens on SSRN

The 60-Second Version#↑ Contents

WhatNumber
10-year Treasury yield, Friday5.23%, highest since 2007
Wednesday's move+13 bp, largest in ~18 months
Rise since the start of 2026+108 bp (from 4.15%)
Estimated loss on a par 10-year since Januaryabout 8% of price
Episode risk when losses cluster (simulation)~3× higher, with VaR/ES flat
Best episode-forecasting market of six testedTreasuries (AUC 0.64 at 99%)
Did the richer CTER model beat simple EWMA?No

The seven things that matter:

  1. This week was a run, not a shock. A hot services PMI, a weak 5-year auction, a 13 bp jump, a soft 7-year auction and a 5.23% close all fed each other over five sessions.
  2. VaR and Expected Shortfall ignore the order of losses. Three bad days in a row and three bad days spread over a quarter produce the same VaR and the same ES.
  3. Tail Episode Risk (TER) measures the run. It groups consecutive extreme-loss days into episodes, sums how far each day exceeded the threshold, and takes the tail of the worst episode over a 10-day horizon.
  4. Clustering can triple episode risk while VaR and ES stay flat. In simulation, 99% TER rose from 1.51 to 4.54 as serial dependence rose. VaR and ES barely moved.
  5. The two views can disagree on which scenario is riskier. A sustained liquidity collapse scored 12.00 on episode severity against 3.50 for scattered spikes that VaR and ES ranked higher.
  6. Bond-market episodes are the most predictable of six markets tested. Long-duration Treasuries held an AUC of 0.64 to 0.67 even in the deep tail. Equities, currencies and Bitcoin faded toward a coin flip or below.
  7. The honest result: more features did not help. The full CTER conditioning set forecast episode severity worse than a simple EWMA volatility model at the 95% and 97.5% thresholds, and no better at 99%.

What's in this report#


1. The Week in Numbers#↑ Contents

Market figures below are as reported by CNBC, CNN Business, NBC News and others for the week of September 21 to 25, 2026. Full links are in Methodology and Sources.

IndicatorLevel / moveWhy it matters
10-year Treasury yield5.23% (Friday), up from 4.15% at the start of 2026Highest since 2007; the benchmark for mortgages and corporate borrowing
Wednesday's 10-year move+13 bp to ~5.10%Largest one-day jump in ~18 months
2-year Treasury yield~4.87% to 4.90%Highest since 2024; signals expected Fed tightening
30-year Treasury yield~5.4%Highest since before the 2008 financial crisis
Fed funds target3.75% to 4.00%First hike in three years, on September 16
October hike odds (CME FedWatch )~64%Markets expect more tightening
S&P Global services PMI 58.7 (from 56.5)Strongest in nearly five years
5-year and 7-year auctions Weak / soft demandSupply pressure on the curve
WTI–10yr correlation (1-month)0.96 (BMO)Oil and rates moving almost in lockstep
30-year mortgage rateabove 7.2% in daily lender surveysA one-year high for household borrowing costs
Michigan 1-yr inflation expectations4.6% (from 4.0%)Highest since June
S&P 500 (Friday close)7,743.41Winning week after a three-day losing streak

1.1 How the week's losses fed on each other#

Read in order, the week was a chain, and every link raised the odds of the next:

  1. Hot growth data. The S&P Global services PMI printed 58.7, its strongest in nearly five years.
  2. Rate-hike odds rise. Markets priced roughly a 64% chance of another Fed hike in October.
  3. A weak 5-year auction showed buyers demanding more yield to absorb new supply.
  4. The 10-year jumps 13 bp on Wednesday, the largest one-day move in about 18 months.
  5. A soft 7-year auction the next day added pressure rather than relief.
  6. The 10-year reaches 5.23% on Friday.

Running underneath all six steps was oil. With the Iran war keeping crude elevated, the one-month correlation between WTI and the 10-year yield reached 0.96, according to BMO. Higher oil fed inflation expectations, which fed hike odds, which fed yields. The last step then becomes the first step of whatever comes next.

1.2 What that means for a bond portfolio#

A 10-year Treasury near a 5.2% yield has a modified duration of about 7.7. For a par bond paying semiannually, modified duration is (1 − 1.026⁻²⁰) ÷ 0.052 ≈ 7.72.

MoveYield changeApprox. price impactLoss on $10B of 10-yr-equivalent
Start of 2026 → now≈ +108 bp≈ −8%≈ $800M
Early September → now≈ +45 bp≈ −3.4%≈ $340M

The arithmetic: −7.72 × 1.08% = −8.3%, plus a small convexity offset of about +0.4%, gives roughly −8%. For the 45 bp move, −7.72 × 0.45% = −3.5%, plus about +0.1% convexity, gives roughly −3.4%. These are first-order illustrative estimates, not reported figures.

Equities held up surprisingly well. The S&P 500 recovered from a three-day losing streak to finish the week higher. That resilience is why this story is easy to miss. Nothing crashed, but pressure built up over several days.


2. The Blind Spot in VaR and Expected Shortfall#↑ Contents

VaR and Expected Shortfall measure how bad individual days are, not how bad days line up. Both are calculated from the marginal distribution of losses, the set of daily outcomes without regard to the order they arrive in.

The two are the main tools of market risk measurement. VaR gives the loss threshold you are unlikely to exceed at a chosen confidence level. Expected Shortfall averages the losses beyond that threshold. Under Basel's Fundamental Review of the Trading Book , ES is now the core measure of regulatory market-risk capital.

A simple example shows the problem. Take two six-day loss paths with a tail threshold of 2:

Day123456VaR / ESWorst-episode severity
Path A (dispersed)404040Identical2
Path B (clustered)444000Identical6

Both paths contain three losses of 4 and three zeros, so their VaR and ES are identical. Path A has three separate one-day episodes, each with an excess of 4 − 2 = 2. Path B has one continuous episode with a cumulative excess of 2 + 2 + 2 = 6, three times larger.

Bar chart comparing two six-day loss paths with identical VaR and ES: the dispersed path has a worst episode severity of 2, the clustered path 6

This matters because the damage that breaks institutions is usually sequential:

  • Funding and margin stress builds with consecutive losses, since collateral calls compound before assets can be sold.
  • Forced selling feeds on itself when each day's loss triggers the next day's liquidation.
  • Depositor and investor confidence reacts to how long bad news lasts, not only to one headline.

Silicon Valley Bank in 2023 is the clearest recent example. The bank did not fail because of one bad bond-market day. It failed because a long run of rising rates built up unrealized losses on its securities portfolio faster than any daily risk number showed, until depositors noticed.


3. What Is Tail Episode Risk (TER)?#↑ Contents

Tail Episode Risk (TER) is the upper-tail quantile of the worst connected run of extreme losses over a fixed horizon. Where VaR asks how large one bad day can be, TER asks how severe the worst unbroken streak of bad days can become. It was proposed in the working paper Conditional Tail Episode Risk (SSRN, September 2026).

3.1 How TER is built, in three steps#

  1. Set a tail threshold. Any day whose loss exceeds a high quantile, such as the 95th percentile, counts as extreme.
  2. Group extreme days into episodes. An episode is an unbroken run of consecutive extreme days. Its severity is the sum of how far each day in the run exceeded the threshold.
  3. Take the worst episode, then its tail. Over a horizon (10 trading days in the study), find the most severe episode, called S*. TER is the upper-tail quantile of that worst-episode severity.

In formula terms, with daily losses L, threshold q at level α, tail level β, horizon H and information at time t:

TER  = Q_β [ max over episodes j of  Σ (L_t+k − q_t,α)⁺  for k in j  |  information at t ]

CTER = P( an episode occurs in the next H days | X_t )      <- how likely?
     × ES_β( S*_t,H | an episode occurs, X_t, Z_t )          <- how bad?

3.2 Conditional TER: how likely, and how bad?#

Conditional Tail Episode Risk (CTER) splits episode risk into the probability of entering an extreme-loss episode and the expected severity of that episode if it happens, both conditioned on current market conditions. Those are the two questions a risk committee actually asks:

How likely are we to enter an extreme-loss episode, given current conditions? > If we do, how bad could it get?

A risk report that today says "VaR 5%, ES 7%" could add an episode probability and an episode severity next to them.

3.3 How TER compares with VaR, ES and drawdown#

MeasureQuestion it answersDepends on order of losses?Sensitive to clustering?Status
VaRHow large is the loss threshold?NoNoExisting
ESHow severe is the marginal tail?NoNoExisting (Basel FRTB)
Aggregate excessHow much total exceedance occurs in an extreme event?Event-levelYesPrior art (Anderson, 1994)
Conditional Expected Drawdown How deep can a peak-to-trough decline get?YesYesRelated
TERHow severe can the worst connected run of extreme losses become?YesYesProposed
CTERGiven today's conditions, how likely and how severe is that run?YesYesProposed

What the paper does not claim. Summing threshold exceedances is established extreme value theory , and the mathematics of extreme clusters is well developed. The contribution is turning the worst episode over a finite horizon into a conditional forecast target for financial risk, splitting it into probability and severity, and testing it with modern forecast-evaluation standards. The paper also proves TER is not a disguised drawdown measure: the two can rank the same paths in opposite order.


4. What the Research Found, Good and Bad#↑ Contents

The framework was tested three ways: controlled simulation, deliberately constructed stress scenarios, and 26 years of real data (2000 to 2026) across six markets: the S&P 500, Nasdaq, long-duration Treasuries (TLT ), EUR/USD, USD/JPY and Bitcoin.

#QuestionResultVerdict
1Does clustering raise episode risk when VaR/ES stay flat?99% TER rises ~3× while VaR/ES barely moveYes
2Can VaR/ES and TER disagree on which scenario is riskier?Liquidity collapse: VaR 3.5 but S* 12.0Yes
3Does episode severity identify real crises?Flags 2020, 2008, 2025, 2011, 2000–02Yes
4Can episodes be predicted in advance?Weak overall; bonds are the exception (AUC 0.64 to 0.67)Mixed
5Does CTER's richer feature set beat a simple volatility model?No. Worse at 95% and 97.5%, no difference at 99%No

4.1 Clustering can triple episode risk while VaR and ES stay flat#

The paper simulated return series with the same marginal distribution and increased only the serial dependence ρ.

Serial dependence ρVaR 95%ES 95%TER 95%TER 99%
0.01.6332.0590.9811.508
0.31.6452.0531.0101.737
0.71.6552.0461.0652.580
0.91.6682.0931.0604.541

Line chart: as serial dependence rises from 0 to 0.9, 99% TER climbs from 1.51 to 4.54 while 95% VaR and ES stay flat

VaR and ES barely moved, while 99% TER rose 3.0× (4.541 ÷ 1.508). With fat-tailed, variance-standardized Student-t innovations, the pattern held: 99% TER rose from 3.1 to 5.3.

Takeaway: loss clustering, the defining feature of a week like this one, can sharply increase episode risk while the headline risk numbers stay the same.

4.2 VaR/ES and TER can rank the same scenarios in opposite order#

The paper constructed six 60-day stress paths with a common, fixed threshold (q = 2.5).

ScenarioVaR 95%ES 95%Worst-episode severity S*Which looks riskier?
Isolated crash1.424.006.50
Tail thickening (isolated spikes)6.006.003.50VaR/ES say this one
Liquidity collapse (sustained run)3.503.5012.00TER says this one
Volatility explosion (cluster)6.427.6825.62Both agree
Broad contagion (all elevated)3.394.012.40Neither captures it well
Escalating deterioration3.584.527.80

Paired bar chart of six stress scenarios comparing 95% Expected Shortfall with worst-episode severity; liquidity collapse scores 12.00 on episode severity against 3.50 on ES

VaR and ES say the scattered-spikes scenario is riskier. TER says the sustained liquidity collapse is 3.4 times worse by episode severity (12.00 ÷ 3.50). Neither is simply correct. They capture different dimensions of stress, and a risk system tracking only one misses the other.

4.3 The measure identifies real crises#

Applied to realized S&P 500 losses from 2000 to 2026 (10-day horizon), the most severe distinct episodes line up with the periods practitioners remember as crises:

RankEpisodeContext
1March 2020COVID crash (10-day worst-episode severity 11.3%)
2October to November 2008Global financial crisis
3March 2025Selloff
4August 2011U.S. downgrade and euro crisis
52000 to 2002Dot-com unwind

4.4 Bonds are where episodes are most predictable#

The paper tested whether current conditions could forecast whether an episode would occur over the next 10 days. The conditioning set was the VIX, realized volatility, drawdown, the Chicago Fed NFCI , the St. Louis Fed Financial Stress Index , and recent episode behavior. AUC measures that ability.

AssetAUC @ 95%AUC @ 97.5%AUC @ 99%Read
S&P 5000.5630.4710.405Fades in the tail
Nasdaq0.5630.5190.423Fades in the tail
TLT (20+ yr Treasuries)0.6720.6640.644Holds up in the tail
EUR/USD0.5970.6010.467Fades at 99%
USD/JPY0.5990.5810.473Fades at 99%
Bitcoin0.3770.4270.403Below random
Cross-asset mean0.560.540.47Weak overall

Dot plot of out-of-sample AUC by asset at three thresholds; long-duration Treasuries stay near 0.65 while equities, currencies and Bitcoin fall below 0.5 at the 99% threshold

In equities, currencies and Bitcoin, predictability faded as the threshold rose. In long-duration Treasuries it held up, even in the deep tail.

This is consistent with the week's pattern. Treasury stress tends to build through observable steps: inflation data, auction results, Fed communication and term premium . That gives risk managers something to monitor. Bond episodes build in ways that show up in the data, and our measures should be designed to see them.

4.5 The forecasting layer did not beat a simple baseline#

This is the result most papers would bury, and the most important one to state plainly.

The paper tested whether the full CTER conditioning set improved episode-severity forecasts over a simple EWMA volatility baseline . Both used the same gradient-boosted estimator and the same walk-forward protocol with a purge gap , so there was no look-ahead.

ThresholdMAE difference (CTER − baseline)95% block-bootstrap CI Diebold–Mariano p Verdict
95%+0.00123[+0.00073, +0.00179]below 0.001CTER significantly worse
97.5%+0.00072[+0.00034, +0.00114]below 0.001CTER significantly worse
99%+0.00017[−0.00008, +0.00041]0.077No significant difference

Interval chart of forecast error differences: CTER is significantly worse than the EWMA baseline at the 95% and 97.5% thresholds and not significantly different at 99%

The result held across horizons from 5 to 60 days, across five threshold levels, and under four scoring approaches: point MAE, TER pinball loss , a Fissler–Ziegel severity score, and a parametric extreme-value model.

Robustness checkRange testedDid conditioning help?
HorizonH = 5, 10, 20, 60 daysNo; the gap widens at longer horizons
Thresholdα = 90% to 99%No; significantly worse everywhere except 99%
Scoring ruleMAE, pinball, Fissler–ZiegelNo systematic gain under any
EstimatorGradient boosting and conditional GPD Better on typical episodes, worse on deep-tail ones

Conclusion: over 2000 to 2026, adding more market-state information did not improve episode forecasts over a simple volatility model. More variables added estimation noise without enough extra signal.


5. Why a Negative Result Is Still Useful#↑ Contents

A risk measure can capture something real and still be hard to forecast. The paper separates those two questions and answers them differently:

QuestionAnswerEvidence
Does TER measure something VaR and ES miss?YesTheory, simulation, stress tests, crisis history
Does CTER's conditional layer forecast better than simple models?Not yetOut-of-sample tests across six markets, 2000 to 2026

Put as a pipeline, risk functional → target → estimator → forecast → decision, the research supports the first two links. The estimator, forecast and decision links are still open.

This distinction matters beyond academia. Financial institutions are adopting AI-driven risk models quickly, and a model with more features and more complexity can look more impressive while forecasting worse. Under the Federal Reserve's SR 11-7 model risk guidance, complexity is not evidence of quality. This result shows why: a sophisticated conditioning layer lost to EWMA, a model from the 1990s.

The framework is also designed to be auditable:

ObjectElicitable? How to backtest it
TERYesPinball loss plus a coverage test; model-free, like VaR
CTER (as one number)No (proved by counterexample)Evaluate in parts
CTER: episode probabilityYesBrier score / AUC
CTER: episode severityYes (jointly)Fissler–Ziegel score, as with VaR/ES

That gives supervisors and model validators a clear evaluation protocol rather than a black box.


6. Why Episode Risk Matters for the U.S. Financial System#↑ Contents

Episode risk is not only a quant's concern. Weeks like this one test four parts of the U.S. economy.

ChannelThis week's evidenceWhy episode risk matters
1. Bank balance sheetsUnrealized losses on bank securities portfolios grow as yields climbUnrealized losses and deposit flight both build over episodes, as at SVB
2. Housing affordability30-year mortgage rates above 7.2% in daily lender surveysA sustained rise feeds into what millions of households pay; one volatile day that reverses does not
3. Public financeWeak 5-year and 7-year auctionsConsecutive weak auctions raise the cost of financing the national debt in steps
4. Cross-market contagionWTI–10yr correlation of 0.96Energy and rates now fail together, so a single-asset measure misses joint episodes

The fourth channel points directly to the next step in this research. In the stress tests, the broad contagion scenario, where all assets are moderately elevated, produced low single-asset episode severity (S* = 2.40), even though it is exactly the kind of systemic stress regulators care about. For more on how oil and the dollar have moved together since 2022, see our DXY, gold and oil analysis.


7. What Risk Managers Can Do Now#↑ Contents

#ActionWhy
1Report episodes alongside days. Add realized worst-episode severity over rolling 10-day windows to VaR/ES reports.It can be computed after the fact with no model
2Backtest the clustered-loss periods. Check model behavior across connected runs of losses (2008, 2020, 2022, this month).Individual breach counts can look acceptable while a run of losses breaks you
3Start simple. Use EWMA as the benchmark to beat.The evidence shows it is hard to beat for episode severity
4Watch the bond market closely. Track the auction calendar, inflation releases and Fed communication.Fixed income is where episodes are most predictable
5Stress-test joint moves. Design scenarios where rates and energy shock together.This year's data shows they move together
6Treat any new risk measure as a model under validation, including this one.Threshold, horizon and episode definition are choices that need governance

8. Limitations and What Comes Next#↑ Contents

  • Forecasting is unsolved. TER measures something real, but the conditional CTER layer did not beat EWMA out of sample. Until a model does, the simple baseline should be the default.
  • Design choices matter. The threshold level, the 10-day horizon and the rule that an episode ends on the first non-extreme day are all choices. Different choices give different numbers, which is why they need governance.
  • Single-asset only. The current framework measures episodes in one market at a time, which is why the broad-contagion scenario scored low.
  • Working paper. The research has not yet been peer reviewed.

The most important extension is a multivariate TER, which measures episodes when several markets fail together. It is the right tool for a period when geopolitics, energy and interest rates move as one.


Frequently Asked Questions#↑ Contents

Why did the 10-year Treasury yield rise above 5% in September 2026? A run of connected pressures rather than one shock. A strong S&P Global services PMI (58.7), rising odds of another Fed hike in October (about 64% on CME FedWatch), weak 5-year and 7-year auctions, and higher oil prices tied to the Iran war pushed the 10-year from about 4.97% to 5.10% on September 23 and to 5.23% by Friday, September 25, its highest level since 2007.

What is Tail Episode Risk (TER)? Tail Episode Risk is a risk measure proposed by Niraj Neupane, CA (ICAI), in a 2026 SSRN working paper. It groups consecutive days of extreme losses into episodes, measures each episode's severity as the sum of how far each day exceeded a tail threshold, and takes the upper-tail quantile of the worst episode over a fixed horizon, such as 10 trading days.

What is Conditional Tail Episode Risk (CTER)? CTER is the conditional version of TER. It multiplies the probability of entering an extreme-loss episode, given current market conditions, by the expected severity of that episode if it occurs. It answers the two questions a risk committee asks: how likely is a run of extreme losses, and how bad could it get?

How is TER different from Value-at-Risk and Expected Shortfall? VaR and Expected Shortfall are calculated from the distribution of daily losses and ignore the order in which losses arrive. Three bad days in a row and three bad days spread across a quarter give the same VaR and ES. TER depends on that order, so it rises when losses cluster. In the paper's simulation, 99% TER roughly tripled as clustering increased while VaR and ES barely moved.

Is TER the same as maximum drawdown? No. Drawdown measures the peak-to-trough fall in portfolio value, including ordinary days. TER counts only days beyond an extreme-loss threshold and only while they run consecutively. The paper proves the two can rank the same loss paths in opposite order.

Can tail-loss episodes be predicted in advance? Only weakly in most markets. Across six markets from 2000 to 2026, the average AUC at the 99% threshold was 0.47, no better than a coin flip. Long-duration Treasuries were the exception, with AUC between 0.64 and 0.67 at every threshold, because bond stress tends to build through observable steps such as auctions, inflation data and Fed communication.

Did the CTER model beat a simple volatility model? No. Using the same estimator and a walk-forward test with no look-ahead, the full CTER conditioning set forecast episode severity significantly worse than a simple EWMA volatility baseline at the 95% and 97.5% thresholds, and no differently at 99% (p = 0.077). The result held across horizons of 5 to 60 days and four scoring methods.

How much has a 10-year Treasury lost in price in 2026? Roughly 8%. The 10-year yield rose about 108 basis points from 4.15% at the start of 2026 to 5.23%. With a modified duration of about 7.7, that implies a price decline of about 8% after a small convexity adjustment, or about $800 million on $10 billion of 10-year-equivalent exposure. This is an illustrative estimate, not a reported figure.

What should risk managers do differently after a week like this? Report realized worst-episode severity over rolling 10-day windows alongside VaR and ES, backtest models across clustered-loss periods such as 2008, 2020 and 2022, keep EWMA as the benchmark any new model must beat, and stress-test scenarios in which interest rates and energy prices shock together.

Where can I read the full Conditional Tail Episode Risk paper? The working paper is available free on SSRN at abstract 7502499. It has not been peer reviewed.


Conclusion and Key Takeaways#↑ Contents

The Fed meets again on October 27 to 28. The 10-year may stabilize near current levels or keep rising. Either way, this week showed that tail risk has two dimensions:

DimensionMeasured by
How large a single loss can beVaR, ES
How losses line up over timeTER, CTER
  1. This week's Treasury selloff was an episode: five connected sessions that took the 10-year from below 5% to 5.23%, each manageable on its own.
  2. VaR and ES were built for the first dimension and cannot see the second.
  3. TER measures the second dimension, and in simulation it tripled when losses clustered while VaR and ES stayed flat.
  4. Bonds are where episodes are most predictable, which makes this market the right place to start monitoring them.
  5. A richer forecasting model did not beat EWMA. Forecasting episodes is still an open problem, and simple baselines deserve respect.
  6. The next step is a multivariate version, because the risks we face now arrive across markets at the same time.

Measuring that second dimension is a practical question for bank safety, household borrowing costs and the stability of the U.S. Treasury market. The next stress event will likely look like this week did: several consecutive days that each seem manageable on their own.

SSRN Working Paper, September 2026 · Abstract 7502499

Conditional Tail Episode Risk: A Path-Dependent Framework for Extreme-Loss Episodes Beyond Value-at-Risk and Expected Shortfall

Niraj Neupane

Read the full paperFree · Opens on SSRN

Methodology and Sources#↑ Contents

This article applies the author's own working paper to the Treasury market of September 21 to 25, 2026. All research figures (simulation, stress scenarios, crisis ranking, AUC and forecast-comparison results) are as reported in the paper. Market figures are secondary, as reported by the news sources below. The portfolio-loss figures in section 1.2 are illustrative duration-based estimates by Financial Gurkha, with the arithmetic shown. Charts are drawn by Financial Gurkha from the paper's tables.

Primary source

Market data (week of September 21 to 25, 2026)

  • CNBC, "10-year Treasury yield rockets to 19-year high" (Sept 23, 2026). cnbc.com
  • CNBC, "The 10-year Treasury yield is at its highest in nearly two decades. How we got here" (Sept 26, 2026). cnbc.com
  • CNBC, Treasury yields, Warsh, Bessent and the national debt (Sept 24, 2026). cnbc.com
  • CNBC, Fed hike and the 10-year above 5% (Sept 16, 2026). cnbc.com
  • CNBC, 10-year at its highest since 2007 (Sept 15, 2026). cnbc.com
  • CNN Business, "10-year Treasury yield hits 5.1% for first time in 19 years" (Sept 23, 2026). cnn.com
  • NBC News, Treasury yields and oil (Sept 23, 2026). nbcnews.com
  • CNBC, Stock market live updates (Sept 25, 2026). cnbc.com
  • Yahoo Finance, Stock market today (Sept 25, 2026). finance.yahoo.com
  • Charles Schwab, market update (Sept 25, 2026). schwab.com
  • BNN Bloomberg / CP24 (Sept 25, 2026). cp24.com
  • Forbes Advisor, "Mortgage Rates Today: September 25, 2026 – 30-Year Rate Hits One-Year High." forbes.com

Research

  • Anderson, C. W. (1994). The aggregate excess measure of severity of extreme events. Journal of Research of NIST, 99, 555–561.
  • Fissler, T., & Ziegel, J. F. (2016). Higher order elicitability and Osband's principle. Annals of Statistics, 44, 1680–1707.
  • Goldberg, L. R., & Mahmoud, O. (2017). Drawdown: from practice to theory and back again. Mathematics and Financial Economics, 11, 275–297.
  • Patton, A. J., Ziegel, J. F., & Chen, R. (2019). Dynamic semiparametric models for expected shortfall (and Value-at-Risk). Journal of Econometrics, 211, 388–413.
  • Board of Governors of the Federal Reserve System / OCC. SR 11-7: Guidance on Model Risk Management.

About the author#

Niraj Neupane, CA (ICAI), is a quantitative researcher and founder of Korvane (AI-powered trade, risk and validation) and Calderyn Institute (quant finance AI engineering education). His working papers on tail-episode risk and machine-learning Value-at-Risk are available on SSRN. ORCID: 0009-0003-7026-7026. Read more on his author profile.

Views expressed are the author's own and do not represent any employer.

Financial Gurkha · Order Ticket

FG-3861

Done reading? Let's play a game.

Pick a desk. The House Fulfills.

Financial Gurkha — Independent Financial Market Research from Wall Street, New York