Thirteen validation winters, 2013/14 to 2025/26, scored on the October–March windstorm season, on a network of 3,002 stations across nine north-west European countries drawn from a daily hindcast library of about 5,300 European stations. Predictions are scored against two records built from different data: a station observation archive reaching back to the 1930s, and a 1950–2025 windstorm event catalogue we built from reanalysis. Every skill number is measured against a no-skill baseline made from the same climatology, so what is reported is what the forecast adds. No loss data is used in fitting or scoring anywhere; loss figures appear only as context. This is the short view for risk and analytics teams: what the system predicts, how well it has done over the scored record, and the alarms that follow. The instrument-level validation, the specification and threshold workings, the glossary and the season-by-season detail are in the full technical validation report, available on request.
Insured windstorm outcomes are not one quantity. Fifty-three winters of station observations separate cleanly into two families, and a winter can be extreme in one while sitting near normal in the other. That is the single most important structural fact in this programme, because it decides how many calls the system needs to make.
Wind activity, peak severity, storm clustering and station peak wind all move together: they correlate +0.73 to +0.97 with each other over 53 winters. They also carry the strongest link to the independently built catalogue extremity index. This is the axis the industry has historically priced: the 1990-type winter.
The worst 7-day burst and clumping correlate +0.79 with each other, but far more loosely with the first family: burst +0.47 to +0.56, clumping +0.26 to +0.37. Against the catalogue index they manage only +0.24 and +0.18. A concentrated fortnight inside an otherwise ordinary season shares some ground with a relentless season, but not enough for either to stand in for the other, and the catalogue does not see it at all.
| activity | severity | clustering | station peak | burst | clumping | catalogue index | |
|---|---|---|---|---|---|---|---|
| activity | +0.97 | +0.95 | +0.79 | +0.53 | +0.29 | +0.62 | |
| severity | +0.97 | +0.93 | +0.87 | +0.56 | +0.37 | +0.60 | |
| clustering | +0.95 | +0.93 | +0.73 | +0.48 | +0.26 | +0.52 | |
| station peak | +0.79 | +0.87 | +0.73 | +0.47 | +0.31 | +0.55 | |
| burst | +0.53 | +0.56 | +0.48 | +0.47 | +0.79 | +0.24 | |
| clumping | +0.29 | +0.37 | +0.26 | +0.31 | +0.79 | +0.18 | |
| catalogue index | +0.62 | +0.60 | +0.52 | +0.55 | +0.24 | +0.18 |
Observed station measures against each other and against the catalogue extremity index, 53 winters (the index column covers 41). The two blocks in the top-left and bottom-right are the two families; the weak corner between them is the point. Green is positive, red negative, intensity is strength.
| observed station measure | r vs the catalogue extremity index | p | winters |
|---|---|---|---|
| activity | +0.62 | 0.0001 | 41 |
| severity | +0.60 | 0.0000 | 41 |
| clustering | +0.52 | 0.0004 | 41 |
| station peak | +0.55 | 0.0001 | 41 |
| burst | +0.24 | 0.1310 | 41 |
| clumping | +0.18 | 0.2601 | 41 |
Permutation-tested over the winters where both series exist. Activity, severity, station peak and clustering all track the catalogue index at p≤0.0004 across 41 winters: two records built from different data agreeing on which winters were bad. Burst and clumping do not, and that is not a failure: they are measuring the other family.
Dudley, Eunice and Franklin arrived between 16 and 21 February 2022: three damaging storms in six days, EUR 3.85bn of industry loss. All three fell inside the season, so this is not a case of looking at the wrong months.
On family one the winter was unremarkable: activity +0.30σ, peak severity +0.45σ, station clustering +0.28σ, and the catalogue index put it at the 23.6th percentile of 75 years. On family two it was extreme: worst 7-day burst +2.91σ and clumping +2.18σ across October–March, the second highest burst of the thirteen behind 2018/19.
The forecast side split the same way. The October wind call read −0.81σ, so nothing in the family-one instruments saw it coming. The August concentration index read +0.56, its second-highest reading of the thirteen: six months before the fortnight.
And the week itself can be located from the station data alone. Asked which seven days carried most of the network’s exceedance that winter, with no knowledge of named storms, the answer is 16 to 22 February 2022, carrying 4.04 times the exceedance an evenly spread season would put in a week. The map below shows which stations took that hit: a corridor from Ireland through to Denmark, with Scandinavia untouched.
The record is split by issuance, because the two issuances predict different families and one table invites the wrong comparison. The first row of toggles picks the issuance; the second picks the geography, with pooled Europe, the product, as the default. Click any observed column for what predicted it and the comparison chart.
Signals are standardised on the thirteen-winter issuance record; observed columns are σ against the frozen 1972–2011 station reference. On weighting: every observed column drawn from station data is a network mean on the same population kernel as the signal, so forecast and outcome are weighted identically. The only unweighted column is the catalogue extremity percentile. That percentile is missing 2022/23 and does not yet reach 2025/26, so those two cells are blank, which is why every catalogue figure in this report is scored on eleven winters rather than thirteen. This table is the source. Every alarm threshold is computed from exactly these columns, so any figure in the alarm tables can be reproduced from the numbers on this screen with nothing more than a spreadsheet.
On the composite columns. Both issuances carry the three instruments and the composite built from them, and the composite competes for the best-predictor highlight like any other column, because the report makes claims about it. On the August issuance the composite is the concentration index, the product for family two. On the October issuance it is the family-one composite, which is shown precisely because it loses to rain: combining helps on one family and dilutes on the other, and this is where that is visible in the data. One thing to expect. The best-predictor figure is the plain correlation of the two columns as printed, so it can be checked against this table directly. The findings quote partial correlations, with the no-skill climatology removed from both sides. The two occasionally rank differently, most visibly on October–March burst, where the two are close: wind 0.71 against the index 0.70 on the plain numbers, and the index +0.73 against wind +0.70 once each is scored against its own Floor. Neither gap is meaningful on thirteen winters. The panel says which measure it is reporting whenever a published figure exists for the column.
Two families, and two different kinds of severe winter inside the first one, mean the system needs three alarms rather than one. The 53-winter observed record shows why with unusual clarity, and it also shows which of the three we can currently build.
| winter | activity σ | 7-day burst σ | catalogue extremity pct | severity alarm would fire | concentration alarm would fire |
|---|---|---|---|---|---|
| 1973/74 | +0.76 | +0.92 | 40.0 | Tier 1 | : |
| 1982/83 | +0.71 | +0.55 | 72.7 | Tier 1 | : |
| 1983/84 | +0.77 | +1.68 | 90.9 | Tier 1 | fires |
| 1989/90 | +1.61 | +2.09 | 95.5 | Tier 2 | fires |
| 1993/94 | +0.68 | +0.51 | 81.8 | Tier 1 | : |
| 1994/95 | +1.29 | +0.14 | 95.5 | Tier 2 | : |
| 2001/02 | +0.94 | +1.04 | 70.0 | Tier 1 | : |
| 2006/07 | +0.91 | +0.59 | 67.3 | Tier 1 | : |
| 2013/14 | +0.92 | +0.22 | 61.8 | Tier 1 | : |
| 2015/16 | +0.66 | +1.42 | 76.4 | Tier 1 | : |
| 2018/19 | −0.45 | +3.05 | 25.5 | : | fires |
| 2019/20 | +1.23 | +1.21 | 89.1 | Tier 1 | : |
| 2021/22 | +0.30 | +2.81 | 23.6 | : | fires |
Every winter in 53 that either the severity ladder or the concentration threshold would flag, with the two side by side. Severity alarm: Tier 1 is observed activity at or above +0.65σ, Tier 2 is activity and station clustering both at or above +1.25σ, on the frozen 1972–2011 scale: 11 Tier 1 firings and 2 Tier 2 across 53 winters. Concentration alarm: observed worst 7-day burst at or above +1.5σ. These are the observed definitions, which is what makes the comparison possible over 53 winters; the prediction-side thresholds are a separate job, noted below. One note on the scale. Every observed σ quoted in this report sits on a frozen pre-2012 reference, so that winters are compared against a baseline that does not include them. What changes between exhibits is the season window: this ladder is built on November to March, the record table above on October to March. The same winter therefore carries two readings, both frozen: 2018/19’s burst is +3.05σ here and +3.18σ there. Neither is more correct, and the report does not mix the two windows within a table.
Two families and two mechanisms inside family one, the winter that is bad for the current climate, and the winter that is exceptional against any climate, make three alarms the natural design. They are not equally ready, and the difference matters more than the design.
August concentration index → observed Oct–Mar burst, +0.73, p=0.009, and → clumping +0.63, p=0.029. Every cell survives dropping any single winter. Fires for the family-two winter from two months before the season opens. This is the alarm that would have caught 2021/22 and 2018/19, the two most concentrated winters of the record.
October conviction → catalogue extremity index. Rain reads +0.5 to +0.6, p between 0.06 and 0.13 on the eleven winters the catalogue indexes. This is the winter that is bad relative to today’s climate, and the strongest family-one relationship in the programme. Fires at an October rain reading of +0.50.
Three legs, each on its own target and its own firing level. Any one fires the alarm. Concentration: August index at +0.63 index units, against an observed Oct–Mar burst of +2.0σ. Extremity: October rain at +0.77, against catalogue extremity at or above the 95th percentile. Breadth: October wind conviction read as a network mean at +0.61, against observed activity at or above +1.23σ. Both the concentration and extremity legs sit above the firing level of the alarm whose instrument they share, so any winter firing them has already fired Alarm 1 or Alarm 2. The breadth leg is the only one that reaches a winter neither has flagged, and it does so in 2% of winters. Detection: 68% for a 1989/90 winter, 45% for 1994/95, 51% for 1983/84. The derivation is below.
| leg | instrument | target | threshold | leg fires |
|---|---|---|---|---|
| concentration | August index | Oct–Mar burst ≥ +2.0σ | +0.63 | 4.8% |
| extremity | October rain | catalogue extremity ≥ 95th pct | +0.77 | 11.0% |
| breadth | October wind, network mean | observed activity ≥ +1.23σ | +0.61 | 5.7% |
Thresholds in native signal units. Any one leg fires the alarm. Detection is the probability that a winter delivering the stated observed values produces a signal above the level: 68% for 1989/90, 45% for 1994/95, 51% for 1983/84 and 42% for 2019/20. The three legs carry correlations of +0.73, +0.57 and +0.30 against their own targets. Firing rates are read from the modelled signal distribution, not from the thirteen-winter tally, so they are not comparable with the counts in the sweep tables above.
When a winter is loaded, the predicted instruments run red together and the observed dials verify together. The verified quantity is how much of the map runs red, winter by winter: not which individual station is hit, which the system does not claim.
Interactive: predicted instruments on the left, observed dials on the right; any dial can be read alone or combined with any others, and the winters step with the chips, the arrow keys or Play. Strongest signal shows the deepest reading any active dial gives each station, which is how an alarm set behaves; Averaged draws the per-station mean instead, which is the aggregation a composite actually uses. Coverage: all thirteen scored winters plus the season ahead. Matched runs puts each dial on the issuance that suits its own target, wind, rain and warmth on the October run over December to February and the index on the August run over October to March. The October station arrays cover the thirteen scored winters; they grey out for 2026/27, for which the store holds no October initialisation. August only puts all four on the August issuance over October to March, on one construction and every winter including the season ahead.
Lead time. The concentration call becomes available at the August issuance and not before: the June and July runs read −0.22 and −0.07 against the observed burst, August reads +0.73, September +0.44 and October +0.29. A separate reading of the core winter against the catalogue follows in October. The industry windstorm season runs October to March, so the August initialisation lands roughly two months before the season opens and four to seven months before the months that usually carry the damage. The core-winter reading arrives as the season starts. The step from July to August is discontinuous rather than gradual, which is a caveat on thirteen winters as much as a property of the system: the workings show the whole curve.
Seasonal storm totals. Against observed activity the forecast does not beat the climate trend at any lead tested. The trend itself is informative and ours is better than a straight line, but that is the climate engine rather than season picking. It is the open item behind Alarm 3, the exceptional-season call.
Peak severity. Not predicted at any lead, on any window. A single exceptional storm sits outside what a seasonal system can see.
Which stations get hit. Tested directly: same-winter geographic pattern agreement does not beat cross-winter agreement. The map’s claim is breadth, not placement.
The sample. Thirteen winters of predictions, scored against 53 winters of station observations and 75 years of catalogue index. Thirteen is what the hindcast covers and we are not extending it, so the work went into making those thirteen hard to argue with: partial correlations against a no-skill baseline built from the same climatology, permutation tests rather than parametric p-values, leave-one-winter-out refitting on every headline cell, and a robustness check that rebuilds the targets without population weighting. The chain also rests on far more than thirteen years: the observed measures track the catalogue index across 41 winters, and the day-level accuracy behind the instruments is verified over 68 million day-predictions.
Issued from the August 2026 run, on the October–March season. Concentration is the measure the system predicts at this range; the December–February reading against the catalogue index needs the October initialisation and does not exist yet.
| instrument | 2026/27 reading σ | lowest of the thirteen | highest of the thirteen |
|---|---|---|---|
| wind | −0.59 | −0.47 | +0.75 |
| rain | −1.13 | −0.69 | +0.78 |
| warmth | −0.46 | −0.63 | +0.86 |
The season call is one number, but the forecast is issued month by month. The same August run resolved into all six months of the season, each scored against that month at that lead in the thirteen previous August runs. The eighteen cells below average to the three instrument readings above, which in turn average to the index.
Seventeen of the eighteen cells read below their thirteen-winter mean and none reaches −2σ. Most emphatic: rain in December (−1.64σ). The single cell above its mean is warmth in October (+0.08σ), which is close enough to normal to be read as neutral. The measured skill is for the season-shape measure, not for individual months, so the month resolution is context rather than six separate forecasts.
What it says. The volume-and-structure risk for the extended season reads below normal on the measure the system predicts best at this range.
What it does not say. Nothing about how many storms will arrive or how severe the worst one will be, and nothing yet about the core winter against the catalogue index: that needs October.
What would change it. The October issuance is a genuinely different observation rather than a refinement: it reads the autumn state that conditions the core winter. In 2019/20 the August run gave no warning of a winter the October wind call went on to place at its record maximum.
Run agreement, in three numbers rather than a table. Splitting the thirteen winters at the median spread between issuances, the agreeing half missed the outcome by 0.87σ and the divergent half by 1.22σ. Spread correlates +0.28 with error, and +0.45 with how far the outcome landed from normal in either direction, so divergence marks a season with something in it rather than a reason to discount the call. For 2026/27 the runs disagree, which is why the reading above is provisional and why the September and October issuances are worth waiting for. The full issuance-by-issuance table is in the workings tab.
The same run carries a 24-month horizon, so it does produce readings for October 2027 to March 2028 at leads of 14 to 19 months. They come in uniformly low: every month of every instrument between −1.67σ and −0.49σ. We do not issue that as a view. The uniformity is the reason: a real seasonal signal discriminates between months and between instruments, and this one does not, which is the signature of a run relaxing towards its own climatology. There is also no skill test at that range; the longest lead with a measured result anywhere in this document is four months. The second winter becomes forecastable when its own late-summer initialisation arrives.
Independent of any model output: what the 75-year teleconnection record says about winters like the one ahead. Context for the call, never a call on its own.
| Drivers aligned active | winters | storm count | clustered events | aggregate severity |
|---|---|---|---|---|
| 0 | 11 | 7.4 | 2.2 | 1296 |
| 1 | 26 | 8.4 | 2.2 | 1450 |
| 2 | 29 | 11.6 | 3.1 | 2150 |
| 3 | 9 | 14.1 | 4.7 | 2745 |
Escalation with alignment (p=0.004, trend-robust; storm count and aggregate severity rise monotonically, clustered events are flat between 0 and 1 aligned drivers then rise): the interplay, not any single index, is the signal. ENSO does not lead for Europe: it is tested and carried as a negative, not used. Caveats: 75-year averages, regime-dependent; in the recent slice the relationship reverses (see the note below); long-record context, never a season-caller.
| Winter | Storms | Alignment | 75-yr percentile | z NATI | z AMO | z PDO |
|---|---|---|---|---|---|---|
| 1987/88 | 87J (Oct 1987)* | -0.43 | 26th | +0.91 | -0.21 | +0.59 |
| 1989/90 | Daria, Vivian, Herta | +0.16 | 56th | +1.00 | -1.11 | -0.36 |
| 1990/91 | Undine; run-on season | +0.76 | 87th | -0.10 | -1.02 | -1.17 |
| 1999/00 | Anatol, Lothar, Martin (Dec 1999) | +1.21 | 99th | -1.29 | -0.50 | -1.84 |
| 2006/07 | Kyrill (Jan 2007) | -0.30 | 36th | +0.91 | +0.54 | -0.56 |
| 2009/10 | Xynthia (Feb 2010) | -0.85 | 8th | +1.90 | +0.55 | +0.09 |
| 2013/14 | Serial UK winter | +0.29 | 64th | +0.22 | -0.95 | -0.15 |
| 2021/22 | Dudley, Eunice, Franklin | +0.90 | 91st | -1.34 | +0.45 | -1.82 |
| 2023/24 | High-latitude cluster year | -0.90 | 5th | +1.14 | +2.34 | -0.79 |
Grey: the long record; blue: hindcast-era winters; red: major insured-loss winters; amber: index-extreme, modest-loss. One-signed reading: alignment grades serial storm-family winters (99th/91st/87th percentiles) and does not flag single-monster winters – clustering-axis context, selected-by-loss so it motivates, never validates.