MetSwift · Claros European winter windstorm signal

What the system predicts, how well, and what it says about the winter ahead

Thirteen validation winters, 2013/14 to 2025/26, scored on the October–March windstorm season, on a network of 3,002 stations across nine north-west European countries drawn from a daily hindcast library of about 5,300 European stations. Predictions are scored against two records built from different data: a station observation archive reaching back to the 1930s, and a 1950–2025 windstorm event catalogue we built from reanalysis. Every skill number is measured against a no-skill baseline made from the same climatology, so what is reported is what the forecast adds. No loss data is used in fitting or scoring anywhere; loss figures appear only as context. This is the short view for risk and analytics teams: what the system predicts, how well it has done over the scored record, and the alarms that follow. The instrument-level validation, the specification and threshold workings, the glossary and the season-by-season detail are in the full technical validation report, available on request.

Findings Winter 2026/27

What we are trying to predict

Insured windstorm outcomes are not one quantity. Fifty-three winters of station observations separate cleanly into two families, and a winter can be extreme in one while sitting near normal in the other. That is the single most important structural fact in this programme, because it decides how many calls the system needs to make.

Family one: how severe the season was overall

Wind activity, peak severity, storm clustering and station peak wind all move together: they correlate +0.73 to +0.97 with each other over 53 winters. They also carry the strongest link to the independently built catalogue extremity index. This is the axis the industry has historically priced: the 1990-type winter.

Family two: whether the damage arrived together

The worst 7-day burst and clumping correlate +0.79 with each other, but far more loosely with the first family: burst +0.47 to +0.56, clumping +0.26 to +0.37. Against the catalogue index they manage only +0.24 and +0.18. A concentrated fortnight inside an otherwise ordinary season shares some ground with a relentless season, but not enough for either to stand in for the other, and the catalogue does not see it at all.

activityseverityclusteringstation peakburstclumpingcatalogue index
activity+0.97+0.95+0.79+0.53+0.29+0.62
severity+0.97+0.93+0.87+0.56+0.37+0.60
clustering+0.95+0.93+0.73+0.48+0.26+0.52
station peak+0.79+0.87+0.73+0.47+0.31+0.55
burst+0.53+0.56+0.48+0.47+0.79+0.24
clumping+0.29+0.37+0.26+0.31+0.79+0.18
catalogue index+0.62+0.60+0.52+0.55+0.24+0.18

Observed station measures against each other and against the catalogue extremity index, 53 winters (the index column covers 41). The two blocks in the top-left and bottom-right are the two families; the weak corner between them is the point. Green is positive, red negative, intensity is strength.

Definitions: the observed measures, precisely
  • Wind activity. Days per station-winter above that station’s own era-relative 85th-percentile daily threshold, averaged across the network on the population kernel, expressed as σ against the frozen 1972–2011 reference. The canonical “how windy was the season” series.
  • Peak severity. The strongest daily wind each station records all season, network mean. Note the catalogue index uses a daily mean wind while station data carries the daily maximum, so the two severity measures are not like-for-like and are never compared raw.
  • Storm clustering (station pairs). Separate exceedance episodes at the same station arriving within 72 hours of one another, counted per station and then averaged. Temporal clustering, measured locally.
  • Station peak wind. Absolute daily maximum per station, network mean. Sensitive to single storms.
  • Worst 7-day burst. The most intense seven-day window of network-wide exceedance anywhere in the winter. Asks whether the season’s damage piled into one short stretch.
  • Clumping. Variance-to-mean ratio of weekly network exceedance counts across the season. High when the winter arrives in a few concentrated spells rather than evenly.
  • The concentration index, and how to read its scale. The product itself. For each instrument and each of the six months, the population-weighted conviction-day rate is standardised on its own thirteen-winter record, the six months are averaged, and the three instruments are then averaged. So eighteen standardised numbers, each with one equal vote. It is built out of standard deviations but it is not itself on a standard-deviation scale: averaging eighteen partly-correlated standardised numbers shrinks the spread, and the index has a standard deviation of about 0.38 across the thirteen winters against roughly 0.45 for each single instrument. So a reading of −0.73 is about 2.0 index standard deviations below normal, not two thirds of one, and the +0.5 alarm threshold sits at about 1.3. Index readings in this report are quoted as index units and never carry a sigma sign; observed outcomes are quoted in sigma against the frozen 1972–2011 reference and always do.
  • The windstorm catalogue, and the four measures it carries. Our own 1950–2025 windstorm event catalogue, built from reanalysis fields rather than from the station archive. It shares no data with the station instruments or the station outcomes, which is what makes it a genuine cross-check, but it is ours, not a third party’s, and it is not an external validation. Four measures come off it, and they keep the names they were given at the start of the programme:
    • Catalogue events: the number of windstorm events identified in the winter. A count.
    • Catalogue clustering: how many of those events fall within 72 hours of another. The catalogue’s own clustering axis, and the closest independent analogue to our concentration family.
    • Catalogue severity: the strongest single event wind of the winter, in m/s. Daily-mean based, where station data carries the daily maximum, so the two severity measures are never compared raw.
    • Catalogue extremity index: the blended index, quoted as a percentile of the 75-year record. This is the measure the report scores against, and the one referred to elsewhere simply as the catalogue index.
Concentration and clustering are not the same thing: how they differ
  • Clustering asks a local question: at one station, did exceedances arrive in quick succession? It is counted per station, then averaged over three thousand stations and the whole season.
  • Concentration (burst and clumping) asks a seasonal, network-wide question: did the network’s exceedance pile into one short window of the winter?
  • The two can disagree sharply. A winter of many small paired events spread evenly from November to March clusters strongly and concentrates weakly. A winter where one region is hit three times in a fortnight and the rest of the season is quiet concentrates strongly and clusters weakly, because a short run affecting part of the network is diluted by twenty quiet weeks and three thousand stations when it is averaged.
  • 2021/22 is that second case. Station clustering read +0.28σ, essentially normal. Worst 7-day burst read +2.91σ, the second highest of the thirteen behind 2018/19.
  • This document uses concentration throughout for family two and reserves clustering for the station-pair measure.

Why family one is worth predicting: the long-record link

observed station measurer vs the catalogue extremity indexpwinters
activity+0.620.000141
severity+0.600.000041
clustering+0.520.000441
station peak+0.550.000141
burst+0.240.131041
clumping+0.180.260141

Permutation-tested over the winters where both series exist. Activity, severity, station peak and clustering all track the catalogue index at p≤0.0004 across 41 winters: two records built from different data agreeing on which winters were bad. Burst and clumping do not, and that is not a failure: they are measuring the other family.

The results

Family two: season concentration, October–March, from the August initialisation

The record, scored in both directions. Two winters were flagged at or above +0.5 on the index, and both were the two most concentrated of the thirteen: observed burst +3.18σ and +2.91σ, averaging +3.04σ. No false alarm. Going the other way, five winters delivered observed burst at or above +1.0σ: the index caught the two largest and read the other three between −0.25 and −0.02, close to normal rather than wrong-signed. So the index discriminates the extremes well and the middle poorly: it separates a +3σ volley winter from an ordinary one, and it does not rank a +1σ winter above an average one. For a capital-planning use that is the right way round; for anything needing a full ranking it is not.

2021/22: the case that shows why two families need two calls

Dudley, Eunice and Franklin arrived between 16 and 21 February 2022: three damaging storms in six days, EUR 3.85bn of industry loss. All three fell inside the season, so this is not a case of looking at the wrong months.

On family one the winter was unremarkable: activity +0.30σ, peak severity +0.45σ, station clustering +0.28σ, and the catalogue index put it at the 23.6th percentile of 75 years. On family two it was extreme: worst 7-day burst +2.91σ and clumping +2.18σ across October–March, the second highest burst of the thirteen behind 2018/19.

The forecast side split the same way. The October wind call read −0.81σ, so nothing in the family-one instruments saw it coming. The August concentration index read +0.56, its second-highest reading of the thirteen: six months before the fortnight.

And the week itself can be located from the station data alone. Asked which seven days carried most of the network’s exceedance that winter, with no knowledge of named storms, the answer is 16 to 22 February 2022, carrying 4.04 times the exceedance an evenly spread season would put in a week. The map below shows which stations took that hit: a corridor from Ireland through to Denmark, with Scandinavia untouched.

The thirteen-winter record

The record is split by issuance, because the two issuances predict different families and one table invites the wrong comparison. The first row of toggles picks the issuance; the second picks the geography, with pooled Europe, the product, as the default. Click any observed column for what predicted it and the comparison chart.

Signals are standardised on the thirteen-winter issuance record; observed columns are σ against the frozen 1972–2011 station reference. On weighting: every observed column drawn from station data is a network mean on the same population kernel as the signal, so forecast and outcome are weighted identically. The only unweighted column is the catalogue extremity percentile. That percentile is missing 2022/23 and does not yet reach 2025/26, so those two cells are blank, which is why every catalogue figure in this report is scored on eleven winters rather than thirteen. This table is the source. Every alarm threshold is computed from exactly these columns, so any figure in the alarm tables can be reproduced from the numbers on this screen with nothing more than a spreadsheet.

On the composite columns. Both issuances carry the three instruments and the composite built from them, and the composite competes for the best-predictor highlight like any other column, because the report makes claims about it. On the August issuance the composite is the concentration index, the product for family two. On the October issuance it is the family-one composite, which is shown precisely because it loses to rain: combining helps on one family and dilutes on the other, and this is where that is visible in the data. One thing to expect. The best-predictor figure is the plain correlation of the two columns as printed, so it can be checked against this table directly. The findings quote partial correlations, with the no-skill climatology removed from both sides. The two occasionally rank differently, most visibly on October–March burst, where the two are close: wind 0.71 against the index 0.70 on the plain numbers, and the index +0.73 against wind +0.70 once each is scored against its own Floor. Neither gap is meaningful on thirteen winters. The panel says which measure it is reporting whenever a published figure exists for the column.

Alarms

Two families, and two different kinds of severe winter inside the first one, mean the system needs three alarms rather than one. The 53-winter observed record shows why with unusual clarity, and it also shows which of the three we can currently build.

The observed ladder, and what it cannot see

winteractivity σ7-day burst σcatalogue extremity pctseverity alarm would fireconcentration alarm would fire
1973/74+0.76+0.9240.0Tier 1:
1982/83+0.71+0.5572.7Tier 1:
1983/84+0.77+1.6890.9Tier 1fires
1989/90+1.61+2.0995.5Tier 2fires
1993/94+0.68+0.5181.8Tier 1:
1994/95+1.29+0.1495.5Tier 2:
2001/02+0.94+1.0470.0Tier 1:
2006/07+0.91+0.5967.3Tier 1:
2013/14+0.92+0.2261.8Tier 1:
2015/16+0.66+1.4276.4Tier 1:
2018/19−0.45+3.0525.5: fires
2019/20+1.23+1.2189.1Tier 1:
2021/22+0.30+2.8123.6: fires

Every winter in 53 that either the severity ladder or the concentration threshold would flag, with the two side by side. Severity alarm: Tier 1 is observed activity at or above +0.65σ, Tier 2 is activity and station clustering both at or above +1.25σ, on the frozen 1972–2011 scale: 11 Tier 1 firings and 2 Tier 2 across 53 winters. Concentration alarm: observed worst 7-day burst at or above +1.5σ. These are the observed definitions, which is what makes the comparison possible over 53 winters; the prediction-side thresholds are a separate job, noted below. One note on the scale. Every observed σ quoted in this report sits on a frozen pre-2012 reference, so that winters are compared against a baseline that does not include them. What changes between exhibits is the season window: this ladder is built on November to March, the record table above on October to March. The same winter therefore carries two readings, both frozen: 2018/19’s burst is +3.05σ here and +3.18σ there. Neither is more correct, and the report does not mix the two windows within a table.

The ladder works on the family it was built for. Its two Tier 2 winters, 1989/90 and 1994/95, sit at the 95.5th percentile of the catalogue index. Its Tier 1 winters run from the 40th to the 90.9th, mostly high. This is the alarm for the winter that grinds: the type the market priced through the 1980s and 1990s.
And it is blind to the other family. The two most concentrated winters in the record, 2018/19 (burst +3.05σ) and 2021/22 (burst +2.81σ), fire neither tier of the severity ladder. Their activity readings are −0.45σ and +0.30σ, and the catalogue index puts them at the 25.5th and 23.6th percentiles. They are also the two most recent extremes in the series. A severity ladder on its own would have called both winters quiet, which is the whole case for a second alarm.

What the architecture should therefore be: three alarms, at three stages of maturity

Two families and two mechanisms inside family one, the winter that is bad for the current climate, and the winter that is exceptional against any climate, make three alarms the natural design. They are not equally ready, and the difference matters more than the design.

Alarm 1: concentration evidenced

August concentration index → observed Oct–Mar burst, +0.73, p=0.009, and → clumping +0.63, p=0.029. Every cell survives dropping any single winter. Fires for the family-two winter from two months before the season opens. This is the alarm that would have caught 2021/22 and 2018/19, the two most concentrated winters of the record.

Alarm 2: winter extremity, Dec–Feb specified

October conviction → catalogue extremity index. Rain reads +0.5 to +0.6, p between 0.06 and 0.13 on the eleven winters the catalogue indexes. This is the winter that is bad relative to today’s climate, and the strongest family-one relationship in the programme. Fires at an October rain reading of +0.50.

Alarm 3: the exceptional season specified

Three legs, each on its own target and its own firing level. Any one fires the alarm. Concentration: August index at +0.63 index units, against an observed Oct–Mar burst of +2.0σ. Extremity: October rain at +0.77, against catalogue extremity at or above the 95th percentile. Breadth: October wind conviction read as a network mean at +0.61, against observed activity at or above +1.23σ. Both the concentration and extremity legs sit above the firing level of the alarm whose instrument they share, so any winter firing them has already fired Alarm 1 or Alarm 2. The breadth leg is the only one that reaches a winter neither has flagged, and it does so in 2% of winters. Detection: 68% for a 1989/90 winter, 45% for 1994/95, 51% for 1983/84. The derivation is below.

What the two working alarms deliver when they are run together. A winter counts as one to call if it delivered either an Oct–Mar burst at or above +1.0σ or a catalogue extremity index at or above the 60th percentile. Firing either alarm at a +0.5 reading, across the thirteen winters: three hits, two false alarms, five correct quiet calls, three misses. The hits are 2018/19 and 2021/22, the two most concentrated winters of the record, and 2015/16 on an October rain call of +0.68. The false alarms are 2017/18 and 2023/24; the misses are 2013/14, 2014/15 and 2019/20. Three of the six firing decisions sit within 0.06 of the threshold. What the period did not contain is the 1989/90 type, which is why the third alarm is scoped from the long observed record rather than from this sample.

Alarm 3 → the three legs, their targets and their firing levels

leginstrumenttargetthresholdleg fires
concentrationAugust indexOct–Mar burst ≥ +2.0σ+0.634.8%
extremityOctober raincatalogue extremity ≥ 95th pct+0.7711.0%
breadthOctober wind, network meanobserved activity ≥ +1.23σ+0.615.7%

Thresholds in native signal units. Any one leg fires the alarm. Detection is the probability that a winter delivering the stated observed values produces a signal above the level: 68% for 1989/90, 45% for 1994/95, 51% for 1983/84 and 42% for 2019/20. The three legs carry correlations of +0.73, +0.57 and +0.30 against their own targets. Firing rates are read from the modelled signal distribution, not from the thirteen-winter tally, so they are not comparable with the counts in the sweep tables above.

The map

When a winter is loaded, the predicted instruments run red together and the observed dials verify together. The verified quantity is how much of the map runs red, winter by winter: not which individual station is hit, which the system does not claim.

Interactive: predicted instruments on the left, observed dials on the right; any dial can be read alone or combined with any others, and the winters step with the chips, the arrow keys or Play. Strongest signal shows the deepest reading any active dial gives each station, which is how an alarm set behaves; Averaged draws the per-station mean instead, which is the aggregation a composite actually uses. Coverage: all thirteen scored winters plus the season ahead. Matched runs puts each dial on the issuance that suits its own target, wind, rain and warmth on the October run over December to February and the index on the August run over October to March. The October station arrays cover the thirteen scored winters; they grey out for 2026/27, for which the store holds no October initialisation. August only puts all four on the August issuance over October to March, on one construction and every winter including the season ahead.

What the system does and does not do

Lead time. The concentration call becomes available at the August issuance and not before: the June and July runs read −0.22 and −0.07 against the observed burst, August reads +0.73, September +0.44 and October +0.29. A separate reading of the core winter against the catalogue follows in October. The industry windstorm season runs October to March, so the August initialisation lands roughly two months before the season opens and four to seven months before the months that usually carry the damage. The core-winter reading arrives as the season starts. The step from July to August is discontinuous rather than gradual, which is a caveat on thirteen winters as much as a property of the system: the workings show the whole curve.

Seasonal storm totals. Against observed activity the forecast does not beat the climate trend at any lead tested. The trend itself is informative and ours is better than a straight line, but that is the climate engine rather than season picking. It is the open item behind Alarm 3, the exceptional-season call.

Peak severity. Not predicted at any lead, on any window. A single exceptional storm sits outside what a seasonal system can see.

Which stations get hit. Tested directly: same-winter geographic pattern agreement does not beat cross-winter agreement. The map’s claim is breadth, not placement.

The sample. Thirteen winters of predictions, scored against 53 winters of station observations and 75 years of catalogue index. Thirteen is what the hindcast covers and we are not extending it, so the work went into making those thirteen hard to argue with: partial correlations against a no-skill baseline built from the same climatology, permutation tests rather than parametric p-values, leave-one-winter-out refitting on every headline cell, and a robustness check that rebuilds the targets without population weighting. The chain also rests on far more than thirteen years: the observed measures track the catalogue index across 41 winters, and the day-level accuracy behind the instruments is verified over 68 million day-predictions.

Winter 2026/27

Issued from the August 2026 run, on the October–March season. Concentration is the measure the system predicts at this range; the December–February reading against the catalogue index needs the October initialisation and does not exist yet.

Concentration index, Oct 2026 – Mar 2027
−0.73
The lowest of the fourteen readings, just below the thirteen-winter low of −0.57. On the thirteen scored winters, no reading below zero was followed by a top-two concentration season.
Instrument readings
wind −0.59 · rain −1.13 · warmth −0.46
All three below their thirteen-winter means. Wind and rain both read below anything in the scored record (previous lows −0.47 and −0.69); warmth sits inside it. Rain is the most emphatic.
Confidence
Provisional
August is the first issuance with demonstrated skill, so this is the first reading that counts. The June and July runs for 2026/27 read −0.84 and +0.06, but those issuances score −0.11 against the outcome over thirteen winters and are not evidence either way. September and October are the next real information.
instrument2026/27 reading σlowest of the thirteenhighest of the thirteen
wind−0.59−0.47+0.75
rain−1.13−0.69+0.78
warmth−0.46−0.63+0.86

Where in the season the quiet reading sits

The season call is one number, but the forecast is issued month by month. The same August run resolved into all six months of the season, each scored against that month at that lead in the thirteen previous August runs. The eighteen cells below average to the three instrument readings above, which in turn average to the index.

-2σ-1σ+1σ+2σ−0.32−0.48+0.08Oct−0.85−0.76−0.13Nov−0.66−1.64−0.71Dec−0.72−1.23−0.93Jan−0.27−1.39−0.61Feb−0.70−1.30−0.47Marwindrainwarmtheach bar is that month’s 2026/27 reading against the same month at the same lead in the thirteen previous August issuances

Seventeen of the eighteen cells read below their thirteen-winter mean and none reaches −2σ. Most emphatic: rain in December (−1.64σ). The single cell above its mean is warmth in October (+0.08σ), which is close enough to normal to be read as neutral. The measured skill is for the season-shape measure, not for individual months, so the month resolution is context rather than six separate forecasts.

How to use this

What it says. The volume-and-structure risk for the extended season reads below normal on the measure the system predicts best at this range.

What it does not say. Nothing about how many storms will arrive or how severe the worst one will be, and nothing yet about the core winter against the catalogue index: that needs October.

What would change it. The October issuance is a genuinely different observation rather than a refinement: it reads the autumn state that conditions the core winter. In 2019/20 the August run gave no warning of a winter the October wind call went on to place at its record maximum.

Run agreement, in three numbers rather than a table. Splitting the thirteen winters at the median spread between issuances, the agreeing half missed the outcome by 0.87σ and the divergent half by 1.22σ. Spread correlates +0.28 with error, and +0.45 with how far the outcome landed from normal in either direction, so divergence marks a season with something in it rather than a reason to discount the call. For 2026/27 the runs disagree, which is why the reading above is provisional and why the September and October issuances are worth waiting for. The full issuance-by-issuance table is in the workings tab.

What can be said about 2027/28

The same run carries a 24-month horizon, so it does produce readings for October 2027 to March 2028 at leads of 14 to 19 months. They come in uniformly low: every month of every instrument between −1.67σ and −0.49σ. We do not issue that as a view. The uniformity is the reason: a real seasonal signal discriminates between months and between instruments, and this one does not, which is the signature of a run relaxing towards its own climatology. There is also no skill test at that range; the longest lead with a measured result anywhere in this document is four months. The second winter becomes forecastable when its own late-summer initialisation arrives.

Driver context from the long records

Independent of any model output: what the 75-year teleconnection record says about winters like the one ahead. Context for the call, never a call on its own.

The 75-year driver escalation (independent of any model output)

Drivers aligned activewintersstorm countclustered eventsaggregate severity
0117.42.21296
1268.42.21450
22911.63.12150
3914.14.72745

Escalation with alignment (p=0.004, trend-robust; storm count and aggregate severity rise monotonically, clustered events are flat between 0 and 1 aligned drivers then rise): the interplay, not any single index, is the signal. ENSO does not lead for Europe: it is tested and carried as a negative, not used. Caveats: 75-year averages, regime-dependent; in the recent slice the relationship reverses (see the note below); long-record context, never a season-caller.

The reversal in the recent slice. Over 75 years the escalation with driver alignment is orderly, but lately it inverts: 2023/24 produced twelve catalogue events with seven clustered while sitting at the 5th percentile of driver alignment, and 2021/22 sat at the 91st and produced seven. The long-record relationship is real and is why the drivers are carried as context, but it has not held over the last decade, and nothing in the findings tab depends on it.

The 75-year driver anchor – every famous winter placed on the alignment scale

1987/881989/901990/911999/002006/072009/102013/142021/222023/24low alignment -1.32high alignment +1.21
WinterStormsAlignment75-yr percentilez NATIz AMOz PDO
1987/8887J (Oct 1987)*-0.4326th+0.91-0.21+0.59
1989/90Daria, Vivian, Herta+0.1656th+1.00-1.11-0.36
1990/91Undine; run-on season+0.7687th-0.10-1.02-1.17
1999/00Anatol, Lothar, Martin (Dec 1999)+1.2199th-1.29-0.50-1.84
2006/07Kyrill (Jan 2007)-0.3036th+0.91+0.54-0.56
2009/10Xynthia (Feb 2010)-0.858th+1.90+0.55+0.09
2013/14Serial UK winter+0.2964th+0.22-0.95-0.15
2021/22Dudley, Eunice, Franklin+0.9091st-1.34+0.45-1.82
2023/24High-latitude cluster year-0.905th+1.14+2.34-0.79

Grey: the long record; blue: hindcast-era winters; red: major insured-loss winters; amber: index-extreme, modest-loss. One-signed reading: alignment grades serial storm-family winters (99th/91st/87th percentiles) and does not flag single-monster winters – clustering-axis context, selected-by-loss so it motivates, never validates.

Forward teleconnection view

Reserved for the sub-model forward teleconnection projections: the independent driver-based outlook for 2026/27, with its own verification record, to sit alongside the instrument-based view above. To be supplied.
September 2026