Claros European winter windstorm signal – evidence and forward view
MetSwift · Claros European winter windstorm signal · prepared for Aon

What the system predicts, how well, and what it says about the winter ahead

Thirteen validation winters, 2013/14 to 2025/26, scored on the October–March windstorm season, on a network of 3,002 stations across nine north-west European countries drawn from a daily hindcast library of about 5,300 European stations. Predictions are scored against two records built from different data: a station observation archive reaching back to the 1930s, and a 1950–2025 windstorm event catalogue we built from reanalysis. Every skill number is measured against a no-skill baseline made from the same climatology, so what is reported is what the forecast adds. No loss data is used in fitting or scoring anywhere; loss figures appear only as context. Every technical term used in this report is defined in the glossary at the top of the workings tab.

Findings Winter 2026/27 Workings and further evidence

1  What we are trying to predict

Insured windstorm outcomes are not one quantity. Fifty-three winters of station observations separate cleanly into two families, and a winter can be extreme in one while sitting near normal in the other. That is the single most important structural fact in this programme, because it decides how many calls the system needs to make.

Family one: how severe the season was overall

Wind activity, peak severity, storm clustering and station peak wind all move together: they correlate +0.73 to +0.97 with each other over 53 winters. They also carry the strongest link to the independently built catalogue extremity index. This is the axis the industry has historically priced: the 1990-type winter.

Family two: whether the damage arrived together

The worst 7-day burst and clumping correlate +0.79 with each other, but far more loosely with the first family: burst +0.47 to +0.56, clumping +0.26 to +0.37. Against the catalogue index they manage only +0.24 and +0.18. A concentrated fortnight inside an otherwise ordinary season shares some ground with a relentless season, but not enough for either to stand in for the other, and the catalogue does not see it at all.

activityseverityclusteringstation peakburstclumpingcatalogue index
activity+0.97+0.95+0.79+0.53+0.29+0.62
severity+0.97+0.93+0.87+0.56+0.37+0.60
clustering+0.95+0.93+0.73+0.48+0.26+0.52
station peak+0.79+0.87+0.73+0.47+0.31+0.55
burst+0.53+0.56+0.48+0.47+0.79+0.24
clumping+0.29+0.37+0.26+0.31+0.79+0.18
catalogue index+0.62+0.60+0.52+0.55+0.24+0.18

Observed station measures against each other and against the catalogue extremity index, 53 winters (the index column covers 41). The two blocks in the top-left and bottom-right are the two families; the weak corner between them is the point. Green is positive, red negative, intensity is strength.

Definitions: the observed measures, precisely
  • Wind activity. Days per station-winter above that station’s own era-relative 85th-percentile daily threshold, averaged across the network on the population kernel, expressed as σ against the frozen 1972–2011 reference. The canonical “how windy was the season” series.
  • Peak severity. The strongest daily wind each station records all season, network mean. Note the catalogue index uses a daily mean wind while station data carries the daily maximum, so the two severity measures are not like-for-like and are never compared raw.
  • Storm clustering (station pairs). Separate exceedance episodes at the same station arriving within 72 hours of one another, counted per station and then averaged. Temporal clustering, measured locally.
  • Station peak wind. Absolute daily maximum per station, network mean. Sensitive to single storms.
  • Worst 7-day burst. The most intense seven-day window of network-wide exceedance anywhere in the winter. Asks whether the season’s damage piled into one short stretch.
  • Clumping. Variance-to-mean ratio of weekly network exceedance counts across the season. High when the winter arrives in a few concentrated spells rather than evenly.
  • The concentration index, and how to read its scale. The product itself. For each instrument and each of the six months, the population-weighted conviction-day rate is standardised on its own thirteen-winter record, the six months are averaged, and the three instruments are then averaged. So eighteen standardised numbers, each with one equal vote. It is built out of standard deviations but it is not itself on a standard-deviation scale: averaging eighteen partly-correlated standardised numbers shrinks the spread, and the index has a standard deviation of about 0.38 across the thirteen winters against roughly 0.45 for each single instrument. So a reading of −0.73 is about 2.0 index standard deviations below normal, not two thirds of one, and the +0.5 alarm threshold sits at about 1.3. Index readings in this report are quoted as index units and never carry a sigma sign; observed outcomes are quoted in sigma against the frozen 1972–2011 reference and always do.
  • The windstorm catalogue, and the four measures it carries. Our own 1950–2025 windstorm event catalogue, built from reanalysis fields rather than from the station archive. It shares no data with the station instruments or the station outcomes, which is what makes it a genuine cross-check, but it is ours, not a third party’s, and it is not an external validation. Four measures come off it, and they keep the names they were given at the start of the programme:
    • Catalogue events: the number of windstorm events identified in the winter. A count.
    • Catalogue clustering: how many of those events fall within 72 hours of another. The catalogue’s own clustering axis, and the closest independent analogue to our concentration family.
    • Catalogue severity: the strongest single event wind of the winter, in m/s. Daily-mean based, where station data carries the daily maximum, so the two severity measures are never compared raw.
    • Catalogue extremity index: the blended index, quoted as a percentile of the 75-year record. This is the measure the report scores against, and the one referred to elsewhere simply as the catalogue index.
Concentration and clustering are not the same thing: how they differ
  • Clustering asks a local question: at one station, did exceedances arrive in quick succession? It is counted per station, then averaged over three thousand stations and the whole season.
  • Concentration (burst and clumping) asks a seasonal, network-wide question: did the network’s exceedance pile into one short window of the winter?
  • The two can disagree sharply. A winter of many small paired events spread evenly from November to March clusters strongly and concentrates weakly. A winter where one region is hit three times in a fortnight and the rest of the season is quiet concentrates strongly and clusters weakly, because a short run affecting part of the network is diluted by twenty quiet weeks and three thousand stations when it is averaged.
  • 2021/22 is that second case. Station clustering read +0.28σ, essentially normal. Worst 7-day burst read +2.91σ, the second highest of the thirteen behind 2018/19.
  • This document uses concentration throughout for family two and reserves clustering for the station-pair measure.

Why family one is worth predicting: the long-record link

observed station measurer vs the catalogue extremity indexpwinters
activity+0.620.000141
severity+0.600.000041
clustering+0.520.000441
station peak+0.550.000141
burst+0.240.131041
clumping+0.180.260141

Permutation-tested over the winters where both series exist. Activity, severity, station peak and clustering all track the catalogue index at p≤0.0004 across 41 winters: two records built from different data agreeing on which winters were bad. Burst and clumping do not, and that is not a failure: they are measuring the other family.

2  The instruments, and how well we predict them

Claros issues daily exceedance probabilities months ahead for wind, rain and warmth. The seasonal instrument is the count of conviction days: station-days where the forecast probability reaches 1.5× the base rate against that station’s own rebuilt calendar-day threshold. Below, each instrument’s conviction rate against what those same stations actually recorded.

WIND CONVICTION vs OBSERVED WIND
+0.43
thirteen winters, October issuance, DJF, same stations both sides, both axes standardised on the same thirteen winters so the comparison is like for like. p=0.14
RAIN CONVICTION vs OBSERVED RAIN
+0.53
thirteen winters, October issuance, DJF. p=0.06
WARMTH CONVICTION vs OBSERVED WARMTH
+0.44
thirteen winters, October issuance, DJF. p=0.13

Wind

-2σ-1σ+1σ+2σ2013/142014/152015/162016/172017/182018/192019/202020/212021/222022/232023/242024/252025/26bars: forecast conviction rate, October issuance · line: what those same stations actually recorded · one shared σ axis

Rain

-2σ-1σ+1σ+2σ2013/142014/152015/162016/172017/182018/192019/202020/212021/222022/232023/242024/252025/26bars: forecast conviction rate, October issuance · line: what those same stations actually recorded · one shared σ axis

Warmth

-2σ-1σ+1σ+2σ2013/142014/152015/162016/172017/182018/192019/202020/212021/222022/232023/242024/252025/26bars: forecast conviction rate, October issuance · line: what those same stations actually recorded · one shared σ axis

What the instrument layer is worth

At the day level these instruments are thoroughly verified. A flagged day exceeds at about twice the base rate, measured across 68 million day-predictions, and that accuracy holds at every lead tested: 2.00× at one to three months and 1.99× at thirteen to twenty-four. Accuracy that does not decay with lead is unusual in seasonal forecasting and it is the foundation everything above is built on.

At the seasonal level the same instruments point the same way. Wind at +0.43 and rain at +0.53 are respectable correlations for a seasonal signal; with thirteen winters they are indicative rather than conclusive taken one at a time (p=0.14 and p=0.06), and warmth reads +0.44 alongside them.

The value lands one step further on. What these instruments predict best is not how many exceedance days a winter contains but how those days are distributed through it, and there the same three read +0.46 to +0.70 individually, with wind and rain both clearing significance on their own (p=0.012 and p=0.040). Combined into the index they reach +0.73 (p=0.009). The instruments are the mechanism; the measures in section 4 are the product.

Is this predicting windstorms by predicting something else?

It is a fair challenge and it deserves a direct answer, because a reader could reasonably suspect that three instruments and several measures were searched until something correlated.

The three instruments are not three candidate predictors. They are three readings of a single engine: all of the analogue weighting derives from one set of sea-level-pressure pattern relationships. The instruments correlate +0.62 (wind with rain) and +0.81 (wind with warmth), and once wind is in the model rain adds +0.14 (p=0.69) and warmth −0.06 (p=0.85). So there was never a pool of independent predictors to pick from: there is one signal, and the question is only which observable it is best read through. That is a property of how each observable relates to windstorm activity, not a choice we made.

And the measures are not interchangeable either. Section 1 shows they fall into two families with a weak corner between them, established over 53 winters of observations before any forecast was scored. Predicting rain days in order to say something about how a windstorm season is structured is a real inference, and it earns its place only because the structural link between the measures is independently documented and because the result survives permutation, leave-one-out and the removal of population weighting.

What would make it cherry-picking is choosing the season window, the initialisation month or the combination rule after seeing which choice scored best. The workings tab records each of those decisions and the evidence that fixed it, including the one place where the better-scoring option was rejected for exactly this reason.

3  What each dial buys

One line per instrument: the strongest measured link it carries, and what it is therefore for. Each points at a call the system could make; section 6 turns them into alarms.

Wind conviction days. +0.70 against the observed Oct–Mar burst from the August initialisation, p=0.015: the only single instrument that clears significance on the season measure. +0.53 against clumping (p=0.072) points the same way without reaching it. Its historical importance is family one, observed wind activity is the measure that tracks the catalogue index best over 41 winters at +0.62, but at thirteen winters the forecast does not beat the climate trend on activity totals. What it buys today: the strongest single instrument for extended-season concentration. What it is still for: the exceptional-season alarm, once the calibration step in section 6 is done.
Rain conviction days. +0.60 against Oct–Mar burst from August (p=0.063, just short) and +0.5 to +0.6 against the catalogue extremity index from October (p 0.06 to 0.13, on the eleven winters the catalogue covers). What it buys: the best performer in the hindcast on both families, and the instrument to read if conditions continue as they have. Serial storm families drag rain shields across the network before the wind field resolves, which is the mechanism behind the second number, and over 53 winters observed rain days track station clustering at +0.59. What it is not: the observable that defines a severe catalogue winter. Over 41 winters observed rain days manage +0.23 against that index while observed wind activity reaches +0.62.
Warmth conviction days. +0.46 against Oct–Mar burst from August (p=0.130) and +0.55 against clumping (p=0.064). Neither clears on its own, and its own fidelity is the weakest of the three. What it buys: regime context. What it is not: an independent confirmation: the three instrument correlates +0.81 with wind, and once wind is in the model it adds −0.29 (p=0.32). Treat the three as readings of one signal rather than three votes.
The combined concentration index. +0.73 against observed Oct–Mar burst, p=0.009, from the August initialisation. What it buys: the product for family two, and the only construction here that clears significance on both concentration measures (+0.73 burst, +0.63 clumping). The two winters it flagged were the two most concentrated of the thirteen, averaging +3.04σ observed. What it is for: the concentration alarm, and the forward view.

4  The results

Family two: season concentration, October–March, from the August initialisation

Claros runs monthly, so the first question is which run to use. Against the concentration outcome the late-summer runs are the usable ones: August +0.73 and September +0.44, with the two combined reading +0.73. June and July are both slightly negative at −0.11. Bootstrapping the differences over thirteen winters, August cannot be separated from September (90% interval on the gap −0.17 to +0.80; September comes out ahead in 12% of resamples) but both can be separated from July (July beats August in 2.6% of resamples). There is no known meteorological barrier between August and September either, so the honest reading is that the late-summer window carries the signal and the split between its two runs is inside the noise. The report quotes August because it is the earlier of the two and therefore the more useful, not because it is measurably better. This also settles a coordination question: the teleconnection sub-model is validated on September, and on this evidence that is an equally defensible choice.

measureinstrumentpartial rpleave-one-winter-out range
Oct–Mar burstwind+0.700.012+0.59 to +0.78
rain+0.600.040+0.45 to +0.71
warmth+0.460.130+0.33 to +0.54
the index+0.730.009+0.63 to +0.84
Oct–Mar clumpingwind+0.540.070+0.33 to +0.83
rain+0.530.076+0.41 to +0.65
warmth+0.550.064+0.43 to +0.81
the index+0.630.029+0.47 to +0.81

Partial correlations with the no-skill baseline controlled out, thirteen winters, full station network, scored across the whole October–March season on both sides. The product is the index, and it clears significance on both concentration measures: +0.73 against the worst 7-day burst (p=0.009) and +0.63 against clumping (p=0.029). Among the single instruments wind and rain clear on burst (+0.70, p=0.012 and +0.60, p=0.040); warmth (+0.46) points the same way without reaching it on its own, which is what you would expect from three readings of one model run rather than three independent tests. The leave-one-winter-out range recomputes each cell thirteen times, dropping a different winter, and every concentration cell stays positive.

The weighting is not where the result comes from. The instruments are population-weighted, and the outcome can be built either the same way or as a plain station average that treats every station alike. Rebuilding the October–March outcome unweighted and rescoring the same index: burst holds: +0.71 weighted against +0.68 unweighted, still clearing significance (p=0.017). Clumping does not: +0.61 weighted falls to +0.47 unweighted (p=0.127). So the burst result is a property of the forecast and survives however the network is averaged, while the clumping result depends on weighting towards where people live. That is why burst is the headline measure and clumping is labelled the weaker of the two throughout.
-2σ-1σ+1σ+2σ+3σ2013/142014/152015/162016/172017/182018/192019/202020/212021/222022/232023/242024/252025/262026/27bars: the concentration index from the August initialisation · line: observed worst 7-day burst · green = flagged, orange = the season ahead

The August index against the observed worst 7-day burst, October–March, thirteen winters, r=+0.73 (p=0.009). Its two highest readings are the two most concentrated winters of the record. The orange bar is the season ahead, which reads the lowest of the fourteen.

The record, scored in both directions. Two winters were flagged at or above +0.5 on the index, and both were the two most concentrated of the thirteen: observed burst +3.18σ and +2.91σ, averaging +3.04σ. No false alarm. Going the other way, five winters delivered observed burst at or above +1.0σ: the index caught the two largest and read the other three between −0.25 and −0.02, close to normal rather than wrong-signed. So the index discriminates the extremes well and the middle poorly: it separates a +3σ volley winter from an ordinary one, and it does not rank a +1σ winter above an average one. For a capital-planning use that is the right way round; for anything needing a full ranking it is not.

2021/22: the case that shows why two families need two calls

Dudley, Eunice and Franklin arrived between 16 and 21 February 2022: three damaging storms in six days, EUR 3.85bn of industry loss. All three fell inside the season, so this is not a case of looking at the wrong months.

On family one the winter was unremarkable: activity +0.30σ, peak severity +0.45σ, station clustering +0.28σ, and the catalogue index put it at the 23.6th percentile of 75 years. On family two it was extreme: worst 7-day burst +2.91σ and clumping +2.18σ across October–March, the second highest burst of the thirteen behind 2018/19.

The forecast side split the same way. The October wind call read −0.81σ, so nothing in the family-one instruments saw it coming. The August concentration index read +0.56, its second-highest reading of the thirteen: six months before the fortnight.

And the week itself can be located from the station data alone. Asked which seven days carried most of the network’s exceedance that winter, with no knowledge of named storms, the answer is 16 to 22 February 2022, carrying 4.04 times the exceedance an evenly spread season would put in a week. Section 7 maps which stations took that hit: a corridor from Ireland through to Denmark, with Scandinavia untouched.

Family one: winter extremity, December–February, from the October initialisation

The concentration call is issued in August, four to seven months ahead, and it answers a broad question about the shape of the season. The October initialisation answers a narrower and more valuable one: how extreme the core of the winter will be against the catalogue. It arrives later, it is a sharper reading, and it is scored against the one record built from entirely different data to the station archive.

days pooled across the three monthseach month standardised first
instrumentpartial rppartial rp
wind+0.380.277+0.440.212
rain+0.660.035+0.570.090
warmth+0.390.262+0.490.153
the index+0.530.114+0.560.096

October initialisation, December–February, against the catalogue extremity index, over the eleven winters the catalogue covers: the prediction record runs to thirteen winters, but 2022/23 and 2025/26 are not yet indexed, so every catalogue figure in this report is scored on eleven. Two ways of combining the three months are shown because both are defensible and they agree closely: the pooled and standardised versions of each instrument correlate +0.89 to +0.94 with one another, so the gap between the two columns is sampling noise on eleven winters, not a choice that changes the answer. Read rain as +0.5 to +0.6 with p between 0.06 and 0.13, and the index as +0.46 to +0.47, p 0.17 to 0.18, jackknife +0.20 to +0.80.

On this family the composite does not help, and that is worth stating rather than hiding. Combining the three instruments is what makes the concentration call work: on October–March burst the index reads +0.73 against the best single instrument’s +0.70, so the three add information to each other. Here the opposite happens. Rain carries the relationship on its own at +0.5 to +0.6, and averaging wind and warmth into it pulls the composite down to +0.46 to +0.47, because on this target those two are close to noise (+0.31 to +0.39) and the average dilutes rather than reinforces. So the family-one call is rain, not the index, and the index column is in the record table below so the dilution is visible rather than asserted. The same three instruments combine usefully for one family and not the other, which is a statement about the two families rather than about the instruments.

Rain for the winters we have been having; wind for the winters the record contains

Rain is the best performer in the hindcast. Across the prediction record it is the strongest single instrument against the catalogue index at +0.5 to +0.6, and it is also the strongest against extended-season concentration. If conditions continue broadly as they have through the validation decade, rain is the instrument to read. It is not quite significant on the eleven winters the catalogue covers, and two or three more October runs would settle that.

Wind is what the long record says a severe winter is made of. Over 41 winters, observed wind activity tracks the catalogue index at +0.62 (p=0.0001). Observed warm days manage +0.34 and observed rain days +0.23. Chained through what we can actually forecast, wind offers about +0.28 to the catalogue (fidelity +0.45 against a long-record link of +0.62) where rain offers about +0.09. So the two records disagree about which observable matters, and the disagreement carries information: rain measures better over thirteen recent winters, wind measures better over fifty-three.

Which is why both belong in the alarm set. The validation decade has been a period of moderate windstorm activity by historical standards. Rain is calibrated to that period and performs well in it. The winters that defined the market’s view of this peril, 1989/90 and 1994/95, are wind-activity extremes of a kind the prediction record does not contain, and wind is the instrument tied to them. Reading only rain risks being well tuned to a regime that changes; reading only wind discards the better recent performer. Section 6 sets out the alarms that follow, and the one that is not yet buildable.

5  The thirteen-winter record

The record is split by issuance, because the two issuances predict different families and one table invites the wrong comparison. The first row of toggles picks the issuance; the second picks the geography, with pooled Europe, the product, as the default. Click any observed column for what predicted it and the comparison chart.

Signals are standardised on the thirteen-winter issuance record; observed columns are σ against the frozen 1972–2011 station reference. On weighting: every observed column drawn from station data is a network mean on the same population kernel as the signal, so forecast and outcome are weighted identically. The only unweighted column is the catalogue extremity percentile. That percentile is missing 2022/23 and does not yet reach 2025/26, so those two cells are blank, which is why every catalogue figure in this report is scored on eleven winters rather than thirteen. This table is the source. Every alarm threshold in section 6 is computed from exactly these columns, so any figure in the sweep tables can be reproduced from the numbers on this screen with nothing more than a spreadsheet.

On the composite columns. Both issuances carry the three instruments and the composite built from them, and the composite competes for the best-predictor highlight like any other column, because the report makes claims about it. On the August issuance the composite is the concentration index, the product for family two. On the October issuance it is the family-one composite, which is shown precisely because it loses to rain: section 4 sets out why combining helps on one family and dilutes on the other, and this is where that is visible in the data. One thing to expect. The best-predictor figure is the plain correlation of the two columns as printed, so it can be checked against this table directly. The findings quote partial correlations, with the no-skill climatology removed from both sides. The two occasionally rank differently, most visibly on October–March burst, where the two are close: wind 0.71 against the index 0.70 on the plain numbers, and the index +0.73 against wind +0.70 once each is scored against its own Floor. Neither gap is meaningful on thirteen winters. The panel says which measure it is reporting whenever a published figure exists for the column.

6  Alarms

Two families, and two different kinds of severe winter inside the first one, mean the system needs three alarms rather than one. The 53-winter observed record shows why with unusual clarity, and it also shows which of the three we can currently build.

The observed ladder, and what it cannot see

winteractivity σ7-day burst σcatalogue extremity pctseverity alarm would fireconcentration alarm would fire
1973/74+0.76+0.9240.0Tier 1:
1982/83+0.71+0.5572.7Tier 1:
1983/84+0.77+1.6890.9Tier 1fires
1989/90+1.61+2.0995.5Tier 2fires
1993/94+0.68+0.5181.8Tier 1:
1994/95+1.29+0.1495.5Tier 2:
2001/02+0.94+1.0470.0Tier 1:
2006/07+0.91+0.5967.3Tier 1:
2013/14+0.92+0.2261.8Tier 1:
2015/16+0.66+1.4276.4Tier 1:
2018/19−0.45+3.0525.5: fires
2019/20+1.23+1.2189.1Tier 1:
2021/22+0.30+2.8123.6: fires

Every winter in 53 that either the severity ladder or the concentration threshold would flag, with the two side by side. Severity alarm: Tier 1 is observed activity at or above +0.65σ, Tier 2 is activity and station clustering both at or above +1.25σ, on the frozen 1972–2011 scale: 11 Tier 1 firings and 2 Tier 2 across 53 winters. Concentration alarm: observed worst 7-day burst at or above +1.5σ. These are the observed definitions, which is what makes the comparison possible over 53 winters; the prediction-side thresholds are a separate job, noted below. One note on the scale. Every observed σ quoted in this report sits on a frozen pre-2012 reference, so that winters are compared against a baseline that does not include them. What changes between exhibits is the season window: this ladder is built on November to March, the record table in section 5 on October to March. The same winter therefore carries two readings, both frozen: 2018/19’s burst is +3.05σ here and +3.18σ there. Neither is more correct, and the report does not mix the two windows within a table.

The ladder works on the family it was built for. Its two Tier 2 winters, 1989/90 and 1994/95, sit at the 95.5th percentile of the catalogue index. Its Tier 1 winters run from the 40th to the 90.9th, mostly high. This is the alarm for the winter that grinds: the type the market priced through the 1980s and 1990s.
And it is blind to the other family. The two most concentrated winters in the record, 2018/19 (burst +3.05σ) and 2021/22 (burst +2.81σ), fire neither tier of the severity ladder. Their activity readings are −0.45σ and +0.30σ, and the catalogue index puts them at the 25.5th and 23.6th percentiles. They are also the two most recent extremes in the series. A severity ladder on its own would have called both winters quiet, which is the whole case for a second alarm.

What the architecture should therefore be: three alarms, at three stages of maturity

Two families and two mechanisms inside family one, the winter that is bad for the current climate, and the winter that is exceptional against any climate, make three alarms the natural design. They are not equally ready, and the difference matters more than the design.

Alarm 1: concentration evidenced

August concentration index → observed Oct–Mar burst, +0.73, p=0.009, and → clumping +0.63, p=0.029. Every cell survives dropping any single winter. Fires for the family-two winter from two months before the season opens. This is the alarm that would have caught 2021/22 and 2018/19, the two most concentrated winters of the record.

Alarm 2: winter extremity, Dec–Feb developing

October conviction → catalogue extremity index. Rain reads +0.5 to +0.6, p between 0.06 and 0.13. This is the winter that is bad relative to today’s climate, and it is the strongest family-one relationship anywhere in the programme: the right size for a useful alarm, on a sample too short to establish it. Two more October issuances would settle it either way.

Alarm 3: the exceptional season next step

The 1989/90 and 1994/95 type: the grinding winter the observed activity ladder catches at the 95.5th percentile of the catalogue index. On the observed side this is the best-defined alarm in the programme. The prediction side needs a step the other two alarms do not: the signal has to be scaled against the observed variability of the target before a threshold means anything, because the placed calls are conservative in spread by construction. That calibration is the next piece of work rather than a closed question.

What the search for the exceptional season found, and what it did not settle. Wind is the instrument that ought to reach this family: observed wind activity is the measure that tracks the catalogue index best over 41 winters, at +0.62. So we looked properly: three instruments × four issuances (July to October) × two forecast windows × five family-one targets (the activity/severity/clustering composite, observed activity alone, the catalogue extremity index, catalogue event count, catalogue clustering): 120 cells, Floor-controlled, permutation-tested. No raw cell clears p<0.05 with a positive sign. The strongest is the rain-to-index cell behind Alarm 2 at +0.60 (p=0.06); the best wind cell is +0.44 (p=0.15). Searching 120 cells is also the reason a lone p<0.05 would have been suspect: finding none is a cleaner result than finding one. What this establishes is that no untransformed signal reads this family directly on thirteen winters. It does not settle the calibrated route, which scales the signal against the target’s observed variability rather than correlating the two raw series, and which is the work Alarm 3 is waiting on.
What the two working alarms deliver when they are run together. A winter counts as one to call if it delivered either an Oct–Mar burst at or above +1.0σ or a catalogue extremity index at or above the 60th percentile. Firing either alarm at a +0.5 reading, across the thirteen winters: three hits, two false alarms, five correct quiet calls, three misses. The hits are 2018/19 and 2021/22, the two most concentrated winters of the record, caught by concentration; and 2015/16, caught on winter extremity off an October rain call of +0.68. The false alarms are 2017/18, where rain read +1.19 against a catalogue at the 35.5th percentile, and 2023/24, where the index read +0.52 against a burst of +0.49 and the catalogue at the 52.7th. Of the three misses, 2013/14 was marginal on every measure at the 61.8th percentile, 2014/15 delivered a +1.40σ burst with no instrument above +0.5, and 2019/20 reached the 89.1st catalogue percentile with October rain at +0.35, just short of firing. Three of those six calls sit within 0.06 of the threshold, which is the calibration problem in section 6 rather than a skill result. The correlations are the stable quantity on thirteen winters; a firing level read off them is not, and shifting the threshold by a twentieth of a sigma reorders half the tally. What the period did not contain is the 1989/90 type, which is why the third alarm is scoped from the long observed record rather than from this sample.

Where the thresholds would sit

For the two alarms that have a route, every threshold scored. Alarm 3 is absent because the calibration that would set its threshold has not been run.

August concentration index → observed Oct–Mar burst ≥ +1.0σ

thresholdfirescaughtfalse alarmsmissedquiet & rightcatch rate
+0.006421680%
+0.203213740%
+0.353213740%
+0.503213740%
+0.65000580%
+0.80000580%
+1.00000580%

Thirteen winters. The concentration alarm has no false alarm at all from +0.5 upward, at the cost of missing the marginal winters: it catches 2 of the 5 winters that delivered a +1.0σ burst, and both of the two that delivered a +2.9σ one.

October rain → catalogue extremity index ≥ 60th pct

thresholdfirescaughtfalse alarmsmissedquiet & rightcatch rate
+0.0073404100%
+0.205231567%
+0.355231567%
+0.503122633%
+0.652112733%
+0.80101370%
+1.00101370%

Eleven winters, because the catalogue does not yet index 2022/23 or 2025/26. The rain alarm is flat across most of its range: 2 of 3 caught with 2 false alarms at every threshold from +0.20 to +0.65. A flat response over that span is what a real but imprecise signal looks like on eleven winters.

October wind → observed Oct–Mar activity ≥ +0.5σ

thresholdfirescaughtfalse alarmsmissedquiet & rightcatch rate
+0.006331675%
+0.205232650%
+0.354133625%
+0.50202470%
+0.65101480%
+0.80101480%
+1.00101480%

Shown for completeness: the wind-to-activity alarm catches 2 of 4 with 2 false alarms at the two lowest thresholds and 1 of 4 above them, which is why it is not proposed in this form. This is the cell the calibration work behind Alarm 3 has to improve on.

The calibration step, which is what stands between these tables and an operational alarm. An imperfect forecast is always narrower in spread than the thing it predicts: a signal that correlates +0.7 with an outcome moves about 70% as far as the outcome does. We leave that compression in rather than inflating the forecast back out to match, because inflating it would make every reading look more confident than the evidence supports. The consequence is that a σ on the signal side and a σ on the observed side are not the same unit, so the thresholds above, which are read off the observed distributions, cannot be transferred to the signal as they stand. The work outstanding is to derive signal-side thresholds from the joint distribution of signal and outcome, with stated detection and false-alarm rates and a deliberately asymmetric objective, since for a capital-planning use a miss costs materially more than a false alarm. That study is the same one Alarm 3 waits on, and it has not been run. Until it is, read the thresholds here as the shape of the trade-off rather than the alarm specification.

7  The map

When a winter is loaded, the predicted instruments run red together and the observed dials verify together. The verified quantity is how much of the map runs red, winter by winter: not which individual station is hit, which the system does not claim.

Interactive: predicted instruments on the left, observed dials on the right; any dial can be read alone or combined with any others, and the winters step with the chips, the arrow keys or Play. Strongest signal shows the deepest reading any active dial gives each station, which is how an alarm set behaves; Averaged draws the per-station mean instead, which is the aggregation a composite actually uses. Coverage: all thirteen scored winters plus the season ahead. Matched runs puts each dial on the issuance that suits its own target, wind, rain and warmth on the October run over December to February and the index on the August run over October to March. The October station arrays cover the thirteen scored winters; they grey out for 2026/27, for which the store holds no October initialisation. August only puts all four on the August issuance over October to March, on one construction and every winter including the season ahead.

The networks in play, and they are not the same on both sides. Every result in this report pairs a forecast network with an observed one, and the two are deliberately different. The forecast side is the selection of methodology 2.3, the top stations by population weight to 85% of cumulative weight per country: 1,041 stations for wind, 1,033 for rain and 1,039 for warmth, the difference being stations absent from one variable’s archive. The Floor is read at the same stations as the signal it controls. The observed side is the fixed panel of 3,002 stations, which is what carries the long record back to 1973/74 and what every observed sigma is measured against; the rain and warmth station dials sit on the sub-panels that hold those variables, about 1,050 and 1,045 stations. Both sides are population-weighted on the same kernel, so a forecast and its outcome are weighted identically even though the panels differ.

Two exhibits use different networks again, because the question they answer is different. The regional panel draws every station in each region on both sides, since it asks whether the signal works region by region rather than what the population-weighted number is. The wider-Europe test uses a latitude-stratified equal-weighted sample outside the core nine, so no population kernel is assumed where the product does not operate. The fidelity dials are the one place where both sides sit on the forecast selection, because there the question is whether the forecast and the observation agree at the same stations. Each is labelled at the point of use.

The concentration index, station by station

The fourth prediction dial is the index itself, on the same October–March season and the same August issuance as the number in section 4. Not a proxy and not a shorter window: the same quantity, taken apart.

The mean of the map is the number. The index is a population-weighted network average, so it does not decompose into per-station index readings, but it does decompose exactly into per-station contributions. Each station’s contribution is its own weighted share of the network exceedance rate, measured against its own normal, in index units. Those contributions sum to the index by construction. The map shows each station’s contribution scored against its own thirteen-winter record, so red means a station pushed the reading up harder than it usually does, and the panel prints the weighted mean beside it. A station absent from one instrument contributes nothing to that instrument’s network rate, so its contribution there is zero rather than missing, and with that convention the contributions sum to the index exactly: the largest gap across all fourteen winters is 0.004.

Two things to expect when it is selected. The lead label changes to August issue, Oct to Mar, leads 2 to 7 months, and shows both entries if the selection mixes issuances with the October instruments. And the map draws every station in the nine-country network, 2,555 of them for the index, so it is no sparser than the other dials. Three station sets are in play here and they should not be confused. The index number quoted throughout this report is the population-weighted average over the 1,041-station selection of section 2.3, which is 85% of population weight per country. The exact decomposition spans the 1,178 stations that appear in at least one instrument’s selection, and their contributions sum to the published index to within 0.004 across all fourteen winters. The map draws the whole nine-country network, 2,555 stations for the index, because a thousand-station picture is not a picture of Europe; its whole-network mean sits about 0.05 from the published reading and ranks the thirteen winters with it at +0.99, so the map can be read on the whole network without misrepresenting the number.

Where the volley landed: burst participation

Every other dial on that map is a season total. Burst participation is not, and it is the only one that shows the measure the system actually predicts as a picture rather than a number.

How it is built. For each winter, find the seven-day window in which the network as a whole recorded most of its exceedance. Then ask, of each station separately, what share of that station’s whole-season exceedance fell inside those seven days, and score that share against the same station’s own 1974–2011 record. Red means this station took an unusually concentrated hit in that one week. It sits on the same frozen reference as the other observed dials, so it composites with them like any other.

The method locates the storms without being told about them. Nothing in the construction knows about named storms, catalogue events or loss. Run on 2021/22, the window it returns is 16 to 22 February 2022: Dudley, Eunice and Franklin, which arrived between the 16th and the 21st. That week carried 4.04 times the exceedance an evenly spread season would put in seven days, the most concentrated week of the thirteen. On 2018/19 it returns 9 to 15 March 2019 at 3.90 times. Those are the two winters the August concentration index flagged, and this is the third independent construction to pick them out.

What the picture adds. Load 2021/22 with burst participation alone. It runs deep red across Ireland, Britain, the Low Countries, northern Germany and Denmark, and blue across Scandinavia, while the season-total dials beside it read close to normal almost everywhere. That is the difference between the two families rendered geographically: a corridor took most of a winter’s damage in six days, and no season total sees it. It also shows what the system does not claim. The forecast predicts that a winter will be concentrated, not where the corridor will fall, and the map is the honest way to show both at once.

Over the full 53-winter record 2021/22 ranks 4th and 2018/19 6th on the network mean of this measure, so neither is unprecedented: the validation decade contained two strongly concentrated winters, not the two worst on record. Across the thirteen scored winters the participation ratio correlates +0.81 with the worst 7-day burst the report predicts, which is the check that the map and the headline measure are two views of one thing rather than two measures.

8  What the system does and does not do

Lead time. The concentration call becomes available at the August issuance and not before: the June and July runs read −0.22 and −0.07 against the observed burst, August reads +0.73, September +0.44 and October +0.29. A separate reading of the core winter against the catalogue follows in October. Aon’s windstorm season runs October to March, so the August initialisation lands roughly two months before the season opens and four to seven months before the months that usually carry the damage. The core-winter reading arrives as the season starts. The step from July to August is discontinuous rather than gradual, which is a caveat on thirteen winters as much as a property of the system: the workings show the whole curve.

Seasonal storm totals. Against observed activity the forecast does not beat the climate trend at any lead tested. The trend itself is informative and ours is better than a straight line, but that is the climate engine rather than season picking. It is the open item behind Alarm 3, the exceptional-season call.

Peak severity. Not predicted at any lead, on any window. A single exceptional storm sits outside what a seasonal system can see.

Which stations get hit. Tested directly: same-winter geographic pattern agreement does not beat cross-winter agreement. The map’s claim is breadth, not placement.

The sample. Thirteen winters of predictions, scored against 53 winters of station observations and 75 years of catalogue index. Thirteen is what the hindcast covers and we are not extending it, so the work went into making those thirteen hard to argue with: partial correlations against a no-skill baseline built from the same climatology, permutation tests rather than parametric p-values, leave-one-winter-out refitting on every headline cell, and a robustness check that rebuilds the targets without population weighting. The chain also rests on far more than thirteen years: the observed measures track the catalogue index across 41 winters, and the day-level accuracy behind the instruments is verified over 68 million day-predictions.

Winter 2026/27

Issued from the August 2026 run, on the October–March season. Concentration is the measure the system predicts at this range; the December–February reading against the catalogue index needs the October initialisation and does not exist yet.

Concentration index, Oct 2026 – Mar 2027
−0.73
The lowest of the fourteen readings, just below the thirteen-winter low of −0.57. On the thirteen scored winters, no reading below zero was followed by a top-two concentration season.
Instrument readings
wind −0.59 · rain −1.13 · warmth −0.46
All three below their thirteen-winter means. Wind and rain both read below anything in the scored record (previous lows −0.47 and −0.69); warmth sits inside it. Rain is the most emphatic.
Confidence
Provisional
August is the first issuance with demonstrated skill, so this is the first reading that counts. The June and July runs for 2026/27 read −0.84 and +0.06, but those issuances score −0.11 against the outcome over thirteen winters and are not evidence either way. September and October are the next real information.
instrument2026/27 reading σlowest of the thirteenhighest of the thirteen
wind−0.59−0.47+0.75
rain−1.13−0.69+0.78
warmth−0.46−0.63+0.86

Where in the season the quiet reading sits

The season call is one number, but the forecast is issued month by month. The same August run resolved into all six months of the season, each scored against that month at that lead in the thirteen previous August runs. The eighteen cells below average to the three instrument readings above, which in turn average to the index.

-2σ-1σ+1σ+2σ−0.32−0.48+0.08Oct−0.85−0.76−0.13Nov−0.66−1.64−0.71Dec−0.72−1.23−0.93Jan−0.27−1.39−0.61Feb−0.70−1.30−0.47Marwindrainwarmtheach bar is that month’s 2026/27 reading against the same month at the same lead in the thirteen previous August issuances

Seventeen of the eighteen cells read below their thirteen-winter mean and none reaches −2σ. Most emphatic: rain in December (−1.64σ). The single cell above its mean is warmth in October (+0.08σ), which is close enough to normal to be read as neutral. The measured skill is for the season-shape measure, not for individual months, so the month resolution is context rather than six separate forecasts.

How to use this

What it says. The volume-and-structure risk for the extended season reads below normal on the measure the system predicts best at this range.

What it does not say. Nothing about how many storms will arrive or how severe the worst one will be, and nothing yet about the core winter against the catalogue index: that needs October.

What would change it. The October issuance is a genuinely different observation rather than a refinement: it reads the autumn state that conditions the core winter. In 2019/20 the August run gave no warning of a winter the October wind call went on to place at its record maximum.

Run agreement, in three numbers rather than a table. Splitting the thirteen winters at the median spread between issuances, the agreeing half missed the outcome by 0.87σ and the divergent half by 1.22σ. Spread correlates +0.28 with error, and +0.45 with how far the outcome landed from normal in either direction, so divergence marks a season with something in it rather than a reason to discount the call. For 2026/27 the runs disagree, which is why the reading above is provisional and why the September and October issuances are worth waiting for. The full issuance-by-issuance table is in the workings tab.

What can be said about 2027/28

The same run carries a 24-month horizon, so it does produce readings for October 2027 to March 2028 at leads of 14 to 19 months. They come in uniformly low: every month of every instrument between −1.67σ and −0.49σ. We do not issue that as a view. The uniformity is the reason: a real seasonal signal discriminates between months and between instruments, and this one does not, which is the signature of a run relaxing towards its own climatology. There is also no skill test at that range; the longest lead with a measured result anywhere in this document is four months. The second winter becomes forecastable when its own late-summer initialisation arrives.

Driver context from the long records

Independent of any model output: what the 75-year teleconnection record says about winters like the one ahead. Context for the call, never a call on its own.

The 75-year driver escalation (independent of any model output)

Drivers aligned activewintersstorm countclustered eventsaggregate severity
0117.42.21296
1268.42.21450
22911.63.12150
3914.14.72745

Escalation with alignment (p=0.004, trend-robust; storm count and aggregate severity rise monotonically, clustered events are flat between 0 and 1 aligned drivers then rise): the interplay, not any single index, is the signal. ENSO does not lead for Europe: it is tested and carried as a negative, not used. Caveats: 75-year averages, regime-dependent; in the recent slice the relationship reverses (see the note below); long-record context, never a season-caller.

The reversal in the recent slice. Over 75 years the escalation with driver alignment is orderly, but lately it inverts: 2023/24 produced twelve catalogue events with seven clustered while sitting at the 5th percentile of driver alignment, and 2021/22 sat at the 91st and produced seven. The long-record relationship is real and is why the drivers are carried as context, but it has not held over the last decade, and nothing in the findings tab depends on it.

The 75-year driver anchor – every famous winter placed on the alignment scale

1987/881989/901990/911999/002006/072009/102013/142021/222023/24low alignment -1.32high alignment +1.21
WinterStormsAlignment75-yr percentilez NATIz AMOz PDO
1987/8887J (Oct 1987)*-0.4326th+0.91-0.21+0.59
1989/90Daria, Vivian, Herta+0.1656th+1.00-1.11-0.36
1990/91Undine; run-on season+0.7687th-0.10-1.02-1.17
1999/00Anatol, Lothar, Martin (Dec 1999)+1.2199th-1.29-0.50-1.84
2006/07Kyrill (Jan 2007)-0.3036th+0.91+0.54-0.56
2009/10Xynthia (Feb 2010)-0.858th+1.90+0.55+0.09
2013/14Serial UK winter+0.2964th+0.22-0.95-0.15
2021/22Dudley, Eunice, Franklin+0.9091st-1.34+0.45-1.82
2023/24High-latitude cluster year-0.905th+1.14+2.34-0.79

Grey: the long record; blue: hindcast-era winters; red: major insured-loss winters; amber: index-extreme, modest-loss. One-signed reading: alignment grades serial storm-family winters (99th/91st/87th percentiles) and does not flag single-monster winters – clustering-axis context, selected-by-loss so it motivates, never validates.

Forward teleconnection view

Reserved for the sub-model forward teleconnection projections: the independent driver-based outlook for 2026/27, with its own verification record, to sit alongside the instrument-based view above. To be supplied.

Workings and further evidence

Everything that supports the findings without being one: the glossary, the lead-time curve, the window tests, the reference comparison, the geography breakdowns, the run-by-run record and the season-by-season roster.

Glossary: every term this report leans on

Conviction day. A station-day on which Claros put the probability of exceedance at 1.5 times or more the station’s nominal base rate for that calendar day. Wind uses a Q85 threshold (base rate 15%, so the gate is 22.5%); rain and warmth use Q90 (base rate 10%, gate 15%). Realised winter exceedance runs nearer 18%, so in practice the wind gate is about 1.25 times the realised rate rather than 1.5.

Conviction rate. The share of station-days in a month that were conviction days, weighted by population as described below. This is the raw quantity every instrument is built from.

Population weighting. Each station carries a weight from the population living near it, so the network average leans towards where people and property are rather than treating a Highland gauge and a Rotterdam gauge as equals. It is a proxy for where damage concentrates, not client exposure data, and it is a proxy we test against: every headline result is rebuilt as a plain unweighted station average, and section 5 reports what moves.

Detrended calendar-day threshold. Each station’s exceedance threshold is computed on its own record with the climate trend removed, so a warming or windier baseline does not itself generate conviction days. This is what makes a conviction day a statement about this winter rather than about the decade.

The no-skill baseline (the Floor). For every forecast quantity we build the same quantity from climatology alone, with no forecast information in it. Every correlation in this report is a partial correlation: the Floor is regressed out of both the signal and the outcome first, and what is reported is the relationship that survives. So a number here is what the forecast adds over knowing the climatology, not the total predictability of the outcome.

Month standardisation. Each calendar month’s conviction rate is converted to a z-score against that same month at that same lead across the thirteen scored winters, and the months are then averaged. Without it, a month with naturally large swings dominates the season average and the quieter months contribute nothing. Section 5.5 of the methodology sets out the test that decides when it is required.

Instrument composite, and the concentration index. An instrument composite is the six standardised months of one variable averaged into one number for the season. The concentration index is the three instrument composites (wind, rain, warmth) averaged. Note the scale: it is built out of standard deviations but is not itself on a σ scale, because averaging correlated-but-not-identical components shrinks the spread. Its own standard deviation across the thirteen winters is about 0.38, so a reading of −0.73 is roughly 2.0 of its own standard deviations below normal.

Signal compression. Any imperfect forecast is narrower in spread than the outcome it predicts. We leave that compression in rather than stretching the forecast back out, which is why signal-side and observed-side σ are not interchangeable and why the alarm thresholds need their own calibration step.

Worst 7-day burst. The most exceedance-heavy seven-day window anywhere in the season, network-wide, expressed as a z-score. A seasonal maximum, so it answers “how bad was the worst fortnight” rather than “how bad was the season”.

Clumping. How unevenly the season’s exceedance was spread through time, network-wide. High clumping means the season’s activity arrived in a few concentrated spells rather than steadily.

Station clustering. A different question from the two above, asked locally: at one station, did exceedances arrive in quick succession? Counted per station, then averaged. A season can cluster strongly and concentrate weakly, and the reverse.

Station composite. Per station-winter, the highest of that station’s three z-scores for wind activity, storm clustering and peak severity, then averaged across the network: how bad the winter was at each station on whichever axis hit it hardest. Used for description and for the map, never as a prediction target, because it is built from the same station data as the instruments.

The catalogue. A 1950–2025 windstorm event catalogue we built from reanalysis. It is ours, not a vendor product, and it shares no data with the station archive, which is what makes it a genuinely separate scoring record. The catalogue extremity index ranks each winter against the other 75; catalogue events is the raw count; catalogue clustering counts events falling within 72 hours of another; catalogue severity is the strongest single event wind that winter, on a daily-mean basis.

The frozen 1972–2011 reference. A fixed forty-winter baseline used for every observed measure quoted as a σ, so that winters are compared against a common scale instead of each being measured against a moving average that includes itself. Burst participation uses 1974–2011 and the network indices use winters to 2011; all three end before the Claros hindcast library begins, so no scored winter contributes to the yardstick it is scored against. Claros readings are standardised on the 2013–2025 hindcast record, the only history the model has. The fidelity dials in section 3 are the one exhibit where the observed side is standardised on the thirteen winters instead, so that both axes of that chart share a footing; a correlation is unaffected by the location and scale of either variable, so the reading is the same either way. The one thing that varies between observed exhibits is the season window, October to March or November to March, and the report labels which and does not mix them inside a table.

Placed call. A forecast reading for a specific season and a specific geography, as it was issued, with no later adjustment. Every forward number in this report is a placed call.

Permutation test. Rather than assuming a distribution, we shuffle the signal series thousands of times and count how often chance reproduces the observed relationship. The signal and its Floor baseline are shuffled together, using the same index order, because shuffling the signal alone breaks their pairing and overstates significance.

Jackknife. Refitting the same relationship thirteen times, each time dropping one winter, and reporting the range. A result that depends on one winter shows up as a range that crosses zero.

Driver alignment, NATI, AMO, PDO. Teleconnection context from the long records, not model output. NATI is a North Atlantic circulation index; AMO and PDO are the Atlantic and Pacific multidecadal ocean modes. Driver alignment is a percentile ranking of how favourably these stood for European windstorm activity in a given winter. Carried as context only: nothing in the findings tab depends on it.

Aggregate severity. In the map panel, the sum of event severity across a winter rather than its maximum: a measure of total rather than peak.

When the concentration call arrives

0.00.20.40.60.8−0.22June−0.07July+0.73August+0.44September+0.30Octoberpartial correlation of each issuance’s concentration index with the observed worst 7-day burst, thirteen winters · each issuance’s own Floor removed

Partial correlation of each issuance’s concentration index with the observed worst 7-day burst over the thirteen winters, each issuance scored against its own Floor. June reads −0.22, July −0.07, August +0.73, September +0.44 and October +0.30. The product uses August alone, and the reason is lead time rather than a claim that August is the best month: August and September are not separable on this sample (difference +0.27, p=0.52), and combining them gives +0.73, which is within noise of August by itself. August is chosen because it is the earlier of the two indistinguishable issuances and therefore the more useful. The July to August step is the only one the sample comes close to supporting: +0.82, p=0.042. Every other step on the curve is well inside noise, September and October included, so the curve is read as a rough shape with an onset, not month by month. Every issuance on the curve is extracted station by station with signal and Floor taken on the same station visit, so each point removes a baseline built from exactly the same stations rather than a nearby approximation to them.

How the season window was settled

The report is scored October–March because that is the season the market uses. Getting there required one decision that is not obvious, so the whole audit trail is here rather than summarised.

Step zero: why the core winter is not enough, from the observations alone

Before any forecast is involved: how much of the October–March season does the December–February core actually account for? Regressing each Oct–Mar measure on the same DJF measure over 51 winters of observations gives activity R²=0.77, worst 7-day burst R²=0.28, clumping R²=0.05. So a core-winter view is a fair proxy for how many windy days a season contains, and a poor one for how those days are arranged: roughly three quarters of the variance in the measure this report predicts sits outside the core season. That is measured from the observed record, with no reference to any forecast or correlation, and it is the reason the season has to be the full one.

Step one: the two candidate targets are nearly the same thing

The observed worst 7-day burst over October–March correlates +0.97 with the same measure over November–March. Extending the measured season barely changes what is being measured, because a seasonal maximum can only move one way when a month is added. So the target choice is not where the difficulty is.

Step two, how the forecast months are combined is where it is

construction of the forecast sideindex vs Oct–Mar burstwind vs Oct–Mar burstused?
Pool raw station-days across Oct–Marn/a+0.06no, October swamps the total
Drop October from the forecast side, keep it in the target+0.83+0.82no, a window chosen after seeing the result
Standardise each month, then average (all six months)+0.71+0.68yes

Why the weakest of the three is the one used. Pooling raw days fails for a measurable reason: the October wind conviction rate carries roughly 6.7× the level and 23× the standard deviation of the November–March rate, so October is 17% of the days and dominates the mean, a day-pooled October–March signal correlates −0.04 with its own November–March component. Standardising each month first fixes that, and the fix is justified by the variance ratio, which is measurable without reference to any correlation. The middle option scores best and is not used, because choosing which months feed the forecast after seeing which choice scores best is how a spurious result gets made. The test we apply to any methodology choice in this programme is whether it could have been justified before the correlation was looked at. The variance fix passes that test; dropping October does not.

Step three, the earlier window tests, kept for reference

measureforecast windowOctSepAugJul
catalogue indexDJF+0.56+0.20+0.10+0.22
Nov–Mar burstDJF+0.43+0.26+0.44+0.17
catalogue indexNDJFM+0.36−0.03+0.04+0.02
Nov–Mar burstNDJFM+0.35+0.31+0.65+0.03

These grids are from the window-selection work and are deliberately on other windows, DJF and November–March, because their purpose was to establish the rule that a measure should be predicted from a forecast covering the same months. That rule is why the report is scored October–March throughout, and why the one exception (the catalogue index, on a core-winter window) is flagged where it appears.

What the reference contributes

Claros has two working parts that can be measured separately. The climate reference, an era-relative, station-specific, calendar-day-specific threshold, rebuilt each year, is what makes any discrimination possible; the analogue weighting supplies the year-specific signal on top of it. This table isolates the first by holding the forecast fixed and swapping only the reference it is judged against.

targetissuedour rebuilt referencefixed 30-year normalgain
catalogue extremity indexOct+0.54−0.25+0.79
observed activityOct+0.56−0.56+1.12
Nov–Mar burstAug+0.58−0.32+0.91
Nov–Mar clumpingAug+0.57−0.11+0.68

Identical forecast, identical stations, identical winters. These are raw correlations, so the Nov–Mar burst figure is on the earlier window and basis rather than the no-skill baseline controlled out. Against a fixed 30-year normal the signal flags 17.5 days per station-winter instead of 0.14: the climate has drifted away from the stale reference, the signal saturates and all discrimination is lost. The rebuilt reference is a precondition for season picking, not a refinement of it.

Should the network be wider than nine countries?

The nine-country scope is the area European windstorm loss actually falls in, so it needs little defending. It was still worth testing rather than asserting, because “why only these countries?” is the first question a reviewer asks.

Six non-core regions were built from the wider library, Iberia, Alpine and central Europe, Finland and Iceland, Italy, the Baltics, south-east Europe, and then aggregated, because one large network is a different test from six small ones: a signal that was regime-scale and simply under-sampled by the core would show up as the network grew. It does the opposite. Against the concentration outcome, the nine-country core reads +0.77; adding Finland and Iceland gives +0.70; adding the Baltics as well, +0.55; the whole continent, +0.36. And everything except the core, aggregated into one 676-station network, reads −0.03. That monotone decline with dilution is the signature of adding noise to a real signal rather than of a signal needing more room.

And the skill is not the only thing that thins out. Measured in absolute wind speed rather than in each station’s own percentiles, the core network is the windiest part of Europe bar one. Taking the core as 1.00, the typical daily-max wind at its 85th percentile runs at 0.84 in Iberia, 0.82 in Italy, 0.73 in south-east Europe, 0.72 across the Alps and central Europe and 0.67 in the Baltics, with the 99th percentile of observed wind giving the same ordering. So the regions where no skill is found are also the regions with materially less windstorm to find, which is the more useful way round to state it: this is not a boundary drawn where the model happens to work, it is drawn where European windstorm happens. The single exception proves the point. Finland and Iceland are the one non-core group windier than the core, at 1.08, and they are also the one non-core group where any signal survived at all.

910 stations, 26 per country on a latitude-stratified sample, equal-weighted so no population kernel is assumed outside the core, August initialisation, thirteen winters, outcomes rebuilt from each network’s own daily record. Both wind and rain, against burst, clumping and activity: 48 non-core combinations, of which two reach p<0.05: one positive, one negative, in different regions on different measures with nothing else nearby. The individual regional numbers are not reproduced here because at 26 stations per country they are not strong enough to read one at a time; the aggregate is the test that answers the question. Two limits worth stating: this is decisive for wind and weaker for rain, which is the more weighting-sensitive instrument and this test uses no weighting; and no catalogue exists outside the core, so no non-core result could be checked against anything but the station data behind it.

Geography inside the core

Two separate regional breakdowns were built inside the nine countries, on different station selections and weightings. The point they make together is not about any one country.

Pooling is not a compromise, it is where the skill is. On the first breakdown the pooled nine-country call scores +0.57 against its own outcome, and that is higher than every individual country measured against its own outcome: Benelux +0.49, Great Britain +0.48, Denmark and France +0.45, Sweden +0.24, Germany +0.22, Ireland −0.05, Norway −0.18. The whole is better than any of its parts, which is the same finding as the scope test above read from the other direction. There is a footprint at which this signal is measurable, and cutting below it loses skill just as surely as diluting above it does.

Which is why no country figure is quoted. The two breakdowns disagree sharply once they are cut that fine: Benelux reads +0.49 on one basis and −0.49 on the other, Denmark +0.45 against −0.34, Great Britain +0.48 against +0.04. Germany and France agree. The difference is station selection and weighting inside each geography rather than different observations, and at 150 to 600 stations per country a handful of them moves a correlation a long way. Neither breakdown is reported as a result, and the report makes no claim about any country or sub-region. Reconciling the two is scheduled work; it does not change anything in the findings, all of which are the pooled call.

Issuance-by-issuance record

Each cell is the concentration index from that issuance alone, standardised on that issuance’s own thirteen-winter scored record and applied unchanged to 2026/27. Spread is the standard deviation across the issuances present. 2026/27 has three because September and October do not exist yet; on those same three months its spread ranks 4th of 14. The headline statistics from this table are quoted in the forward tab.

Season by season

wintercatalogue eventsclusteringseverity m/sextremity pct7-day burstsummer index callOct wind calldriver alignmentnotable stormsmarket context
2013/1416823.3 (52 mph)61.8−0.07+0.08+0.23+0.29 (64th)Xaver, Dirk (Dec); serial UK storms Jan-Feb (Anne, Christina)
2014/1510123.0 (51 mph)30.9+1.40+0.02−0.23-0.16 (40th)Elon-Felix (Jan); Mike, Ole
2015/168326.7 (60 mph)76.4+1.31−0.24+0.09-0.31 (35th)Desmond (Dec, flood-led); Gertrude, Imogen (Feb)
2016/175022.1 (49 mph)20.0−1.65−0.34−0.35-0.04 (48th)Egon (Jan); Doris (Feb)
2017/186323.6 (53 mph)35.5+0.50−0.05+0.62+0.35 (73rd)Burglind, Eleanor, Friederike (Jan); Emma (late Feb)
2018/196022.1 (49 mph)25.5+3.18+0.61+1.05+0.27 (62nd)Alfrida (Jan); quiet core
2019/2012523.3 (52 mph)89.1+1.07+0.00+0.38+0.13 (53rd)Ciara-Sabine, Dennis, Jorge (serial February)
2020/217324.1 (54 mph)35.5−0.50−0.44+0.48+0.42 (74th)Bella (Dec); wind-quiet, cold spells
2021/227121.8 (49 mph)23.6+2.91+0.58−0.72+0.90 (91st)Barra (Dec); Malik-Corrie (Jan); Dudley, Eunice, Franklin (Feb triple)Dudley+Eunice+Franklin EUR 3.85bn industry (PERILS)
2022/234028.3 (63 mph)n/a−0.61−0.57−0.21+0.33 (70th)Otto (Feb); mild season
2023/2412729.8 (67 mph)52.7+0.49+0.52−0.31-0.90 (5th)Pia, Gerrit (Dec); Henk, Isha, Jocelyn (Jan). Ciaran was November (autumn)Ciaran EUR 2.04bn industry (Nov, autumn; PERILS)
2024/2510322.9 (51 mph)35.5−1.70−0.13−0.63-0.06 (44th)Darragh (Dec); Eowyn (Jan, Irish wind records); HerminiaEowyn EUR 765m industry final (PERILS)
2025/26n/a−0.65−0.04−0.39

The thirteen hindcast winters. Catalogue events, clustering and severity come from our reanalysis-built catalogue; the 7-day burst is the station measure; the summer index call and October wind call are the forecasts as issued. 2022/23 and 2025/26 are not yet indexed in the catalogue, which is why every catalogue figure in this report is scored on eleven winters rather than thirteen. Catalogue severity is daily-mean based while station data carries the daily maximum, so the two severity measures are never compared raw; the record table’s definitions cover what that column does and does not discriminate.

Method

The season. October to March, on both sides of every comparison in the findings. That is the windstorm season as the market defines it and as named-storm contract wording follows it. Earlier internal work in this programme was scored on December–February and on November–March; those windows are not used here.

The forecast. Claros weights historical analogues on 28 teleconnection drivers and their interactions to issue daily exceedance probabilities months ahead. A conviction day is a station-day where that probability reaches 1.5× the nominal base rate against the station’s own rebuilt calendar-day threshold (wind: 85th percentile, nominal base 15%; rain and warmth: 90th, base 10%). The seasonal instrument is the population-weighted count of conviction days across the network.

How the six months are combined, and why it matters. Each month’s conviction rate is standardised on its own thirteen-winter record first, and the six standardised months are then averaged. The instrument index is the average of the three instruments built that way. This is not cosmetic. The alternative, pooling the raw station-days across all six months, fails, because the months are not on comparable scales: the October wind conviction rate has roughly 6.7 times the level and 23 times the standard deviation of the November–March rate, so in a day-pooled mean October is 17% of the days and dominates the total. Measured directly, a day-pooled October–March wind signal correlates −0.04 with its own November–March component: the rest of the season is not diluted, it is erased. Standardising first gives every month one equal voice and lets no month swamp the others. The justification is a property of the data, measured before any correlation was looked at, and no month is dropped from either side.

The thresholds. Era-relative and station-specific: each station’s own percentile for that calendar day, detrended, so an exceedance day means the same thing in 1975 as in 2020. Measured against these thresholds the realised winter exceedance rate is nearer 18% than the nominal 15%, so the conviction gate sits slightly lower than 1.5× realised. The rule is frozen and applied identically to every winter, and a higher bar was tested up to 8× and performs worse.

The baseline. Every skill number is a partial correlation with the no-skill track controlled out: the same climatology, same stations, same annual vintages, no analogue weighting. That removes the shared climate trend from both sides, so what remains is year-specific information.

Significance. By permutation. We shuffle which forecast is paired with which winter, twenty thousand times, and count how often a random pairing does as well as the real one. If a fifth of the shuffles beat it, the result is worth nothing; if one in three hundred does, p=0.003.

One detail matters and is easy to get wrong. Every headline number is measured after removing what the no-skill baseline, the same climatology with no analogue weighting, already explains. So when the winters are shuffled, each forecast has to keep its own baseline attached to it. If the baseline stays put while the forecast moves, the shuffled pairs are no longer a fair picture of chance: they have been stripped of a relationship the real pairs still have, so chance looks worse than it is and the result looks better than it is. Keeping each forecast and its baseline together is the strict version, it is what is used here, and it agrees with the textbook formula on every cell to within 0.006.

Wind is the common factor, and the three instruments are not three independent votes. On the construction used here, wind correlates +0.62 with rain and +0.81 with warmth at the August initialisation, while rain and warmth correlate only +0.24 with each other. At October all three are tighter, +0.63 to +0.87. So wind carries most of what the other two carry: once wind is in the model, rain adds +0.14 (p=0.69) and warmth adds −0.06 (p=0.85), neither significant. Two consequences, pulling in opposite directions. The concentration result must not be presented as three independent confirmations, because it is substantially one signal read three ways. But the index still earns its place over wind alone, because averaging partly-independent readings of one signal reduces noise even when none of them adds fresh information: the index scores +0.73 against wind’s +0.70.

Testing many things at once, and what we did about it. This is the risk that matters most in work like this, so it is worth being concrete rather than reassuring. If you test twenty unrelated things at the conventional threshold, roughly one will look significant purely by luck. This programme has tested several hundred instrument-and-measure combinations. So a single good-looking number, produced on its own, would be worth very little, and none is quoted that way.

Three specific checks stand behind the two results the report does quote, and each one is a thing a reader can ask for and re-run:

  • Shuffle the winters. Re-pair each forecast with a randomly chosen winter, twenty thousand times, and count how often a random pairing does as well as the real one. For the index against the burst measure, 11 shuffles in 1,000 matched it.
  • Remove each winter in turn. Refit the result thirteen times, each time with a different season left out. If one exceptional winter were carrying the answer, one of those thirteen refits would collapse. The lowest is +0.59 and the highest +0.79, so none does.
  • Take the population weighting out of the outcome. Rebuild it as a plain station average, so the answer cannot be an artefact of weighting towards where people live. Burst goes from +0.71 to +0.68 and stays significant; clumping goes from +0.61 to +0.47 and does not. So burst passes and clumping fails, and clumping is labelled the weaker measure everywhere it appears as a result.

What is quoted, and what is not. Quoted as findings: the two index rows in section 4, the August index against October–March burst (+0.73) and against clumping (+0.63), because both pass all three checks. Carried as a developing result rather than a finding: October rain against the catalogue index, +0.5 to +0.6 with p between 0.06 and 0.13 on the eleven winters the catalogue covers. It does not clear the first check on that sample, and section 4 sets out why it is still the call worth watching. Not quoted at all: every single-instrument cell that reaches p<0.05 in one configuration and not another.

And one search that found nothing, which is the more useful evidence. Looking for a route to the exceptional-season family, we ran the full grid, three instruments across four initialisation months, two forecast windows and five different targets, 120 combinations, and recorded all of it rather than the best cell. Nothing reached significance with a positive sign. The point is not that 120 is a large number; it is that when this method is pointed at something that is not there, it returns nothing. That is the only real evidence available that the numbers it does return are not manufactured by searching.

Where this is tested, and where it is not. The hindcast library covers 5,300 stations across roughly forty European countries. Every scored number in this report is the nine-country north-west European network, Great Britain, Ireland, France, Germany, Netherlands, Belgium, Denmark, Norway, Sweden, 3,002 stations, with a 1,041-station population-weighted selection carrying the headline signal. Nothing here is a claim about any other part of Europe, or about any country or sub-region taken on its own. Whether the scope should be wider was tested rather than assumed; the test is in the workings tab.

Weighting. A 20-city population kernel (120 km Gaussian) over the north-west European core, used as the closest available proxy for where insured value sits. It is not client exposure data. Signal and outcome are weighted identically, and the burst result survives removing the weighting from the target altogether.

v8 · 24 Aug 2026