Thirteen validation winters, 2013/14 to 2025/26, scored on the October–March windstorm season, on a network of 3,002 stations across nine north-west European countries drawn from a daily hindcast library of about 5,300 European stations. Predictions are scored against two records built from different data: a station observation archive reaching back to the 1930s, and a 1950–2025 windstorm event catalogue we built from reanalysis. Every skill number is measured against a no-skill baseline made from the same climatology, so what is reported is what the forecast adds. No loss data is used in fitting or scoring anywhere; loss figures appear only as context. Every technical term used in this report is defined in the glossary at the top of the workings tab.
Insured windstorm outcomes are not one quantity. Fifty-three winters of station observations separate cleanly into two families, and a winter can be extreme in one while sitting near normal in the other. That is the single most important structural fact in this programme, because it decides how many calls the system needs to make.
Wind activity, peak severity, storm clustering and station peak wind all move together: they correlate +0.73 to +0.97 with each other over 53 winters. They also carry the strongest link to the independently built catalogue extremity index. This is the axis the industry has historically priced: the 1990-type winter.
The worst 7-day burst and clumping correlate +0.79 with each other, but far more loosely with the first family: burst +0.47 to +0.56, clumping +0.26 to +0.37. Against the catalogue index they manage only +0.24 and +0.18. A concentrated fortnight inside an otherwise ordinary season shares some ground with a relentless season, but not enough for either to stand in for the other, and the catalogue does not see it at all.
| activity | severity | clustering | station peak | burst | clumping | catalogue index | |
|---|---|---|---|---|---|---|---|
| activity | +0.97 | +0.95 | +0.79 | +0.53 | +0.29 | +0.62 | |
| severity | +0.97 | +0.93 | +0.87 | +0.56 | +0.37 | +0.60 | |
| clustering | +0.95 | +0.93 | +0.73 | +0.48 | +0.26 | +0.52 | |
| station peak | +0.79 | +0.87 | +0.73 | +0.47 | +0.31 | +0.55 | |
| burst | +0.53 | +0.56 | +0.48 | +0.47 | +0.79 | +0.24 | |
| clumping | +0.29 | +0.37 | +0.26 | +0.31 | +0.79 | +0.18 | |
| catalogue index | +0.62 | +0.60 | +0.52 | +0.55 | +0.24 | +0.18 |
Observed station measures against each other and against the catalogue extremity index, 53 winters (the index column covers 41). The two blocks in the top-left and bottom-right are the two families; the weak corner between them is the point. Green is positive, red negative, intensity is strength.
| observed station measure | r vs the catalogue extremity index | p | winters |
|---|---|---|---|
| activity | +0.62 | 0.0001 | 41 |
| severity | +0.60 | 0.0000 | 41 |
| clustering | +0.52 | 0.0004 | 41 |
| station peak | +0.55 | 0.0001 | 41 |
| burst | +0.24 | 0.1310 | 41 |
| clumping | +0.18 | 0.2601 | 41 |
Permutation-tested over the winters where both series exist. Activity, severity, station peak and clustering all track the catalogue index at p≤0.0004 across 41 winters: two records built from different data agreeing on which winters were bad. Burst and clumping do not, and that is not a failure: they are measuring the other family.
Claros issues daily exceedance probabilities months ahead for wind, rain and warmth. The seasonal instrument is the count of conviction days: station-days where the forecast probability reaches 1.5× the base rate against that station’s own rebuilt calendar-day threshold. Below, each instrument’s conviction rate against what those same stations actually recorded.
At the day level these instruments are thoroughly verified. A flagged day exceeds at about twice the base rate, measured across 68 million day-predictions, and that accuracy holds at every lead tested: 2.00× at one to three months and 1.99× at thirteen to twenty-four. Accuracy that does not decay with lead is unusual in seasonal forecasting and it is the foundation everything above is built on.
At the seasonal level the same instruments point the same way. Wind at +0.43 and rain at +0.53 are respectable correlations for a seasonal signal; with thirteen winters they are indicative rather than conclusive taken one at a time (p=0.14 and p=0.06), and warmth reads +0.44 alongside them.
The value lands one step further on. What these instruments predict best is not how many exceedance days a winter contains but how those days are distributed through it, and there the same three read +0.46 to +0.70 individually, with wind and rain both clearing significance on their own (p=0.012 and p=0.040). Combined into the index they reach +0.73 (p=0.009). The instruments are the mechanism; the measures in section 4 are the product.
It is a fair challenge and it deserves a direct answer, because a reader could reasonably suspect that three instruments and several measures were searched until something correlated.
The three instruments are not three candidate predictors. They are three readings of a single engine: all of the analogue weighting derives from one set of sea-level-pressure pattern relationships. The instruments correlate +0.62 (wind with rain) and +0.81 (wind with warmth), and once wind is in the model rain adds +0.14 (p=0.69) and warmth −0.06 (p=0.85). So there was never a pool of independent predictors to pick from: there is one signal, and the question is only which observable it is best read through. That is a property of how each observable relates to windstorm activity, not a choice we made.
And the measures are not interchangeable either. Section 1 shows they fall into two families with a weak corner between them, established over 53 winters of observations before any forecast was scored. Predicting rain days in order to say something about how a windstorm season is structured is a real inference, and it earns its place only because the structural link between the measures is independently documented and because the result survives permutation, leave-one-out and the removal of population weighting.
What would make it cherry-picking is choosing the season window, the initialisation month or the combination rule after seeing which choice scored best. The workings tab records each of those decisions and the evidence that fixed it, including the one place where the better-scoring option was rejected for exactly this reason.
One line per instrument: the strongest measured link it carries, and what it is therefore for. Each points at a call the system could make; section 6 turns them into alarms.
Claros runs monthly, so the first question is which run to use. Against the concentration outcome the late-summer runs are the usable ones: August +0.73 and September +0.44, with the two combined reading +0.73. June and July are both slightly negative at −0.11. Bootstrapping the differences over thirteen winters, August cannot be separated from September (90% interval on the gap −0.17 to +0.80; September comes out ahead in 12% of resamples) but both can be separated from July (July beats August in 2.6% of resamples). There is no known meteorological barrier between August and September either, so the honest reading is that the late-summer window carries the signal and the split between its two runs is inside the noise. The report quotes August because it is the earlier of the two and therefore the more useful, not because it is measurably better. This also settles a coordination question: the teleconnection sub-model is validated on September, and on this evidence that is an equally defensible choice.
| measure | instrument | partial r | p | leave-one-winter-out range |
|---|---|---|---|---|
| Oct–Mar burst | wind | +0.70 | 0.012 | +0.59 to +0.78 |
| rain | +0.60 | 0.040 | +0.45 to +0.71 | |
| warmth | +0.46 | 0.130 | +0.33 to +0.54 | |
| the index | +0.73 | 0.009 | +0.63 to +0.84 | |
| Oct–Mar clumping | wind | +0.54 | 0.070 | +0.33 to +0.83 |
| rain | +0.53 | 0.076 | +0.41 to +0.65 | |
| warmth | +0.55 | 0.064 | +0.43 to +0.81 | |
| the index | +0.63 | 0.029 | +0.47 to +0.81 |
Partial correlations with the no-skill baseline controlled out, thirteen winters, full station network, scored across the whole October–March season on both sides. The product is the index, and it clears significance on both concentration measures: +0.73 against the worst 7-day burst (p=0.009) and +0.63 against clumping (p=0.029). Among the single instruments wind and rain clear on burst (+0.70, p=0.012 and +0.60, p=0.040); warmth (+0.46) points the same way without reaching it on its own, which is what you would expect from three readings of one model run rather than three independent tests. The leave-one-winter-out range recomputes each cell thirteen times, dropping a different winter, and every concentration cell stays positive.
The August index against the observed worst 7-day burst, October–March, thirteen winters, r=+0.73 (p=0.009). Its two highest readings are the two most concentrated winters of the record. The orange bar is the season ahead, which reads the lowest of the fourteen.
Dudley, Eunice and Franklin arrived between 16 and 21 February 2022: three damaging storms in six days, EUR 3.85bn of industry loss. All three fell inside the season, so this is not a case of looking at the wrong months.
On family one the winter was unremarkable: activity +0.30σ, peak severity +0.45σ, station clustering +0.28σ, and the catalogue index put it at the 23.6th percentile of 75 years. On family two it was extreme: worst 7-day burst +2.91σ and clumping +2.18σ across October–March, the second highest burst of the thirteen behind 2018/19.
The forecast side split the same way. The October wind call read −0.81σ, so nothing in the family-one instruments saw it coming. The August concentration index read +0.56, its second-highest reading of the thirteen: six months before the fortnight.
And the week itself can be located from the station data alone. Asked which seven days carried most of the network’s exceedance that winter, with no knowledge of named storms, the answer is 16 to 22 February 2022, carrying 4.04 times the exceedance an evenly spread season would put in a week. Section 7 maps which stations took that hit: a corridor from Ireland through to Denmark, with Scandinavia untouched.
The concentration call is issued in August, four to seven months ahead, and it answers a broad question about the shape of the season. The October initialisation answers a narrower and more valuable one: how extreme the core of the winter will be against the catalogue. It arrives later, it is a sharper reading, and it is scored against the one record built from entirely different data to the station archive.
| days pooled across the three months | each month standardised first | |||
|---|---|---|---|---|
| instrument | partial r | p | partial r | p |
| wind | +0.38 | 0.277 | +0.44 | 0.212 |
| rain | +0.66 | 0.035 | +0.57 | 0.090 |
| warmth | +0.39 | 0.262 | +0.49 | 0.153 |
| the index | +0.53 | 0.114 | +0.56 | 0.096 |
October initialisation, December–February, against the catalogue extremity index, over the eleven winters the catalogue covers: the prediction record runs to thirteen winters, but 2022/23 and 2025/26 are not yet indexed, so every catalogue figure in this report is scored on eleven. Two ways of combining the three months are shown because both are defensible and they agree closely: the pooled and standardised versions of each instrument correlate +0.89 to +0.94 with one another, so the gap between the two columns is sampling noise on eleven winters, not a choice that changes the answer. Read rain as +0.5 to +0.6 with p between 0.06 and 0.13, and the index as +0.46 to +0.47, p 0.17 to 0.18, jackknife +0.20 to +0.80.
Rain is the best performer in the hindcast. Across the prediction record it is the strongest single instrument against the catalogue index at +0.5 to +0.6, and it is also the strongest against extended-season concentration. If conditions continue broadly as they have through the validation decade, rain is the instrument to read. It is not quite significant on the eleven winters the catalogue covers, and two or three more October runs would settle that.
Wind is what the long record says a severe winter is made of. Over 41 winters, observed wind activity tracks the catalogue index at +0.62 (p=0.0001). Observed warm days manage +0.34 and observed rain days +0.23. Chained through what we can actually forecast, wind offers about +0.28 to the catalogue (fidelity +0.45 against a long-record link of +0.62) where rain offers about +0.09. So the two records disagree about which observable matters, and the disagreement carries information: rain measures better over thirteen recent winters, wind measures better over fifty-three.
Which is why both belong in the alarm set. The validation decade has been a period of moderate windstorm activity by historical standards. Rain is calibrated to that period and performs well in it. The winters that defined the market’s view of this peril, 1989/90 and 1994/95, are wind-activity extremes of a kind the prediction record does not contain, and wind is the instrument tied to them. Reading only rain risks being well tuned to a regime that changes; reading only wind discards the better recent performer. Section 6 sets out the alarms that follow, and the one that is not yet buildable.
The record is split by issuance, because the two issuances predict different families and one table invites the wrong comparison. The first row of toggles picks the issuance; the second picks the geography, with pooled Europe, the product, as the default. Click any observed column for what predicted it and the comparison chart.
Signals are standardised on the thirteen-winter issuance record; observed columns are σ against the frozen 1972–2011 station reference. On weighting: every observed column drawn from station data is a network mean on the same population kernel as the signal, so forecast and outcome are weighted identically. The only unweighted column is the catalogue extremity percentile. That percentile is missing 2022/23 and does not yet reach 2025/26, so those two cells are blank, which is why every catalogue figure in this report is scored on eleven winters rather than thirteen. This table is the source. Every alarm threshold in section 6 is computed from exactly these columns, so any figure in the sweep tables can be reproduced from the numbers on this screen with nothing more than a spreadsheet.
On the composite columns. Both issuances carry the three instruments and the composite built from them, and the composite competes for the best-predictor highlight like any other column, because the report makes claims about it. On the August issuance the composite is the concentration index, the product for family two. On the October issuance it is the family-one composite, which is shown precisely because it loses to rain: section 4 sets out why combining helps on one family and dilutes on the other, and this is where that is visible in the data. One thing to expect. The best-predictor figure is the plain correlation of the two columns as printed, so it can be checked against this table directly. The findings quote partial correlations, with the no-skill climatology removed from both sides. The two occasionally rank differently, most visibly on October–March burst, where the two are close: wind 0.71 against the index 0.70 on the plain numbers, and the index +0.73 against wind +0.70 once each is scored against its own Floor. Neither gap is meaningful on thirteen winters. The panel says which measure it is reporting whenever a published figure exists for the column.
Two families, and two different kinds of severe winter inside the first one, mean the system needs three alarms rather than one. The 53-winter observed record shows why with unusual clarity, and it also shows which of the three we can currently build.
| winter | activity σ | 7-day burst σ | catalogue extremity pct | severity alarm would fire | concentration alarm would fire |
|---|---|---|---|---|---|
| 1973/74 | +0.76 | +0.92 | 40.0 | Tier 1 | : |
| 1982/83 | +0.71 | +0.55 | 72.7 | Tier 1 | : |
| 1983/84 | +0.77 | +1.68 | 90.9 | Tier 1 | fires |
| 1989/90 | +1.61 | +2.09 | 95.5 | Tier 2 | fires |
| 1993/94 | +0.68 | +0.51 | 81.8 | Tier 1 | : |
| 1994/95 | +1.29 | +0.14 | 95.5 | Tier 2 | : |
| 2001/02 | +0.94 | +1.04 | 70.0 | Tier 1 | : |
| 2006/07 | +0.91 | +0.59 | 67.3 | Tier 1 | : |
| 2013/14 | +0.92 | +0.22 | 61.8 | Tier 1 | : |
| 2015/16 | +0.66 | +1.42 | 76.4 | Tier 1 | : |
| 2018/19 | −0.45 | +3.05 | 25.5 | : | fires |
| 2019/20 | +1.23 | +1.21 | 89.1 | Tier 1 | : |
| 2021/22 | +0.30 | +2.81 | 23.6 | : | fires |
Every winter in 53 that either the severity ladder or the concentration threshold would flag, with the two side by side. Severity alarm: Tier 1 is observed activity at or above +0.65σ, Tier 2 is activity and station clustering both at or above +1.25σ, on the frozen 1972–2011 scale: 11 Tier 1 firings and 2 Tier 2 across 53 winters. Concentration alarm: observed worst 7-day burst at or above +1.5σ. These are the observed definitions, which is what makes the comparison possible over 53 winters; the prediction-side thresholds are a separate job, noted below. One note on the scale. Every observed σ quoted in this report sits on a frozen pre-2012 reference, so that winters are compared against a baseline that does not include them. What changes between exhibits is the season window: this ladder is built on November to March, the record table in section 5 on October to March. The same winter therefore carries two readings, both frozen: 2018/19’s burst is +3.05σ here and +3.18σ there. Neither is more correct, and the report does not mix the two windows within a table.
Two families and two mechanisms inside family one, the winter that is bad for the current climate, and the winter that is exceptional against any climate, make three alarms the natural design. They are not equally ready, and the difference matters more than the design.
August concentration index → observed Oct–Mar burst, +0.73, p=0.009, and → clumping +0.63, p=0.029. Every cell survives dropping any single winter. Fires for the family-two winter from two months before the season opens. This is the alarm that would have caught 2021/22 and 2018/19, the two most concentrated winters of the record.
October conviction → catalogue extremity index. Rain reads +0.5 to +0.6, p between 0.06 and 0.13. This is the winter that is bad relative to today’s climate, and it is the strongest family-one relationship anywhere in the programme: the right size for a useful alarm, on a sample too short to establish it. Two more October issuances would settle it either way.
The 1989/90 and 1994/95 type: the grinding winter the observed activity ladder catches at the 95.5th percentile of the catalogue index. On the observed side this is the best-defined alarm in the programme. The prediction side needs a step the other two alarms do not: the signal has to be scaled against the observed variability of the target before a threshold means anything, because the placed calls are conservative in spread by construction. That calibration is the next piece of work rather than a closed question.
For the two alarms that have a route, every threshold scored. Alarm 3 is absent because the calibration that would set its threshold has not been run.
| threshold | fires | caught | false alarms | missed | quiet & right | catch rate |
|---|---|---|---|---|---|---|
| +0.00 | 6 | 4 | 2 | 1 | 6 | 80% |
| +0.20 | 3 | 2 | 1 | 3 | 7 | 40% |
| +0.35 | 3 | 2 | 1 | 3 | 7 | 40% |
| +0.50 | 3 | 2 | 1 | 3 | 7 | 40% |
| +0.65 | 0 | 0 | 0 | 5 | 8 | 0% |
| +0.80 | 0 | 0 | 0 | 5 | 8 | 0% |
| +1.00 | 0 | 0 | 0 | 5 | 8 | 0% |
Thirteen winters. The concentration alarm has no false alarm at all from +0.5 upward, at the cost of missing the marginal winters: it catches 2 of the 5 winters that delivered a +1.0σ burst, and both of the two that delivered a +2.9σ one.
| threshold | fires | caught | false alarms | missed | quiet & right | catch rate |
|---|---|---|---|---|---|---|
| +0.00 | 7 | 3 | 4 | 0 | 4 | 100% |
| +0.20 | 5 | 2 | 3 | 1 | 5 | 67% |
| +0.35 | 5 | 2 | 3 | 1 | 5 | 67% |
| +0.50 | 3 | 1 | 2 | 2 | 6 | 33% |
| +0.65 | 2 | 1 | 1 | 2 | 7 | 33% |
| +0.80 | 1 | 0 | 1 | 3 | 7 | 0% |
| +1.00 | 1 | 0 | 1 | 3 | 7 | 0% |
Eleven winters, because the catalogue does not yet index 2022/23 or 2025/26. The rain alarm is flat across most of its range: 2 of 3 caught with 2 false alarms at every threshold from +0.20 to +0.65. A flat response over that span is what a real but imprecise signal looks like on eleven winters.
| threshold | fires | caught | false alarms | missed | quiet & right | catch rate |
|---|---|---|---|---|---|---|
| +0.00 | 6 | 3 | 3 | 1 | 6 | 75% |
| +0.20 | 5 | 2 | 3 | 2 | 6 | 50% |
| +0.35 | 4 | 1 | 3 | 3 | 6 | 25% |
| +0.50 | 2 | 0 | 2 | 4 | 7 | 0% |
| +0.65 | 1 | 0 | 1 | 4 | 8 | 0% |
| +0.80 | 1 | 0 | 1 | 4 | 8 | 0% |
| +1.00 | 1 | 0 | 1 | 4 | 8 | 0% |
Shown for completeness: the wind-to-activity alarm catches 2 of 4 with 2 false alarms at the two lowest thresholds and 1 of 4 above them, which is why it is not proposed in this form. This is the cell the calibration work behind Alarm 3 has to improve on.
When a winter is loaded, the predicted instruments run red together and the observed dials verify together. The verified quantity is how much of the map runs red, winter by winter: not which individual station is hit, which the system does not claim.
Interactive: predicted instruments on the left, observed dials on the right; any dial can be read alone or combined with any others, and the winters step with the chips, the arrow keys or Play. Strongest signal shows the deepest reading any active dial gives each station, which is how an alarm set behaves; Averaged draws the per-station mean instead, which is the aggregation a composite actually uses. Coverage: all thirteen scored winters plus the season ahead. Matched runs puts each dial on the issuance that suits its own target, wind, rain and warmth on the October run over December to February and the index on the August run over October to March. The October station arrays cover the thirteen scored winters; they grey out for 2026/27, for which the store holds no October initialisation. August only puts all four on the August issuance over October to March, on one construction and every winter including the season ahead.
The networks in play, and they are not the same on both sides. Every result in this report pairs a forecast network with an observed one, and the two are deliberately different. The forecast side is the selection of methodology 2.3, the top stations by population weight to 85% of cumulative weight per country: 1,041 stations for wind, 1,033 for rain and 1,039 for warmth, the difference being stations absent from one variable’s archive. The Floor is read at the same stations as the signal it controls. The observed side is the fixed panel of 3,002 stations, which is what carries the long record back to 1973/74 and what every observed sigma is measured against; the rain and warmth station dials sit on the sub-panels that hold those variables, about 1,050 and 1,045 stations. Both sides are population-weighted on the same kernel, so a forecast and its outcome are weighted identically even though the panels differ.
Two exhibits use different networks again, because the question they answer is different. The regional panel draws every station in each region on both sides, since it asks whether the signal works region by region rather than what the population-weighted number is. The wider-Europe test uses a latitude-stratified equal-weighted sample outside the core nine, so no population kernel is assumed where the product does not operate. The fidelity dials are the one place where both sides sit on the forecast selection, because there the question is whether the forecast and the observation agree at the same stations. Each is labelled at the point of use.
The fourth prediction dial is the index itself, on the same October–March season and the same August issuance as the number in section 4. Not a proxy and not a shorter window: the same quantity, taken apart.
The mean of the map is the number. The index is a population-weighted network average, so it does not decompose into per-station index readings, but it does decompose exactly into per-station contributions. Each station’s contribution is its own weighted share of the network exceedance rate, measured against its own normal, in index units. Those contributions sum to the index by construction. The map shows each station’s contribution scored against its own thirteen-winter record, so red means a station pushed the reading up harder than it usually does, and the panel prints the weighted mean beside it. A station absent from one instrument contributes nothing to that instrument’s network rate, so its contribution there is zero rather than missing, and with that convention the contributions sum to the index exactly: the largest gap across all fourteen winters is 0.004.
Two things to expect when it is selected. The lead label changes to August issue, Oct to Mar, leads 2 to 7 months, and shows both entries if the selection mixes issuances with the October instruments. And the map draws every station in the nine-country network, 2,555 of them for the index, so it is no sparser than the other dials. Three station sets are in play here and they should not be confused. The index number quoted throughout this report is the population-weighted average over the 1,041-station selection of section 2.3, which is 85% of population weight per country. The exact decomposition spans the 1,178 stations that appear in at least one instrument’s selection, and their contributions sum to the published index to within 0.004 across all fourteen winters. The map draws the whole nine-country network, 2,555 stations for the index, because a thousand-station picture is not a picture of Europe; its whole-network mean sits about 0.05 from the published reading and ranks the thirteen winters with it at +0.99, so the map can be read on the whole network without misrepresenting the number.
Every other dial on that map is a season total. Burst participation is not, and it is the only one that shows the measure the system actually predicts as a picture rather than a number.
How it is built. For each winter, find the seven-day window in which the network as a whole recorded most of its exceedance. Then ask, of each station separately, what share of that station’s whole-season exceedance fell inside those seven days, and score that share against the same station’s own 1974–2011 record. Red means this station took an unusually concentrated hit in that one week. It sits on the same frozen reference as the other observed dials, so it composites with them like any other.
The method locates the storms without being told about them. Nothing in the construction knows about named storms, catalogue events or loss. Run on 2021/22, the window it returns is 16 to 22 February 2022: Dudley, Eunice and Franklin, which arrived between the 16th and the 21st. That week carried 4.04 times the exceedance an evenly spread season would put in seven days, the most concentrated week of the thirteen. On 2018/19 it returns 9 to 15 March 2019 at 3.90 times. Those are the two winters the August concentration index flagged, and this is the third independent construction to pick them out.
What the picture adds. Load 2021/22 with burst participation alone. It runs deep red across Ireland, Britain, the Low Countries, northern Germany and Denmark, and blue across Scandinavia, while the season-total dials beside it read close to normal almost everywhere. That is the difference between the two families rendered geographically: a corridor took most of a winter’s damage in six days, and no season total sees it. It also shows what the system does not claim. The forecast predicts that a winter will be concentrated, not where the corridor will fall, and the map is the honest way to show both at once.
Over the full 53-winter record 2021/22 ranks 4th and 2018/19 6th on the network mean of this measure, so neither is unprecedented: the validation decade contained two strongly concentrated winters, not the two worst on record. Across the thirteen scored winters the participation ratio correlates +0.81 with the worst 7-day burst the report predicts, which is the check that the map and the headline measure are two views of one thing rather than two measures.
Lead time. The concentration call becomes available at the August issuance and not before: the June and July runs read −0.22 and −0.07 against the observed burst, August reads +0.73, September +0.44 and October +0.29. A separate reading of the core winter against the catalogue follows in October. Aon’s windstorm season runs October to March, so the August initialisation lands roughly two months before the season opens and four to seven months before the months that usually carry the damage. The core-winter reading arrives as the season starts. The step from July to August is discontinuous rather than gradual, which is a caveat on thirteen winters as much as a property of the system: the workings show the whole curve.
Seasonal storm totals. Against observed activity the forecast does not beat the climate trend at any lead tested. The trend itself is informative and ours is better than a straight line, but that is the climate engine rather than season picking. It is the open item behind Alarm 3, the exceptional-season call.
Peak severity. Not predicted at any lead, on any window. A single exceptional storm sits outside what a seasonal system can see.
Which stations get hit. Tested directly: same-winter geographic pattern agreement does not beat cross-winter agreement. The map’s claim is breadth, not placement.
The sample. Thirteen winters of predictions, scored against 53 winters of station observations and 75 years of catalogue index. Thirteen is what the hindcast covers and we are not extending it, so the work went into making those thirteen hard to argue with: partial correlations against a no-skill baseline built from the same climatology, permutation tests rather than parametric p-values, leave-one-winter-out refitting on every headline cell, and a robustness check that rebuilds the targets without population weighting. The chain also rests on far more than thirteen years: the observed measures track the catalogue index across 41 winters, and the day-level accuracy behind the instruments is verified over 68 million day-predictions.
Issued from the August 2026 run, on the October–March season. Concentration is the measure the system predicts at this range; the December–February reading against the catalogue index needs the October initialisation and does not exist yet.
| instrument | 2026/27 reading σ | lowest of the thirteen | highest of the thirteen |
|---|---|---|---|
| wind | −0.59 | −0.47 | +0.75 |
| rain | −1.13 | −0.69 | +0.78 |
| warmth | −0.46 | −0.63 | +0.86 |
The season call is one number, but the forecast is issued month by month. The same August run resolved into all six months of the season, each scored against that month at that lead in the thirteen previous August runs. The eighteen cells below average to the three instrument readings above, which in turn average to the index.
Seventeen of the eighteen cells read below their thirteen-winter mean and none reaches −2σ. Most emphatic: rain in December (−1.64σ). The single cell above its mean is warmth in October (+0.08σ), which is close enough to normal to be read as neutral. The measured skill is for the season-shape measure, not for individual months, so the month resolution is context rather than six separate forecasts.
What it says. The volume-and-structure risk for the extended season reads below normal on the measure the system predicts best at this range.
What it does not say. Nothing about how many storms will arrive or how severe the worst one will be, and nothing yet about the core winter against the catalogue index: that needs October.
What would change it. The October issuance is a genuinely different observation rather than a refinement: it reads the autumn state that conditions the core winter. In 2019/20 the August run gave no warning of a winter the October wind call went on to place at its record maximum.
Run agreement, in three numbers rather than a table. Splitting the thirteen winters at the median spread between issuances, the agreeing half missed the outcome by 0.87σ and the divergent half by 1.22σ. Spread correlates +0.28 with error, and +0.45 with how far the outcome landed from normal in either direction, so divergence marks a season with something in it rather than a reason to discount the call. For 2026/27 the runs disagree, which is why the reading above is provisional and why the September and October issuances are worth waiting for. The full issuance-by-issuance table is in the workings tab.
The same run carries a 24-month horizon, so it does produce readings for October 2027 to March 2028 at leads of 14 to 19 months. They come in uniformly low: every month of every instrument between −1.67σ and −0.49σ. We do not issue that as a view. The uniformity is the reason: a real seasonal signal discriminates between months and between instruments, and this one does not, which is the signature of a run relaxing towards its own climatology. There is also no skill test at that range; the longest lead with a measured result anywhere in this document is four months. The second winter becomes forecastable when its own late-summer initialisation arrives.
Independent of any model output: what the 75-year teleconnection record says about winters like the one ahead. Context for the call, never a call on its own.
| Drivers aligned active | winters | storm count | clustered events | aggregate severity |
|---|---|---|---|---|
| 0 | 11 | 7.4 | 2.2 | 1296 |
| 1 | 26 | 8.4 | 2.2 | 1450 |
| 2 | 29 | 11.6 | 3.1 | 2150 |
| 3 | 9 | 14.1 | 4.7 | 2745 |
Escalation with alignment (p=0.004, trend-robust; storm count and aggregate severity rise monotonically, clustered events are flat between 0 and 1 aligned drivers then rise): the interplay, not any single index, is the signal. ENSO does not lead for Europe: it is tested and carried as a negative, not used. Caveats: 75-year averages, regime-dependent; in the recent slice the relationship reverses (see the note below); long-record context, never a season-caller.
| Winter | Storms | Alignment | 75-yr percentile | z NATI | z AMO | z PDO |
|---|---|---|---|---|---|---|
| 1987/88 | 87J (Oct 1987)* | -0.43 | 26th | +0.91 | -0.21 | +0.59 |
| 1989/90 | Daria, Vivian, Herta | +0.16 | 56th | +1.00 | -1.11 | -0.36 |
| 1990/91 | Undine; run-on season | +0.76 | 87th | -0.10 | -1.02 | -1.17 |
| 1999/00 | Anatol, Lothar, Martin (Dec 1999) | +1.21 | 99th | -1.29 | -0.50 | -1.84 |
| 2006/07 | Kyrill (Jan 2007) | -0.30 | 36th | +0.91 | +0.54 | -0.56 |
| 2009/10 | Xynthia (Feb 2010) | -0.85 | 8th | +1.90 | +0.55 | +0.09 |
| 2013/14 | Serial UK winter | +0.29 | 64th | +0.22 | -0.95 | -0.15 |
| 2021/22 | Dudley, Eunice, Franklin | +0.90 | 91st | -1.34 | +0.45 | -1.82 |
| 2023/24 | High-latitude cluster year | -0.90 | 5th | +1.14 | +2.34 | -0.79 |
Grey: the long record; blue: hindcast-era winters; red: major insured-loss winters; amber: index-extreme, modest-loss. One-signed reading: alignment grades serial storm-family winters (99th/91st/87th percentiles) and does not flag single-monster winters – clustering-axis context, selected-by-loss so it motivates, never validates.
Everything that supports the findings without being one: the glossary, the lead-time curve, the window tests, the reference comparison, the geography breakdowns, the run-by-run record and the season-by-season roster.
Conviction day. A station-day on which Claros put the probability of exceedance at 1.5 times or more the station’s nominal base rate for that calendar day. Wind uses a Q85 threshold (base rate 15%, so the gate is 22.5%); rain and warmth use Q90 (base rate 10%, gate 15%). Realised winter exceedance runs nearer 18%, so in practice the wind gate is about 1.25 times the realised rate rather than 1.5.
Conviction rate. The share of station-days in a month that were conviction days, weighted by population as described below. This is the raw quantity every instrument is built from.
Population weighting. Each station carries a weight from the population living near it, so the network average leans towards where people and property are rather than treating a Highland gauge and a Rotterdam gauge as equals. It is a proxy for where damage concentrates, not client exposure data, and it is a proxy we test against: every headline result is rebuilt as a plain unweighted station average, and section 5 reports what moves.
Detrended calendar-day threshold. Each station’s exceedance threshold is computed on its own record with the climate trend removed, so a warming or windier baseline does not itself generate conviction days. This is what makes a conviction day a statement about this winter rather than about the decade.
The no-skill baseline (the Floor). For every forecast quantity we build the same quantity from climatology alone, with no forecast information in it. Every correlation in this report is a partial correlation: the Floor is regressed out of both the signal and the outcome first, and what is reported is the relationship that survives. So a number here is what the forecast adds over knowing the climatology, not the total predictability of the outcome.
Month standardisation. Each calendar month’s conviction rate is converted to a z-score against that same month at that same lead across the thirteen scored winters, and the months are then averaged. Without it, a month with naturally large swings dominates the season average and the quieter months contribute nothing. Section 5.5 of the methodology sets out the test that decides when it is required.
Instrument composite, and the concentration index. An instrument composite is the six standardised months of one variable averaged into one number for the season. The concentration index is the three instrument composites (wind, rain, warmth) averaged. Note the scale: it is built out of standard deviations but is not itself on a σ scale, because averaging correlated-but-not-identical components shrinks the spread. Its own standard deviation across the thirteen winters is about 0.38, so a reading of −0.73 is roughly 2.0 of its own standard deviations below normal.
Signal compression. Any imperfect forecast is narrower in spread than the outcome it predicts. We leave that compression in rather than stretching the forecast back out, which is why signal-side and observed-side σ are not interchangeable and why the alarm thresholds need their own calibration step.
Worst 7-day burst. The most exceedance-heavy seven-day window anywhere in the season, network-wide, expressed as a z-score. A seasonal maximum, so it answers “how bad was the worst fortnight” rather than “how bad was the season”.
Clumping. How unevenly the season’s exceedance was spread through time, network-wide. High clumping means the season’s activity arrived in a few concentrated spells rather than steadily.
Station clustering. A different question from the two above, asked locally: at one station, did exceedances arrive in quick succession? Counted per station, then averaged. A season can cluster strongly and concentrate weakly, and the reverse.
Station composite. Per station-winter, the highest of that station’s three z-scores for wind activity, storm clustering and peak severity, then averaged across the network: how bad the winter was at each station on whichever axis hit it hardest. Used for description and for the map, never as a prediction target, because it is built from the same station data as the instruments.
The catalogue. A 1950–2025 windstorm event catalogue we built from reanalysis. It is ours, not a vendor product, and it shares no data with the station archive, which is what makes it a genuinely separate scoring record. The catalogue extremity index ranks each winter against the other 75; catalogue events is the raw count; catalogue clustering counts events falling within 72 hours of another; catalogue severity is the strongest single event wind that winter, on a daily-mean basis.
The frozen 1972–2011 reference. A fixed forty-winter baseline used for every observed measure quoted as a σ, so that winters are compared against a common scale instead of each being measured against a moving average that includes itself. Burst participation uses 1974–2011 and the network indices use winters to 2011; all three end before the Claros hindcast library begins, so no scored winter contributes to the yardstick it is scored against. Claros readings are standardised on the 2013–2025 hindcast record, the only history the model has. The fidelity dials in section 3 are the one exhibit where the observed side is standardised on the thirteen winters instead, so that both axes of that chart share a footing; a correlation is unaffected by the location and scale of either variable, so the reading is the same either way. The one thing that varies between observed exhibits is the season window, October to March or November to March, and the report labels which and does not mix them inside a table.
Placed call. A forecast reading for a specific season and a specific geography, as it was issued, with no later adjustment. Every forward number in this report is a placed call.
Permutation test. Rather than assuming a distribution, we shuffle the signal series thousands of times and count how often chance reproduces the observed relationship. The signal and its Floor baseline are shuffled together, using the same index order, because shuffling the signal alone breaks their pairing and overstates significance.
Jackknife. Refitting the same relationship thirteen times, each time dropping one winter, and reporting the range. A result that depends on one winter shows up as a range that crosses zero.
Driver alignment, NATI, AMO, PDO. Teleconnection context from the long records, not model output. NATI is a North Atlantic circulation index; AMO and PDO are the Atlantic and Pacific multidecadal ocean modes. Driver alignment is a percentile ranking of how favourably these stood for European windstorm activity in a given winter. Carried as context only: nothing in the findings tab depends on it.
Aggregate severity. In the map panel, the sum of event severity across a winter rather than its maximum: a measure of total rather than peak.
Partial correlation of each issuance’s concentration index with the observed worst 7-day burst over the thirteen winters, each issuance scored against its own Floor. June reads −0.22, July −0.07, August +0.73, September +0.44 and October +0.30. The product uses August alone, and the reason is lead time rather than a claim that August is the best month: August and September are not separable on this sample (difference +0.27, p=0.52), and combining them gives +0.73, which is within noise of August by itself. August is chosen because it is the earlier of the two indistinguishable issuances and therefore the more useful. The July to August step is the only one the sample comes close to supporting: +0.82, p=0.042. Every other step on the curve is well inside noise, September and October included, so the curve is read as a rough shape with an onset, not month by month. Every issuance on the curve is extracted station by station with signal and Floor taken on the same station visit, so each point removes a baseline built from exactly the same stations rather than a nearby approximation to them.
The report is scored October–March because that is the season the market uses. Getting there required one decision that is not obvious, so the whole audit trail is here rather than summarised.
Before any forecast is involved: how much of the October–March season does the December–February core actually account for? Regressing each Oct–Mar measure on the same DJF measure over 51 winters of observations gives activity R²=0.77, worst 7-day burst R²=0.28, clumping R²=0.05. So a core-winter view is a fair proxy for how many windy days a season contains, and a poor one for how those days are arranged: roughly three quarters of the variance in the measure this report predicts sits outside the core season. That is measured from the observed record, with no reference to any forecast or correlation, and it is the reason the season has to be the full one.
The observed worst 7-day burst over October–March correlates +0.97 with the same measure over November–March. Extending the measured season barely changes what is being measured, because a seasonal maximum can only move one way when a month is added. So the target choice is not where the difficulty is.
| construction of the forecast side | index vs Oct–Mar burst | wind vs Oct–Mar burst | used? |
|---|---|---|---|
| Pool raw station-days across Oct–Mar | n/a | +0.06 | no, October swamps the total |
| Drop October from the forecast side, keep it in the target | +0.83 | +0.82 | no, a window chosen after seeing the result |
| Standardise each month, then average (all six months) | +0.71 | +0.68 | yes |
Why the weakest of the three is the one used. Pooling raw days fails for a measurable reason: the October wind conviction rate carries roughly 6.7× the level and 23× the standard deviation of the November–March rate, so October is 17% of the days and dominates the mean, a day-pooled October–March signal correlates −0.04 with its own November–March component. Standardising each month first fixes that, and the fix is justified by the variance ratio, which is measurable without reference to any correlation. The middle option scores best and is not used, because choosing which months feed the forecast after seeing which choice scores best is how a spurious result gets made. The test we apply to any methodology choice in this programme is whether it could have been justified before the correlation was looked at. The variance fix passes that test; dropping October does not.
| measure | forecast window | Oct | Sep | Aug | Jul |
|---|---|---|---|---|---|
| catalogue index | DJF | +0.56 | +0.20 | +0.10 | +0.22 |
| Nov–Mar burst | DJF | +0.43 | +0.26 | +0.44 | +0.17 |
| catalogue index | NDJFM | +0.36 | −0.03 | +0.04 | +0.02 |
| Nov–Mar burst | NDJFM | +0.35 | +0.31 | +0.65 | +0.03 |
These grids are from the window-selection work and are deliberately on other windows, DJF and November–March, because their purpose was to establish the rule that a measure should be predicted from a forecast covering the same months. That rule is why the report is scored October–March throughout, and why the one exception (the catalogue index, on a core-winter window) is flagged where it appears.
Claros has two working parts that can be measured separately. The climate reference, an era-relative, station-specific, calendar-day-specific threshold, rebuilt each year, is what makes any discrimination possible; the analogue weighting supplies the year-specific signal on top of it. This table isolates the first by holding the forecast fixed and swapping only the reference it is judged against.
| target | issued | our rebuilt reference | fixed 30-year normal | gain |
|---|---|---|---|---|
| catalogue extremity index | Oct | +0.54 | −0.25 | +0.79 |
| observed activity | Oct | +0.56 | −0.56 | +1.12 |
| Nov–Mar burst | Aug | +0.58 | −0.32 | +0.91 |
| Nov–Mar clumping | Aug | +0.57 | −0.11 | +0.68 |
Identical forecast, identical stations, identical winters. These are raw correlations, so the Nov–Mar burst figure is on the earlier window and basis rather than the no-skill baseline controlled out. Against a fixed 30-year normal the signal flags 17.5 days per station-winter instead of 0.14: the climate has drifted away from the stale reference, the signal saturates and all discrimination is lost. The rebuilt reference is a precondition for season picking, not a refinement of it.
The nine-country scope is the area European windstorm loss actually falls in, so it needs little defending. It was still worth testing rather than asserting, because “why only these countries?” is the first question a reviewer asks.
Six non-core regions were built from the wider library, Iberia, Alpine and central Europe, Finland and Iceland, Italy, the Baltics, south-east Europe, and then aggregated, because one large network is a different test from six small ones: a signal that was regime-scale and simply under-sampled by the core would show up as the network grew. It does the opposite. Against the concentration outcome, the nine-country core reads +0.77; adding Finland and Iceland gives +0.70; adding the Baltics as well, +0.55; the whole continent, +0.36. And everything except the core, aggregated into one 676-station network, reads −0.03. That monotone decline with dilution is the signature of adding noise to a real signal rather than of a signal needing more room.
And the skill is not the only thing that thins out. Measured in absolute wind speed rather than in each station’s own percentiles, the core network is the windiest part of Europe bar one. Taking the core as 1.00, the typical daily-max wind at its 85th percentile runs at 0.84 in Iberia, 0.82 in Italy, 0.73 in south-east Europe, 0.72 across the Alps and central Europe and 0.67 in the Baltics, with the 99th percentile of observed wind giving the same ordering. So the regions where no skill is found are also the regions with materially less windstorm to find, which is the more useful way round to state it: this is not a boundary drawn where the model happens to work, it is drawn where European windstorm happens. The single exception proves the point. Finland and Iceland are the one non-core group windier than the core, at 1.08, and they are also the one non-core group where any signal survived at all.
910 stations, 26 per country on a latitude-stratified sample, equal-weighted so no population kernel is assumed outside the core, August initialisation, thirteen winters, outcomes rebuilt from each network’s own daily record. Both wind and rain, against burst, clumping and activity: 48 non-core combinations, of which two reach p<0.05: one positive, one negative, in different regions on different measures with nothing else nearby. The individual regional numbers are not reproduced here because at 26 stations per country they are not strong enough to read one at a time; the aggregate is the test that answers the question. Two limits worth stating: this is decisive for wind and weaker for rain, which is the more weighting-sensitive instrument and this test uses no weighting; and no catalogue exists outside the core, so no non-core result could be checked against anything but the station data behind it.
Two separate regional breakdowns were built inside the nine countries, on different station selections and weightings. The point they make together is not about any one country.
Pooling is not a compromise, it is where the skill is. On the first breakdown the pooled nine-country call scores +0.57 against its own outcome, and that is higher than every individual country measured against its own outcome: Benelux +0.49, Great Britain +0.48, Denmark and France +0.45, Sweden +0.24, Germany +0.22, Ireland −0.05, Norway −0.18. The whole is better than any of its parts, which is the same finding as the scope test above read from the other direction. There is a footprint at which this signal is measurable, and cutting below it loses skill just as surely as diluting above it does.
Which is why no country figure is quoted. The two breakdowns disagree sharply once they are cut that fine: Benelux reads +0.49 on one basis and −0.49 on the other, Denmark +0.45 against −0.34, Great Britain +0.48 against +0.04. Germany and France agree. The difference is station selection and weighting inside each geography rather than different observations, and at 150 to 600 stations per country a handful of them moves a correlation a long way. Neither breakdown is reported as a result, and the report makes no claim about any country or sub-region. Reconciling the two is scheduled work; it does not change anything in the findings, all of which are the pooled call.
Each cell is the concentration index from that issuance alone, standardised on that issuance’s own thirteen-winter scored record and applied unchanged to 2026/27. Spread is the standard deviation across the issuances present. 2026/27 has three because September and October do not exist yet; on those same three months its spread ranks 4th of 14. The headline statistics from this table are quoted in the forward tab.
| winter | catalogue events | clustering | severity m/s | extremity pct | 7-day burst | summer index call | Oct wind call | driver alignment | notable storms | market context |
|---|---|---|---|---|---|---|---|---|---|---|
| 2013/14 | 16 | 8 | 23.3 (52 mph) | 61.8 | −0.07 | +0.08 | +0.23 | +0.29 (64th) | Xaver, Dirk (Dec); serial UK storms Jan-Feb (Anne, Christina) | |
| 2014/15 | 10 | 1 | 23.0 (51 mph) | 30.9 | +1.40 | +0.02 | −0.23 | -0.16 (40th) | Elon-Felix (Jan); Mike, Ole | |
| 2015/16 | 8 | 3 | 26.7 (60 mph) | 76.4 | +1.31 | −0.24 | +0.09 | -0.31 (35th) | Desmond (Dec, flood-led); Gertrude, Imogen (Feb) | |
| 2016/17 | 5 | 0 | 22.1 (49 mph) | 20.0 | −1.65 | −0.34 | −0.35 | -0.04 (48th) | Egon (Jan); Doris (Feb) | |
| 2017/18 | 6 | 3 | 23.6 (53 mph) | 35.5 | +0.50 | −0.05 | +0.62 | +0.35 (73rd) | Burglind, Eleanor, Friederike (Jan); Emma (late Feb) | |
| 2018/19 | 6 | 0 | 22.1 (49 mph) | 25.5 | +3.18 | +0.61 | +1.05 | +0.27 (62nd) | Alfrida (Jan); quiet core | |
| 2019/20 | 12 | 5 | 23.3 (52 mph) | 89.1 | +1.07 | +0.00 | +0.38 | +0.13 (53rd) | Ciara-Sabine, Dennis, Jorge (serial February) | |
| 2020/21 | 7 | 3 | 24.1 (54 mph) | 35.5 | −0.50 | −0.44 | +0.48 | +0.42 (74th) | Bella (Dec); wind-quiet, cold spells | |
| 2021/22 | 7 | 1 | 21.8 (49 mph) | 23.6 | +2.91 | +0.58 | −0.72 | +0.90 (91st) | Barra (Dec); Malik-Corrie (Jan); Dudley, Eunice, Franklin (Feb triple) | Dudley+Eunice+Franklin EUR 3.85bn industry (PERILS) |
| 2022/23 | 4 | 0 | 28.3 (63 mph) | n/a | −0.61 | −0.57 | −0.21 | +0.33 (70th) | Otto (Feb); mild season | |
| 2023/24 | 12 | 7 | 29.8 (67 mph) | 52.7 | +0.49 | +0.52 | −0.31 | -0.90 (5th) | Pia, Gerrit (Dec); Henk, Isha, Jocelyn (Jan). Ciaran was November (autumn) | Ciaran EUR 2.04bn industry (Nov, autumn; PERILS) |
| 2024/25 | 10 | 3 | 22.9 (51 mph) | 35.5 | −1.70 | −0.13 | −0.63 | -0.06 (44th) | Darragh (Dec); Eowyn (Jan, Irish wind records); Herminia | Eowyn EUR 765m industry final (PERILS) |
| 2025/26 | – | – | – | n/a | −0.65 | −0.04 | −0.39 | – | – |
The thirteen hindcast winters. Catalogue events, clustering and severity come from our reanalysis-built catalogue; the 7-day burst is the station measure; the summer index call and October wind call are the forecasts as issued. 2022/23 and 2025/26 are not yet indexed in the catalogue, which is why every catalogue figure in this report is scored on eleven winters rather than thirteen. Catalogue severity is daily-mean based while station data carries the daily maximum, so the two severity measures are never compared raw; the record table’s definitions cover what that column does and does not discriminate.
The season. October to March, on both sides of every comparison in the findings. That is the windstorm season as the market defines it and as named-storm contract wording follows it. Earlier internal work in this programme was scored on December–February and on November–March; those windows are not used here.
The forecast. Claros weights historical analogues on 28 teleconnection drivers and their interactions to issue daily exceedance probabilities months ahead. A conviction day is a station-day where that probability reaches 1.5× the nominal base rate against the station’s own rebuilt calendar-day threshold (wind: 85th percentile, nominal base 15%; rain and warmth: 90th, base 10%). The seasonal instrument is the population-weighted count of conviction days across the network.
How the six months are combined, and why it matters. Each month’s conviction rate is standardised on its own thirteen-winter record first, and the six standardised months are then averaged. The instrument index is the average of the three instruments built that way. This is not cosmetic. The alternative, pooling the raw station-days across all six months, fails, because the months are not on comparable scales: the October wind conviction rate has roughly 6.7 times the level and 23 times the standard deviation of the November–March rate, so in a day-pooled mean October is 17% of the days and dominates the total. Measured directly, a day-pooled October–March wind signal correlates −0.04 with its own November–March component: the rest of the season is not diluted, it is erased. Standardising first gives every month one equal voice and lets no month swamp the others. The justification is a property of the data, measured before any correlation was looked at, and no month is dropped from either side.
The thresholds. Era-relative and station-specific: each station’s own percentile for that calendar day, detrended, so an exceedance day means the same thing in 1975 as in 2020. Measured against these thresholds the realised winter exceedance rate is nearer 18% than the nominal 15%, so the conviction gate sits slightly lower than 1.5× realised. The rule is frozen and applied identically to every winter, and a higher bar was tested up to 8× and performs worse.
The baseline. Every skill number is a partial correlation with the no-skill track controlled out: the same climatology, same stations, same annual vintages, no analogue weighting. That removes the shared climate trend from both sides, so what remains is year-specific information.
Significance. By permutation. We shuffle which forecast is paired with which winter, twenty thousand times, and count how often a random pairing does as well as the real one. If a fifth of the shuffles beat it, the result is worth nothing; if one in three hundred does, p=0.003.
One detail matters and is easy to get wrong. Every headline number is measured after removing what the no-skill baseline, the same climatology with no analogue weighting, already explains. So when the winters are shuffled, each forecast has to keep its own baseline attached to it. If the baseline stays put while the forecast moves, the shuffled pairs are no longer a fair picture of chance: they have been stripped of a relationship the real pairs still have, so chance looks worse than it is and the result looks better than it is. Keeping each forecast and its baseline together is the strict version, it is what is used here, and it agrees with the textbook formula on every cell to within 0.006.
Wind is the common factor, and the three instruments are not three independent votes. On the construction used here, wind correlates +0.62 with rain and +0.81 with warmth at the August initialisation, while rain and warmth correlate only +0.24 with each other. At October all three are tighter, +0.63 to +0.87. So wind carries most of what the other two carry: once wind is in the model, rain adds +0.14 (p=0.69) and warmth adds −0.06 (p=0.85), neither significant. Two consequences, pulling in opposite directions. The concentration result must not be presented as three independent confirmations, because it is substantially one signal read three ways. But the index still earns its place over wind alone, because averaging partly-independent readings of one signal reduces noise even when none of them adds fresh information: the index scores +0.73 against wind’s +0.70.
Testing many things at once, and what we did about it. This is the risk that matters most in work like this, so it is worth being concrete rather than reassuring. If you test twenty unrelated things at the conventional threshold, roughly one will look significant purely by luck. This programme has tested several hundred instrument-and-measure combinations. So a single good-looking number, produced on its own, would be worth very little, and none is quoted that way.
Three specific checks stand behind the two results the report does quote, and each one is a thing a reader can ask for and re-run:
What is quoted, and what is not. Quoted as findings: the two index rows in section 4, the August index against October–March burst (+0.73) and against clumping (+0.63), because both pass all three checks. Carried as a developing result rather than a finding: October rain against the catalogue index, +0.5 to +0.6 with p between 0.06 and 0.13 on the eleven winters the catalogue covers. It does not clear the first check on that sample, and section 4 sets out why it is still the call worth watching. Not quoted at all: every single-instrument cell that reaches p<0.05 in one configuration and not another.
And one search that found nothing, which is the more useful evidence. Looking for a route to the exceptional-season family, we ran the full grid, three instruments across four initialisation months, two forecast windows and five different targets, 120 combinations, and recorded all of it rather than the best cell. Nothing reached significance with a positive sign. The point is not that 120 is a large number; it is that when this method is pointed at something that is not there, it returns nothing. That is the only real evidence available that the numbers it does return are not manufactured by searching.
Where this is tested, and where it is not. The hindcast library covers 5,300 stations across roughly forty European countries. Every scored number in this report is the nine-country north-west European network, Great Britain, Ireland, France, Germany, Netherlands, Belgium, Denmark, Norway, Sweden, 3,002 stations, with a 1,041-station population-weighted selection carrying the headline signal. Nothing here is a claim about any other part of Europe, or about any country or sub-region taken on its own. Whether the scope should be wider was tested rather than assumed; the test is in the workings tab.
Weighting. A 20-city population kernel (120 km Gaussian) over the north-west European core, used as the closest available proxy for where insured value sits. It is not client exposure data. Signal and outcome are weighted identically, and the burst result survives removing the weighting from the target altogether.