OSA1.5 | Machine Learning in Weather and Climate
Machine Learning in Weather and Climate
Including EMS Young Scientist Conference Award
Conveners: Noelia Otero Felipe, Sam Allen, Miguel-Ángel Fernández-Torres, Rodrigo Almeida, Richard Müller, Bernhard Reichert, Dennis Schulze, Gert-Jan Steeneveld, Roope Tervo | Co-convener: Angela Meyer
Orals Thu1
| Thu, 10 Sep, 09:00–10:30 (CEST)|Room Mission 2
Orals Thu2
| Thu, 10 Sep, 11:00–12:45 (CEST)|Room Mission 2
Orals Thu3
| Thu, 10 Sep, 14:30–16:30 (CEST)|Room Mission 2
Posters PS-Thu4
| Attendance Thu, 10 Sep, 16:30–18:00 (CEST) | Display Wed, 09 Sep, 14:00–Fri, 11 Sep, 13:00|TransitZone, P83–96
Thu, 09:00
Thu, 11:00
Thu, 14:30
Thu, 16:30
This session focuses on machine learning (ML) methodology applied in the meteorological context. The topics include model architectures, training strategies, uncertainty quantification, evaluation and validation schemes underpinning reliable ML for the Earth system. We aim to bring together contributors from meteorology, climate science, computer science, and applied mathematics who are advancing the theoretical and methodological foundations of ML for weather and climate. We also welcome approaches applied to weather extremes across time scales, with a strong emphasis on uncertainty quantification and probabilistic prediction in operational settings. In particular, we encourage studies that bridge AI-based forecasts with impact-based forecasting, risk assessment, and decision support, including applications to floods, droughts, heatwaves, storms, atmospheric rivers, and compound or cascading hazards.

We invite contributions on topics including, but not limited to:

* Novel model architectures with potential to be applied in meteorology/climatology.
* Novel applications of ML architectures for geophysical data.
* Training strategies and objectives
- including e.g. loss functions, self-supervision, pre-training and fine-tuning, transfer learning, and data augmentation, ...
* Integration of physical knowledge
- physics-informed and hybrid models, constraints and regularisation, stability and robustness, ...
* Uncertainty quantification and reliability
- probabilistic ML, ensembles, Bayesian approaches, decision-relevant evaluation, ...
* Evaluation strategies and evaluation studies
- Intercomparison of different architectures, comparison with physical methods, benchmark strategies.
* Interpretability, explainability and fairness
- methods to understand, diagnose and stress-test ML models...
* Human aspect -- how AI changes our work, organisations, and culture?
* ML and hybrid approaches for extreme event prediction
* Evaluation of AI forecasts for rare and high-impact events
* Integration of AI methods into operational workflows: Case studies demonstrating operational feasibility and societal benefits
* Translation of probabilistic AI forecasts into impact-based warnings and user-oriented products

Orals Thu1: Thu, 10 Sep, 09:00–10:30 | Room Mission 2

Chairpersons: Bernhard Reichert, Richard Müller
09:00–09:30
09:30–09:45
|
EMS2026-2
|
Onsite presentation
Shifa Mathbout, George Boustras, Pierantonios Papazoglou, Joan Albert Lopez Bustins, and Javier Martin Vide
This study analyses how climate stress, violent conflict, and socioeconomic instability interact to shape patterns of forced migration from the Middle East and North Africa (MENA) to the European Union (EU). Using high-resolution climate, conflict, and socioeconomic data covering the period 2000–2023, we develop an integrated empirical framework to identify the primary determinants of cross-border asylum flows. A machine-learning approach based on a Random Forest Model (RFM) is employed and benchmarked against the traditional Gravity Model (GM). By capturing nonlinear effects and complex interactions among drivers, the RFM substantially outperforms the GM, explaining more than 53% of the observed variation in migration flows.
 
The Middle East and North Africa (MENA) region has become a focal point of intersecting climate stress, violent conflict, and large-scale human displacement, with significant implications for regional stability and migration governance beyond its borders. This study examines how climate variability interacts with conflict and socioeconomic fragility to influence forced migration from MENA countries to the European Union (EU). Using high-resolution climate, conflict, and socioeconomic data spanning the period 2000–2023, the analysis develops an integrated empirical framework to identify the key drivers shaping cross-border asylum flows. 
To capture the complexity of migration dynamics, the study applies a machine-learning approach using a Random Forest Model (RFM) and compares its predictive performance with that of the conventional Gravity Model (GM). The results demonstrate that the RFM substantially outperforms the GM, explaining over 53% of the observed variation in asylum applications. This improvement reflects the model’s ability to account for nonlinear relationships and interactions among environmental, political, and economic variables.
 
The findings reveal that violent conflict and economic instability are the primary determinants of forced migration from the MENA region. Climate-related stressors, particularly prolonged droughts, rising temperatures, and declining agricultural productivity, do not independently trigger large-scale displacement. Instead, they function as threat multipliers that intensify existing vulnerabilities, weaken livelihoods, and exacerbate conflict-related pressures, thereby increasing the likelihood of forced mobility. These results underscore that migration surges emerge from the convergence of environmental stress, political instability, and socioeconomic deprivation rather than from climatic factors in isolation. 
 
By highlighting the compound nature of climate-related displacement, this study contributes to ongoing debates on climate security and migration governance. The findings emphasise the need for integrated policy responses that combine short-term humanitarian assistance with long-term investments in climate adaptation, conflict prevention, and economic resilience. A key limitation of the analysis is its exclusive focus on asylum applications to the EU, which excludes internal and regional displacement dynamics and relies on country-level predictive modelling that cannot establish causal relationships or capture subnational variation in climate–conflict interactions.

How to cite: Mathbout, S., Boustras, G., Papazoglou, P., Lopez Bustins, J. A., and Martin Vide, J.: Climate Stress, Conflict Dynamics, and Forced Migration from the MENA Region: Implications for Security and Sustainable Development, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-2, https://doi.org/10.5194/ems2026-2, 2026.

09:45–10:00
|
EMS2026-15
|
EMS Young Scientist Conference Award
|
Onsite presentation
Esma Nur Demirtaş and Barış Önol

Snow cover distribution in mountainous regions is characterized by high spatio-temporal variability, particularly in the form of ephemeral snow cover, which undergoes rapid accumulation and ablation cycles. In the Eastern Black Sea region of Turkey, these dynamics are increasingly influenced by climate change, leading to significant shifts in snowmelt timing and accelerated early-season runoff. Accurately monitoring these processes is constrained by the spatial resolution of regional reanalysis datasets. While the CERRA-Land dataset provides a robust and high-quality long-term climatic record, its 5.5 km spatial resolution is primarily optimized for regional-scale dynamics, which inherently limits the characterization of sub-grid orographic effects and the precise delineation of snow lines in topographically complex terrains. This study implements a novel deep learning framework based on Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN) to downscale CERRA-Land snow cover data to a 500 m resolution. The model is trained using multi-platform MODIS (Terra and Aqua) satellite imagery as the high-resolution ground truth, enabling the capture of diurnal variations and reducing cloud-cover interference. A fundamental aspect of this methodology is the integration of auxiliary topographic features, including elevation, slope, and aspect derived from Digital Elevation Models (DEM). Given the extreme vertical gradients of the Eastern Black Sea, these topographic variables are essential for the model to learn altitude-dependent snow distribution patterns and correct for complex shading effects. Quantitative evaluations demonstrate that the proposed framework significantly outperforms traditional spatial interpolation methods. Specifically, the topography-informed ESRGAN model reduced the Root Mean Square Error (RMSE) from a baseline of 49.20% to 17.40% and improved the Peak Signal-to-Noise Ratio (PSNR) from 6.16 dB to 15.20 dB, successfully reconstructing sharp textural details. Beyond its performance in spatial reconstruction, this method offers significant computational efficiency. By providing a high-fidelity and rapid downscaling mechanism, the framework can be seamlessly applied to historical climate reconstructions and future climate change simulations. Consequently, this research provides a robust foundation for high-resolution hydrological modeling, enabling better-informed water resource management and flood risk assessment under shifting climatic conditions.

How to cite: Demirtaş, E. N. and Önol, B.: Deep Learning Based Super-Resolution Spatial Downscaling of Snow Cover in the Eastern Black Sea Region, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-15, https://doi.org/10.5194/ems2026-15, 2026.

10:00–10:15
|
EMS2026-25
|
EMS Young Scientist Conference Award
|
Onsite presentation
Assaf Shmuel, Niklas Schwind, Kai Kornhuber, Ron Milo, and Carl-Friedrich Schleussner

Discernible differences in global climate responses under varying greenhouse gas emission scenarios are commonly assumed to emerge only after 20 to 30 years.  Here we show that mitigation benefits are detectable within a decade (9±6 years) over the global land area when high-resolution gridded climate data are analysed using a machine learning approach. Specifically, we train an ensemble of gradient-boosted decision tree models on CMIP6 simulations to distinguish between low- and intermediate-emissions scenarios (SSP1-2.6 and SSP2-4.5, respectively) using monthly near-surface air temperature fields. By retaining spatial information, we uncover regional warming signals that remain hidden when relying on global averages and identify the regions in which these signals first emerge using an explainability framework. As a performance baseline, we replace the machine learning approach with a logistic regression model using only global mean surface air temperature, which yields emergence timescales of about 30 years, consistent with previous studies. The spatial pattern of the timing of emergence shows pronounced regional contrasts, with the Tropics standing out as the earliest emerging regions. Even when restricting our analysis to subregions, we find a detectable signal to emerge over the land area of the four highest emitting countries in 13 (±6) years. These results demonstrate that detectable climate benefits of greenhouse gas mitigation appear much earlier than previously recognised and suggest that high emitting countries would also experience near-term benefits from bending the emissions curve. Demonstrating that mitigation produces a discernible climate response within a decade provides a clearer scientific basis for maintaining and accelerating ambitious emissions-reduction efforts.

How to cite: Shmuel, A., Schwind, N., Kornhuber, K., Milo, R., and Schleussner, C.-F.: Climate mitigation benefits emerge within a decade, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-25, https://doi.org/10.5194/ems2026-25, 2026.

10:15–10:30
|
EMS2026-136
|
Onsite presentation
Yuejian Zhu and Xiaohe An

Accurate prediction of tropical cyclones (TCs) is critical for disaster mitigation, yet data-driven AI weather models often struggle with robustness and sensitivity to initial perturbations. This study conducts a comprehensive sensitivity analysis of an AI-based weather model (Pangu-Weather) to evaluate its resilience to geographic location and initial perturbations in TC positions and intensities, framed against seasonal average background conditions. Experiments focus on TCs in the tropical Atlantic and North Western Pacific basins, the feedback of the experiments could help us to understand the dynamic properties and physical reasons of AI-based weather models furtherly and the size of the initial perturbation for numerical design of future global ensemble systems.

Key findings reveal that TC track and intensity predictions are highly sensitive to initial perturbations in both the Atlantic and North Western Pacific basins, though larger perturbations are required for the Atlantic basin than for the North Western Pacific basin. It is confirmed that a data-driven AI model does follow up the dynamics and physics mostly. Notably, small symmetric perturbations (intensity: ±1–3 hPa from the mean surface pressure anomaly) do not generate meaningful error growth (spread) in the Atlantic basin—contrary to typical numerical model behavior—but they suffice for the North Western Pacific. The position perturbations (±50 km) also deviate from expected model responses for both basins. Meanwhile, initial spinup and geographic adjustment are discussed as well. These results may align with prior studies suggesting that AI models lack explicit physical constraints and dynamic conservation. Future work could explore targeted perturbations to further assess their impact on TC forecasts.

How to cite: Zhu, Y. and An, X.: How Initial Perturbations Affect AI-Driven Tropical Cyclone Predictions?, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-136, https://doi.org/10.5194/ems2026-136, 2026.

Orals Thu2: Thu, 10 Sep, 11:00–12:45 | Room Mission 2

Chairpersons: Richard Müller, Bernhard Reichert
11:00–11:15
|
EMS2026-149
|
Onsite presentation
haixia qi

Numerical models often systematically underestimate the intensity of heavy precipitation and struggle to accurately represent the spatial pattern of rainbands over MLYR. This study converts the Spatial Fractional Skill Score (FSS) into a differentiable loss function and the traditional loss function Mean Squared Error (MSE) trained the U-Net deep learning architecture to construct a multi-model ensemble post-processing model that optimizes for spatial pattern and extreme intensity. For heavy precipitation (≥ 25 mm/d), a two-stage data-augmentation fine-tuning strategy is employed to further calibrate the multi-model ensemble outputs. We evaluated the performance of four models, ensemble mean (MEAN), and the multi-model ensemble U-Net(FSS) and U-Net(MSE) for daily precipitation forecasts. The results demonstrate that within 24–240 h lead time, U-Net(FSS) reduces the  averaged RMSE by 3–7% compared to the best individual model, with the improvement increasing as the forecast lead time extends. For heavy precipitation(≥25 mm/d), the FSS of U-Net improvesby approximately 10% over the best individual model at lead times of 24–72 h, maintaining an improvement of over 5% through 168 h. For extreme precipitation (95th percentile), the U-Net sustains TS around 0.28 at  168–240h lead times, improving by 10–20% compared to MEAN. A case study of rainstorms in summer 2024 reveals that the U-Net outperforms both individual models and MEAN in capturing the spatial structure of rainbands and the intensity of extreme centers at lead times of 24, 168, and 240 h. The multi-model ensemble U-Net shows significant potential for optimizing the prediction of both the spatial and intensity of heavy precipitation forecasting.The applicability and improvement in heavy precipitation forecasting were discussed to inform future applications of multi-model ensemble techniques.

Key words: TIGGE,Multi-Model Ensemble Forecasting,Spatial FSS Loss Function,U-Net Deep Learning Model,Heavy Precipitation Forecasting

How to cite: qi, H.: Multimodel Ensemble Heavy Precipitation Forecast with U-Net Deep Learning Model Integrating the Spatial FSS Loss Function, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-149, https://doi.org/10.5194/ems2026-149, 2026.

11:15–11:30
|
EMS2026-297
|
Onsite presentation
Arundhuti Banerjee and David Daou

Reliable and efficient flood mapping techniques are critical for disaster response and risk assessment. Satellite-based Earth observation, particularly Synthetic Aperture Radar (SAR) data from Sentinel-1, provides an effective means of monitoring flood events over large spatial scales, independent of cloud cover and daylight conditions. When combined with advanced deep learning algorithms, SAR imagery enables the generation of accurate and timely flood maps that support disaster mitigation and emergency response. In this study, we benchmark multiple flood segmentation architectures commonly used in Earth observation, including convolutional neural network (CNN)-based U-Net models with different encoder backbones and transformer-based vision architectures such as SegFormer. A multiclass segmentation model is initially trained on a Sentinel-1 flood dataset derived from NASA observations and subsequently fine-tuned on the Sen1Floods11 dataset to evaluate cross-dataset generalization and model robustness. Beyond segmentation accuracy, we investigate explainable deep learning approaches to better understand how models interpret SAR flood signatures. Specifically, we employ Gradient-weighted Class Activation Mapping (Grad-CAM) and entropy-based uncertainty estimation to analyze model attention and prediction confidence. Our results indicate that transformer-based architectures produce more spatially coherent flood predictions compared to traditional CNN models. In several test scenes, SegFormer captures large, continuous inundated regions more effectively, whereas CNN-based U-Net models tend to rely on localized radar backscatter patterns, often resulting in fragmented segmentation outputs. These findings underline the importance of architectural choice for robust flood mapping and demonstrate the added value of explainability in building trust for operational deployment.

How to cite: Banerjee, A. and Daou, D.: Explainable Flood Mapping from Sentinel-1 Imagery using CNNs and Vision Transformers, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-297, https://doi.org/10.5194/ems2026-297, 2026.

11:30–11:45
|
EMS2026-311
|
Onsite presentation
Miltiadis Kofinas and Dim Coumou

Machine learning-based weather prediction (MLWP) models have revolutionized weather modelling by providing forecasts that are orders of magnitude cheaper than numerical weather prediction (NWP) models, while achieving performance that matches or often surpasses state-of-the-art NWP models, including the ECMWF Integrated Forecasting Systems (IFS). Existing MLWP methods, however, are largely re-purposed computer vision architectures and not natively designed for Earth system data; thus, they suffer from important drawbacks. Notably, they still rely on re-analysis data, i.e. they implicitly still require NWP models to generate training data, which consequently leads them to inherit the biases of the underlying numerical models. Furthermore, they operate on dense--often planar--grids that introduce geometric distortions, while requiring constant resolution across the globe. Finally, most MLWP architectures are predominantly local-first, relying on mechanisms such as windowed attention or multi-mesh message passing, which can hinder their ability to model teleconnections and other long-range interactions. Neural fields--continuous fields parameterized by neural networks--have recently emerged as state-of-the-art representations for spatio-temporal modalities, offering an expressive alternative to grid-based representations. Neural fields are continuous in space and time, enabling arbitrary resolutions, and eliminating the need for grid-based input data. Moreover, they are fully differentiable, allowing the incorporation of physics-based loss functions to promote physical consistency. In this work, we introduce Neural Gaia, a neural field that encodes atmospheric data as a function of space--specifically latitude, longitude, and pressure level--and time, and use it to make predictions of future states of the atmosphere. Neural Gaia leverages the qualities of neural fields to model both local and long-range atmospheric phenomena without reliance on dense grids, paving the way for more flexible, accurate, and physically consistent ML-based weather prediction.

How to cite: Kofinas, M. and Coumou, D.: Neural Gaia: Neural Fields for Climate Modelling and Weather Forecasting, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-311, https://doi.org/10.5194/ems2026-311, 2026.

11:45–12:00
|
EMS2026-388
|
Onsite presentation
Francesco Pinto, Luca Lanzilao, Paco Lopez Dekker, and Angela Meyer

We present the first satellite-based wind forecast model, WindCastNet, and demonstrate that it outperforms state-of-the-art regional weather forecast models by lead times of up to ~2.5 hours in forecasts of 10-meter wind speed. Our model, WindCastNet, is a space-time aware convolutional long short-term memory network (ConvLSTM) that enables short-term wind nowcasting over offshore regions using exclusively satellite scatterometer observations (ASCAT and HSCAT) as input. Unlike geostationary satellites and ground radar, satellite scatterometers provide observations that are sparse and irregular in space and time, covering variable spatial domains and providing only a moderate training set size. 

WindCastNet overcomes these challenges by encoding spatial coverage masks, geographic coordinates, and inter-observation time intervals as explicit input channels, enabling the model to remain robust under varying observational configurations. From a training perspective, the limited availability of satellite data - approximately 13,000 usable overpasses between 2021 and 2025 - is addressed through a two-stage strategy: pretraining on ERA5 reanalysis fields to learn overall dynamics of 10-meter wind fields, followed by fine-tuning on scatterometer measurements. This transfer learning approach proves essential to achieve skillful forecasts, as demonstrated by the training diagnostics. Although the model is optimized for a 3-hour horizon, its recurrent architecture allows inference at longer lead times. 

Evaluation against HARMONIE-AROME (MEPS) over the North Sea shows 55% lower RMSE at 1h and 10% lower at 2h lead time, while HARMONIE-AROME exhibits lower forecast RMSE beyond 3h lead time. Compared to persistence, the model achieves 65%, 32%, and 21% RMSE reduction at 1h, 2h, and 3h, respectively. 
We further characterize the WindCastNet behavior across different meteorological regimes and discuss improvement potential, particularly at domain boundaries and for phenomena originating outside the training domain. Finally, we dicuss introducing physical constraints, outline pathways toward uncertainty, quantification and probabilistic extensions as natural next steps for operational integration. 

How to cite: Pinto, F., Lanzilao, L., Lopez Dekker, P., and Meyer, A.:  Satellite-based Offshore Wind Nowcasting: addressing spatio-temporal irregularities of scatterometer observations with Deep Learning , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-388, https://doi.org/10.5194/ems2026-388, 2026.

12:00–12:15
|
EMS2026-451
|
Onsite presentation
Fabian Schubert, Reinhold Hess, and Cristina Primo

Nowcasting (NWC) and numerical weather prediction (NWP) systems provide high-resolution forecasts, but their performance depends on lead time as well as on the underlying observations, model configurations, and physical parameterisations across spatial and temporal scales. The ensemble nowcasting system STEPS, operated by DWD, delivers highly accurate precipitation forecasts for the first few hours, whereas the regional ICON ensemble variants ICON-RUC-EPS and ICON-D2-EPS outperform STEPS at longer lead times. These systems differ in their physical parameterisations, computational cost, and forecast ranges (up to +48h for ICON-D2-EPS and +14h for ICON-RUC-EPS). Combining STEPS with these NWP models offers the potential to produce a seamless probabilistic forecast up to 48h that outperforms each individual system.

Within the SINFONY 3.0 project at DWD, we develop a machine learning approach to generate a calibrated, seamlessly blended probabilistic forecast of hourly precipitation. Radar observations from DWD’s network are used for training. The method builds on recent advances in machine learning-based post-processing: Grönquist et al. [1] use deep U-Nets for bias and spread estimation, Rempel et al. [2] focus on the blending of ensemble nowcasting and ensemble NWP  and Primo et al. [3] generate calibrated probabilistic distributions using a neural network and include additional contextual data features such as seasonal and orographic parameters to improve predictions.

We introduce a U-Net-based neural network architecture with context-dependent modulation that adaptively weights the contributing forecast systems. This allows the relative importance of each input source to be adjusted dynamically based on lead time, seasonality and orography. The inclusion of contextual information therefore supports calibration, temporal consistency, and overall forecast skill.

[1] Peter Grönquist, Chengyuan Yao, Tal Ben-Nun, Nikoli Dryden, Peter Dueben, Shigang Li, and Torsten Hoefler. Deep learning for post-processing ensemble weather forecasts. Philos. Trans. A Math. Phys. Eng. Sci., 379(2194):20200092, April 2021.
[2] Martin Rempel, Peter Schaumann, Reinhold Hess, Volker Schmidt, and Ulrich Blahak. Adaptive blending of probabilistic precipitation forecasts with emphasis on calibration and temporal forecast consistency. Artificial Intelligence for the Earth Systems, 1(4), October 2022.
[3] Cristina Primo, Benedikt Schulz, Sebastian Lerch, and Reinhold Hess. Comparison of model output statistics and neural networks to postprocess wind gusts. In: A. Ott, W. Reichel, J. Warwicker, editors. Applications of Mathematics in Sciences, Engineering, and Economics Cham: Springer Nature Switzerland; 153–180; 2026.

How to cite: Schubert, F., Hess, R., and Primo, C.: Context-Aware Blending for Seamless Probabilistic Precipitation Forecasting, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-451, https://doi.org/10.5194/ems2026-451, 2026.

12:15–12:30
|
EMS2026-484
|
Onsite presentation
Seppo Pulkkinen, Heikki Myllykoski, Calum Baugh, and Marc Berenguer

We present probabilistic warning tools for flood hazard and risk based on deep learning techniques and weather radar observations. This work has been done in the EU-funded INLINE project. The scope is on nowcasting heavy rainfall and associated flash floods in short time ranges (0-3 hours) and at high spatial and temporal resolutions (1-2 km and 5-15 minutes). Pan-European precipitation nowcasts are produced from OPERA rain rate composites using a convolutional neural network based on the SimVP architecture. Training the network with alternative loss functions together with three post-processing techniques are applied to improve the utility of the nowcasts. Underestimation of localized heavy precipitation is reduced by using a cumulative distribution transformation. Stochastic post-processing is then applied to produce ensemble members that reproduce the lost spatial variability. The ensembles are generated by Fourier-filtering and transforming white noise fields to reproduce the spatiotemporal correlation structure and the distribution of forecast errors. Finally, exceedance probabilities estimated from the ensembles are calibrated by using logistic regression. We present a verification study to determine the maximum time ranges of useful skill across different spatial scales and accumulation periods and to show that the proposed methodology outperforms simple extrapolation-based techniques in most cases. We also give examples of verification metrics where this is not the case. Precipitation rates from deterministic nowcasts are translated into color-coded hazard levels by using user-specified thresholds. The hazard levels are further translated into flood risk levels by using exposure information (i.e. population or critical infrastructure). In addition, we present different methods to translate ensemble nowcasts into probability-aware predictions of hazard and risk and give discussion of their practical utility. Practical use of the proposed methodology is demonstrated by using major flood events during the years 2024 and 2025 that affected multiple European countries.

How to cite: Pulkkinen, S., Myllykoski, H., Baugh, C., and Berenguer, M.: Probabilistic precipitation and flood risk nowcasting on pan-European scale by using deep learning, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-484, https://doi.org/10.5194/ems2026-484, 2026.

12:30–12:45
|
EMS2026-487
|
Onsite presentation
Wout Dewettinck, Dieter Van den Bleeken, Michiel Van Ginderachter, Leon Adriaensen, Hans Van De Vyver, Steven Caluwaerts, and Piet Termonia

Recent advances in data-driven weather and climate modelling have challenged the traditional dominance of physics-based numerical prediction systems. Machine-learning-based models, such as AIFS, have demonstrated competitive or superior performance for a range of standard forecast metrics. However, these evaluations primarily focus on large-scale variables under typical conditions, while the representation of local, high-impact extreme events remains insufficiently explored.

This study investigates the ability of data-driven models to represent extreme precipitation and to generalise beyond the conditions encountered during training. A graph neural network model, trained using the Anemoi framework, is developed based on output from a convection-permitting (4 km) regional climate simulation with the ALARO model over Western Europe. Model performance is quantified using precipitation return levels derived from annual maxima across multiple accumulation durations, ranging from hourly to multi-day timescales. This experimental design enables a systematic comparison between data-driven and physics-based representations of extremes.

To explicitly assess generalisation, a second model is trained on a modified dataset in which the upper tail of the precipitation distribution is selectively masked. By comparing return levels from both data-driven models with those from the reference simulation, we evaluate the extent to which such models can reproduce extremes that are absent from the training data.

This framework provides a controlled setting to assess the robustness of AI-based models for high-impact precipitation events and their ability to generalise to unseen conditions. This is particularly relevant for extreme precipitation, which is inherently underrepresented in training data and is expected to evolve under climate change, potentially leading to conditions outside the historical training distribution.

How to cite: Dewettinck, W., Van den Bleeken, D., Van Ginderachter, M., Adriaensen, L., Van De Vyver, H., Caluwaerts, S., and Termonia, P.: Can AI Weather Models Extrapolate Extremes? Evaluating Generalisation to Unseen Precipitation Events, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-487, https://doi.org/10.5194/ems2026-487, 2026.

Orals Thu3: Thu, 10 Sep, 14:30–16:30 | Room Mission 2

Chairpersons: Noelia Otero Felipe, Roope Tervo
14:30–14:45
|
EMS2026-550
|
Onsite presentation
Cristian Lussana, Even Marius Nordhagen, Rolf Heilemann Myhre, Thomas N. Nipen, Amélie Neuville, Ivar A. Seierstad, and Christian Lessig

WeatherGenerator is a pan-European initiative that integrates advanced machine learning architectures with high-performance computing to build an open, kilometer-scale foundation model of the coupled Earth system. The project is structured into four thematic areas; the application presented here represents one of the twenty-two developed within Theme 3 by project partners.

The Norwegian Meteorological Institute (MET Norway) plans to use the foundation model to reconstruct several decades of atmospheric variables over Scandinavia, ideally covering the most recent 40 years. The main goal is to apply the WeatherGenerator framework for climate monitoring purposes. The reconstructed datasets comprise near-surface variables as well as atmospheric variables across multiple pressure levels. This approach combines diverse data sources (including in situ observations, reanalysis datasets, and numerical model outputs) by exploiting the foundation model to produce coherent, high-resolution fields for both weather and climate monitoring. A target spatial resolution of a few kms is achieved through data fusion methods and a dedicated tail network trained to generate gridded analyses at this scale.

This contribution aims to present MET Norway’s results from the initial phase of the WeatherGenerator project. The emphasis is therefore on preliminary findings, outlining the methodological framework and highlighting the potential of these approaches for high-resolution climate monitoring.

------------------------------------------

Note: The WeatherGenerator project (grant agreement No101187947) is funded by the European Union. Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the Commission. Neither the European Union nor the granting authority can be held responsible for them.

How to cite: Lussana, C., Nordhagen, E. M., Heilemann Myhre, R., Nipen, T. N., Neuville, A., Seierstad, I. A., and Lessig, C.: Use of the WeatherGenerator foundation model for atmospheric reconstruction over Scandinavia, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-550, https://doi.org/10.5194/ems2026-550, 2026.

14:45–15:00
|
EMS2026-567
|
Onsite presentation
Çağlar Küçük, Brigitta Goger, Pascal Gfäller, Irene Schicker, and Alexander Kann

Machine learning has been transforming weather prediction at an unprecedented pace, driven by advances in modelling approaches, rapid progress enabled by open-source collaboration, and growing availability of high-quality atmospheric datasets across global to regional scales. One prominent example at the crossroads of these developments is the anemoi framework, which supports both model development and practical application in data-driven weather prediction. 

In this contribution, we present our experiences applying data-driven weather prediction using the anemoi framework for the Greater Alpine Region centred on Austria. Specifically, we use a multi-stage training pipeline to reduce training costs while learning atmospheric dynamics from reanalysis datasets with varying spatiotemporal coverage and resolution using a deterministic architecture. In addition, we extend this transfer learning approach to increase temporal resolution of the forecasting model and to fine-tune it using other datasets complementing reanalysis datasets.

Our data-driven models achieve consistently better scores compared to operational numerical weather prediction models running in-house, although capturing extreme values and spatial structures of predicted fields at high resolution remains a challenge. We analyse our models with increasing temporal resolution from 6- to 3- hours and show that error growth with increasing lead time is not a critical issue within the 72-hours lead time required for our models. We also discuss options for extending the set of predicted variables by fine-tuning with complementary datasets and present initial results. Finally, we provide an initial assessment of our model for real-time predictions over the convective season of 2026 to discuss the path towards operationalisation. These findings aim to support future design choices and contribute to a clearer understanding of the capabilities and limitations of data-driven weather prediction models.

How to cite: Küçük, Ç., Goger, B., Gfäller, P., Schicker, I., and Kann, A.: Data-Driven Weather Prediction for the Greater Alpine Region Using the Anemoi Framework, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-567, https://doi.org/10.5194/ems2026-567, 2026.

15:00–15:15
|
EMS2026-577
|
Onsite presentation
Robin Richardson, Bart Schilperoort, Peter Kalverla, and Gijs van den Oord

The success and wide-spread availability of machine learning approaches in earth and environmental sciences has resulted in a proliferation of deep learning models adapted for weather prediction. Neural networks have proven highly successful for a multitude of data-driven tasks such as bias correction, downscaling and nowcasting. 

Inspired by recent advances in generative modeling of textual data through large language models, the EU WeatherGenerator project aims to develop the leading European AI foundation model for weather and atmospheric climate modeling. This model is pre-trained with petabytes of multi-modal data (reanalyses, station observations, satellite products, etc.) necessitating training on powerful computing clusters, including Europe’s first exascale-class supercomputers. 

In pursuit of this “foundation model” status, and in order to evaluate the flexibility and added value of the model, it is necessary to integrate the WeatherGenerator to a wide range of existing forecast pipelines. The pilot application we present here is nowcasting of heavy precipitation in Western Africa. Dense networks of (openly available) automated weather stations and radars cannot be found in every region of the world. In sub-Sahara Western Africa, a lack of available radar data means that nowcasting is mostly limited to available satellite products, complicating efforts to predict the risk of flash floods from high-intensity precipitation. 

We investigate fine-tuning the pre-trained WeatherGenerator to SEVIRI output, training a tail network that predicts rainfall retrieval from the MSG-CPP product. We also explore transfer learning with WeatherGenerator, using a decoder trained to EURADCLIM over the European continent with SEVIRI input and assessing its accuracy over the target region. 

Finally, we present the details of our upcoming service call, through which the Netherlands eScience Center plans to bring the WeatherGenerator technology to potential stakeholders. The service call will be open to applicants from a broad range of sectors including the European research community, public institutions and industry, and provide for both (short term) deployment support as well as (longer term) explorative research applications. 

How to cite: Richardson, R., Schilperoort, B., Kalverla, P., and van den Oord, G.: The WeatherGenerator foundation model - pilot application and service call , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-577, https://doi.org/10.5194/ems2026-577, 2026.

15:15–15:30
|
EMS2026-582
|
Onsite presentation
Aram Farhad Shafiq Salihi, Sophie Buurman, Even Norhagen, Mario Santa Cruz, Michiel Van Ginderachter, and Thomas Nipen

The domain of weather forecasting is currently undergoing a significant transformation driven by advances in machine learning. State of the art data-driven weather models have demonstrated performance that surpasses the state of the art traditional numerical weather prediction (NWP) models, while operating at a fraction of the computational cost (Bouallegue et al., 2024). 

While global models like AIFS (Lang et al.,2024), GraphCast, and Pangu have gained a lot of attention, different flavours of high-resolution regional modeling have emerged  developed by different meteorological institutes. Stretched-grid is a global model with an increased spatial resolution and dynamics over a region of interest, a novel approach to regional modelling. This capability is demonstrated to be highly competitive and even surpass state of the art regional NWP for certain variables (Nipen et al., 2025; Nordhagen et al., 2025).

Building on the concept of stretched-grid and generalizing the idea to incorporate more high-resolution data, while avoiding intermediate fine-tuning and transfer learning steps, one can achieve high resolution predictions for any domain in Europe.  In this work, we propose a probabilistic multi-domain model, introducing a new way of training across domains and resolutions, by utilizing the concept of dynamical graphs. This is done by alternating between different global and regional data and its corresponding graph across different spatial regions, grid-types and resolutions. This capability is enabled through a flexible encoder, processor and decoder architecture. The idea is to make the model less biased towards certain grid types, terrain and dynamics induced by the training data, enabling the model to generalize across resolutions and climate zones. 

The kilometre scale model has been trained on analyses from four high resolution NWP models covering various parts of Europe, including AROME-Artic (2.5km), the MetCoOp Ensemble Prediction System (MEPS, 2.5km), UWC-West (2km) and the Austrian Reanalysis (ARA, 2.5km). We evaluate the performance of this model on another domain that was left out of the training, and show improved generalizability when compared to a model trained on fewer domains.  With one year of verification, we show improvements across parameters such as 2m temperature, mean sea level pressure, 10m wind speed and total precipitation.

How to cite: Salihi, A. F. S., Buurman, S., Norhagen, E., Santa Cruz, M., Van Ginderachter, M., and Nipen, T.: Multi-domain: A dynamic way of training across domains and resolutions, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-582, https://doi.org/10.5194/ems2026-582, 2026.

15:30–15:45
|
EMS2026-588
|
Onsite presentation
Serban Vadineanu, Bastien François, Sophie Buurman, Jasper Wijnands, and Maurice Schmeits

Global machine learning weather prediction (MLWP) models have recently achieved impressive predictive performance, in some cases outperforming global numerical weather prediction (NWP) models at lower computational costs. Initial approaches focused on developing deterministic MLWP models. Although these models perform well on deterministic metrics, they often produce overly smooth forecasts due to the use of MSE loss and do not capture uncertainty. To address these limitations, recent efforts have shifted toward developing ensemble MLWP models. For example, ECMWF has developed AIFS ENS, a variant of its Artificial Intelligence Forecasting System (AIFS), which is capable of producing skillful global ensemble forecasts. However, most of the global MLWP ensemble models operate at coarse spatial resolution, limiting their ability to provide actionable information at fine spatial scales. This capability is crucial for national meteorological institutes such as KNMI, where high-resolution probabilistic forecasts are required for risk assessment and weather warnings. Building on recent advances in stretched-grid techniques, this study develops a pre-trained European ensemble model by combining a stretched-grid framework with ensemble generation methods. The model is trained on ERA5 and CERRA reanalysis datasets (o96 (1 degree) and 5.5 km resolution, respectively), leveraging their extensive multi-decade archives. We investigate how different CRPS loss function designs (such as multi-scale and spectral variants) affect the quality of probabilistic forecasts, with a focus on calibration, sharpness, and mitigating spatial smoothness. The resulting pre-trained model is designed to serve as a computationally efficient starting point for subsequent stretched-grid finetuning on higher-resolution reanalysis datasets, allowing meteorological institutes to streamline the development of MLWP ensemble models. This work supports the mission of KNMI to deliver skilful probabilistic forecasts for weather warnings, and contributes toward enabling operational, data-driven, high-resolution ensemble forecasting.

How to cite: Vadineanu, S., François, B., Buurman, S., Wijnands, J., and Schmeits, M.: Pretraining high-resolution European ML weather ensembles with a stretched-grid framework, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-588, https://doi.org/10.5194/ems2026-588, 2026.

15:45–16:00
|
EMS2026-643
|
Onsite presentation
Konstantinos V. Varotsos, Jun She, Margaux Emma Hilt, Gianmaria Sannino, Andrea Orlandi, Kostas Tsiaras, and Christos Giannakopoulos

Regional ocean simulations require high-resolution atmospheric forcing, whereas CMIP6 climate projections are typically too coarse in both space and time for direct application. In the MOIRAI project, we developed a spatio-temporal statistical downscaling framework to produce multi-decadal, high-resolution 3-hourly atmospheric forcing over the Mediterranean, North Sea, and Arctic domains. The methodology consists of two main steps. First, daily climate model fields were spatially downscaled to the target regional grids and bias-adjusted using high-resolution reanalysis datasets as reference. Second, the daily bias-adjusted fields were temporally disaggregated to 3-hourly resolution using machine-learning models trained on reanalysis data for 1985–2014. The framework was applied to key forcing variables, including 2 m temperature, relative humidity, mean sea level pressure, 10 m wind, precipitation, and radiative fluxes. A range of machine-learning methods was evaluated, including support vector machines, random forests, k-nearest neighbours, XGBoost, and neural networks. As the overall predictive skill was broadly similar among methods, XGBoost was selected for implementation because it provided the best compromise between performance, computational efficiency, and scalability for long transient climate simulations.

The results show that the framework reproduces realistic sub-daily variability and coherent spatial patterns, particularly for variables with a pronounced diurnal cycle, such as temperature, relative humidity, and radiative fluxes. The analysis also highlights clear differences in reconstruction skill among variables. Predictors at daily resolution provide stronger constraints for variables dominated by the day–night cycle, whereas sub-daily reconstruction is more challenging for variables such as mean sea level pressure, whose variability is largely driven by synoptic-scale dynamics rather than diurnal forcing. In these cases, the method preserves the large-scale daily structure, but the exact hour-to-hour evolution may differ from the reference reanalysis.

Overall, the proposed framework offers a computationally efficient approach for transforming coarse-resolution global climate projections into high-resolution atmospheric forcing suitable for regional marine and coastal climate applications.

Funding This work was supported by the European Union through the MOIRAI project, Grant Agreement No. 101180994.

How to cite: Varotsos, K. V., She, J., Hilt, M. E., Sannino, G., Orlandi, A., Tsiaras, K., and Giannakopoulos, C.: A machine-learning framework for spatio-temporal statistical downscaling of CMIP6 climate projections, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-643, https://doi.org/10.5194/ems2026-643, 2026.

16:00–16:15
|
EMS2026-626
|
Onsite presentation
Tamara Happé, Jasper Wijnands, Paolo Scussolini, Peter Pfleiderer, and Dim Coumou

Due to climate change, heatwaves are becoming more frequent and intense, with western Europe experiencing the strongest trends in the Northern Hemisphere mid-latitudes. Part of the temperature trends are caused by circulation changes. Here we deploy Deep Learning techniques to classify European heatwaves based on their atmospheric circulation and to study their associated changes over time. We train a Variational Autoencoder (VAE) to reduce the dimensionality of the heatwave samples, on large ensemble climate model simulations. We show that the VAE generalizes well to observed heatwave circulations in ERA5 reanalysis, without the need for transfer learning. We then use probabilistic clustering techniques to find physically distinct heatwave types in ERA5 reanalysis. The circulation features relevant for heatwaves in ERA5 are consistent with the climate model heatwaves. Regression analysis reveals that the Atlantic Plume type of heatwaves are becoming more frequent over time, while the Atlantic High heatwaves are becoming less frequent. Then, we introduce new and straightforward interpretability methods to study the latent space, including feature importance identification for the clustering model. We investigate which circulation features are associated with the most important nodes in the latent space and how the latent space changes over time. For example, we find that the Atlantic Low heatwave shows a deepening of the low pressure system off the Atlantic coast over time. Each heatwave type is undergoing unique changes in their circulation, highlighting the necessity to study each heatwave type separately. Our method can furthermore be used to boost specific aspects of extreme events, and we illustrate how heatwave circulation could change in the future if the current trends persist, with in some cases an intensification of features.

How to cite: Happé, T., Wijnands, J., Scussolini, P., Pfleiderer, P., and Coumou, D.: Interpretable Deep Learning for European heatwave dynamics, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-626, https://doi.org/10.5194/ems2026-626, 2026.

16:15–16:30
|
EMS2026-650
|
Onsite presentation
Eliot Walt, Miltiadis Kofinas, Nikolaj Mücke, Efstratios Gavves, and Dim Coumou

In recent years, deep learning-based weather prediction (DLWP) systems have gained significant traction, rivalling their physics-based counterparts for a fraction of the computational cost. Nevertheless, the potential of current DLWPs is hindered by important limitations. For instance, they primarily rely on numerical weather prediction (NWP) model outputs as sources of training data and usually offer a very limited set of prognostic variables compared to standard NWPs. Furthermore, they cannot easily be coupled with other models and modalities. Perhaps most importantly, the majority of the DLWPs are deterministic. However, given the inherently chaotic nature of atmospheric dynamics, DLWP outputs are often unusable as they fail to capture the range of plausible futures. Classical probabilistic methods, such as input perturbation, are useful to a certain degree, but may fail to produce enough spread. Conversely, training natively stochastic DLWP models from scratch is often prohibitively expensive. In this work, we explore the idea of converting a deterministic DLWP into a probabilistic model. We present Xaurora, an operational stochastic DLWP based on Aurora, a 1.3 billion-parameter, transformer-based Earth system foundation model. Using the stochastic interpolant framework, we fine-tune Aurora as a drift prediction model using low-rank adaptation (LoRA). We compare our model with state-of-the-art NWPs and DLWPs on a variety of weather and climate modelling tasks and achieve competitive performance. Our main contribution is to demonstrate that large-scale DLWPs can be effectively converted into stochastic models at a fraction of the cost of training from scratch. Xaurora also paves the way for further research into the intersection of weather and climate modelling and generative models, which could help to overcome the current limitations of DLWPs and move beyond the forecasting-only paradigm.

How to cite: Walt, E., Kofinas, M., Mücke, N., Gavves, E., and Coumou, D.: Xaurora: Probabilistic Weather Forecasting with Foundation Models and Stochastic Interpolants, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-650, https://doi.org/10.5194/ems2026-650, 2026.

Posters: Thu, 10 Sep, 16:30–18:00 | TransitZone

Display time: Wed, 9 Sep, 14:00–Fri, 11 Sep, 13:00
Chairpersons: Rodrigo Almeida, Roope Tervo
P83
|
EMS2026-67
Bu-Yo Kim, Hae-Jung Koo, Joo Wan Cha, and Seungbum Kim

Dead fuel moisture (FM) strongly influences wildfire ignition and spread, but direct measurements remain sparse and difficult to obtain over large mountainous regions. This highlights the need for spatially continuous FM estimation using satellite and ground-based observations. Gangwon State, located in northeastern South Korea, is characterized by predominantly forested and mountainous terrain with complex topographic variability, making high-resolution FM estimation particularly important. We developed tree-based machine learning (ML) models to estimate 10-h FM in Gangwon State, South Korea, using AWS meteorological observations, GK-2A satellite radiance, and temporal and topographic variables. These input variables were designed to reflect atmospheric conditions, local characteristics, and diurnal variations relevant to FM dynamics. Random forest, extreme gradient boosting, and light gradient boosting models were trained and optimized using five-fold cross-validation. The models achieved high accuracy for 10-min interval estimates, with RMSE values of 1.20–1.41% and R2 values of 0.94–0.96. The extreme gradient boosting model showed the best overall performance, while the ensemble of the three models provided stable and competitive estimates. The models also performed well in identifying the wildfire-risk threshold of FM < 10% (equitable threat score = 0.78; accuracy = 0.95). The classification results further indicate that the models can reliably distinguish critically dry fuel conditions associated with elevated wildfire risk. These findings suggest that integrating geostationary satellite data with ground observations and tree-based ML can support high-resolution FM estimation and wildfire risk applications in complex terrain.

 

Acknowledgments: This work was funded by the Korea Meteorological Administration Research and Development Program “Research on Weather Modification and Cloud Physics” under Grant (KMA2018-00224).

How to cite: Kim, B.-Y., Koo, H.-J., Cha, J. W., and Kim, S.: Estimating 10-h Dead Fuel Moisture in Gangwon State, South Korea, using GK-2A Satellite Data and Ground Observations with Tree-Based Machine Learning, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-67, https://doi.org/10.5194/ems2026-67, 2026.

P84
|
EMS2026-124
Johannes Marian Landmann, Roman Attinger, Tiago Hungerland, Gabriela Aznar Siguan, and Ulrich Hamann

Towering Cumulus clouds (TCUs) pose a significant challenge to aviation because they are associated with strong vertical motion and should therefore be avoided during all flight phases. Beyond the meteorological challenges, TCU nowcasting is further complicated by the large variety of existing definitions, which differ across user groups and operational purposes. In the context of TCU nowcasting for aviation, it is important to identify a minimum viable definition that satisfies the needs of different users, ranging from air traffic controllers to pilots and meteorological forecasters. A common distinction is that TCUs do not produce lightning, while storms are classified as Cumulonimbus as soon as lightning activity begins.  However, current operational convection forecasts at MeteoSwiss are trained on lightning-based Cumulonimbus/thunderstorm labels and therefore miss TCUs entirely. 

In this work, we explore an impact-oriented nowcasting approach that explicitly targets TCUs as observed by aircraft onboard weather radars. Our approach is to retrain an existing recurrent convolutional neural network (RCNN), originally designed for lightning and hail prediction, on radar-derived TCU labels in a transfer-learning framework.  The training dataset is generated by a TCU detection algorithm that is already used operationally by MeteoSwiss for METAR (METeorological Aerodrome Report) products at the airports of Geneva and Zurich. In this product, TCUs are identified from Swiss radar images using two-dimensional maximum radar reflectivity fields combined with spatial Fourier bandpass filtering and, optionally, three-dimensional radar information. 

To obtain spatially consistent training data, we extend the currently airport-centered automatic METAR processing to cover the entire Swiss domain. In an initial step, the computationally expensive 3D component of the algorithm is omitted, allowing us to focus on a more efficient 2D radar-based classification approach. This approach enables the generation of multi-year TCU datasets from radar archives that can serve as training data for machine learning models. 

Based on these radar-derived TCU fields, we retrain the RCNN to produce spatially explicit probabilistic forecasts of TCU occurrence over Switzerland, with a temporal resolution of 10 minutes and lead times of up to 4 hours. We present the full data pipeline, from radar and lightning archives to TCU ground-truth generation and model training, and we discuss the potential of combining radar-based classification with deep recurrent networks. This approach bridges the gap between radar observations, operational aviation definitions of convection, and probabilistic machine learning-based impact forecasting. 

How to cite: Landmann, J. M., Attinger, R., Hungerland, T., Aznar Siguan, G., and Hamann, U.: Radar-based ground truth for deep learning Towering Cumulus (TCU) nowcasting: towards a minimum definition for aviation , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-124, https://doi.org/10.5194/ems2026-124, 2026.

P85
|
EMS2026-211
Yamin Hu, Dangfu Yang, Xinru Liu, and Shengjun Liu

Convolutional neural networks (CNNs) have been widely studied and found to obtain favorable results in statistical downscaling to derive high-resolution climate variables from large-scale coarse general circulation models (GCMs). However, there is a lack of research exploring the predictor selection for CNN modeling. This paper presents an effective and efficient greedy elimination algorithm to address this problem. The algorithm has three main steps: predictor importance attribution, predictor removal, and CNN retraining, which are performed sequentially and iteratively. The importance of individual predictors is measured by a gradient-based importance metric computed by a CNN backpropagation technique, which was initially proposed for CNN interpretation. The algorithm is tested on the CNN-based statistical downscaling of monthly precipitation with 20 candidate predictors and compared with a correlation analysis-based approach. Linear models are implemented as benchmarks. The experiments illustrate that the predictor selection solution can reduce the number of input predictors by more than half, improve the accuracy of both linear and CNN models, and outperform the correlation analysis method. Although the RMSE (root-mean-square error) is reduced by only 0.8%, only 9 out of 20 predictors are used to build the CNN, and the FLOPs (Floating Point Operations) decrease by 20.4%. The results imply that the algorithm can find subset predictors that correlate more to the monthly precipitation of the target area and seasons in a nonlinear way. It is worth mentioning that the algorithm is compatible with other CNN models with stacked variables as input and has the potential for nonlinear correlation predictor selection.

How to cite: Hu, Y., Yang, D., Liu, X., and Liu, S.: Predictor Selection for CNN-based Statistical Downscaling of Monthly Precipitation, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-211, https://doi.org/10.5194/ems2026-211, 2026.

P86
|
EMS2026-372
Tim Radke, Johanna Baehr, and Marc Rautenhaus

Reliable detection of atmospheric features such as tropical cyclones, atmospheric rivers, and atmospheric surface fronts is important for weather forecasting and climate analysis. Traditionally, this detection relies on expert knowledge or rule‑based approaches. Artificial neural networks (ANNs) offer a new method for detection. While they detect atmospheric features fast and often accurate, they are black‑box systems, making it difficult to assess whether they identify atmospheric features based on physically meaningful patterns or spurious correlations in the data. Such physically implausible patterns can lead to reduced performance on unseen data even though the ANN performs well in a test setup. Explainable artificial intelligence methods aim to open this black box. Previous studies have adapted Layer‑wise Relevance Propagation (LRP), one of these methods, for explaining ANNs trained for atmospheric feature‑detection. While these studies showed that the extraction of learned patterns is feasible, they also demonstrated that interpreting especially the structure of large‑scale detection patterns remains challenging. To address this, we adapt Concept Relevance Propagation (CRP) for the detection of atmospheric features. Unlike LRP, CRP decomposes the ANN’s detection into concepts, each representing a pattern used by the ANN. Using CRP, we analyze an ANN trained to detect tropical cyclones and atmospheric rivers examining both patterns found detecting singular features and patterns found across the dataset. As an example, for atmospheric rivers, CRP reveals that, in addition to known patterns such as elongated moisture bands and strong winds, the ANN also considers surrounding dry‑air regions - an aspect not previously found. This adaptation of CRP improves the interpretability of ANNs for atmospheric feature detection, enabling a clearer assessment of whether the networks rely on physically plausible patterns or on implausible correlations in the data.

How to cite: Radke, T., Baehr, J., and Rautenhaus, M.: Improving on the explanation of patterns learned by artificial neural networks trained to detect atmospheric features, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-372, https://doi.org/10.5194/ems2026-372, 2026.

P87
|
EMS2026-373
Spatiotemporal bias correction of CAMS PM₂.₅ forecasts over Europe via ensemble machine learning
(withdrawn)
Tetiana Vovk and Maciej Kryza
P88
|
EMS2026-392
Physically Constrained Windblow Detection Using Open Earth Observation Data
(withdrawn)
Muhammad Bin Arif, Conor Sweeney, Michela Bertolotto, and Liz Gavin
P89
|
EMS2026-474
Linna Zhao and Shu Lu

    Reliable forecasts of maximum air temperatures are essential to prevent heat-related disasters, efficiently mitigate the damages caused by high-temperature disasters and appropriately respond to them. Deviations usually exist in the prediction of near-surface elements from numerical models due to complex factors such as atmospheric dynamic processes, physical processes, local topography and geomorphology. In particular, the deviation between the prediction and observation of daily maximum temperature is relatively larger when the weather changes drastically. Therefore, it is still a challenge to realize refined and accurate forecasts for daily maximum temperature. 
    Given the complexity of the interactions between the atmosphere and the Earth’s surface, no single machine learning method can consistently and effectively eliminate biases in numerical weather prediction (NWP) models. This is primarily because different regression variables in machine learning can significantly affect forecast performance. Furthermore, due to the randomness of initial weight parameters, neural network models generally have high variances. Consequently, individual machine learning models, particularly deep learning networks, often suffer from the “bias-variance” trade-off, and the ensemble machine learning can be an effective way to address this issue. A successful way to reduce the high variance of neural network models is to train multiple models instead of a single model and integrate the predictions of these models, i.e., integration learning. Integration learning not only reduces the variance of predictions, but also yields better predictions than any single model. Using ensemble models can effectively reduce the variance and enhance the model generalization ability. 
    Here, we proposed a stacking ensemble model named FLT, which consists of a fully connected neural network with embedded layers (ED-FCNN), a long short-term memory (LSTM) network and a temporal convolutional network (TCN) to overcome the high variance of a single neural network and to improve prediction of maximum air temper-ature. The case study of daily maximum temperature forecast evaluated with observation of al-most 2400 weather stations shows substantial improvement over that of single neural network model, ECMWF-IFS and statistical post-processing model. The FLT model can more effectively improve the forecast bias of the ECMWF-IFS model than that of any of the above single neural network model, with the RMSE reduced by 52.36% and the accuracy of temperature forecast in-creased by 43.12% compared with the ECMWF-IFS model. The average RMSEs of the FLT model decreases by 8.39%, 1.50%, 2.96% and 16.03%, respectively, compared with ED-FCNN, LSTM, TCN and the decaying average method.
    This study demonstrates that ensemble learning models constructed using stacked generalization methods can effectively reduce the forecasting bias associated with single neural network models, enhance the models’ generalization ability, and improve the accuracy of maximum temperature forecasts.

How to cite: Zhao, L. and Lu, S.: Ensemble Deep Learning for Improving Numerical Weather Prediction of Daily Maximum Air Temperature , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-474, https://doi.org/10.5194/ems2026-474, 2026.

P90
|
EMS2026-479
Tsvetomira Krikoryan, Alessia Hu, Sem Vijverberg, Isa Rethans, Steven van den Tol, Jannes van Ingen, and Dim Coumou

The rapid growth of weather-dependent renewable energy increases Europe’s vulnerability to grid volatility, making accurate seasonal-to-subseasonal (S2S) forecasts increasingly important. Traditional Numerical Weather Prediction systems, however, often struggle to maintain useful accuracy beyond a 10-day horizon. To address this predictability gap, we leverage recent advancements in open-source AI Foundation Models (FoMos), since they offer major advantages in computational efficiency, scalability, and adaptability compared to traditional approaches.

In this contribution we focus on demonstrating the potential of fine-tuning ECMWF’s AIFS (Artificial Intelligence Forecasting System) for improved prediction of European Weather Regimes. AIFS-ENS offers particurlarly high skill in large-scale flow. By attaching a second Decoder (called the ‘diagnostic head’) to the Processor, we demonstrate a lightweight fine-tuning strategy. The diagnostic head can be swapped out for different targets, offering flexibility to let any AI model focus on a specific region, variable(s), lead-time, and extremes. In addition, it allows for add additional input features which are relevant for the forecasting task. In this case it can allow for more explicit information on the MJO dynamics which are important for the modeling the Weather Regimes. By focusing on large-scale regime-level dynamics, we show how AI models can better exploit sources of long-range predictability compared to traditional approaches. We evaluate the ability of fine-tuned models to represent European weather regimes at sub-seasonal lead times and assess their potential to improve predictive skill in this challenging forecast range. This work is part of the broader EIC-backed TAILOR project, which aims to develop a Finetuning-As-A-Service (FAAS) platform and we discuss how these advances form a key building block within the broader TAILOR vision.

How to cite: Krikoryan, T., Hu, A., Vijverberg, S., Rethans, I., van den Tol, S., van Ingen, J., and Coumou, D.: Fine-Tuning AI Foundation Models for Subseasonal Weather Forecasting, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-479, https://doi.org/10.5194/ems2026-479, 2026.

P91
|
EMS2026-489
Jan-Philip Kraayvanger and Julian Quinting

 1. Introduction

Reliable and computationally affordable forecasts of renewable energy production values are necessary for effective grid management and energy market integration and thus for a fast and sustainable transition of the power sector. State-of-the-art Machine Learning based weather prediction (MLWP) models are getting cheaper and better continuously nowadays, making them the perfect option to provide the needed weather forecasts. On the other hand, they lack the variables for solar power generation (solar capacity factor or irradiance or at least cloud cover). This study aims to answer the question of whether MLWP is suitable for deriving solar energy values from weather forecasts, provide a suitable post-processing pipeline and compare the results for different MLWP models.

2. Methodology

A comprehensive ML-based post-processing technique is developed to predict the solar capacity factor using weather data from forecasts or reanalysis datasets. In addition to basic calculation and data processing steps, the methodology consists of a Convolutional Neural Networks (CNN) trained on ERA5 and the “C3S operational energy dataset”. From ERA5 only the variables wind, humidity, pressure, and temperature were used in the training, making the model suitable for use with MLWP data. From the energy dataset, the solar capacity factor is used as ground truth.

With this architecture, weather forecasts of MLWP models (currently Pangu-Weather, soon NeuralGCM and GraphCast) are used to predict the solar capacity factor for up to 10 days lead time.

3. Current (and Upcoming) Results

Compared to a simple persistence baseline, the CNN applied to Pangu-Weather consistently yields a lower RMSE, with the error reduction ranging from approximately 51% for a one-day lead time to 11% for a lead time of 10 days. Similar results can be achieved by comparing the approach against a climatology baseline.

Future work will include comparing the performance across different MLWP model forecasts to identify the optimal models for energy sector predictions.

4. Conclusions This research demonstrates that sophisticated ML-based post-processing is essential for transforming raw weather model outputs into actionable, reliable capacity factor forecasts for the energy sector. The enhanced skill provides a vital tool for energy trading and grid operators to manage risk, optimize renewable energy resource deployment, and bolster grid resilience in the face of climate variability.

How to cite: Kraayvanger, J.-P. and Quinting, J.: Post-Processing of ML-Based Weather Prediction for Solar Capacity Factor Forecasting, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-489, https://doi.org/10.5194/ems2026-489, 2026.

P92
|
EMS2026-548
Brigitta Goger, Çağlar Küçük, Pascal Gfäller, Irene Schicker, and Alexander Kann

Data-driven weather prediction is rapidly emerging as a transformative approach in modern forecasting. While these models often achieve improved traditional skill scores, such as root mean square error (RMSE) for near-surface variables like 2m temperature, they frequently exhibit overly smoothed spatial structures. This smoothing of spatial fields limits physical interpretability and reduces skill in representing localized extreme events, particularly convective phenomena associated with heavy precipitation. 

In this study, we investigate whether augmenting training datasets with adding vertical velocity (w) as a diagnostic variable can improve the representation and predictability of convective extremes. Since w represents convective updrafts, using the variable as a diagnostic can increase the similarity of data-driven models to the traditional, physics-based NWP models. Using the anemoi framework in a stretched-grid, limited-area configuration, we train models on ERA5 data and the Austrian regional re-analysis (ARA) at a horizontal grid spacing of 2.5km. Vertical velocity at multiple atmospheric levels is introduced as an additional diagnostic predictor. 

Model performance is evaluated using case studies of a strong summertime convective events, a known challenge for both physics-based and data-driven forecasting systems. We compare simulations from (i) a baseline data-driven model, (ii) the new  version including vertical velocity, and (iii) our physics-based numerical weather prediction system (CLAEF-AA). We will analyse whether incorporating vertical velocity improves the representation of convective structures and provides added value for extreme event prediction. We discuss implications for model design and outline future directions for integrating physically meaningful diagnostics into data-driven weather prediction systems. 

How to cite: Goger, B., Küçük, Ç., Gfäller, P., Schicker, I., and Kann, A.: The impact of vertical velocity as a variable on predictability of convective events in data-driven weather prediction , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-548, https://doi.org/10.5194/ems2026-548, 2026.

P93
|
EMS2026-570
Maicon Hieronymus, Richard Müller, and Ulrich Blahak

For many applications and regions, radar coverage remains a challenge in weather prediction. Remote regions, oceanic areas, and mountain-shadowed terrain suffer from sparse or absent radar observations, and even where infrastructure exists, data accessibility and downtime create further gaps. Satellite imagery offers a promising opportunity to fill these spatial and temporal voids. Meteosat Third Generation (MTG) satellites provide unprecedented spatial resolution and temporal revisit rates across Europe and Africa, making them particularly well-suited for learning continuous radar-like reflectivity fields.

We present a UNet-based approach to synthesize 2D radar composites from MTG satellite channels and lightning data (LINET). Our architecture employs wavelet decomposition for multi-scale feature extraction, going beyond standard lowpass filtering to better capture the spatial structure of precipitation. To preserve the sharpness of convective features, we employ a loss function that explicitly emphasizes sharp edges in the target and penalizes the synthesized output accordingly. The bottleneck combines Swin-style local window attention with a global pooled attention branch, enabling the model to simultaneously capture fine-grained local structure and mesoscale spatial context. We also employ an efficient channel attention mechanism for very wide feature maps.

A central focus of this work is the geographic transferability of the trained model. Using radar observations from Europe, we systematically evaluate cross-regional generalization by training and validating on subsets of countries and testing on held-out regions. This design allows us to assess how well learned satellite-to-radar mappings transfer across different climatic regimes and precipitation characteristics, with implications for deploying such models in radar-sparse or radar-free regions globally.

How to cite: Hieronymus, M., Müller, R., and Blahak, U.: Cross-regional generalization of satellite-based radar synthesis with deep learning, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-570, https://doi.org/10.5194/ems2026-570, 2026.

P94
|
EMS2026-659
Icíar Lloréns Jover, Francesco Zanetta, and Cornelia Schwierz

This poster presents the library obsweatherscale and showcases it

obsweatherscale is an open-source Python library for machine learning-based probabilistic interpolation and regression of surface weather variables using Gaussian Processes (GPs). Built on GPyTorch, the library provides a modular and extensible framework for constructing neural-augmented Gaussian Processes models that incorporate trainable mean and kernel functions that accept arbitrary input features.  Key features include plug-and-play data transformations, support for uncertainty quantification,  GPU acceleration via the deep-learning framework PyTorch (https://pytorch.org/projects/pytorch/), as well as training and inference routines.

The library was developed to generate high-resolution wind maps over Switzerland, where complex alpine terrain and sparse observations challenge traditional methods. The aim was to create detailed wind maps that closely align with observations. Leveraging comprehensive orography maps, topographic descriptors, and numerical model outputs, we use obsweatherscale to downscale hourly ICON reanalysis wind fields from 1 km to sub-kilometer resolution integrating the wide variety of predictors. This showcases the model's ability to improve spatial detail and observational consistency, as well as provide calibrated uncertainty estimates.

It may be useful to note, that obsweatherscale can be applied to any surface parameter and additional applications: Beyond wind downscaling, obsweatherscale generalizes to a range of meteorological applications, including bias correction of model outputs and probabilistic spatial interpolation of observational datasets.

 

References: 

Lloréns Jover, I., & Zanetta, F (2024). obsweatherscale: observation-conditioned ML downscaling of surface weather fields. GitHub repository: https://github.com/MeteoSwiss/obsweatherscale 

Zanetta, F., Nerini, D., Buzzi, M., & Moss, H. (2025). Efficient modeling of sub-kilometer surface wind with Gaussian processes and neural networks. Artificial Intelligence for the Earth Systems. 

 

How to cite: Lloréns Jover, I., Zanetta, F., and Schwierz, C.: obsweatherscale: a Python library for ML-based probabilistic interpolation and downscaling of surface weather fields, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-659, https://doi.org/10.5194/ems2026-659, 2026.

P95
|
EMS2026-738
Sara Top, Jonas Kittner, Charles Pierce, Sara Speelman, Luise Wolf, Benjamin Bechtel, and Lesley De Cruz

Machine learning (ML) and artificial intelligence (AI) have become integral to modern weather and climate science, and their rapid evolution is driving new opportunities for urban climate research. Across the urban climate community, ML methods of varying complexities are being used for tasks ranging from fast point-based predictions to neighborhood- or city‑scale urban climate modelling for one or multiple (bio)meteorological variables. For instance, novel ML and AI approaches support scenario generation for impact studies such as quantifying the effects of urban vegetation, assessing thermal comfort under different warming pathways, or identifying locations for future cool spaces. In addition, ML techniques are also being applied to provide boundary conditions for micro‑scale models. 

While the field is expanding quickly, it remains highly fragmented. AI/ML in urban climate research spans diverse methods, scales, datasets, and scientific aims, making it difficult to define a common research agenda or benchmarks. The lack of an established interdisciplinary network between AI and urban climate communities further increases the risk of duplicated efforts and missed opportunities for coordinated progress. 

The newly founded AI4UrbanClimate working group addresses this gap by establishing an international community focused on AI/ML applications in urban climate research. Our goals are to (1) bring people together with similar research interests, (2) develop a common understanding of the scope and landscape of ML-based urban climate research, (3) review existing work across domains, and (4) identify persistent challenges and priorities and initiate and coordinate collective action - such as benchmarking datasets, standardized evaluation metrics, and best‑practice guidelines. 

As an initial step, we are mapping ongoing activities across the community to enable the development of meaningful benchmarks and identify where joint efforts could accelerate research. We will present the first outcomes from the initiative’s activities, including insights on who is currently involved and thoughts that were shared during the kick‑off meeting and online social event. We will also outline how you can join the AI4UrbanClimate network and contribute to building this emerging community. 

How to cite: Top, S., Kittner, J., Pierce, C., Speelman, S., Wolf, L., Bechtel, B., and De Cruz, L.: Building a community and research agenda for machine learning in urban climate: the AI4UrbanClimate initiative , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-738, https://doi.org/10.5194/ems2026-738, 2026.

P96
|
EMS2026-537
Marcos Martínez-Roig, Cesar Azorín-Molina, Mario Santa Cruz Lopez, Ana Prieto Nemesio, Christian Lessig, Maria Teresa García Gálvez, Fernando Belinchón Martín, Cristina Toledano Lozano, Javier Martínez Amaya, Jose Luis Casado Rubio, and Maria Yolanda Luna Rico

The rapid emergence of Data-Driven Weather Prediction (DDWP) models has revolutionized global forecasting, with ECMWF’s Artificial Intelligence Forecasting System (AIFS) [1] showing the ability to capture complex dynamics and demonstrating competitive skill against traditional Numerical Weather Prediction (NWP) models. While models such as the AIFS are highly effective, most are designed for global applications and trained on datasets like ERA5. Despite its extensive temporal coverage, the spatial resolution of ERA5 is often insufficient to resolve the fine-scale atmospheric processes critical for limited-area modeling where kilometer scale resolution is essential, e. g., to properly capture extreme events. Their adaptation to convection-permitting scales remains a significant frontier, especially in regions with complex topography, such as the Iberian Peninsula. This study addresses these limitations by presenting an evaluation of a high-resolution regional AI forecasting model for the Iberian Peninsula, that bridges the gap between global and regional systems and builds on the ANEMOI open-source framework.

Our approach utilizes the Anemoi framework to train specialized models for the Iberian Peninsula, leveraging the valuable global context of ERA5 for limited-area modeling. Central to our methodology is a tiered training strategy designed to harmonize the extensive temporal record of ERA5 with the high spatial resolution of AEMET’s HARMONIE-AROME [2] dataset (2.5 km spatial resolution). Following initial hyperparameter tuning on ERA5 O96 and pretraining on O320 grids, we perform fine-tuning with HARMONIE to better capture localized weather phenomena that global models typically fail to resolve. To ensure spatial continuity and provide external boundary conditions, we employ a stretched grid as done in Bris [3]. This technique integrates high-resolution data over the target domain with lower-resolution information from the surrounding global region, creating a robust prediction system that mitigates training limitations arising from the comparatively shorter temporal extent of regional datasets.

Initial deterministic experiments in this setup revealed that while the model captures general atmospheric patterns, it struggles to represent the magnitude and frequency of localized extreme events. To address this, our research has transitioned toward the application of probabilisticforecasting architectures available in Anemoi framework. Specifically, this study evaluates the performance of Generative Diffusion Models in comparison to the Continuous Ranked Probability Score (CRPS) optimized probabilistic approach developed at AEMET.

Ultimately, this work serves as a practical demonstration of how the Anemoi ecosystem can be implemented to establish high-resolution, computationally efficient probabilistic regional forecasting workflows within a National Meteorological Service context, such as AEMET. Beyond this implementation, our primary scientific objective is to assess the validity of Generative Diffusion Models as a robust approach to accurately represent uncertainty and capture the localized extreme events over the complex terrain of the Iberian Peninsula, which remain a challenge for both traditional NWP and current data-driven systems. 

How to cite: Martínez-Roig, M., Azorín-Molina, C., Santa Cruz Lopez, M., Prieto Nemesio, A., Lessig, C., García Gálvez, M. T., Belinchón Martín, F., Toledano Lozano, C., Martínez Amaya, J., Casado Rubio, J. L., and Luna Rico, M. Y.: Regional Data-Driven Weather Forecasting: Preliminary Results of a Probabilistic model using HARMONIE-AROME 2.5 km Data for the Iberian Peninsula., EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-537, https://doi.org/10.5194/ems2026-537, 2026.