OSA1.4 | Verification for AI- and physics-based weather prediction: new challenges, techniques and observations
Verification for AI- and physics-based weather prediction: new challenges, techniques and observations
Conveners: Estíbaliz Gascón, Bastien François, Sabrina Wahl
Orals Mon2
| Mon, 07 Sep, 11:00–13:00 (CEST)|Room Quest
Orals Mon3
| Mon, 07 Sep, 14:30–16:00 (CEST)|Room Quest
Posters PS-Tue4
| Attendance Tue, 08 Sep, 16:30–18:00 (CEST) | Display Mon, 07 Sep, 08:00–Tue, 08 Sep, 18:00|TransitZone, P96–97
Mon, 11:00
Mon, 14:30
Tue, 16:30
This session addresses recent advances and emerging challenges in the verification of numerical weather prediction (NWP), artificial intelligence–based weather prediction (AIWP), and climate modelling systems across a broad range of spatial and temporal scales. Contributions are welcome across the full verification spectrum, from methodological research to operational practice and user-oriented applications.

The scope encompasses established verification approaches for physical NWP models, as well as novel methodologies required for AIWP and hybrid systems, including their extension to subseasonal-to-seasonal and climate applications
• Use and interpretation of new and emerging observational datasets for verification, including non-traditional observations, impact-based data, and applications related to high-impact and user-oriented services such as warnings for hazardous weather.
• Advances in verification approaches tailored to different modelling systems (physical, artificial intelligence-based and hybrid models, and climate models), including suitable metrics, techniques, and effective communication of forecast skill and uncertainty.
• Methodological innovations such as spatial, temporal, and object-based verification, extremes verification, process-based evaluation, and probabilistic methods.
• Verification strategies adapted to high-resolution, convection permitting, ensemble, and variable-resolution modelling systems.
• Verification approaches aimed at supporting decision-making and end-user needs, including sector-specific verification, evaluation of risk-relevant events, and applications bridging weather and climate services.

Orals Mon2: Mon, 7 Sep, 11:00–13:00 | Room Quest

Chairpersons: Estíbaliz Gascón, Bastien François, Sabrina Wahl
11:00–11:30
|
EMS2026-633
|
solicited
|
Onsite presentation
Romain Pic and the Bridging The Gap

Forecast verification is a critical component in the development and evaluation of weather prediction models. Spatial weather fields present unique challenges for verification due to their complex structures and the inherent spatial dependencies. Two main approaches have emerged to tackle these challenges: spatial verification methods designed to mitigate the double penalty effect and consistent/proper scores motivated by theoretical decision-making principles.

Bridging The Gap is a research project aimed at connecting these two communities to foster collaboration and advance the field of spatial verification for ensemble forecasts. Bridging The Gap is an activity of the Joint Working Group on Forecast Verification Research (JWGFVR) under the World Meteorological Organization (WMO) World Weather Research Programme (WWRP). The project seeks to explore the links between spatial verification methods and consistent/proper scores, develop new methodologies, and address current challenges of spatial verification (e.g., ensemble forecasts and AI-based weather prediction models).

In June 2026, the project will hold its first dedicated workshop at the University of Reading (UK). The event will bring together researchers from both communities for contributed and invited talks, and closed discussion sessions aimed at identifying future research directions and kickstarting a joint review article. We provide an overview of the current state of spatial verification for ensemble forecasts, summarize the key themes and outcomes of the workshop, and outline open research questions at the intersection of spatial verification, proper scoring rules, and the evaluation of ensemble and AI-based forecasts.

Acknowledgements. The Bridging The Gap workshop was made possible through the support of the World Weather Research Programme of the World Meteorological Organization.

How to cite: Pic, R. and the Bridging The Gap: Bridging The Gap : project overview and workshop outcomes, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-633, https://doi.org/10.5194/ems2026-633, 2026.

11:30–11:45
|
EMS2026-170
|
Onsite presentation
Francesco Pasquini, Michiel Baatsen, Bastien François, Natalie Theeuwes, and Maurice Schmeits

Weather forecasting has traditionally relied on Numerical Weather Prediction (NWP) models, which simulate weather
by solving the governing fluid equations. Recently, the emergence of Deep Learning Weather Prediction (DLWP)
models has opened a new era in weather forecasting, offering a data-driven alternative to classical NWP approaches.
Regional DLWP models such as the stretched-grid model Bris developed by Met Norway, have demonstrated perfor-
mance on par with, or even slightly better than regional NWP models across a range of standard forecast metrics.
By overcoming the coarse horizontal resolution that constrained earlier global data-driven models, the operational
use of regional DLWP systems now appears increasingly promising. Nevertheless, the performance of such models
during extreme events is generally inferior to that of regional NWP models, and comprehensive evaluations of their
ability to generate physically realistic forecasts are still lacking.

Here, we present a study comparing the physical consistency of the deterministic version of Bris with the control
run of the operational MetCoOp Ensemble Prediction System (MEPS) in forecasting the severe extratropical cyclone
Poly, which hit the Netherlands on 5 July 2023. We examine whether Bris accurately represents deviations from
key atmospheric balances and whether it reproduces expected dynamics of the storm. We show that, despite its
relatively good performance in terms of RMSE, Bris struggles to capture important mesoscale features of the event
and that it significantly disrupts several atmospheric balances. This unrealistic disruption is mainly linked to the
fine-scale noise evidenced in its output fields, which leads to incorrect and unrealistic spatial gradients. The analysis
of the amplitude spectra also reveals that, in the stretched-grid DDM, fine-scale noise coexists with a seemingly
competing smoothing of the same meteorological variables at larger scales. This tendency for large-scale smoothing
is commonly observed in DLWP models trained with MSE loss and we show that it has notable consequences when
forecasting extreme windstorms such as Poly. These results raise critical questions for improving AI-based models
to better represent extreme events and how to ensure physical consistency in their predictions.

How to cite: Pasquini, F., Baatsen, M., François, B., Theeuwes, N., and Schmeits, M.: Assessing the ability of a stretched-grid deep-learning weather prediction model to capture physical balances, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-170, https://doi.org/10.5194/ems2026-170, 2026.

11:45–12:00
|
EMS2026-469
|
Onsite presentation
Britta Seegebrecht, Sabrina Wahl, Stefanie Hollborn, George Craig, Erik Pavel, Wael Almikaeel, Sebastian Buschow, Martin G. Schultz, Christian Lessig, Ilaria Luise, Anas Al-Iahham, and Juergen Gall

Data-driven weather prediction models based on artificial intelligence (AI) have rapidly advanced in recent years and are frequently reported to outperform traditional physics-based numerical weather prediction (NWP) models for selected verification scores. However, optimization with respect to a specific loss function can adversely affect other metrics, potentially leading to unrealistic forecast characteristics, such as overly smooth spatial structures when mean-squared or mean-absolute error–based loss functions are used.

In Bonavita & Geer, 2026 an orthogonal decomposition of the Mean Squared Error (MSE) into Information Error and Noise Error is introduced to unravel different strategies for minimizing this commonly used accuracy metric. These insights allow for a more meaningful interpretation of the MSE as accuracy measure.

Additionally, a scale dependent analysis of model performance can help to reveal systematic differences between AI and NWP models specifically on smaller scales such as the effective resolution.

The combination of both approaches – the decomposition of the scale dependent MSE into Information and Noise Error – has been derived and applied to different AI and NWP models.

First results show the potential of this method to disentangle different effects allowing for a fairer, more comprehensive comparison between AI and NWP weather prediction models.

The analysis is partly based on forecasts from the Weather Prediction Model Intercomparison Project (WP MIP), which provides a collection of NWP and AI-model forecasts from multiple national weather services and research institutions.

The work is conducted within the RAINA project, which aims to develop a foundation model for the atmosphere with a particular focus on reliable, high-resolution forecasts of extreme wind and precipitation events.

Bonavita, M. & Geer, A.J. (2026) Forecast verification using information and noise. Quarterly Journal of the Royal Meteorological Society, e70109. Available from: https://doi.org/10.1002/qj.70109

How to cite: Seegebrecht, B., Wahl, S., Hollborn, S., Craig, G., Pavel, E., Almikaeel, W., Buschow, S., Schultz, M. G., Lessig, C., Luise, I., Al-Iahham, A., and Gall, J.: Leveraging scale dependent accuracy measures for a fair comparison between AI and NWP weather prediction models, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-469, https://doi.org/10.5194/ems2026-469, 2026.

12:00–12:15
|
EMS2026-292
|
Onsite presentation
Zied Ben Bouallegue

We are interested in the ability of weather models to predict extreme events. Nowadays, weather forecasts are generated not only using theory‑driven (or physics‑based) models, but also data‑driven (or machine‑learning), and hybrid (nudging‑based) models. The rapid development of data‑driven models and their adoption in operational weather forecasting have triggered investigations and discussions into how well these models compare to theory‑driven models when it comes to extreme events. This work aims to contribute to this effort.

The idea followed here consists in focusing on the extrema (maximum and minimum) of weather forecast fields (e.g., temperature anomaly). The general framework described in [1] is used to assess three aspects of the forecast separately.

  • First, the physical realism of the extrema is checked to answer the question: “Is the extremum consistent with our understanding of the physics of the Earth system?”
  • Second, the structural realism is assessed in terms of block‑extrema represented as parametric distributions, addressing the question: “Do the characteristics of the forecastextrema distribution resemble those of the observations?”
  • Finally, the functional realism is measured to summarise the ability of the model to produce forecasts close to the observations. Here again, the focus is on block-extrema, and we discuss how these results differ from focusing on block-averages instead.

The Weather Prediction Model Intercomparison Project (WP-MIP) offers a convenient framework for our investigations. The corresponding dataset resembles forecasts of a variety of weather parameters from different contributing institutions. We show results for a subset of deterministic forecasts from theory-driven, data-driven, and hybrid models. We also suggest an extension of our approach to ensemble forecasts.

[1] Zied Ben Bouallègue (2026), “What is a realistic forecast?" Assessing data-driven weather forecasts, a journey from verification to falsification. Preprint, https://doi.org/10.48550/arXiv.2602.00622

How to cite: Ben Bouallegue, Z.: On the realism of forecast extrema , EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-292, https://doi.org/10.5194/ems2026-292, 2026.

12:15–12:30
|
EMS2026-379
|
Onsite presentation
Leonard Knirsch, Lea Beusch, Jonas Bhend, and Christoph Spirig

Historically, weather warnings at MeteoSwiss are verified manually by a team of forecasters. This method effectively leverages expert knowledge and has proven successful over the years. However, the granularity of the results and consistencyof this manual verification is limited and could only be increased at substantial costs. We have therefore developed an automatic verification framework that allows to evaluate warning quality at increased granularity both in space and time, i.e. at individual regions and for different lead times.

Our automatic verification algorithm first creates two sets of reference warnings to verify the forecaster issued warnings against. Both use a gridded precipitation data set from combined radar and rain gauges observations to assign a warning level to each one of Switzerland’s 159 warning regions (varying in size between 105 and 475 km2). The first set determines the optimal warning level and duration for each warning region. The second determines the optimal aggregated warning level, following forecasters’ established guidelines on minimal warning areas (1000 km2) and homogeneous durations across regions for warning events, analogously to how forecasters themselves aggregate weather forecast information when issuing warnings. Hence, while the first one leads to patchy warning maps that closely correspond to measurements, the second one results in smooth warning maps. These are further away from measurements, but more desirable from a communication perspective and account better for predictability limitations. In a second step, our algorithm compares the issued warnings with both sets of reference warnings: it classifies each hour as hit, miss, false alarm or correct rejection per region. Additionally, mechanisms for selecting the degree of tolerance regarding warning level and temporal accuracy are also implemented.

In this contribution, we present key findings from applying our algorithm to all rain warnings that were issued by MeteoSwiss between 2015 and 2025 and reflect on how we can use them to further improve our warnings. Additionally, we highlight how byproducts of our algorithm can be used to support the manual verification efforts. Finally, we will give an outlook on how we envision to include this algorithm in our operational verification suite.

How to cite: Knirsch, L., Beusch, L., Bhend, J., and Spirig, C.: Automated Verification of Rain Warnings at MeteoSwiss, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-379, https://doi.org/10.5194/ems2026-379, 2026.

12:30–12:45
|
EMS2026-614
|
Onsite presentation
Dennis Schulze, Evelyn Müller, and Jan Hoffmann

MeteoIQ operates an independent forecast verification service (verify.meteoiq.com) that evaluates the performance of a wide range of meteorological forecast products using in situ station observations as ground truth. The dataset currently spans several years, beginning in 2020. It includes hourly forecasts for key surface parameters such as temperature, dew point, wind speed, wind gusts, precipitation, and cloud cover. In addition to raw numerical weather prediction (NWP) model outputs, the service also assesses post-processed forecasts from multiple providers.

With the recent emergence of AI-based forecasting systems, there is increasing interest in understanding how these models perform relative to established approaches. In this contribution, we present a comparative analysis of NWP, AI-based, and traditionally post-processed forecasts using a consistent verification framework. Standard metrics such as bias, root mean square error, and categorical scores are used to quantify performance across different parameters, lead times, and meteorological conditions.

We place particular emphasis on identifying systematic differences in forecast characteristics. For example, AI-based models may exhibit distinct error structures, temporal consistency, or skill variations under specific weather regimes compared to physics-based models and statistical post-processing methods. We also explore how forecast performance varies geographically and seasonally, and how these differences impact end-user applications.

The analysis is based on a harmonized dataset that ensures comparability across providers, including consistent temporal and spatial matching with observations. By leveraging a multi-year archive, we are able to assess both average performance and variability over time.

This contribution aims to provide an objective, service-oriented perspective on the current capabilities of AI-based forecasts in comparison to traditional methods. The results are intended to support users and providers in understanding the strengths and limitations of different forecasting approaches, and to inform the integration of new model types into operational decision-making processes.

How to cite: Schulze, D., Müller, E., and Hoffmann, J.: Comparing NWP, AI and traditional post-processed forecasts with station observations, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-614, https://doi.org/10.5194/ems2026-614, 2026.

12:45–13:00

Orals Mon3: Mon, 7 Sep, 14:30–16:00 | Room Quest

Chairpersons: Estíbaliz Gascón, Bastien François, Sabrina Wahl
14:30–14:45
|
EMS2026-351
|
Onsite presentation
Ryan Poole

In the first part of this talk, we highlight the importance of data standardization, both within an organisation and across organisations. As both physics and machine‑learning bases weather and climate models advance rapidly, the diversity of available datasets, metadata conventions, and grid geometries grow with it. While this expansion offers new scientific advancements, it also creates challenges for comparing outputs across different models and organisations. Differences in vertical coordinate definitions, horizontal resolutions, and variable naming can hinder consistent analysis, reduce reproducibility, and introduce ambiguity when comparing datasets. 

The rise of machine‑learning (ML) workflows further increases the need for reliable, standardised input data. ML systems are highly sensitive to inconsistencies, and even minor metadata issues can degrade performance and take time to resolve. By embedding data standardisation directly into preprocessing, we can reduce manual manipulation of the data, lower the risk of silent errors, and support the creation of high‑quality datasets suitable for training and evaluation. 

In the second part of this talk, we discuss how the Met Office handles standardization: StaGE, The Standard Gridding Engine. This a python library designed to streamline the preparation and standardisation of meteorological datasets. StaGE provides a unified framework for harmonising metadata, enforcing naming conventions, and ensuring that key attributes such as units, coordinates etc, are clear and consistent across different datasets. This allows us to output standardised data for customers, as well as other departments within the Met Office.

Of the various submodules in StaGE, a key component is its robust regridding functionality, which enables model datasets, which can often have model specific vertical levels and sit on staggered horizontal grids, to be mapped onto a common, predefined set of vertical levels (height and pressure for example) and unstaggered horizontal grids. As models increasingly adopt diverse geometries and resolutions, this functionality is essential for meaningful comparisons.

Overall, StaGE provides an efficient and scientifically robust foundation for working with the expanding range of meteorological data, enhancing interoperability and enabling clearer, more reliable analysis across models and methods.

How to cite: Poole, R.: StaGE: The Standard Gridding Engine, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-351, https://doi.org/10.5194/ems2026-351, 2026.

14:45–15:00
|
EMS2026-583
|
Onsite presentation
Nevio Babić, Jadran Jurković, and Vinko Šoljan

The presence of cumulonimbus (CB) and thunderstorm areas always has a strong influence on aviation. Observations, and especially forecasts, of CB areas are important in air traffic flow management (ATM). Eumetnet provides a unique product, the Cross-Border Convection Forecast (CBCF), which supports planning of air traffic flows across Europe for Eurocontrol and other local air traffic flow management systems. CBCF is issued from May to October on the European domain for the current and following five three-hourly periods between 06 and 21 UTC. Forecasters in each state simultaneously produce forecast (FCST) polygons based on a risk matrix that reveals the probability and convective mode (isolated ISOL, clustered CLST, or widespread WSPR). In addition to CBCF, for internal air traffic management purposes, the meteorological watch office in Croatia issues a similar internal ATM forecast (ATMF) covering the Croatian domain, valid for all days (h24) in three-hourly intervals. After some efforts to subjectively verify ATMF, we developed a novel verification approach with an emphasis on the convective mode.
For the purposes of this analysis, to define observed CB (OBS) polygons we rely solely on 1-min lightning detection data. To obtain OBS polygons in an unbiased and objective manner, we employ HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise), a popular unsupervised machine learning algorithm often used in image segmentation. Since this algorithm also relies on a set of optional, but purpose-dependent thresholds, we will demonstrate the sensitivity of final verification scores on these parameters.
On the one hand, we demonstrate verification results in a traditional parameter space bounded by so-called precision and recall. On the other hand, we report verification of FCST polygons also by the Dice similarity coefficient which, unlike precision and recall, can be viewed as a combined quality score combining areas of both OBS and FCST polygons. 
Preliminary verification scores indicate progressively better forecasting as convective mode organization increases from ISOL to WSPR. In other words, this particular method of verification does not seem plausible when attempting to verify rather small ISOL polygons.
We presented a method for verifying forecasted CB areas, with particular emphasis on the treatment of observed organisational convection mode and verification scores. This approach could also be applied to the verification of other novel Eumetnet products, such as CBCF and eGAFOR. 

How to cite: Babić, N., Jurković, J., and Šoljan, V.: Verifying thunderstorm forecast polygons using a novel machine learning approach, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-583, https://doi.org/10.5194/ems2026-583, 2026.

15:00–15:15
|
EMS2026-765
|
Onsite presentation
Iris Odak, Sara Anđela Perić, and Jakov Lozuk

Forecast verification is essential for evaluating numerical weather prediction performance, particularly when categorical measures are applied across events of varying frequency. Traditional verification metrics often exhibit strong dependence on climatological frequency, limiting their interpretability for rare but high-impact phenomena. This study assesses the applicability of both conventional and rare-event categorical verification measures across the full spectrum of event frequencies using theoretical cases, with a focus on wind forecasts. Wind speed and direction observations at 10 m, representing both strong-wind and weak-wind regimes, are used to verify operational forecasts from two similar model configurations.
Results from theoretical contingency tables and real forecast data reveal the main strengths and substantial limitations of widely used measures, expectedly including pronounced base-rate dependence for some. Unexpectedly, the Extremal Dependence Index (EDI) provides a consistent assessment of forecast quality for both rare and frequent events. The findings demonstrate that although EDI was originally developed for rare-event verification, it shows a broader range of applicability across different event frequencies than might be expected, even though retaining some of the known limitations common to categorical measures.

Building on this, the results can be naturally extended toward a more integrated representation of forecast performance. By visualizing these aspects within a polar framework (performance rose), the circular nature of wind direction is explicitly accounted for, enabling a more intuitive interpretation compared to Cartesian representations. Such an approach allows multiple attributes (e.g., distribution, bias, and accuracy) to be conveyed simultaneously, while highlighting the importance of a balanced and targeted selection of information to avoid information overload, potentially leading to loss of clarity.

How to cite: Odak, I., Perić, S. A., and Lozuk, J.: The Extremal Dependence Index (EDI) – More Than a Tool for Rare Event Forecast Evaluation?, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-765, https://doi.org/10.5194/ems2026-765, 2026.

15:15–15:30
|
EMS2026-46
|
Onsite presentation
Abhilasha Sevta

Climate change has intensified extreme heat events, particularly in semi-arid regions where agriculture, water resources, and livelihoods are highly climate-sensitive. The reliable temperature projections are therefore essential for effective climate risk assessment and adaptation planning. In this context, General Circulation Models (GCMs) are widely used to project future climate, but their performance varies across regions and across different temperature characteristics. This study proposes a robust framework to identify the most reliable CMIP6 GCMs for temperature projections in semi-arid regions. The proposed methodology evaluates GCM performance against observed data for both mean temperature and extreme temperature indices defined by the Expert Team on Climate Change Detection and Indices (ETCCDI), categorised into intensity, duration, and frequency indices. Ten statistical evaluation metrics are used to quantify GCM performance. To minimise subjectivity and dependence on individual metrics, a leave-one-out approach is applied to determine the optimal combination of evaluation metrics. The GCMs are then ranked using two multi-criteria decision-making methods, namely the Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) and Višekriterijumsko Kompromisno Rangiranje (VIKOR), to enhance the robustness of model selection. The framework is applied to the semi-arid region of North-Western India using CMIP6 GCMs for the historical period, with ERA5 used as the reference observational dataset. The results indicate that GISS-E2-1-G, KACE-1-0-G, EC-EARTH3, UKESM1-0-LL, and NORESM2-MM consistently demonstrate better performance across the region based on the optimal set of evaluation metrics (RMSE, MAE, NSE, and R²). However, no single model performs best across all temperature characteristics because of regional heterogeneity. The proposed framework provides a transparent and reproducible approach for robust GCM selection and improves the assessment of temperature extremes in semi-arid regions.

Keywords: Temperature extremes; Semi-arid region; CMIP6; TOPSIS; VIKOR

How to cite: Sevta, A.: Evaluation and Ranking of CMIP6 GCMs for Temperature Extremes in the Semi-Arid Region of NorthWestern India, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-46, https://doi.org/10.5194/ems2026-46, 2026.

15:30–15:45
|
EMS2026-598
|
Online presentation
Ron McTaggart-Cowan, Linus Magnusson, Inna Polichtchouk, Duncan Ackerley, Martin Koehler, and Barbara Casati

Rapid progress in the field of machine-learning for weather prediction has led to the emergence of algorithms whose forecasting skill can exceed that of traditional physically based models.  This development represents an opportunity to improve the quality of forecasting services provided by operational centers, particularly given the speed at which machine-learning based models generate predictions.  Despite the clear promise of these systems, questions remain about the ability of the current generation of machine-learning models to generate physically consistent predictions of the full suite of required forecast fields under all conditions.  Answering these questions will require careful comparisons between the well-understood physically based models, current state-of-the-art machine-learning models, and the hybrid models that combine elements of these two archetypes.  The Weather Prediction Model Intercomparison Project (WP-MIP) is a World Meteorological Organization-supported initiative whose initial goal is to create a centralized database of physically based, machine-learning and hybrid models to enable a distributed assessment and evaluation effort.  More information and instructions for contributing to the project can be found on the project website: https://www.wcrp-esmo.org/activities/wp-mip.  The first instance of WP-MIP focuses on global deterministic predictions using both center-specific and common initializations to facilitate sensitivity studies.  Forecasts contributed by institutions across six continents will be used to develop AI-ready verification techniques that highlight the strengths and weaknesses of each class of prediction system, with the goal of establishing best-practice guidance to model developers and national weather centers.  The broad engagement of the operational and forecast-evaluation communities in WP-MIP will ensure that the project’s results are highly relevant to the development and deployment of next-generation weather prediction systems.

How to cite: McTaggart-Cowan, R., Magnusson, L., Polichtchouk, I., Ackerley, D., Koehler, M., and Casati, B.: WP-MIP: An Artificial Intelligence, Hybrid and Physically Based Model Intercomparison Project for Weather Prediction, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-598, https://doi.org/10.5194/ems2026-598, 2026.

15:45–16:00
|
EMS2026-754
|
Online presentation
Juan Carlos Tufino, Adrian Huerta, Aldo Moya, Alan Llacza, Alexis Ibañez, Jorge Llamocca, and Marco Bojorquez

Heavy rainfall events in northern Peru, primarily associated with the Coastal El Niño phenomenon, are a major cause of devastating floods and severe socioeconomic impacts. To address a deficiency in operational forecasting, this study presents the first systematic verification-based assessment of the predictive skill of two regional numerical weather prediction (NWP) models used by SENAMHI: the WRF and Eta models. The analysis focuses on the most intense precipitation events in the Piura region during February–March 2023, corresponding to the most severe Coastal El Niño since 2017. Predictive skill of 24-hour and 48-hour forecasts is evaluated for each episode.

Both models were forced with identical initial and boundary conditions from the GFS global model. Sensitivity experiments with WRF identified an optimal configuration combining Thompson microphysics, New Tiedtke cumulus, and YSU planetary boundary layer schemes, which exhibited the lowest bias and RMSE in preliminary analysis.

Beyond traditional metrics, we employ spatial verification using the Fractions Skill Score (FSS) to assess precipitation pattern accuracy, and extreme event metrics including the Symmetric Extremal Dependence Index (SEDI) and Extreme Dependency Score (EDS). Results show that the optimized WRF configuration consistently outperforms the operational Eta setup, with lower errors validated against satellite estimates and observations from 27 weather stations, more accurately reproducing the spatial distribution and intensity of precipitation.

However, both models exhibit systematic limitations. WRF tends to overestimate precipitation over mid-to-high Andean slopes and slightly underestimate it in coastal areas. The Eta model underestimates precipitation at high-elevation stations and produces more widespread overestimation across the coastal plain. These persistent biases are quantified spatially, providing a benchmark for future model development. These findings identify a more robust configuration and diagnose regional biases, offering a basis for improving operational verification, extreme rainfall forecasting, and early warning systems in the Piura region.

How to cite: Tufino, J. C., Huerta, A., Moya, A., Llacza, A., Ibañez, A., Llamocca, J., and Bojorquez, M.: Verification of Physics-Based NWP for Hazardous Rainfall Forecasting in Piura, Peru: Parametrization Sensitivity of WRF and Eta Using Spatial and Extreme-Event Metrics, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-754, https://doi.org/10.5194/ems2026-754, 2026.

Posters: Tue, 8 Sep, 16:30–18:00 | TransitZone

Display time: Mon, 7 Sep, 08:00–Tue, 8 Sep, 18:00
Chairpersons: Estíbaliz Gascón, Bastien François, Sabrina Wahl
P96
|
EMS2026-721
Duško Mrkonjić, Katarina Veljović Koračin, and Zorica Podraščanin

The aim of this work was to set up an optimal configuration of moist physics in the Weather Research and Forecasting (WRF) model by assessing the sensitivity of precipitation intensity forecasts over most of the Balkans, with particular reference to Bosnia and Herzegovina (BiH). This system is supposed to be an operational forecasting tool at the BiH entity, Republic Hydrometeorological Institute of Republic of Srpska. The study focuses on cases when precipitation resulted from the activity of a large cyclone in the Mediterranean centered over Sardinia and Sicily during the winter period. Simulations were conducted for five cases of precipitation that affected BiH, with precipitation also recorded at stations in Montenegro and Greece. The WRF model was nested into the Global Forecast System (GFS). Initial and boundary conditions were available at approximately 13 km horizontal resolution. Forecasts were produced on a 10 km horizontal resolution and a vertical grid with 45 levels, with the top at a 20 hPa level. The model start time for all cases was 00 UTC. Sensitivity of the model results to the microphysical schemes (Ferrier, Thompson, and Thompson scheme with aerosols) was tested. Forecasts were classified according to actual precipitation intensity, and skill scores were calculated for different precipitation thresholds. Quantitative evaluation of model results was performed based on recorded daily accumulations. The statistical metrics included root mean square error and bias as standard measures, and sensitivity, specificity, and true skill statistics, as binary measures. Forecasted synoptic fields of pressure, temperature, and geopotential were compared with fields from GFS analyses.

How to cite: Mrkonjić, D., Veljović Koračin, K., and Podraščanin, Z.: Sensitivity of WRF precipitation intensity forecasts to microphysical parametrizations, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-721, https://doi.org/10.5194/ems2026-721, 2026.

P97
|
EMS2026-434
Elena Guk

Standard verification approaches for AI weather prediction (AIWP) models often rely on scores derived from data-dense regions, while verification across the polar areas remains limited. Surface-level performance in data-sparse regions demands targeted evaluation using in-situ observations. We address this gap by evaluating GraphCast-predicted near-surface wind speeds against both ERA5 and in-situ observations at Antarctic stations. This study applies a three-way comparison between observations, reanalysis, and AIWP predictions to Antarctic stations, where such evaluation is currently absent; and extends the ERA5 vs in-situ observations (Obs) Antarctic wind evaluation by Rakoczy et al. [1].

We compare monthly mean 10 m wind speed from GraphCast T+24h forecasts, ERA5 reanalysis and SCAR Met READER station observations [2] at four Antarctic low-altitude stations for 2018–2020. Near-surface winds play a key role in Antarctic climate, influencing sea ice formation, precipitation, boundary layer stability and ice shelf dynamics [3]. Accurate representation in both reanalysis and forecast systems is therefore relevant to climate research as well as operational applications. Katabatic winds, channelled by ice sheet topography at scales below the model grid, may represent a particular challenge for AIWP and NWP models.

The three comparisons show a pattern. ERA5 vs Obs errors are substantial at all four stations, indicating that ERA5 shows substantial biases against observations. GraphCast vs Obs errors are systematically larger than ERA5 vs Obs at three of four stations, suggesting GraphCast may amplify ERA5's errors when evaluated against reality. GraphCast vs ERA5 errors are small at all Antarctic stations, indicating GraphCast closely replicates ERA5 patterns. The combination of these three comparisons is a noticeable preliminary finding: GraphCast faithfully reproduces ERA5 at Antarctic stations, but ERA5 deviates substantially from observations.

The station sample is too small to support strong conclusions, but these preliminary findings suggest that in-situ polar observations can expose verification blind spots not captured by reanalysis-based AIWP benchmarks, and that the fidelity of an AIWP model to ERA5 data is different from its ability to accurately predict in-situ observations. This study is ongoing; extensions to additional stations, longer records, and further AIWP models are planned prior to the conference.

[1] Rakoczy, B. C., Bromwich, D. H., & Wang, S. (2026). Evaluation of ERA5 Near-Surface Winds Over Antarctica: Spatial Variability, Biases, and Large-Scale Influences. Journal of Climate (published online ahead of print 2026), e250215, Article e250215. https://doi.org/10.1175/JCLI-D-25-0215.1

[2] https://legacy.bas.ac.uk/met/READER/data.html (accessed 25.03.2026) https://doi.org/10.5285/569d53fb-9b90-47a6-b3ca-26306e696706 

[3] Davrinche, C., Orsi, A., Amory, C., Kittel, C., and Agosta, C. (2025). Future changes in Antarctic near-surface winds: regional variability and key drivers under a high-emission scenario, The Cryosphere, 19, 6023–6042, https://doi.org/10.5194/tc-19-6023-2025 

How to cite: Guk, E.: GraphCast faithfully reproduces ERA5 near-surface wind speed at Antarctic stations - but ERA5 itself deviates substantially from in-situ observations: a preliminary three-way evaluation, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-434, https://doi.org/10.5194/ems2026-434, 2026.