- MeteoIQ, Berlin, Germany (info@meteoiq.com)
MeteoIQ operates an independent forecast verification service (verify.meteoiq.com) that evaluates the performance of a wide range of meteorological forecast products using in situ station observations as ground truth. The dataset currently spans several years, beginning in 2020. It includes hourly forecasts for key surface parameters such as temperature, dew point, wind speed, wind gusts, precipitation, and cloud cover. In addition to raw numerical weather prediction (NWP) model outputs, the service also assesses post-processed forecasts from multiple providers.
With the recent emergence of AI-based forecasting systems, there is increasing interest in understanding how these models perform relative to established approaches. In this contribution, we present a comparative analysis of NWP, AI-based, and traditionally post-processed forecasts using a consistent verification framework. Standard metrics such as bias, root mean square error, and categorical scores are used to quantify performance across different parameters, lead times, and meteorological conditions.
We place particular emphasis on identifying systematic differences in forecast characteristics. For example, AI-based models may exhibit distinct error structures, temporal consistency, or skill variations under specific weather regimes compared to physics-based models and statistical post-processing methods. We also explore how forecast performance varies geographically and seasonally, and how these differences impact end-user applications.
The analysis is based on a harmonized dataset that ensures comparability across providers, including consistent temporal and spatial matching with observations. By leveraging a multi-year archive, we are able to assess both average performance and variability over time.
This contribution aims to provide an objective, service-oriented perspective on the current capabilities of AI-based forecasts in comparison to traditional methods. The results are intended to support users and providers in understanding the strengths and limitations of different forecasting approaches, and to inform the integration of new model types into operational decision-making processes.
How to cite: Schulze, D., Müller, E., and Hoffmann, J.: Comparing NWP, AI and traditional post-processed forecasts with station observations, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-614, https://doi.org/10.5194/ems2026-614, 2026.