EMS Annual Meeting Abstracts
Vol. 23, EMS2026-297, 2026, updated on 22 Jun 2026
https://doi.org/10.5194/ems2026-297
EMS Annual Meeting 2026
© Author(s) 2026. This work is distributed under
the Creative Commons Attribution 4.0 License.
Oral | Thursday, 10 Sep, 11:15–11:30 (CEST)| Room Mission 2
Explainable Flood Mapping from Sentinel-1 Imagery using CNNs and Vision Transformers
Arundhuti Banerjee and David Daou
Arundhuti Banerjee and David Daou
  • UNU EHS

Reliable and efficient flood mapping techniques are critical for disaster response and risk assessment. Satellite-based Earth observation, particularly Synthetic Aperture Radar (SAR) data from Sentinel-1, provides an effective means of monitoring flood events over large spatial scales, independent of cloud cover and daylight conditions. When combined with advanced deep learning algorithms, SAR imagery enables the generation of accurate and timely flood maps that support disaster mitigation and emergency response. In this study, we benchmark multiple flood segmentation architectures commonly used in Earth observation, including convolutional neural network (CNN)-based U-Net models with different encoder backbones and transformer-based vision architectures such as SegFormer. A multiclass segmentation model is initially trained on a Sentinel-1 flood dataset derived from NASA observations and subsequently fine-tuned on the Sen1Floods11 dataset to evaluate cross-dataset generalization and model robustness. Beyond segmentation accuracy, we investigate explainable deep learning approaches to better understand how models interpret SAR flood signatures. Specifically, we employ Gradient-weighted Class Activation Mapping (Grad-CAM) and entropy-based uncertainty estimation to analyze model attention and prediction confidence. Our results indicate that transformer-based architectures produce more spatially coherent flood predictions compared to traditional CNN models. In several test scenes, SegFormer captures large, continuous inundated regions more effectively, whereas CNN-based U-Net models tend to rely on localized radar backscatter patterns, often resulting in fragmented segmentation outputs. These findings underline the importance of architectural choice for robust flood mapping and demonstrate the added value of explainability in building trust for operational deployment.

How to cite: Banerjee, A. and Daou, D.: Explainable Flood Mapping from Sentinel-1 Imagery using CNNs and Vision Transformers, EMS Annual Meeting 2026, Utrecht, Netherlands, 6–11 Sep 2026, EMS2026-297, https://doi.org/10.5194/ems2026-297, 2026.