- Free Institut for Planetary science, Köln, Germany (h.schmerling@uni-koeln.de)
Advances in exoplanet detection methods have steadily increased the number of known planets. With more than 6000 confirmed by early 2026, exoplanet catalogs now enable increasingly powerful statistical studies of planetary populations. However, every detection technique — transits, radial velocity, microlensing, direct imaging — carries its own observational biases, and because the true underlying planetary population is unknown, these biases cannot themselves be fully characterized. Most demographic analyses have relied on classical statistical approaches, while data-driven, unsupervised machine-learning methods have been used less frequently for exploratory population studies. Here, we explore the Extrasolar Planets Encyclopaedia dataset using a range of unsupervised learning algorithms. We first select a subset of system features and complete missing entries by comparing several imputation strategies, ranging from simple statistical fillers to a feature-prediction model trained on the catalog itself. We then apply outlier-detection methods to identify objects with parameter combinations inconsistent with the bulk of the sample, and finally apply multiple clustering algorithms — in both an unweighted form and a weighted variant intended to mitigate selection effects — to search for latent structure. To assess how strongly observational selection shapes the results, we run this pipeline on four versions of the data: the full catalog as listed, a Kepler subset that has undergone careful bias mitigation, the full catalog with a deliberately bias correction applied to all entries, and a fully synthetic dataset constructed under known input distributions. The pipeline yields: (i) a feature-prediction engine that can infer previously missing system parameters with precision particularly high for stellar features, (ii) a set of catalog entries whose reported parameters may warrant re-examination, and (iii) clusters that reproduce known demographic patterns while also suggesting additional structure among small planets orbiting M-dwarf stars.
How to cite: Schmerling, H., Hribar, R., Grziwa, S., and Pätzold, M.: Clustering the Exoplanet Database; Unraveling Hidden Patterns in Exoplanet Populations using UnsupervisedMachine Learning Techniques, Europlanet Science Congress 2026, The Hague, The Netherlands, 7–11 Sep 2026, EPSC2026-215, https://doi.org/10.5194/epsc2026-215, 2026.