- CNRS, Observatoire de Paris, LIRA, Meudon, France (baptiste.cecconi@obspm.fr)
Language models are powerful tools for processing linguistic data, i.e., unstructured data. In recent years, interest in language models within astrophysics has grown rapidly, driven by the increasing number of scientific publications and the expansion of astronomical terminology. A key objective is to create an environment that facilitates knowledge accessibility, often involving the alignment of data with controlled vocabularies to enable efficient search.
However, only a limited number of publications comply with existing standards, and not all semantic artifacts are interoperable. As a result, they require a standardization process, which can be tedious to perform manually on large datasets. The NASA Astrophysics Data System is engaged in such efforts across multiple areas.
This need to process unstructured data in astronomy, combined with recent breakthroughs in artificial intelligence driven by the emergence of large language models, has led to the development of domain-specific models. These include fine-tuned BERT-based models such as AstroBERT (2022), AstroLLaMA (2023), and Pathfinder (2024). Most of these models are trained to predict the next token in large corpora of astronomical literature, allowing them to acquire domain-specific knowledge and function as powerful scientific assistants.
At Paris Observatory, we worked on two projects involving unstructured data: the alignment of observation facilities across multiple semantic artifacts, and the assignment of UAT keywords to uncategorized papers.
We leveraged a range of natural language processing techniques and language models, in both supervised and unsupervised settings, to perform various tasks. These include LLM-based decision-making and label or definition generation through prompting; transformer-based models and TF-IDF vectorization for ranking-based recommendation systems, such as entity alignment and keyword recommendation; and linear regression and graph neural networks for classification.
We discuss the advantages and limitations of these methods and illustrate them through two case studies.
How to cite: Fretel, L., Cecconi, B., and Louis, C.: Language Models and Natural Language Processing applications in Astronomy: two Case Studies, Europlanet Science Congress 2026, The Hague, The Netherlands, 7–11 Sep 2026, EPSC2026-1219, https://doi.org/10.5194/epsc2026-1219, 2026.