Sustainable AI for Astronomical Archival Data Discovery in the Petabyte Era
2026-11-02 –, Banquet Hall

The rapid adoption of generative AI is transforming the way scientists interact with astronomy science archives and platforms. However, the growing dependence on large foundation models raises important questions about sustainability, cost, and long-term operational viability. In the Data Science and Archives Division at the European Space Astronomy Centre (ESAC), we are exploring different ways to use AI to support our users. Among others, we are exploring how smaller open source language models combined with retrieval-augmented generation (RAG) and domain-specific tooling can enhance astronomical data discovery while minimizing computational overhead. Rather than treating state of the art frontier models as the default solution, we are investigating if smaller models can be effectively deployed for many archive-support tasks with considerably lower infrastructure requirements.

The ESAC astronomy science archives present a particularly attractive use case for this approach. In legacy missions, the number of specialists available to support users decreases with time, while the scientific value of these missions remains high for decades. By combining lightweight large language models (LLM) with RAG pipelines built from mission documentation, archive interfaces user guides, and selected scientific publications, it is possible to provide conversational interfaces that preserve mission knowledge that can be used by scientists to discover and use relevant datasets. Similarly, for missions in operations, such as Euclid, AI assistants could help scientists navigate its complex data model, learn the details of hundreds of distinct data products, and generate ADQL example queries and notebooks.

We argue that this approach could complement other alternatives and represents a more sustainable path for scientific archives. Running open-source LLMs with “only” a few billion parameters on local GPU infrastructure would allow institutions to reuse existing resources, reduce dependence on commercial AI subscriptions, and limit the environmental and financial costs associated with large-scale cloud inference. This approach also maximizes the value of open-source models whose development has already required significant community investment. In this presentation we will showcase some of the experimental work we are doing at ESAC using open source models running on our local infrastructure, for example, enhancing the existing ESASky chatbot with generative AI capabilities and building RAG pipelines for mission documentation and archival data discovery.

Astronomer and Data Scientist with over 20 years of experience supporting astronomy space missions from ESA, NASA and JAXA. Currently working as Innovation Lead in SSC Space and R&D engineer at the ESAC Science Data Center working on world-class space missions like Euclid, JWST and HST. Some of these missions are producing peta-byte scale data sets and we are developing science archives and platforms with the tools that scientists need to discover and consume this data.