BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.adass.org//adass2026//speaker//UMDVNL
BEGIN:VTIMEZONE
TZID:AWST
BEGIN:STANDARD
DTSTART:20000101T000000
RRULE:FREQ=YEARLY;BYMONTH=1;UNTIL=20051231T160000Z
TZNAME:AWST
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
END:STANDARD
BEGIN:STANDARD
DTSTART:20070325T040000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=3
TZNAME:AWST
TZOFFSETFROM:+0900
TZOFFSETTO:+0800
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20071028T030000
RRULE:FREQ=YEARLY;BYDAY=4SU;BYMONTH=10;UNTIL=20081025T190000Z
TZNAME:AWDT
TZOFFSETFROM:+0800
TZOFFSETTO:+0900
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-adass2026-QG8EJW@pretalx.adass.org
DTSTART;TZID=AWST:20261102T140000
DTEND;TZID=AWST:20261102T141500
DESCRIPTION:The rapid adoption of generative AI is transforming the way sci
 entists interact with astronomy science archives and platforms. However\, 
 the growing dependence on large foundation models raises important questio
 ns about sustainability\, cost\, and long-term operational viability. In t
 he Data Science and Archives Division at the European Space Astronomy Cent
 re (ESAC)\, we are exploring different ways to use AI to support our users
 . Among others\, we are exploring how smaller open source language models 
 combined with retrieval-augmented generation (RAG) and domain-specific too
 ling can enhance astronomical data discovery while minimizing computationa
 l overhead. Rather than treating state of the art frontier models as the d
 efault solution\, we are investigating if smaller models can be effectivel
 y deployed for many archive-support tasks with considerably lower infrastr
 ucture requirements.\n\nThe ESAC astronomy science archives present a part
 icularly attractive use case for this approach. In legacy missions\, the n
 umber of specialists available to support users decreases with time\, whil
 e the scientific value of these missions remains high for decades. By comb
 ining lightweight large language models (LLM) with RAG pipelines built fro
 m mission documentation\, archive interfaces user guides\, and selected sc
 ientific publications\, it is possible to provide conversational interface
 s that preserve mission knowledge that can be used by scientists to discov
 er and use relevant datasets. Similarly\, for missions in operations\, suc
 h as Euclid\, AI assistants could help scientists navigate its complex dat
 a model\, learn the details of hundreds of distinct data products\, and ge
 nerate ADQL example queries and notebooks.\n\nWe argue that this approach 
 could complement other alternatives and represents a more sustainable path
  for scientific archives. Running open-source LLMs with “only” a few b
 illion parameters on local GPU infrastructure would allow institutions to 
 reuse existing resources\, reduce dependence on commercial AI subscription
 s\, and limit the environmental and financial costs associated with large-
 scale cloud inference. This approach also maximizes the value of open-sour
 ce models whose development has already required significant community inv
 estment. In this presentation we will showcase some of the experimental wo
 rk we are doing at ESAC using open source models running on our local infr
 astructure\, for example\, enhancing the existing ESASky chatbot with gene
 rative AI capabilities and building RAG pipelines for mission documentatio
 n and archival data discovery.
DTSTAMP:20261001T101732Z
LOCATION:Banquet Hall
SUMMARY:Sustainable AI for Astronomical Archival Data Discovery in the Peta
 byte Era - Marcos López-Caniego
URL:https://pretalx.adass.org/adass2026/talk/QG8EJW/
END:VEVENT
END:VCALENDAR
