BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.adass.org//adass2026//talk//ATU73E
BEGIN:VTIMEZONE
TZID:AWST
BEGIN:STANDARD
DTSTART:20000101T000000
RRULE:FREQ=YEARLY;BYMONTH=1;UNTIL=20051231T160000Z
TZNAME:AWST
TZOFFSETFROM:+0800
TZOFFSETTO:+0800
END:STANDARD
BEGIN:STANDARD
DTSTART:20070325T040000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=3
TZNAME:AWST
TZOFFSETFROM:+0900
TZOFFSETTO:+0800
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20071028T030000
RRULE:FREQ=YEARLY;BYDAY=4SU;BYMONTH=10;UNTIL=20081025T190000Z
TZNAME:AWDT
TZOFFSETFROM:+0800
TZOFFSETTO:+0900
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-adass2026-ATU73E@pretalx.adass.org
DTSTART;TZID=AWST:20261104T093000
DTEND;TZID=AWST:20261104T094500
DESCRIPTION:The Nancy Grace Roman Space Telescope will produce approximatel
 y 20 PB over its five-year mission\, requiring data reduction systems that
  are elastic\, reproducible\, cost-aware\, and tightly integrated with the
  Roman Research Nexus science platform. We present the architecture and op
 erational model for the Space Telescope Science Institute’s Roman data p
 rocessing pipeline. This fully cloud-based data processing system is part 
 of the larger distributed ground systems developed for NASA's latest flags
 hip astrophysics telescope.\n\nThe Roman pipeline is built around containe
 rized processing workloads orchestrated by Kubernetes and Apache Airflow\,
  with fully reproducible infrastructure through Infrastructure as Code. Th
 e design supports multiple concurrent versions of the Roman calibration so
 ftware\, enabling controlled reprocessing\, validation\, and operational f
 lexibility as algorithms and mission needs evolve. The system scales elast
 ically from zero workers when idle to thousands of parallel workers during
  processing campaigns.\n\nWe describe the architectural patterns used to s
 upport petabyte-scale processing\, including workflow abstraction\, storag
 e considerations\, external interfaces for data receipt and distribution\,
  and access patterns for downstream systems. Observed load has exceeded th
 ousands of concurrent jobs without issue. At demonstrated scale\, the syst
 em easily provisioned 60TB of memory across 7\,000 vCPUs achieving remarka
 ble hourly throughput. These figures represent tested operating points rat
 her than architectural limits\; continued load testing is expected to expl
 ore far higher levels of concurrency\, with practical scaling constraints 
 driven primarily by cloud resource availability\, service quotas\, and cos
 t.\n\nThe presentation will also address design tradeoffs for speed\, port
 ability\, scale\, security\, and cost\, along with monitoring and observab
 ility strategies required to operate a large production cloud workflow sys
 tem. We will also present lessons learned from operating at scale in AWS\,
  including service limits\, performance bottlenecks\, and discuss how pipe
 line outputs integrate with the Roman Research Nexus to provide a compute-
 to-data science platform for accessible\, reproducible analysis.
DTSTAMP:20261001T111404Z
LOCATION:Banquet Hall
SUMMARY:Building a Cloud-Native\, Petabyte-Scale Pipeline for the Roman Spa
 ce Telescope - John Glorioso
URL:https://pretalx.adass.org/adass2026/talk/ATU73E/
END:VEVENT
END:VCALENDAR
