2026-11-04 –, Banquet Hall
The Nancy Grace Roman Space Telescope will produce approximately 20 PB over its five-year mission, requiring data reduction systems that are elastic, reproducible, cost-aware, and tightly integrated with the Roman Research Nexus science platform. We present the architecture and operational model for the Space Telescope Science Institute’s Roman data processing pipeline. This fully cloud-based data processing system is part of the larger distributed ground systems developed for NASA's latest flagship astrophysics telescope.
The Roman pipeline is built around containerized processing workloads orchestrated by Kubernetes and Apache Airflow, with fully reproducible infrastructure through Infrastructure as Code. The design supports multiple concurrent versions of the Roman calibration software, enabling controlled reprocessing, validation, and operational flexibility as algorithms and mission needs evolve. The system scales elastically from zero workers when idle to thousands of parallel workers during processing campaigns.
We describe the architectural patterns used to support petabyte-scale processing, including workflow abstraction, storage considerations, external interfaces for data receipt and distribution, and access patterns for downstream systems. Observed load has exceeded thousands of concurrent jobs without issue. At demonstrated scale, the system easily provisioned 60TB of memory across 7,000 vCPUs achieving remarkable hourly throughput. These figures represent tested operating points rather than architectural limits; continued load testing is expected to explore far higher levels of concurrency, with practical scaling constraints driven primarily by cloud resource availability, service quotas, and cost.
The presentation will also address design tradeoffs for speed, portability, scale, security, and cost, along with monitoring and observability strategies required to operate a large production cloud workflow system. We will also present lessons learned from operating at scale in AWS, including service limits, performance bottlenecks, and discuss how pipeline outputs integrate with the Roman Research Nexus to provide a compute-to-data science platform for accessible, reproducible analysis.
John Glorioso is a Chief Engineer in the Data Management Division at the Space Telescope Science Institute. He is the cloud architect for the Roman Science Data Pipeline and responsible for guiding development efforts. His career has spanned thirty years in software development and architecture with a focus on high throughput data processing.