Sagar Borra

Hold a Masters degree in Computer Science - Focus on Big data and AI.
Working for German Center for Astrophysics as a Data scientist.


Session

11-04
13:45
15min
Bridging Kubernetes and Slurm for Transparent HPC Job Offloading
Sagar Borra

Bridging Kubernetes and Slurm for Transparent HPC Job Offloading

Authors:
S. Borra1, U. Canbolat1, M. Drobek1, L. Haupt1

Affiliation:
1 German Center for Astrophysics (DZA), Görlitz, Germany

Date:
July 30, 2026

Abstract:
Kubernetes has emerged as a control plane for modern scientific platforms, supporting reproducible and high-throughput computing workflows. Additionally, high-performance computing (HPC) systems managed by Slurm remain essential for large-scale, compute-intensive processing. Despite their complementary roles, these environments are typically disconnected, requiring manual workload transitions and resulting in fragmented observability and control.

Several approaches were explored to bridging this gap, including systems such as interLink, Volcano, REANA (Reusable Analyses) and Slinky. In this work, we present interLink, a lightweight integration that extends Kubernetes orchestration into a Slurm-managed HPC cluster. By representing HPC resources within Kubernetes, workloads can be offloaded directly and executed as Slurm jobs without requiring changes to user workflows or tooling. A Kubernetes pod, the smallest deployable unit in Kubernetes, encapsulating one or more containers and their execution context, acts as the interface through which users submit and monitor jobs. interLink ensures that job state, logs, and execution progress are continuously reflected back into this originating pod, preserving a consistent operational view.

This approach enables Kubernetes to function as a unified control plane for both service-oriented and batch-oriented workloads. In particular, a remote user can use Jupyter notebooks interactively, for example, in preparation of batch jobs to be run on large datasets stored in the HPC center. This presentation explains how interLink reduces operational friction, maintains end-to-end observability across system boundaries, and enables seamless integration of high-throughput and HPC workflows within a single environment.

Building and operating science platforms and workflows in the Petabyte Era
Banquet Hall