Radio interferometric pipelines encode expert judgment as fixed heuristics: reasonable defaults for a "typical" dataset. When a dataset doesn't fit the mold, an expert inspects the diagnostics and steps in by hand. That doesn't scale as surveys push toward higher data volumes and less human oversight. LLMs can close that gap. They read the same diagnostics an expert would and reason about them per dataset. But letting a model both reason and act is risky. LLMs hallucinate, drift, and skip steps silently. We built an architecture that strictly separates measurement from reasoning. The model gets room to adapt, but it never touches the data path directly.
An MCP layer wraps the reduction software (CASA, here) and exposes each Measurement Set operation (metadata queries, instrument geometry, calibration, imaging) as an independent tool. Every tool returns structured data with explicit completeness and provenance. None of them interpret their own output or call each other. Reasoning lives outside the tools, in skills: version-controlled, plain-text documents that encode interferometric expertise. They're fed into the LLM's context stage by stage, so it can reason about what the tools hand back.
An orchestration layer walks the model through a sequence of stages that looks like a processing pipeline, except the parameters aren't fixed. Skill-based reasoning lets the model make an informed, per-dataset call instead of falling back on defaults. Every stage's state, its parameter choices, and the sequencing are written out as documentation and a reproducible script. Since reasoning lives outside the MCP layer, the orchestrator doesn't care which model is driving it. Cloud (Claude, Codex) or local and open (Gemma, Qwen): both work. A model equipped with real domain expertise, via skills, beats a one-size-fits-all pipeline, dataset by dataset.
We demonstrate end-to-end inspection and calibration on VLA, GMRT, and ALMA data. We discuss extending the approach to other instruments, facilities, and pipelines.