A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations
This work addresses the challenges of managing high-quality, dynamically evolving annotated data essential for Grade of Automation 3–4 (GoA3–GoA4) train operations, where existing data catalogs suffer from high operational overhead, weak development integration, and documentation drift. To overcome these limitations, the authors propose a lightweight, GitOps-driven data catalog architecture that introduces the Data-as-Code paradigm into railway AI perception systems for the first time. By leveraging CI/CD pipelines and static site generation, the approach enables versioned data management, end-to-end traceability, and automated documentation. This solution significantly reduces operational costs, enhances development integration, ensures regulatory compliance, and automatically produces high-performance dataset overview pages.