Staff Software Engineer, Ray Data

Anyscale
San Francisco2026-09-09

About the job

Ray Data is a Python-native data processing engine and a one-stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting-edge AI frameworks using both multimodal and structured data.

The Ray Data team develops and maintains Ray Data, building the underlying distributed data processing infrastructure that powers modern AI workloads. We are a team of engineers passionate about solving challenging problems in distributed systems, data processing, and performance at scale. We are looking for exceptional engineers to build, optimize, and scale Ray Data for increasingly complex AI workloads, including multimodal data processing and large-scale batch inference.

Responsibilities

- Design, build, and improve the core systems that power Ray Data, with a focus on performance, scalability, and reliability.

- Design and optimize distributed execution across different stages of data pipelines in heterogeneous environments.

- Build data loading and processing solutions for production training and inference workloads.

- Solve challenging problems in distributed execution, scheduling, resource management, data partitioning, fault tolerance, and performance optimization.

- Make system-level architectural decisions and reason through tradeoffs in areas such as resource allocation, execution models, batch vs. streaming workloads, and consistency and availability.

- Work with customers and new-age AI-native companies to understand and solve challenges in scaling their AI workloads.

Qualifications

Minimum

- 6+ years of experience building production-grade software, infrastructure, or developer-facing systems, with strong Python engineering experience.

- 6+ years of experience personally owning core architectural decisions within a distributed data or compute engine, rather than primarily operating or using a platform someone else designed.

- Deep experience with distributed systems internals, such as scheduling, fault tolerance, data partitioning, distributed execution, performance optimization, or database and query engine internals.

- A track record of reasoning through system-level tradeoffs and defending architectural decisions, such as batch vs. streaming, static vs. dynamic resource allocation, or consistency vs. availability.

- Passion for solving the unsolved problems in large-scale AI infrastructure and building systems that enable the next generation of AI applications.

Preferred

No preferred qualifications listed.