About the job
You will work on the Model Factory, building the data and analytics systems underneath it such as our agent trajectory store, experiment configuration registry, experiment metrics, and others. In this role, you will engage directly with researchers to understand their experiments and workflows so you can improve our data and analytics. Since training efficiency is a function of how quickly we learn and draw insights from each run, your work will directly accelerate Poolside's ability to develop models. The work on this team will be a mix of greenfield projects as well as extending existing systems to support larger scale, higher reliability, and more demanding performance requirements.
Responsibilities
- Design, build, and operate the data and analytics systems of our Model Factory https://poolside.ai/research
- Own the entire data lifecycle of experiment data including modelling, ingestion, retrieval, query, analysis, and retention
- Partner closely with researchers to understand their workflows and ensure we're guided by their data and analytics needs
Qualifications
Minimum
- Strong programming skills in Go, Python, or other similar languages
- Proven experience building and operating distributed systems at scale, with a strong understanding of consistency models, queue/stream processing and data pipelines
- Production experience with Kubernetes, cloud infrastructure, observability systems
- Ability to lead ambiguous and cross-functional initiatives, building alignment and driving progress
Preferred
- experience working OLAP databases, timeseries databases, high cardinality datasets, understanding of ML research workflows