About the job
We are seeking a Research Engineer to join our Pre-training team, responsible for developing the next generation of large language models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.
Responsibilities
Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development
Independently lead small research projects while collaborating with team members on larger initiatives
Design, run, and analyze scientific experiments to advance our understanding of large language models
Optimize and scale our training infrastructure to improve efficiency and reliability
Develop and improve dev tooling to enhance team productivity
Contribute to the entire stack, from low-level optimizations to high-level model design
Qualifications
Minimum
Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field
Strong software engineering skills with a proven track record of building complex systems
Expertise in Python and experience with deep learning frameworks (PyTorch preferred)
Familiarity with large-scale machine learning, particularly in the context of language models
Ability to balance research goals with practical engineering constraints
Strong problem-solving skills and a results-oriented mindset
Excellent communication skills and ability to work in a collaborative environment
Care about the societal impacts of your work
Preferred
Work on high-performance, large-scale ML systems
Familiarity with GPUs, Kubernetes, and OS internals
Experience with language modeling using transformer architectures
Knowledge of reinforcement learning techniques
Background in large-scale ETL processes