About the job
We are seeking exceptional researchers who can push the frontier of safety mitigations. You will help derisk frontier models by developing novel safety mitigations, developing and applying new techniques from domains like interpretability, control, and alignment to ensure the safety of OpenAI’s deployed models. You will play a critical role in defining how a safe AI system should look in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI.
Responsibilities
Work on identifying emerging AI safety risks and new methodologies for exploring and mitigating the impact of such risks
Build (and then continuously refine) the evaluations that enable us to assess the extent of these risks; this might include working with domain experts (be it internal or external)
Set research directions and strategies to make our AI systems safer, more aligned, and more robust
Contribute to the development of “best practices” guidelines for AI safety for OpenAI and across the industry
Evaluate and design effective red-teaming pipelines to examine the end-to-end robustness of our safety systems, and identify areas for future improvement
Qualifications
Minimum
2+ years of experience in the field of AI safety, especially in areas like RLHF, human-AI collaboration, interpretability, or control
Ph.D. or other degree in computer science, machine learning, or a related field
4+ years of research engineering experience and proficiency in Python or similar languages
Preferred
Excited about OpenAI’s mission of building safe, universally beneficial AGI and are aligned with OpenAI’s charter
Enthusiasm for long-term AI safety, and have thought deeply about technical paths to safe AGI
Willing to “get your hands dirty” and apply methods from domains such as interpretability, robustness, alignment, and control to make OpenAI’s models safe
Thrive in environments involving large-scale AI systems