About the job
We're looking for a Principal Product Manager to sit inside that process. You'll be embedded directly with applied researchers and ML engineers running RLHF, fine-tuning, and preference-tuning pipelines, reading model outputs, calling out what's wrong, and turning that into concrete training priorities. You're the person who brings the end user's voice into a room that's otherwise all metrics and loss curves.
This is a technical, hands-on PM role, not just a roadmap-and-requirements role. You'll spend real time in transcripts and eval dashboards, and you'll need enough grounding in how post-training actually works (RLHF, DPO/preference optimization, instruction tuning, data curation) to have a credible point of view with the research team.
Responsibilities
Embed in the post-training loop: work day-to-day with research and applied ML teams during fine-tuning and RLHF cycles, reviewing model outputs and giving structured feedback on quality, tone, and behavior
Define what "good" means for Copilot's models: set the priorities for capability, personality, and safety tradeoffs across surfaces (Microsoft 365 Copilot, Copilot Studio agents, Windows) — where "good" is often genuinely ambiguous and hasn't been decided before
Run RLHF/preference data campaigns: partner with data and labeling teams to scope what human preference data gets collected, for which behaviors, and why
Build and maintain behavior evals: translate qualitative judgment calls ("this response felt too hedgy," "this refused when it shouldn't have") into evaluation sets that can be tracked release over release
Own the feedback loop: turn user research, enterprise customer feedback, and production incident learnings into concrete post-training priorities — closing the gap between "users are unhappy with X" and "the next fine-tune addresses X"
Make the tuning tradeoffs explicit: helpfulness vs. caution, consistency vs. personality, latency/cost vs. quality — and drive alignment across research, safety, and product leadership on where the line sits
Represent Responsible AI and enterprise requirements into training priorities — compliance behaviors, refusal policies, and brand voice all have to be reflected in what the model is actually tuned to do.
Qualifications
Minimum
Bachelor's Degree AND 10+ years experience in product/service/program management or software development OR equivalent experience.
Preferred
Bachelor's Degree AND 15+ years experience in product/service/program management or software development OR equivalent experience.
6+ years experience taking a product, feature, or experience to market (e.g., design, addressing product market fit, and launch, internal tool/framework).
8+ years experience improving product metrics for a product, feature, or experience in a market (e.g., growing customer base, expanding customer usage, avoiding customer churn).
8+ years experience disrupting a market for a product, feature, or experience (e.g., competitive disruption, taking the place of an established competing product).