About the job
We are building the evaluation backbone for safe, reliable, and efficient agentic engineering in Microsoft Security. This team will create systems that determine when an AI agent, model, prompt, tool, memory strategy, or orchestration pattern is ready to be used in production security and engineering workflows. The role is ideal for engineers who can operate end to end: understand the workflow, design the benchmark, build the harness, implement validators and graders, run experiments, analyze quality and cost tradeoffs, connect results to production feedback, and help teams make evidence-based release decisions.
Responsibilities
Design and build end-to-end evaluation harnesses for agentic security and engineering workflows, including triage, remediation, repo readiness, escalation, tool use, and scan-to-verified-closure paths.
Create representative benchmark suites and golden datasets that include normal, edge, adversarial, failure-recovery, regression, and production-derived cases.
Implement deterministic validators, automated graders, trace analyzers, result stores, comparison views, and workflow adapters that make evaluations repeatable and actionable.
Measure task success, correctness, safety and policy compliance, failure recovery, latency, tool-call behavior, token usage, total cost per successful outcome, and human-review effort.
Compare models, prompts, tools, memory strategies, policies, and orchestration patterns under consistent conditions and help teams understand quality-versus-efficiency tradeoffs.
Integrate evaluations into engineering workflows, CI/CD, release gates, and decision processes so material agent changes are supported by reproducible evidence before production rollout.
Connect offline evaluation results with production feedback, including accepted and rejected outputs, human overrides, incidents, rollbacks, escaped defects, and customer or service-health signals.
Detect regressions, benchmark drift, evaluator miscalibration, and cases where eval scores improve while real-world outcomes do not.
Partner with agent builders, product teams, security engineers, data scientists, program managers, and leadership to turn evaluation results into release recommendations and autonomy-boundary decisions.
Convert learnings into reusable paved paths: playbooks, templates, onboarding guides, dashboards, scorecards, and reference implementations that can scale across MSec.
Qualifications
Minimum
Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 3+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
OR Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 4+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
OR Bachelor's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 6+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
OR equivalent experience.
1+ year(s) people management experience.
Preferred
Doctorate in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 5+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
OR Master's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 8+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
OR Bachelor's Degree in Statistics, Mathematics, Computer Science, Computer Security, or related field AND 12+ years experience in software development lifecycle, large-scale computing, threat analysis or modeling, cybersecurity, vulnerability research, and/or anomaly detection
OR equivalent experience.
Proven software engineering experience building production systems, developer platforms, test infrastructure, automation frameworks, data pipelines, quality systems, or reliability tooling.
Ability to design and implement evaluation systems end to end, including task definition, dataset creation, harness implementation, scoring, analysis, and operational integration.
Demonstrated coding, debugging, system design, and operational excellence skills.
Experience working with structured data, logs, traces, metrics, APIs, automation workflows, and engineering telemetry.
Experience with LLMs, AI agents, model evaluation, prompt/tool orchestration, automated grading, evaluation harnesses, or benchmark design.
Experience with experimentation, statistical confidence, evaluator calibration, regression analysis, human-review protocols, or quality measurement systems.
Experience with security engineering, vulnerability management, SAST/SCA, SARIF, remediation workflows, secure development lifecycle, or compliance-sensitive systems.