About the job
We're looking for a Software Engineer (L4) to help build and operate the CI/CD pipelines and test infrastructure that our engineering organization depends on every day. You'll work on well-scoped, meaningful pieces of larger platform initiatives, with guidance and mentorship from senior engineers and your manager while growing your ability to independently own projects end-to-end. This role blends software engineering with operations: you'll write code that improves pipeline reliability and test execution, and you'll also be on the front line when builds break, tests flake, the CI/CD system needs troubleshooting or the production system needs observability investments.
Responsibilities
Build and maintain components of our CI/CD pipelines (build orchestration, workflow automation, GitHub Actions) that engineering teams rely on to ship safely and quickly
Develop and support test execution infrastructure, including test scheduling, distributed run coordination, and results reporting
Build and ship AI-powered platform features for CI/CD, metrics, and testing infrastructure, agentic systems and AI-assisted automation for problems like alert triage, incident summarization, and test-health diagnostics
Design and implement the distributed backbone for Commerce Observability agentic AI-driven analysis, inference, and orchestration systems
Contribute to tooling that detects, quarantines, and helps diagnose flaky tests
Partner with test-framework and platform teams to integrate new tools and frameworks into the pipeline
Triage, troubleshoot, and resolve pipeline and test-infrastructure incidents; participate in an on-call rotation for CI/CD and test operations
Qualifications
Minimum
4+ years of professional software engineering experience (or equivalent practical experience)
Proficiency in at least one of JavaScript/TypeScript, or Java
Working knowledge of CI/CD tooling such as GitHub Actions, Jenkins, or CircleCI
Gen AI stack: You have strong interest and experience with the latest GenAI stack (LLMs, RAG, Agents). You have familiarity with AI Observability systems like Braintrust
You've shipped LLM-powered systems that are near-deterministic, scalable, and monitored in production
You build and maintain evaluation frameworks, not just features. You know how to measure whether an AI-assisted system (e.g., alert triage, test-health diagnostics) is actually working, and use results to guide design decisions
You understand model internals and ML methodologies well enough to evaluate and select models with rigor and to recognize when AI is not the right tool for a problem
Solid understanding of software testing concepts (unit, integration, E2E) and how automated testing fits into a delivery pipeline
Strong debugging and troubleshooting skills, and comfort working in an operational, on-call context
Clear written and verbal communication; able to document technical work for other engineers
Ability to execute well-scoped projects with moderate guidance, and a track record of growing scope and independence over time
A genuine interest in developer productivity, platform engineering, or testing infrastructure
Preferred
Experience with E2E or test frameworks such as Playwright, Selenium, Cypress, Jest, or Mocha
Exposure to distributed tracing/observability tools (Datadog, Honeycomb, Splunk, or similar)
Hands-on experience with GenAI-assisted developer tooling
Experience with GraphQL
Prior experience supporting production systems or an on-call rotation
Experience working with a globally distributed team