US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


Lead Site Reliability Engineer - Operations Excellence for AI Platforms

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer focussed on Operations Excellence, you will have the opportunity to shape how we respond to, learn from, and prevent incidents across a complex, high-stakes technology environment.

This is a role where your impact is visible, your voice carries weight, and your work directly influences the stability of services relied upon by millions.

As a Lead Software Engineer - Site Reliability Engineer, Operations Excellence at JPMorganChase within the AI/ML & Data Platforms area , you will serve as a technical leader at the intersection of software engineering and reliability engineering - owning the incident management lifecycle, defining reliability standards, and partnering across Engineering, Product, Infrastructure, and Security to deliver durable operational improvements.

You will bring structure to complexity, clarity to high-urgency situations, and a continuous improvement mindset to everything from alert quality to executive reporting.

Your work will directly strengthen the firm's ability to detect, respond to, and prevent production issues at scale.

Job Responsibilities


* Own and continuously improve the incident management lifecycle, including triage, escalation, stakeholder communications, and recovery, ensuring consistent execution and measurable improvement over time.


* Serve as Incident Commander for major incidents, coaching responders to follow defined processes and site reliability engineering best practices while maintaining clear and timely stakeholder communication.


* Drive operational readiness through drills, game days, and failure-mode exercises that improve response effectiveness, surface gaps, and build team resilience before incidents occur.


* Define, implement, and enforce reliability standards across services, including service level indicators and objectives, error budgets, monitoring coverage, and alert quality.


* Strengthen change and release management practices by establishing readiness checks, progressive delivery standards, rollback procedures, runbook quality, and production hygiene expectations.


* Partner with Engineering, Product, Infrastructure, and Security teams to prioritize reliability work and deliver solutions that address user and operational pain points with lasting impact.


* Lead problem management and root cause analysis by debugging complex production issues, identifying systemic root causes, and driving durable remediation and preventative actions.


* Own reliability reporting and operational governance by tracking key performance indicators - including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate - and producing executive-ready summaries.


* Facilitate operational...




Share Job