US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


Lead Site Reliability Engineer

As a Lead Site Reliability Engineer at JPMorgan Chase within the Public Cloud team, you will blend hands-on engineering with program leadership to promote platform stability, ensure consistent execution across SRE teams, and partner closely with Engineering and Product to deliver measurable improvements in availability, support outcomes, and cost of failure.
Job Responsibilities


* Drive consistency across SRE teams: establish and scale a "gold standard" for SLOs, on-call readiness, runbooks, postmortems, action tracking, and operational readiness across AWS/Azure/GCP.


* Partner with Engineering and Product: embed reliability outcomes into roadmaps and delivery plans; convert incidents, support signals, and error budget trends into prioritized backlog with measurable impact.


* Operational excellence & BPMs: design and run operating cadences (incident/stability reviews, KPI reviews, planning inputs), standardize intake/prioritization, and ensure closed-loop execution.


* Risk governance: align reliability operations to risk/control expectations; operationalize recurring remediations (e.g., configuration drift, repeat findings) through centralized automation without sacrificing velocity.


* Ticket reduction & automation opportunities: identify top drivers of support load and operational toil; build and maintain an automation opportunity pipeline; track adoption and deflection.


* Platform stability metrics: own cross-platform reporting (SLO attainment, incident trends, MTTR/MTTI, change failure rate, ticket deflection, customer impact/cost of failure).


* AI for reliability operations: apply AI/LLMs to improve triage, incident summarization, correlation/symptom mapping, and guardrailed automation; measure accuracy, safety, and outcomes.


* Hands-on leadership: stay close to designs and critical implementations; lead systemic remediation and major incident response improvements.

Required qualifications, skills, and capabilities


* 5+ years in SRE / production engineering / platform reliability / infrastructure operations at enterprise scale


* Demonstrated success driving cross-team standardization and measurable reliability outcomes through influence and operating mechanisms.


* Deep knowledge of SLOs/SLIs, error budgets, observability, incident response, postmortems, and change reliability.


* Strong experience partnering with Engineering and Product leadership to align priorities and deliver results.


* Hands-on experience with AWS and/or Azure and automation/IaC fundamentals (e.g., Python/Go/Bash, Terraform, CI/CD).


* Experience using AI/LLM-enabled approaches in operations (AIOps, AI-assisted troubleshooting, agentic workflows) with appropriate controls and measurement.

Preferred Qualifications


* Improved SLO attainment and reduced Sev1/Sev2 frequency; fewer repeat incidents.


* Reduced MTTR/MTTI and lower change failure rate; reduced customer minutes impacted (cost of fai...




Share Job