Lead Site Reliability Engineer - Operations Excellence for AI Platforms
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer focussed on Operations Excellence, you will have the opportunity to shape how we respond to, learn from, and prevent incidents across a complex, high-stakes technology environment.
This is a role where your impact is visible, your voice carries weight, and your work directly influences the stability of services relied upon by millions.
As a Lead Software Engineer - Site Reliability Engineer, Operations Excellence at JPMorganChase within the AI/ML & Data Platforms area , you will serve as a technical leader at the intersection of software engineering and reliability engineering - owning the incident management lifecycle, defining reliability standards, and partnering across Engineering, Product, Infrastructure, and Security to deliver durable operational improvements.
You will bring structure to complexity, clarity to high-urgency situations, and a continuous improvement mindset to everything from alert quality to executive reporting.
Your work will directly strengthen the firm's ability to detect, respond to, and prevent production issues at scale.
Job Responsibilities
* Own and continuously improve the incident management lifecycle, including triage, escalation, stakeholder communications, and recovery, ensuring consistent execution and measurable improvement over time.
* Serve as Incident Commander for major incidents, coaching responders to follow defined processes and site reliability engineering best practices while maintaining clear and timely stakeholder communication.
* Drive operational readiness through drills, game days, and failure-mode exercises that improve response effectiveness, surface gaps, and build team resilience before incidents occur.
* Define, implement, and enforce reliability standards across services, including service level indicators and objectives, error budgets, monitoring coverage, and alert quality.
* Strengthen change and release management practices by establishing readiness checks, progressive delivery standards, rollback procedures, runbook quality, and production hygiene expectations.
* Partner with Engineering, Product, Infrastructure, and Security teams to prioritize reliability work and deliver solutions that address user and operational pain points with lasting impact.
* Lead problem management and root cause analysis by debugging complex production issues, identifying systemic root causes, and driving durable remediation and preventative actions.
* Own reliability reporting and operational governance by tracking key performance indicators - including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate - and producing executive-ready summaries.
* Facilitate operational...
- Rate: Not Specified
- Location: Jersey City, US-NJ
- Type: Permanent
- Industry: Finance
- Recruiter: JPMorgan Chase Bank, N.A.
- Contact: Not Specified
- Email: to view click here
- Reference: 210747379
- Posted: 2026-08-19 09:17:03 -
- View all Jobs from JPMorgan Chase Bank, N.A.
More Jobs from JPMorgan Chase Bank, N.A.
- Lagermitarbeiter / Sortierer für Briefe (m/w/d)
- Lagermitarbeiter / Sortierer für Briefe (m/w/d)
- Export Processor
- Senior FP&A Cost Analyst
- Customer Service Specialist
- Shipping Coordinator
- Operations Manager
- Sr. Buyer
- Customer Account Coordinator
- Senior LLM Engineer
- Senior LLM Engineer
- Senior LLM Engineer
- Storeroom Leader
- Shift Supervisor
- Senior/Staff Engineer Fiberline & Bleach
- Cable Assembly Technician
- Process and Analytics Engineer
- Senior Accounting Analyst
- Ocean Logistics Procurement Lead
- Optical Process Engineer