US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


Vice President - Senior Manager of Site Reliability Engineering

Elevate your engineering leadership to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.

In this high-impact role, you will guide and shape the future of large-scale data platform reliability - bringing your expertise in site reliability engineering, platform engineering, and AI/ML infrastructure to mentor and lead a team of 8-10 engineers.

As a Senior Manager of Site Reliability Engineering at JPMorganChase within the Chief Data and Analytics Office AI/ML and Data Platforms team, you are the non-functional requirement owner and champion for the applications in your remit.

You will define availability targets, embed reliability principles into product design and testing, and ensure service level indicators and objectives are implemented in production to support secure, scalable, and high-performing analytics and AI/ML workloads.

You act in a blameless, data-driven manner and navigate difficult situations with composure and tact.

Job responsibilities


* Lead, mentor, and develop a team of 8-10 site reliability and platform engineers, fostering a culture of ownership, blameless post-mortems, and continuous improvement through tailored feedback and growth plans


* Own and champion non-functional requirements, availability targets, service level indicators, and service level objectives for services supporting large-scale data platforms and AI/ML workloads, ensuring alignment with stakeholders and production readiness standards


* Drive the design, implementation, and evolution of observability and reliability frameworks across distributed systems and data platform environments, leveraging tools such as Grafana, Dynatrace, Prometheus, Datadog, and Splunk


* Oversee the architecture and operational stability of data platform infrastructure, including Databricks, Spark-based data pipelines, and big data ecosystem tools, ensuring scalability, security, and high performance


* Lead reuse-first adoption of enterprise-authorized AI capabilities within site reliability engineering workflows, establishing team standards for traceability, auditability, and alignment to resiliency and security expectations, with human-in-the-loop validation


* Drive a culture of continual improvement by encouraging real-time feedback loops, conducting regular team debriefs, and applying objective, data-driven post-mortem strategies that enable teams to learn from both successes and failures


* Manage stakeholders and ensure teams deliver projects aligned with compliance standards, risk and security requirements, service level agreements, and business objectives


* Contribute to staffing, budget, and resource planning decisions, including hiring, developing, and recognizing engineering talent across the team


* Champion site reliability engineering culture and principles across the organization, ensuring teams document and share knowledge and innov...




Share Job