US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


AI Ops Engineer

AI Ops Engineer

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work.

We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today's complex world.

Our culture thrives on finding new and better ways to accelerate what's next.

We know varied backgrounds are valued and succeed here.

We have the flexibility to manage our work and personal needs.

We make bold moves, together, and are a force for good.

If you are looking to stretch and grow your career our culture will embrace you.

Open up opportunities with HPE.

Job Description:

Job Description

Job Family Definition:

Designs, develops, troubleshoots, and debugs software programs for enhancements and new product development.

Develops software components such as operating systems, compilers, routers, networks, utilities, databases, and Internet-related tools.

Assesses hardware compatibility and/or influences hardware design decisions.

Management Level Definition:

Applies specialized subject matter expertise to resolve common and occasionally complex technical challenges, recommending alternatives as needed.

May act as a project lead and provide guidance to junior professionals.

Exercises independent judgment and collaborates with others to determine the most effective methods for accomplishing work and achieving objectives.

Responsibilities:


* Design and maintain intelligent monitoring and incident response workflows across both cloud and on-premises systems.


* Build and optimize data pipelines to collect, normalize, and enrich logs, metrics, and events for operational analytics.


* Develop and deploy predictive models to detect anomalies, forecast outages, and reduce mean time to resolution.


* Automate remediation runbooks and integrate alerting with service management platforms to enhance system reliability.


* Collaborate with DevOps, Site Reliability, and security teams to improve platform performance, availability, and operational governance.

Education and Experience Required:


* Bachelor's degree in Computer Science, Information Technology, or a related field.


* 5-8 years of experience in IT operations, reliability engineering, platform engineering, or roles focused on AI-driven operations and automation.

Knowledge and Skills:


* Proficiency in Python for automation, data processing, and operational tooling - required.


* Experience with observability platforms such as Splunk, Datadog, Prometheus, or Grafana - required.


* Familiarity with cloud platform operations (Azure, AWS, or Google Cloud) - preferred.


* Knowledge of machine learning model lifecycle management and deployment - advantageous.


* Experie...




Share Job