US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


AI and Machine Learning Engineer

AI and Machine Learning Engineer

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work.

We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today's complex world.

Our culture thrives on finding new and better ways to accelerate what's next.

We know varied backgrounds are valued and succeed here.

We have the flexibility to manage our work and personal needs.

We make bold moves, together, and are a force for good.

If you are looking to stretch and grow your career our culture will embrace you.

Open up opportunities with HPE.

Job Description:

High Performance Computing, AI and Labs are a critical element of HPE.

We are focused on delivering innovative solutions that accelerate our customers' digital transformation, enabling them to tackle their complex, and data-intensive workloads.

Combining deep expertise and the development of the world's most cutting-edge, high-performance supercomputers, is defining the next era of computing delivering valuable insight & innovation.

Join us and redefine what's next for you.

Responsibilities:



* Installs and configures complex IT infrastructure components (servers, storage, network)


* Develop software scripts and configurations for automating deployment.


* Study and improve the performance of Large Language Models run on HPE GPU servers


* Performs system level analysis of server workloads on various HPE platforms running DL and ML code to include accelerated hardware and high-speed networks like InfiniBand


* Writes white papers and other guidance documents for AI workload and model selection


* Captures and reviews system performance data, logs, traces to understand workload behaviour


* Develops software and scripts that help analyse AI workload performance data


* Communicates technical work well and can provide summaries of work to non-technical colleagues


* Works with software and hardware partners in optimizing systems and resolving performance issues


* Documents and reports issues discovered when testing and evaluating the systems


* Communicates project status and concerns to management in a timely manner


* Provides guidance to less-experienced staff members.


* Runs AI and HPC benchmarks.

Education and Experience Required:



* Master's degree or PhD in Computer Science, Engineering, Information Technology or Systems, or relevant field.


* 5+ years of experience.

Knowledge and Skills:



* 5+ years of experience in Machine Learning/Artificial Intelligence and 5+ years of experience in HPC


* Experience running NCCL, HPL and AI benchmarks.


* Experience working with containers and distributed...




Share Job