Data Reliability Engineer
Key Responsibilities:
• Own and improve the reliability, availability, observability, and operational supportability of Maximus UK's Azure Databricks platform, Azure data services, and associated data pipelines and data products.
• Design and implement monitoring, alerting, health checks, and diagnostics across Azure Databricks, Azure data services, orchestration layers, storage, and downstream consumption, extending these patterns into AWS as the estate grows.
• Define and maintain reliability standards, controls, operational runbooks, and support models that improve the resilience, predictability, and supportability of data services.
• Work closely with data engineering teams to identify, prioritise, and remediate reliability, performance, and data quality issues across Databricks notebooks, jobs, workflows, and other Azure data workloads.
• Establish proactive incident detection, triage, and root cause analysis practices, reducing mean time to detect and mean time to recover for data-related issues.
• Design and implement robust data quality controls, validation frameworks, reconciliation processes, and anomaly detection approaches across the end-to-end data lifecycle.
• Configure and use Azure Purview to provide effective data cataloguing, lineage, ownership, and governance, ensuring reliability and quality controls are visible and auditable.
• Collaborate with platform, cloud, architecture, and security teams to ensure the data estate is secure, resilient, cost-effective, and aligned to enterprise standards and patterns.
• Contribute to the reliability engineering approach for an Azure-first data platform while supporting reusable patterns and operational readiness for data services in AWS.
• Partner with architects and engineers so that new pipelines, data products, and platform services are designed with operability, recoverability, scalability, and observability built in from the start.
• Automate repetitive operational tasks, environment checks, dependency verification, failure handling, and recovery processes to increase efficiency and reduce manual intervention and risk.
• Capture lessons learned, codify reliability patterns and standards, and share best practice to continuously improve reliability, transparency, and engineering discipline across the data function.
Essential Skills - What You'll Bring
• Proven experience in data engineering, platform engineering, site reliability engineering, DataOps, or a closely related role focused on data platform reliability and operations.
• Strong hands-on experience with Azure-based data platforms, particularly Azure Databricks and core Azure data services such as Data Lake Storage, Data Factory/Synapse, and analytical stores, with familiarity of equivalent services in AWS.
• Strong understanding of modern data platform architectures, including data lakes, warehouses or lakehouses, orchestration frameworks, transformation pipelines, streaming ser...
- Rate: Not Specified
- Location: St. George, US-UT
- Type: Permanent
- Industry: Finance
- Recruiter: Maximus
- Contact: Not Specified
- Email: to view click here
- Reference: 40044_UT_Salt Lake City
- Posted: 2026-07-15 10:39:55 -
- View all Jobs from Maximus
More Jobs from Maximus
- Purchasing Assistant
- RN Field Coordinator
- Software Engineer III - Test Automation
- Teacher Opportunities
- Weekend Only - Floor Technician
- Legal Assistant
- Direct Support Professional (DSP) - Sunday - Tuesday 8:00 AM - 8:00 PM
- Direct Support Professional - Debra Drive (Wed, Thur, Sat: 8 am - 8 pm)
- Direct Support Professional - (Sun, Mon, Tues: 8 am - 8 pm)
- Direct Support Professional (Thur, Fri, Sat: 8 am - 10 pm)
- Agent
- Direct Support Professional (Sat & Sun: 3 pm - 11 pm)
- Direct Support Professional (Tues, Thur, Fri, Sat : 7 am - 7 pm)
- Grounds Maintenance Laborer (M-F 6:30am-3:00pm / $2,500 Sign-on Bonus)
- Manager, Safety & Training
- Housekeeper
- Copy Of AI Implementation & Project Manager
- Copy Of Area Manager
- Cashier / Receptionist - Full Time
- HR Intern