Why Work Here
The Role
You will own cloud infrastructure and operations for a platform that processes payroll and tax for hundreds of thousands of businesses. Downtime here is not an inconvenience; it means people don't get paid. You will lead the Cloud Operations team, own our Azure estate end to end, and run the CloudOps workstream of our platform modernization program.
This is a player-coach role for a leader who is equally comfortable in an incident bridge, a FinOps review, and an architecture discussion.
What You'll Own
Cloud infrastructure. Full ownership of our Azure environment: compute (Azure Container Apps, App Service, Functions), data (SQL Server, Cosmos DB), messaging (Azure Service Bus), networking, and identity. Infrastructure is managed as code with Terraform; you will hold the line on that discipline.
Modernization infrastructure. Lead the CloudOps pod supporting our monolith decomposition program: landing zones, container platforms, API management, service mesh patterns (Dapr), change data capture (Debezium), and the CI/CD and observability foundations (OpenTelemetry) new services depend on.
Observability platform transition. Own the migration of our observability stack from Application Insights to Grafana. This includes standing up the Grafana platform, migrating dashboards, alerting, and SLO reporting, driving OpenTelemetry-based instrumentation standards so telemetry stays portable rather than vendor-coupled, and managing the cutover so teams keep full visibility throughout. You will own the timeline and the cost case for the move.
Reliability and incident management. Own SLOs, error budget policy, and incident response. We run a structured observe-hypothesize-test incident methodology, blameless postmortems with tracked action items, and monitoring standards with defined detection SLAs. You will keep that culture healthy and raise the bar.
FinOps. Own the Azure budget, monthly actuals-vs-budget reporting, anomaly detection, and cost optimization. You will partner with engineering leadership on capacity planning and unit economics as the platform footprint grows.
Security and compliance partnership. Work closely with the CISO and CIO organizations on SOC 2, GDPR, encryption architecture (Azure Key Vault), device and access audits, and infrastructure controls for AI-assisted and agentic development workflows.
Team leadership. Lead, grow, and retain a team of cloud and site reliability engineers. Set standards, develop leaders, and keep the team focused on platform outcomes rather than ticket queues.
What We're Looking For
Nice to Have
Why This Role
You will not be keeping the lights on for a static estate. You will be building the operational foundation for a multi-year platform transformation with executive backing, real budget, and a team that takes reliability seriously. The work is visible from the board level down.
About isolved®
isolved is a leading provider of human capital management (HCM) solutions that combines modern technology with expert services and support. Purpose-built for People Heroes™, isolved gives HR, payroll and benefits leaders the tools and insights to streamline operations and deliver employee experiences that matter. isolved People Cloud™ is a connected HCM platform with built-in artificial intelligence (AI) and analytics that brings together HR, payroll, benefits, workforce management and talent management in one experience. Built on a legacy of 40 years in the market, isolved is trusted by more than 200,000 employers and used by 9 million U.S. employees, representing about one in 20 American workers. Visit www.isolvedhcm.com.
#LI-KJ1
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.