Purpose of The Position:
The Director, Site Reliability Engineering – Incident Management owns the enterprise Reliability practice across Byte, KFC, and Taco Bell digital platforms, including its strategy, standards, governance, and delivery. Incident Management is the most visible part of that practice and sits with this role exclusively. This leader also owns the reliability platform, meaning the tooling, products, and capabilities that make reliability real for engineers and for markets, and serves as the accountable face to brands and markets for reliability process implementation, reporting, and the operational relationship. This leader establishes the strategy, governance, and operational excellence required to ensure highly available, resilient, and customer-centric technology services while developing high-performing teams and partnering across Engineering, Product, Infrastructure, and Brand leadership to continuously improve reliability and business continuity. This leader operates with exceptional diligence and care, bridging deep technical detail with human and business context so that executives, engineers, and restaurant teams experience clear, calm, and trustworthy communication in the moments that matter most.
Scope and Magnitude:
~30+ Teams – Byte, KFC and Taco Bell
Global Team - Vietnam, India, Colombia and US
Platform Team – Responsible for all products (Edge, Commerce, POS, KDS, Menu, Portal, etc.)
Position Functions:
Strategic Leadership
Provide strategic leadership for Incident Management and Technical Operations across Byte, KFC, and Taco Bell digital platforms, ensuring high availability, operational excellence, and a consistent customer experience.
Establish the vision, governance, and operating model for enterprise Incident Management, driving standardized processes, tooling, and best practices across multiple brands and technology organizations.
Drive enterprise operational readiness for major product launches, restaurant initiatives, promotions, and seasonal events through effective change management, risk assessment, and cross-functional planning.
Own the enterprise framework and targets for service level objectives and indicators, set in partnership with the engineering teams that own the services, which remain accountable for achieving them. Define and monitor MTTR, incident trends, service health, operational maturity, and executive KPIs, using data to prioritize investments and improve platform resilience.
Own the enterprise Incident Management governance framework, ensuring consistent execution, accountability, and continuous improvement across all brands.
Own enterprise observability strategy, setting the standard and direction for how platform health is measured and seen, with Platform Engineering partnering on the underlying platform and instrumentation.
Own the strategy and roadmap for the reliability platform, including the tooling, products, and self-service capabilities that deliver reliability to engineering teams and to markets.
Provide executive-level communications and operational updates to Digital & Technology leadership, Brand CDTOs/CTOs, and executive stakeholders during major incidents and through regular operational business reviews.
Develop trusted partnerships with executive leaders across Product, Engineering, Infrastructure, Security, Restaurant Operations, and Brand Technology to align operational priorities with business objectives and customer experience goals.
Jointly own Business Continuity and Disaster Recovery with Platform Engineering, covering recovery strategies, resiliency testing, crisis management processes, and operational preparedness across Byte, KFC, and Taco Bell. Decisions are made in a standing joint review, with anything unresolved escalating to the Senior Director, Engineering within one cycle. During a declared continuity event or major incident, the Incident Commander decides in the moment.
Sponsor continuous improvement initiatives that enhance operational resilience, reduce enterprise risk, and strengthen the organization's ability to respond to large-scale operational events.
Technical Leadership
Serve as executive Incident Commander during critical enterprise incidents, providing leadership during major outages while ensuring effective cross-functional coordination across Engineering, Product, Infrastructure, Security, Restaurant Operations, Brand Leadership, and external partners.
Partner with Engineering, Platform, Security, and Architecture leaders to influence technology strategy, improve platform reliability, reduce operational toil, and accelerate automation.
Own the incident management and operational resilience strategy, setting the direction for detection, response, recovery, and preparedness across highly available digital platforms.
Influence architecture and platform design decisions to strengthen resiliency, reduce failure points, and improve the reliability of customer-facing services.
Set enterprise observability standards and direction, partnering with Platform Engineering on the platform and instrumentation that deliver against them.
Own detection, response, and remediation automation, partnering with Platform Engineering, which owns provisioning, deployment, and self-service automation.
Partner with Cloud, Platform, Engineering, Security, and Architecture leaders on cloud and platform engineering decisions that improve scalability, reliability, and operational resilience.
Champion a culture of operational excellence by driving post-incident reviews, root cause analysis, corrective actions, and continuous learning across engineering organizations.
Manage relationships with key technology vendors and managed service providers, ensuring operational performance, accountability, and adherence to service level commitments.
Lead cross-brand initiatives to mature observability, incident response automation, disaster recovery readiness, and operational resilience capabilities.
People Leadership
Lead, coach, and develop a high-performing organization of managers and technical leaders responsible for Incident Management and Technical Operations, building organizational capability, succession plans, and a culture of operational excellence and accountability.
Provide leadership for globally distributed Incident Management teams, ensuring consistent operational standards, clear ownership, and effective collaboration across regions and time zones.
Working Relationships:
Internal
DTLT
Brand CDTOs/CTOs
Engineering and Product Teams
Reliability and Platform Engineering (GRE)
Security
External
Vendors - managed service providers and technology partners (incident tooling, observability, cloud), with accountability for operational performance and SLA adherence
Brand market teams
Specialized or Technical Knowledge/Skills
Preferred Requirements:
Success Metrics:
Salary Range: $164,500 - $193,600 annually + bonus eligibility and stock-based compensation. This is the expected salary range for this position. Ultimately, in determining pay, we'll consider the successful candidate’s location, experience, and other job-related factors.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.