Requirements
â Architect and implement highly reliable, scalable, and cost-effective infrastructure
solutions for mission-critical applications across multi-cloud environments (AWS and
Azure).
â Lead the definition and refinement of service level objectives (SLOs), service level
indicators (SLIs), and error budgets, establishing reliability standards across the
organization.
â Design and implement sophisticated Infrastructure as Code (IaC) solutions using
Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
â Drive automation strategies to eliminate toil, improve operational efficiency, and enable
self-service capabilities for development teams.
â Lead incident response efforts, conduct thorough post-incident reviews, and implement
systemic improvements to prevent recurrence.
â Champion cloud-native architectures and modern reliability practices, serving as a
technical advisor for infrastructure and platform decisions.
â Participate in and help optimize the on-call rotation, ensuring sustainable practices and
effective escalation procedures.
â Establish and maintain comprehensive documentation standards, runbooks, and
knowledge repositories that enable team autonomy and effective incident response.
â Design and implement advanced monitoring, logging, and alerting strategies using
observability platforms to enable proactive issue detection and resolution.
â Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement
sophisticated deployment strategies including blue-green, canary, and progressive
delivery patterns.
â Ensure security, compliance, and governance standards are embedded throughout the
infrastructure lifecycle, implementing security-as-code practices.
â Drive capacity planning, performance optimization, and cost management initiatives
across cloud platforms.
â Collaborate with architecture and security teams to establish platform standards,
reference architectures, and best practices.
Skills, Knowledge and Expertise
â 5+ years of proven experience as a Site Reliability Engineer or similar role, with
demonstrated expertise in designing, implementing, and operating large-scale,
distributed systems.
â Deep expertise in Infrastructure as Code (IaC) with Terraform and Ansible, including
module development, state management, and multi-environment orchestration.
â Extensive hands-on experience with both AWS and Azure cloud platforms, including
advanced services, networking, and security features in both environments.
â Expert-level knowledge of container orchestration with Kubernetes, including
architecture, custom resource definitions (CRDs), operators, service mesh
implementations, and production-scale cluster management.
â Advanced proficiency in Linux system administration, performance tuning, and
troubleshooting complex system-level issues.
â Proven experience implementing GitOps workflows using ArgoCD, Flux, or similar tools,
including advanced deployment patterns and progressive delivery.
â Deep understanding of observability principles and hands-on experience with tools such
as Prometheus, Grafana, Datadog, Azure Monitor, or the ELK stack.
â Expert knowledge of networking concepts, including load balancing, CDNs, DNS, VPNs,
service mesh architectures, and distributed systems communication patterns.
â Strong programming and scripting capabilities in Python, Bash, Go, or PowerShell, with
the ability to develop custom tooling and automation frameworks.
â Extensive experience designing and optimizing CI/CD pipelines using Jenkins, GitLab
CI, Azure DevOps, GitHub Actions, or CircleCI.
â Demonstrated ability to lead incident response, conduct root cause analysis, and drive
systemic reliability improvements.
â Excellent communication and leadership skills with proven ability to influence technical
decisions and collaborate with stakeholders at all levels.
â Current certification in AWS (Solutions Architect Associate/Professional or equivalent)
and Azure (Azure Administrator or Azure Solutions Architect), with practical experience
managing production workloads on both platforms.
Good To Have
â Experience with hybrid and multi-cloud networking strategies, including ExpressRoute,
Direct Connect, and cloud interconnects.
â Knowledge of serverless architectures on AWS (Lambda) and Azure (Functions, Logic
Apps) and their operational considerations.
â Proven experience with disaster recovery planning, business continuity, and
implementing multi-region active-active architectures.
â Understanding of machine learning operations (MLOps), data pipeline orchestration, and
supporting ML workloads in production.
â Experience with service mesh technologies such as Istio, Linkerd, or Consul.
â Familiarity with chaos engineering principles and tools like Chaos Monkey or Gremlin.
â Experience with configuration management at scale and policy-as-code tools like Open
Policy Agent (OPA).
â Knowledge of FinOps principles and cloud cost optimization strategies.