The opportunity
NPCI is looking for a highly skilled DevOps Engineer specializing in automation, Kubernetes, and system reliability engineering, with a strong focus on AI-driven testing and resilience engineering. In this role, you will architect and execute automated reliability testing frameworks, simulate real-world failure scenarios, and ensure high availability, fault tolerance, and scalability of mission-critical payment systems operating at national scale.
This is a high-impact role where you'll work at the intersection of:
- DevOps + SRE + Chaos Engineering + AI-driven testing
- Ensuring platforms are failure-ready, not just failure-resistant
Job details
- Job Title: DevOps Engineer – Reliability & Automation
- Division: Quality Control & Monitoring Systems
- Experience: 3–7 Years
- Employment Type: Full-time
- Location: Hyderabad
- Education: Bachelor’s degree in Computer Science / Engineering or related field
- Preferred Certifications: CKA / CKAD / Cloud DevOps (AWS/Azure/GCP)
- Work Mode: 5 Days Work From Office (WFO)
Key Responsibilities
DevOps & Platform Engineering
• Design, implement, and manage CI/CD pipelines for application delivery and infrastructure provisioning
• Containerize applications using Docker and orchestrate via Kubernetes
• Deploy, manage, and optimize production-grade Kubernetes clusters
Reliability & Resilience Engineering
• Design and implement automated resilience and reliability testing frameworks
• Develop and execute failure simulation strategies (node failures, latency, network partitioning, etc.)
• Build and run fault injection mechanisms for controlled chaos testing
• Identify system weaknesses and improve fault tolerance & recovery strategies
Monitoring, Observability & Incident Readiness
• Implement monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK/EFK, Open Telemetry)
• Define and track SLIs, SLOs, and error budgets
• Enhance observability for distributed systems
Infrastructure, Networking & Security
• Manage Linux systems and troubleshoot performance and reliability issues
• Configure networking components including:
➤ Load balancing
➤ DNS
➤ Firewalls & VPN
• Ensure security, compliance, and best practices across infrastructure
Collaboration & Productivity
• Work closely with SRE, DevOps, QA, and development teams
• Improve developer productivity through automation and self-service platforms
• Contribute to platform reliability culture and engineering excellence