Role purpose:
We’re looking for a DevOps Support Engineer to provide hands-on operational support for cloud-native
and containerised solutions running across modern cloud, Docker, and Kubernetes environments.
This role is primarily focused on troubleshooting issues, monitoring platform health, supporting live
services, and helping ensure production systems remain stable, observable, and reliable.
You’re someone who enjoys solving operational problems, investigating alerts, and keeping services
stable. You are comfortable working through unclear issues methodically, including Docker image,
container, deployment, and Kubernetes workload problems, using monitoring and diagnostic data to
narrow down causes. You communicate clearly with engineers and stakeholders during support
activities and value good runbooks, accurate documentation, and practical improvements that
reduce repeated incidents.
You’ll work closely with engineering, platform, and support teams to investigate incidents, respond to
alerts, analyse logs and metrics, support Docker and Kubernetes deployments, maintain runbooks,
and escalate issues where deeper engineering input is required.
Main duties and key responsibilities:
Production Support & Troubleshooting
• Troubleshoot issues across cloud infrastructure, Docker containers, Kubernetes workloads, CI/CD
processes, networking, and platform services
• Investigate production alerts, service degradation, failed Docker/Kubernetes deployments, and
environment issues
• Use logs, metrics, traces, dashboards, and system evidence to identify likely causes and support
resolution
• Escalate complex issues to engineering or platform teams with clear analysis, evidence, and impact
summary
• Track recurring issues and contribute to operational improvements, runbooks, and knowledge articles
Monitoring & Platform Health
• Monitor platform health using Prometheus, Grafana, VictoriaMetrics, Elasticsearch / Kibana,
Sentry, and OpenTelemetry
• Review alerts, dashboards, logs, container events, and telemetry to detect service issues and
infrastructure risks
• Support alert triage, noise reduction, and monitoring coverage improvements for Docker and
Kubernetes workloads
• Validate that services, containers, deployments, jobs, and platform components are operating
as expected
Incident Response & Service Recovery
• Support incident response activities for production and non-production environments
• Assist with initial triage, impact assessment, workaround identification, and service recovery
• Capture incident timelines, evidence, symptoms, and actions taken during support activities
• Contribute to post-incident reviews by identifying recurring patterns, gaps in monitoring, and
opportunities to improve support processes
Deployment, Pipeline & Environment Support
• Support GitHub CI and Jenkins pipeline issues, failed jobs, deployment errors, Docker image
build failures, and release workflow problems
• Help troubleshoot Docker, Helm, ArgoCD, Kubernetes, and environment configuration issues
• Support infrastructure provisioning and configuration issues involving Terraform, Ansible, and
Packer
• Assist engineering teams with Docker deployment validation, container runtime checks,
rollback support, and environment readiness checks
Operational Maintenance & Continuous Improvement
• Maintain and improve runbooks, support procedures, Docker deployment troubleshooting
guides, and operational documentation
• Identify repeated manual support activities that could be automated or simplified
• Support routine maintenance, health checks, capacity checks, container image checks, and
service validation activities
• Help improve reliability, observability, and supportability of cloud, Docker, and Kubernetes
based solutions
Stakeholder & Engineering Support
• Provide technical support to engineering and operations teams for platform, Docker
deployment, container, and environment issues
• Communicate clearly during incidents, handovers, escalations, and support updates
• Work with engineers to validate fixes, confirm service recovery, and close support actions
• Maintain knowledge articles and known-issue documentation to improve future support
efficiency
Required Skills & Experience
• Experience troubleshooting production or near-production cloud-based and containerised
systems
• Hands-on experience supporting Docker containers, Docker images, and Kubernetes
workloads, ideally on EKS
• Working knowledge of AWS infrastructure and common operational support patterns
• Ability to investigate issues using logs, metrics, traces, dashboards, and alerting tools
• Experience supporting CI/CD pipelines using GitHub CI and/or Jenkins, including Docker
image build and deployment issues
• Working knowledge of Terraform, Ansible, Packer, Docker, Helm, and ArgoCD from a support
and troubleshooting perspective
• Familiarity with observability tools such as Prometheus, Grafana, VictoriaMetrics,
Elasticsearch / Kibana, Sentry, and OpenTelemetry
• Understanding of Linux, networking, VPNs, service isolation, container networking, and Zero
Trust concepts
• Ability to document incidents, known issues, Docker deployment procedures, runbooks, and
support procedures clearly
Desirable Experience
• Experience working in an application, infrastructure, platform, or DevOps support
environment
• Exposure to incident management, service recovery, or operational support processes
• Experience supporting stateful services such as Postgres and Kafka
• Familiarity with load testing outputs, performance symptoms, and capacity-related issues
• Experience maintaining Docker registries, Docker image repositories, artifact repositories, or
internal developer tooling
• Exposure to regulated environments, compliance checks, or audit-support activities
• Ability to identify recurring support issues and suggest practical improvements
Example Tech Stack Exposure
Typical technologies in this environment include:
AWS · Docker · Kubernetes · Terraform · Ansible · Packer · GitHub CI · Jenkins · ArgoCD · Helm ·
Prometheus · Grafana · VictoriaMetrics · Elasticsearch · Kafka · Postgres · OpenTelemetry · Zero Trust
Networking
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.