• Lead
the design and implementation of chaos, failure, and performance test
frameworks targeting distributed backend systems and infrastructure.
• Build
and maintain CI/CD pipelines in Jenkins using Groovy -based DSL and
scripted/declarative pipeline definitions.
• Own
infrastructure -level test environments using Kubernetes — spinning up, tearing
down, and managing test clusters for reliability experiments.
• Execute
and automate chaos engineering scenarios: network partitions, node failures,
pod evictions, resource exhaustion, and latency injection.
• Design
and run performance and load tests to identify bottlenecks, regressions, and
capacity limits across services.
• Develop
failure testing strategies to validate system behavior under degraded
conditions — partial failures, cascading failures, and data corruption
scenarios.
• Define
quality metrics and SLOs for infrastructure reliability; report on test
coverage and failure patterns to engineering leadership.
• Mentor
and guide junior QA engineers; champion a reliability -first quality culture
across the engineering org.
• Strong
programming skills in Python for test automation, tooling, and scripting.
• Proficiency
in Groovy, particularly for Jenkins pipeline development (scripted and
declarative).
• Extensive
hands -on Jenkins experience — pipeline authoring, plugin management, shared
libraries, and CI/CD architecture.
• Deep
Kubernetes knowledge — deploying workloads, managing namespaces, configuring
resource limits, and operating clusters for test environments.
• Proven
experience in chaos engineering: fault injection, resilience testing, and tools
such as Chaos Monkey, Litmus, or Gremlin.
• Hands -on
performance testing experience: load generation, profiling, bottleneck
analysis, and tooling (e.g., Locust, k6, Gatling, JMeter).
• Experience
designing failure test scenarios for distributed systems — covering partial
outages, dependency failures, and data -path degradation.
• Strong
understanding of Linux, networking, and infrastructure fundamentals.
• Experience
with service meshes (Istio, Linkerd) for traffic shaping and fault injection.
• Familiarity
with observability tooling: Prometheus, Grafana, Jaeger, or OpenTelemetry for
validating test outcomes.
• Background
in SRE practices: SLOs, error budgets, incident post -mortems.
• Experience
with cloud infrastructure testing on AWS, GCP, or Azure at the platform level.
• Exposure
to eBPF -based tools for low -level system observability during failure tests.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.