What You Do:
- Own the availability, performance, and incident response for Rapid's production EKS clusters
- Design and operate the full observability stack — metrics, logs, traces — with
Open Telemetry as the foundation - Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
- Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
- Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
- Tune autoscaling, networking, and cost efficiency across AWS workloads
- On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
- 10+ years in SRE, DevOps, or infrastructure engineering roles
- Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
surrounding ecosystem - Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
- Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
- Strong foundation in Linux, networking, and distributed systems fundamentals
- Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out - Startup mindset: you make decisions with incomplete information and iterate quickly