Z

Sr. Observability Engineer Grafana & Prometheus

Job Description - Sr. Observability Engineer Grafana & Prometheus

Description

This role is part of the large-scale integration modernization program, focused on migrating WebMethods-based integrations to AWS-native architecture. The engineer will design and implement a centralized observability platform leveraging AWS Managed Grafana, Prometheus, CloudWatch, and OpenTelemetry across APIs, messaging, MFT, and SAP integrations.
The engineer will contribute to building unified dashboards, SLA monitoring, and AI-driven alerting to provide real-time insights into system health, integration flows, and business KPIs. The role is critical to ensuring reliability, troubleshooting efficiency, and operational intelligence across ~100M monthly transactions and ~1,500 flows.



Responsibilities

Observability Platform Engineering

  • Set up and manage Prometheus (metrics collection) and Grafana dashboards (AWS Managed Services)
  • Build and maintain AWS-native observability platform including metrics, logs, and traces using CloudWatch, X-Ray, and OTEL instrumentation
  • Implement instrumentation across APIs, messaging, MFT, and batch integrations for end-to-end visibility
     

Monitoring & Dashboards

  • Monitor application and infrastructure performance across AWS integration estate
  • Create dashboards for KPIs, system health, SLA tracking, and integration performance
  • Develop master dashboards covering API performance, messaging queues, infra health, and integration status 
     

Alerting & Reliability

  • Configure alerts, thresholds, and proactive monitoring frameworks
  • Enable AI-driven alerting and anomaly detection integrated with ITSM tools
  • Support high availability and reliability through proactive observability practices
     

Troubleshooting & Operations

  • Support troubleshooting using metrics, logs, traces, and dashboards
  • Provide deep visibility into payload tracking, error handling, and integration health
  • Enable root cause analysis and faster resolution through observability insights
     

Standardization & Best Practices

  • Establish monitoring standards, observability patterns, and dashboard blueprints
  • Implement structured logging, correlation IDs, and traceability across distributed systems
  • Contribute to observability maturity including automated alerts, runbooks, and analytics


Qualifications

Required Skills & Competencies

  • Strong experience with Grafana and Prometheus
  • Experience setting up monitoring dashboards, alerts, and observability pipelines
  • Basic knowledge of AWS cloud services (CloudWatch, X-Ray, etc.)
  • Solid understanding of monitoring, logging, and alerting concepts
  • Good troubleshooting and analytical skills
     

Good to Have

  • Knowledge of OpenTelemetry (OTEL) instrumentation
  • Exposure to Kubernetes/EKS monitoring
  • Familiarity with integration monitoring (APIs, messaging, batch flows)
  • Understanding of AI-driven observability and anomaly detection concepts
  • Experience in enterprise integration and cloud modernization programs preferred
     

Qualifications

  • Bachelor’s degree in Engineering or related field
  • 8–12 years of experience

 



Original job Sr. Observability Engineer Grafana & Prometheus posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Sr. Observability Engineer Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.