Z

Observability Engineer Grafana & Prometheus

Job Description - Observability Engineer Grafana & Prometheus

Description

This role supports the implementation of the AWS-native observability platform as part of the integration modernization program. The engineer will work on configuring dashboards, metrics collection, and alerting frameworks using Grafana, Prometheus, and AWS monitoring tools.
The role focuses on enabling real-time monitoring and operational visibility for APIs, messaging, and integration flows. The engineer will collaborate with platform and engineering teams to ensure system health, identify issues proactively, and support troubleshooting using observability data. The position plays a key role in maintaining reliability and performance across the integration ecosystem.



Responsibilities

Observability Platform Engineering

  • Set up and manage Prometheus (metrics collection) and Grafana dashboards (AWS Managed Services)
  • Create dashboards for system health, performance KPIs, and integration monitoring
  • Maintain and enhance dashboards for APIs, messaging, and infrastructure components
     

Monitoring & Dashboards

  • Monitor application and infrastructure performance across AWS environments
  • Track key metrics such as latency, error rates, throughput, and availability
  • Support performance tuning through observability insights
     

Alerting & Reliability

  • Configure alerts, thresholds, and proactive monitoring frameworks
  • Enable AI-driven alerting and anomaly detection integrated with ITSM tools
  • Support high availability and reliability through proactive observability practices
     

Troubleshooting & Operations

  • Support troubleshooting using metrics, logs, traces, and dashboards
  • Provide deep visibility into payload tracking, error handling, and integration health
  • Enable root cause analysis and faster resolution through observability insights
     

Standardization & Best Practices

  • Establish monitoring standards, observability patterns, and dashboard blueprints
  • Implement structured logging, correlation IDs, and traceability across distributed systems
  • Contribute to observability maturity including automated alerts, runbooks, and analytics


Qualifications

Required Skills & Competencies

  • Strong experience with Grafana and Prometheus
  • Experience setting up monitoring dashboards, alerts, and observability pipelines
  • Basic knowledge of AWS cloud services (CloudWatch, X-Ray, etc.)
  • Solid understanding of monitoring, logging, and alerting concepts
  • Good troubleshooting and analytical skills
     

Good to Have

  • Knowledge of OpenTelemetry (OTEL) instrumentation
  • Exposure to Kubernetes/EKS monitoring
  • Familiarity with integration monitoring (APIs, messaging, batch flows)
  • Understanding of AI-driven observability and anomaly detection concepts
  • Experience in enterprise integration and cloud modernization programs preferred
  • Understanding of logging and tracing tools
     

Qualifications

  • Bachelor’s degree in Engineering or related field
  • 6–10 years of experience

 



Original job Observability Engineer Grafana & Prometheus posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Observability Engineer Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.