Logo-of-JPMorgan-Chase-hiring-for-jobs-in-India-on-GrabJobs

Site Reliability Engineer II - Incident Management

Job Description - Site Reliability Engineer II - Incident Management

Description

Join a team where your expertise in site reliability engineering shapes the future of AI/ML data platforms. Grow your skills and impact by collaborating with innovators and driving strategic change.

As a Site Reliability Engineer II at JPMorgan Chase within the AI/ML Data Platforms team, you will build scalable and resilient data solutions. You will engage in root cause analysis, production changes, and operational challenges. Your experience will help mentor team members and foster strategic change. You will partner with colleagues across the global network to deliver market-leading solutions.

Job responsibilities

  • Build and support applications using technologies such as Databricks, Snowflake, AWS, and Kubernetes
  • Coordinate incident management coverage to resolve application issues
  • Collaborate with cross-functional teams to perform root cause analysis and implement production changes
  • Develop and support AI/ML solutions for troubleshooting and incident resolution
  • Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes.

 

Required qualifications, capabilities and skills

  • Formal training or certification on security engineering concepts and 2+ years applied experience
  • Proficient in site reliability culture and principles and able to implement site reliability within an application or platform
  • Proficiency in running production incident calls and managing incident resolution
  • Experience in observability including monitoring, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk
  • Strong understanding of SLI/SLO/SLA and Error Budgets
  • Proficiency in Python or PySpark for AI/ML modeling
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows (e.g., troubleshooting support and runbook drafting) with strong validation habits and awareness of data sensitivity. 
  • Ability to assess AI-assisted operational recommendations for correctness and risk, and apply appropriate controls to maintain resiliency, security, and auditability.
  • Ability to reduce toil by building tools to automate repeated tasks
  • Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery
  • Awareness of risk controls and compliance with departmental and company-wide standards

 

Preferred qualifications, capabilities and skills

  • 4+ years in an SRE or production support role with AWS Cloud, Databricks, Snowflake or similar technologies
  • AWS and Databricks certifications


Original job Site Reliability Engineer II - Incident Management posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Site Reliability Engineer II Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.