A

Site Reliability Engineer I

Job Description - Site Reliability Engineer I

Description

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.



Responsibilities
  • Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
  • Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
  • Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
  • Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
  • Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
  • Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
  • Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
  • Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items
  • Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements
  • Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services
  • Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
  • Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation
  • Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic
  • Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
  • Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
  • Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
  • Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
  • Participate and lead the Development change review and change validation processes
  • Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
  • Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers


Qualifications

Education Qualifications:

  • Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
  • Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
  • Strong knowledge of operating systems and application runtimes such as Java and .NET
  • Knowledge of distributed systems and service‑based architectures from an operations and reliability perspective
  • Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
  • Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
  • Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
  • Knowledge of scripting and automation using languages such as PowerShell and Python
  • Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus

 

Work Experience:

  • Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
  • Experience supporting production systems in large‑scale enterprise environments with a focus on reliability and availability
  • Experience in system administration, infrastructure operations, and network troubleshooting
  • Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
  • Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
  • Knowledge of cloud‑based Site Reliability Engineering (SRE) practices with hands‑on experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
  • Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservices‑based architectures
  • Experience using enterprise monitoring and alerting platforms such as ELF
  • Exposure to AI‑assisted monitoring, automation, or AIOps tools is a plus
    • Proficiency in connecting to and administering servers via SSH (Secure Shell)
  • Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access

 

Licenses & Certifications

  • Certification in at least one programming language or runtime such as Java, .NET, or Python
  • Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
  • Public cloud certification in AWS or GCP is a plus
  • Certification or training related to AI platforms, analytics platforms, or AIOps is a plus

 

Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions. 



Original job Site Reliability Engineer I posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Site Reliability Engineer I Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.