Job Description - Reliability Engineer

Job Description:

  • Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g., alert enrichment, routing, correlation, and operational runbooks).
  • Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
  • Track reliability and availability across critical trading applications and their dependencies. Partner with users, development teams, and IT to pinpoint where service levels are degrading.
  • Triage incoming alerts, issues, and escalations—assessing impact, urgency, and ownership.
  • Determine when incident criteria are met, declare incidents, and act as Incident Commander.
  • Coordinate responders and stakeholders; keep incident calls focused on facts, mitigation, and recovery.
  • Maintain clear timelines, actions, and status updates throughout the incident lifecycle.
  • Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
  • Support post-incident review (PIR) follow-ups and recurring issue reviews.
  • Ensure smooth handovers across EMEA, AMER, and APAC using a single global model: one incident standard and one handover process.

Requirements:

  • Experience in production operations, SRE, NOC/command center, trading operations, or a comparable first-line technical role—ideally in a trading, financial services, or other latency-sensitive environment.
  • Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
  • Clear communication (verbal and written): status updates are understandable to both traders and engineers.
  • Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
  • Solid Linux and networking fundamentals, plus the ability to quickly interpret alerts, logs, dashboards, and symptoms.
  • Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data).
  • Familiarity with incident and observability tooling (e.g., PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search).
  • Scripting/automation skills (Python preferred; Bash and Go are a plus), applied to triage, enrichment, routing, and correlation (not product code).
  • Exposure to containerized/cloud-hosted production environments (Kubernetes, Docker, GCP) is a plus.
Original job Reliability Engineer posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar Reliability Engineer Jobs in Hong Kong

GrabJobs is the no1 job portal in Hong Kong, connecting you to thousands of jobs fast! Find the best jobs in Hong Kong, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.