Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g., alert enrichment, routing, correlation, and operational runbooks).
Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
Track reliability and availability across critical trading applications and their dependencies. Partner with users, development teams, and IT to pinpoint where service levels are degrading.
Triage incoming alerts, issues, and escalations—assessing impact, urgency, and ownership.
Determine when incident criteria are met, declare incidents, and act as Incident Commander.
Coordinate responders and stakeholders; keep incident calls focused on facts, mitigation, and recovery.
Maintain clear timelines, actions, and status updates throughout the incident lifecycle.
Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
Support post-incident review (PIR) follow-ups and recurring issue reviews.
Ensure smooth handovers across EMEA, AMER, and APAC using a single global model: one incident standard and one handover process.
Requirements:
Experience in production operations, SRE, NOC/command center, trading operations, or a comparable first-line technical role—ideally in a trading, financial services, or other latency-sensitive environment.
Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
Clear communication (verbal and written): status updates are understandable to both traders and engineers.
Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
Solid Linux and networking fundamentals, plus the ability to quickly interpret alerts, logs, dashboards, and symptoms.
Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data).
Familiarity with incident and observability tooling (e.g., PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search).
Scripting/automation skills (Python preferred; Bash and Go are a plus), applied to triage, enrichment, routing, and correlation (not product code).
Exposure to containerized/cloud-hosted production environments (Kubernetes, Docker, GCP) is a plus.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in Hong Kong.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in Hong Kong, connecting you to thousands of jobs fast!
Find the best jobs in Hong Kong, apply in 1 click and get a job today!