$9,000 - 10,000 monthly
Develop and roll out new observability and automation solutions, driving increased demand for development capacity for Site Reliability Engineering (AIOps and Agentic AI).
Responsibilities:
AIOps & Agentic AI Engineering
• Develop Python and/or LangGraph scripts and agentic AI workflows to automate operational tasks, diagnostics, and remediation.
• Build and integrate AIOps capabilities for predictive alerting, anomaly detection, and intelligent event correlation across infrastructure and applications.
• Design and operationalise agentic AI solutions that augment incident response and reduce manual intervention.
• Develop self-healing and auto-remediation workflows driven by AI/ML and event correlation insights.
Infrastructure Automation and CI/CD
• Develop infrastructure automation to provision, configure, and manage systems in a consistent, auditable manner.
• Manage automation and AI code through GitHub, applying version control, code review, and collaboration best practices.
• Integrate AIOps and automation components into CI/CD pipelines for continuous testing, validation, and deployment.
Observability, Reliability & Collaboration
• Implement observability and event correlation across metrics, logs, and traces to enable proactive reliability management.
• Champion SRE frameworks and reliability practices, embedding AIOps into the delivery lifecycle.
• Collaborate with Application, Infrastructure, Cloud, and Cyber teams to enable predictive incident prevention and root-cause transparency.
Maintain documentation and runbooks, and mentor engineers on AIOps and automation best practices.
Requirements:
• Possess a degree in Computer Science/Information Technology or related fields.
• Hands-on scripting experience with AIOps and Agentic AI - either Python or LangGraph (both languages will be a strong advantage).
• Strong experience in Python and/or LangGraph scripting experience applied to AIOps and Agentic AI.
• Proven experience in using GitHub and CI/CD pipelines.
• Possess infrastructure automation development experience.
• Possess observability and event correlation experience.
• A good team player with excellent written and verbal communication skills, and strong interpersonal skills to coordinate across diverse stakeholders and vendors.
• Certification in the following areas will be advantageous:
o Python Certified Expert / relevant AI/ML certification
o AWS Certified DevOps Engineer / Azure DevOps Expert
o GitHub Actions / CI/CD certification
o Site Reliability Engineering Foundation / Practitioner (DevOps Institute)
o ITIL v4 Foundation
NTT SINGAPORE PTE. LTD.
NTT Singapore Pte Ltd (NTTS) is the regional headquarters of NTT Communications Corporation (NTT Com) for Asia Pacific Region. Established in 1997, NTT Singapore has more than 10 years of expertise in providing information and communications technology (ICT) solutions worldwide. NTT Singapore...
Read more about the companyCopyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.