Site Reliability Engineer – AI Systems & Platforms
Location: EU Timeline or EST Timeline
Department: Site Reliability Engineering / Platform Engineering
Type: Full-Time
Overview
We are seeking a hands-on Site Reliability / Platform Engineer specializing in AI Systems to help build, secure, and scale our internal AI platforms. In this role, you will design shared platform capabilities, enforce robust security guardrails around LLM usage, integrate enterprise data, and manage AI-related infrastructure and costs. You do not need to be a data scientist, but you must bring strong cloud/systems engineering fundamentals and a solid understanding of how modern AI applications operate at scale.
Key Responsibilities
AI Platform & Infrastructure Engineering
Build and operate reusable platform capabilities for LLM access, prompt/agent execution, RAG architectures, and AI gateways.
Provision and manage supporting cloud infrastructure (AWS), Kubernetes clusters, and compute/storage environments using IaC (Terraform) and GitOps.
Provide self-service tools that allow internal teams to adopt AI safely without duplicating infrastructure.
AI Security, Governance & Administration
Implement technical security controls including IAM, RBAC, encryption, secret management, and network isolation for AI workloads.
Build guardrails against AI-specific risks such as prompt injection, unauthorized data training, data leakage, and excessive permissions.
Administer enterprise AI tools (e.g., Claude Code, Claude Apps, Gemini) including SSO integration, license management, and usage policy configuration.
Data Enablement & Observability
Design secure patterns for connecting AI tools to databases, vector stores, data lakes, and internal knowledge platforms while preserving permissions.
Implement end-to-end monitoring, logging, tracing, and token/API cost tracking using tools like Prometheus, Grafana, and Elastic.
Standardization & Cost Optimization
Define organization-wide standards, reusable code libraries, and architecture templates for AI deployment.
Track and optimize expenditures across LLM APIs, GPU/CPU compute, and commercial licenses through quotas, rate limits, and usage reporting.
Required Qualifications
Platform / SRE Experience: 5+ years of experience in SRE, Platform Engineering, DevOps, or Cloud Systems Engineering.
Cloud & Containers: Strong hands-on experience with AWS and production Kubernetes management.
IaC & Scripting: Proficiency with Terraform/Ansible and automation scripting in Python, Go, or Bash.
Systems & Ops: Strong Linux engineering skills, CI/CD pipelines, GitOps workflows, and observability stacks (Prometheus, Grafana, Loki/Elastic).
Security Fundamentals: Deep knowledge of IAM, RBAC, secrets management, encryption, and audit logging.
AI System Fundamentals: Working understanding of LLM deployment options (commercial APIs, AWS Bedrock, self-hosted), AI security risks (prompt injection, data leakage), AI gateways, vector databases, and token/cost management.
AI Tooling: Practical experience integrating AI coding tools into daily development workflows with full ownership of code quality.
Advantageous Experience
Background in telecommunications or highly distributed, security-sensitive environments.
Hands-on experience with LLMOps, MLOps, vector databases, or data platforms.
Exposure to operating GPU workloads, policy-as-code, or recognized AI governance frameworks.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.