About the Role
We're looking for an AI Platform Engineer to design, deploy, and operate enterprise-grade AI platforms — not to build models, but to make AI work in production.
You'll be responsible for:
Deploying and managing LLM hosting, inference services, and embedding services
Running AI agent orchestration platforms in a regulated on-premises environment
Managing GPU infrastructure and optimizing model performance (quantization, vLLM, batching)
Building reusable APIs, SDKs, and shared AI services that other teams can plug into
Keeping AI systems secure, compliant, and observable
This is a hands-on engineering role. You'll be in the terminal, on Kubernetes, tuning inference pipelines, and shipping production services — not writing research papers or building dashboards.
Who We're Looking For
You must have:
2+ years of hands-on experience deploying and managing AI/LLM platforms in production
Strong Kubernetes skills — you've run AI workloads on K8s, not just deployed a single container
Experience with GPU infrastructure — managing GPU nodes, scheduling, monitoring
LLM deployment experience — inference serving, model hosting, prompt pipelines
Software engineering fundamentals — you write clean APIs, use CI/CD, and containerise your work
Experience with vector databases and RAG architectures
Nice to have:
Experience with agentic AI frameworks (LangChain, LlamaIndex, multi-agent systems)
Model optimisation techniques (quantization, vLLM, batching)
Experience working in regulated or on-premise environments (gov, defence, finance)
Background in Computer Science, Computer Engineering, or related fields
BGC GROUP PTE. LTD.
BGC Group is an international recruitment and manpower outsourcing firm that identifies and delivers human capital solutions that imperative to every successful company’s growth. Having helped 25,000 individuals quickly land rewarding careers in companies that drive industries since our inception...
Read more about the companyCopyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.