Stories have the power to make the ordinary extraordinary. They help us think bigger, see beyond ourselves, and deepen our connection with others.
As one of the world’s leading audiobook and e-book streaming platforms, Storytel brings unlimited listening to millions of users across 25 markets.
Driven by our purpose, “Leading the future of storytelling, we move the world through stories”, Storytel Group inspires and entertains people around the world by blending innovation with tradition. We bring stories to life across various formats for everyone to discover. Anytime. Anywhere.
Ready for your next chapter? We’re looking for a Senior ML/AI Platform Engineer to join our Platform Engineering team!
About the Team
The Platform Engineering team focuses on setting up, operating and redefining observability, infrastructure, CI/CD, runtime, Machine Learning and AI agentic tooling at Storytel. We apply software engineering principles to accelerate software delivery and ensure that application development teams are productive in all aspects of the software delivery lifecycle.
We aim to design and build toolchains and workflows covering the operational necessities of the entire lifecycle of an application and to enable that as self-service for our teams. This means providing paved roads that make it easy for our application developers to focus on building a great experience for our users. We support the teams in making the most of our platform’s capabilities and in providing best-practices knowledge.
Our team’s expertise is very diverse, varying from cloud infrastructure, through security, to development tools, ML platform tooling, AI and more. This role is our dedicated specialist for the ML and AI side of that platform.
About the Role
Take the next step in your career as a Senior ML/AI Platform Engineer at Storytel! At Storytel, engineers are viewed as leaders within the Individual Contributor roles, aiming to create impact and drive initiatives to bring our audio entertainment service to millions of users.
ML and AI have become a strategic part of how Storytel builds product experiences — from recommendations and content understanding to LLM- and agent-powered features. This role owns the paved road that makes that possible. You will design, build and operate the ML and AI platform that our data scientists, ML engineers and product teams build on top of, so that going from an experiment to a reliable production workload is a well-trodden path rather than a bespoke project every time.
Concretely, you will:
Own and evolve our ML training and inference platform on Google Cloud — pipeline orchestration with Metaflow, Argo Workflows and Argo Events, scheduled and batch data workflows with Airflow, and workloads running on GKE.
Build the model lifecycle tooling around experiment tracking, model registry, versioning, promotion and rollback using MLflow, and make reproducibility and lineage the default rather than an afterthought.
Shape our AI and agentic capabilities on the Gemini Enterprise Agent Platform (formerly Vertex AI), including Model Garden for model selection and serving, and ADK (Agent Development Kit) as the framework our teams use to build agents.
Productionise agentic systems: tool use and orchestration patterns, evaluation harnesses, guardrails, cost and latency controls, and the observability needed to operate agents reliably once they are in front of users.
Define how ML and AI workloads are deployed, monitored and secured — serving infrastructure, autoscaling, drift and quality monitoring, model and prompt evaluation, and sensible defaults for cost.
Treat the platform as a product: gather feedback from the teams who use it, remove friction, write the docs and templates, and measure success by how productive other engineers are.
Coach and unblock developers and data scientists in their daily work, and lead initiatives that improve the ML/AI developer experience across the organisation.
Python is the lingua franca of this space and the key language for the role.
You will work with technologies including MLflow, Metaflow, Airflow, Argo Workflows, Argo Events, ArgoCD, Kargo, Kubernetes and Google Kubernetes Engine (GKE), Google Cloud Platform — including the Gemini Enterprise Agent Platform (formerly Vertex AI), Model Garden and ADK — GitHub, GitHub Actions, Grafana, Prometheus, Infrastructure as Code (Terraform, Atlantis), Fastly, and more.
About You
To be successful in this role, we believe that you have:
Significant hands-on experience building, deploying and operating ML workloads on Google Cloud — ideally with the Gemini Enterprise Agent Platform (formerly Vertex AI) — not just as a consumer, but building the platform capabilities other teams rely on.
Strong experience with ML workflow orchestration and MLOps tooling in production. Metaflow, Airflow, Argo Workflows, Argo Events and MLflow are what we use; deep experience with comparable tools is equally welcome.
Practical experience with LLM and agentic systems in production — model selection and serving, prompt and context management, evaluation, guardrails, and observability. Experience with ADK, LangGraph or similar agent frameworks is a strong plus.
Strong Python skills, and the habits of a software engineer: testing, code review, CI, and code you’re happy for someone else to maintain.
A platform engineering mindset. You understand that the job is about empowering other engineers to do their absolute best work, and you measure your success by their productivity and happiness.
A deep care for developer experience. You believe that good DX has a direct, positive impact on team morale, delivery speed, and product quality — and you can point to things you’ve built that prove it.
Strong long-term thinking on architecture. You weigh the maintenance cost of every decision and prefer boring, durable solutions over clever ones that someone else will have to live with.
Pragmatic ambition. You have high standards and a vision for where the platform should go, but you also understand the constraints of working in a small organisation and know how to ship value incrementally.
At least 5 years of professional software engineering experience, with solid hands-on development skills.
Solid experience with container technologies, Kubernetes (ideally GKE), and distributed systems in production.
Working knowledge of Infrastructure as Code and CI/CD, and comfort operating in a cloud environment end to end.
An ability to communicate fluently in English (written and spoken) and to collaborate effectively across teams — this role does not work in isolation.
Bonus points for experience or interest in
Feature stores, online/offline serving, and real-time inference at scale
Evaluation and experimentation infrastructure for both classical ML and LLM applications
GPU and accelerator workloads on Kubernetes — scheduling, quotas, and cost management
Recommendation, search, or content-understanding systems
Building internal developer platforms (IDPs), golden paths, and self-service tooling
Data platform and pipeline work adjacent to ML — warehousing, streaming, data quality
AI security and governance — access control for models and data, PII handling, responsible AI practices
Infrastructure-as-code and configuration management at scale
Working with databases — MySQL, Postgres, etc.
Linux/Unix internals
Join a world of stories
At Storytel, we’re a team of creative story lovers who thrive on collaboration and new ideas. Our workplace is friendly, dynamic, and full of opportunities to experiment and make an impact. We believe in trust, flat hierarchies, and empowering you to grow with us.
Does this sound like your next story? Simply fill out the application form and share your CV or LinkedIn profile – no cover letter needed. Answer a few questions, and you’re all set.
Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.