Descripción del trabajo - Site Reliability | Senior
Senior Site Reliability Engineer (SRE)
Location: Barcelona / Madrid (hybrid) or London Company: Next-generation commodity trading platform — Series B ($42M, Feb 2025)
About the company
We're building the next generation of commodity trading platforms — replacing the fragmented tools traders typically rely on with a single, powerful dashboard. Our product aggregates real-time feeds from across the commodities domain and turns them into intuitive, actionable visualisations in one unified interface.
We're growing fast off the back of a $42M Series B, and this is an early-stage opportunity to build your career with a level of ownership and impact you don't get at bigger, more bureaucratic companies.
The role
Traders act on what they see on our platform. That makes reliability a product feature, not an afterthought — and it's why this role exists.
This is a deliberate 50/50 split. Half your time is backend engineering: designing and building the distributed services and pipelines that move real-time market data through our platform. The other half is reliability engineering: making those systems observable, resilient, and cheap to operate. If you've ever shipped a service and then wished you owned how it ran in production, this is that job.
We've kept the split explicit rather than tidy. Our Platform Engineering team owns the shared platform, tooling and automation; you'll be a close partner to them, and the reliability of your own services is yours.
We're looking for people who thrive in an empowered environment — engineers who are comfortable being given problems to solve rather than solutions to implement. You'll enjoy working at pace, value autonomy, and prefer to ask for forgiveness rather than permission.
This is a hybrid role — a couple of days a week in the office, with flexibility built in.
What you'll be doing
Backend engineering
Design, build and maintain the backend services behind our real-time and analytical data processing
Optimise pipelines and services for low latency, high throughput and scale
Own features end to end, from shaping the approach to running them in production
Contribute to design reviews with a clear view on the trade-offs
Reliability engineering
Own the operational health of your services — define what healthy means, then measure it
Improve observability, monitoring and alerting in Datadog, so problems surface before a trader notices
Take part in incident response, then close the loop on the root cause
Work with our runtime across Lambda, ECS and EKS, and help move more workloads onto Kubernetes
Help operate the data and streaming infrastructure you depend on: Kafka, Flink, Redis/Valkey, RDS, Redshift
Extend our infrastructure as code in AWS CDK, and our CI/CD pipelines
Reduce toil — automate the manual, delete the unnecessary, make the next incident less likely
About you
4+ years as a software or reliability engineer, with production systems you've built and supported
High ownership tendencies
Genuine interest in both halves of this role — you want to write the service and own how it runs
Strong in at least one of Kotlin, Java, Python or TypeScript
A solid working understanding of AWS — compute, networking, storage, IAM — from running things in production
Hands-on with infrastructure as code: AWS CDK, Terraform, CloudFormation or similar
Practical experience running container workloads on Kubernetes: deploying, debugging, tuning
Experience with CI/CD tooling and a clear view of a good delivery lifecycle
A habit of instrumenting what you build — metrics, logging, tracing
Experience being on the hook for production, incident response included, and calm when things are on fire
Experience defining or working to service-level objectives, or a clear sense of how you'd start
A strong urge to own and improve things — to spot what isn't working and fix it
Comfortable with agent-based development tools (we use Claude Code; any equivalent is fine)
A clear communicator and a pragmatic problem-solver
Nice to have
Experience building or operating a platform on Kubernetes, EKS especially
Datadog specifically
Data or streaming systems: Kafka, Flink, Redshift, clustered Postgres
Working to error budgets, and using them to make real prioritisation calls
Security and compliance best practice in cloud environments
Complex distributed environments — high throughput, low latency, large datasets
Any exposure to commodities, energy or financial markets (useful, but not required — we'll teach you the domain)
Todos los anuncios de empleo están sujetos a las :Condiciones de GrabJobs. Permitimos a los usuarios marcar los anuncios que puedan infringir dichas condiciones. Los anuncios de empleo también pueden ser marcados por el equipo de moderación de GrabJobs. Sin embargo, ningún sistema de moderación es perfecto, y marcar un anuncio no garantiza que vaya a ser eliminado.
Sé el primero en recibir las últimas ofertas de trabajo de Others Full-Time en Spain.
Setup your job alert:
Al activar las alertas de empleo, acepto los Terms & Privacy Policy de GrabJobs. Puedo darme de baja de las alertas de empleo en cualquier momento.
Saltar
Has alcanzado el número máximo de alertas de empleo.
GrabJobs es el portal de empleo nº 1 en Spain, que te conecta con miles de empleos clave ¡rápidamente!
Encuentra los mejores trabajos de en Spain, ¡solicita en 1 clic y consigue un trabajo hoy mismo!