Site Reliability | Senior (Cataluña)

Site Reliability | Senior (Cataluña)

23 ago
|
Talent-R
|
Cataluña

23 ago

Talent-R

Cataluña

Senior Site Reliability Engineer (SRE)

Location: Barcelona / Madrid (hybrid) or London Company: Next-generation commodity trading platform — Series B ($42M, Feb 2025)

About the company

We're building the next generation of commodity trading platforms — replacing the fragmented tools traders typically rely on with a single, powerful dashboard. Our product aggregates real-time feeds from across the commodities domain and turns them into intuitive, actionable visualisations in one unified interface.

We're growing fast off the back of a $42M Series B, and this is an early-stage opportunity to build your career with a level of ownership and impact you don't get at bigger, more bureaucratic companies.

The role

Traders act on what they see on our platform. That makes reliability a product feature, not an afterthought — and it's why this role exists.

This is a deliberate 50/50 split. Half your time is backend engineering: designing and building the distributed services and pipelines that move real-time market data through our platform. The other half is reliability engineering: making those systems observable, resilient, and cheap to operate. If you've ever shipped a service and then wished you owned how it ran in production, this is that job.

We've kept the split explicit rather than tidy. Our Platform Engineering team owns the shared platform, tooling and automation; you'll be a close partner to them, and the reliability of your own services is yours.

We're looking for people who thrive in an empowered environment — engineers who are comfortable being given problems to solve rather than solutions to implement. You'll enjoy working at pace, value autonomy,



and prefer to ask for forgiveness rather than permission.

This is a hybrid role — a couple of days a week in the office, with flexibility built in.

What you'll be doing

Backend engineering

- Design, build and maintain the backend services behind our real-time and analytical data processing
- Optimise pipelines and services for low latency, high throughput and scale
- Own features end to end, from shaping the approach to running them in production
- Contribute to design reviews with a clear view on the trade-offs

Reliability engineering

- Own the operational health of your services — define what healthy means, then measure it
- Improve observability, monitoring and alerting in Datadog, so problems surface before a trader notices
- Take part in incident response, then close the loop on the root cause
- Work with our runtime across Lambda, ECS and EKS, and help move more workloads onto Kubernetes
- Help operate the data and streaming infrastructure you depend on: Kafka, Flink, Redis/Valkey, RDS, Redshift
- Extend our infrastructure as code in AWS CDK, and our CI/CD pipelines
- Reduce toil — automate the manual, delete the unnecessary, make the next incident less likely

About you

- 4+ years as a software or reliability engineer, with production systems you've built and supported




- High ownership tendencies
- Genuine interest in both halves of this role — you want to write the service and own how it runs
- Strong in at least one of Kotlin, Java, Python or TypeScript
- A solid working understanding of AWS — compute, networking, storage, IAM — from running things in production
- Hands-on with infrastructure as code: AWS CDK, Terraform, CloudFormation or similar
- Practical experience running container workloads on Kubernetes: deploying, debugging, tuning
- Experience with CI/CD tooling and a clear view of a good delivery lifecycle
- A habit of instrumenting what you build — metrics, logging, tracing
- Experience being on the hook for production, incident response included, and calm when things are on fire
- Experience defining or working to service-level objectives, or a clear sense of how you'd start
- A strong urge to own and improve things — to spot what isn't working and fix it
- Comfortable with agent-based development tools (we use Claude Code; any equivalent is fine)
- A clear communicator and a pragmatic problem-solver

Nice to have

- Experience building or operating a platform on Kubernetes, EKS especially
- Datadog specifically
- Data or streaming systems: Kafka, Flink, Redshift, clustered Postgres
- Working to error budgets, and using them to make real prioritisation calls
- Security and compliance best practice in cloud environments
- Complex distributed environments �� high throughput, low latency, large datasets
- Any exposure to commodities, energy or financial markets (useful, but not required — we'll teach you the domain)

📌 Site Reliability | Senior (Cataluña)
🏢 Talent-R
📍 Cataluña

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability | senior (cataluña) / cataluña

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability | senior (cataluña) / cataluña