18 sep
|
TRLLN
|
Barcelona
? Who we are
TRLLN is building the future of asset intelligence. Our Tracking-as-a-Service (TaaS) platform combines a proprietary IoT mesh network, telemetry, and advanced analytics to deliver real-time visibility over physical assets — track any asset, anywhere — with no gates, readers, or fixed infrastructure.
Born inside IFCO, the general market leader in reusable packaging containers for fresh food, TRLLN is now a standalone subsidiary of IFCO: engineered for the world’s largest reusable-asset network — 400 million assets across 50+ countries — and combining that proven, industrial scale with the agility and focus of a tech company, from our Barcelona HQ.
✅ What environment will you be joining?
As a Senior Site Reliability Engineer at TRLLN, you’ll join the product engineering team that builds and runs the core of our platform: the backend and infrastructure that ingest, process, and serve high-volume IoT telemetry in near real time, on a multi-tenant, cloud-native architecture on Google Cloud — GKE, Cloud Run, Managed Kafka, Pub/Sub, AlloyDB and BigQuery, all managed with Terraform.
You will be our first dedicated reliability role. Today reliability is everyone’s job and nobody’s specialty: the team has strong software engineering foundations, an observability stack we are actively maturing, and availability commitments to customers that we want to back with real SLOs, alerting and runbooks. Your job is to turn that into a reliability practice — and to make the engineers around you better at operating what they build.
Security is a big part of this role. The platform goes through external security assessments and is preparing for the EU Cyber Resilience Act; the operational side of that — cloud security posture, vulnerability and supply-chain management, incident readiness, business continuity — will sit with you.
TRLLN is a startup-like environment backed by a global enterprise: we move fast, iterate with purpose, and care deeply about doing things right, always balancing pragmatism, speed and quality. And we face real scale challenges: seasonal traffic peaks from massive IoT device rollouts, a growing hardware fleet in the field, demanding reliability targets, and a platform engineered to keep growing.
? What we’re looking for
- Strong analytical and problem-solving skills, with a focus on delivery and quality.
- Demonstrated leadership in technical decision-making, mentoring, and team collaboration.
- Excellent communication skills. Able to explain complex systems and trade-offs clearly, both verbally and in writing.
- An enabler mindset. You make the team better.
? What you bring
- Proven experience operating high-volume / high-traffic distributed systems in production — event streaming (Kafka, Pub/Sub or similar), telemetry or time-series data, performance under real load, and capacity planning for traffic peaks.
- Strong reliability engineering background: experience working with SLIs/SLOs and error budgets, alerting design, incident management, blameless postmortems.
- Hands-on with the observability stack: Prometheus / OpenTelemetry / Grafana or a cloud-native equivalent (we use Cloud Monitoring + Managed Prometheus) — metrics, logging and distributed tracing, including their cost.
- Solid Kubernetes, Docker and Terraform (or similar IaC) experience — networking, autoscaling,
workload identity, private clusters. We run on GCP, but experience on any major cloud transfers.
- Production databases: PostgreSQL (AlloyDB / Cloud SQL or equivalent) — performance, backups, restore and failover.
- Platform security, hands-on: cloud security posture (IAM and least privilege, network segmentation, secrets management, encryption / KMS), WAF and rate limiting (Cloud Armor or similar), and remediating findings from external security assessments and pentests.
- Vulnerability management and supply chain: image and dependency scanning, SBOMs, signed and attested builds, patch cadence — plus the operational side of security incident response and business continuity / disaster recovery.
- Compliance as an engineer: comfortable turning regulatory and audit requirements (EU Cyber Resilience Act, ISO 27001 / SOC 2) into controls, evidence and runbooks — you’ll help us get audit-ready.
- CI/CD and GitOps: pipelines as code, progressive delivery, safe rollbacks (we use GitHub Actions).
- You write production code: comfortable reading and contributing to Go and/or JVM (Kotlin / Java) services, not just YAML. Automation and instrumentation are part of the job.
- Comfortable working in both English and Spanish.
Nice to have
- Cloud FinOps / cost optimisation experience.
- Load and chaos testing.
- Exposure to IoT / device fleets or telemetry-heavy products.
- Security certifications (e.g. CKS, GCP Professional Cloud Security Engineer) or prior experience on an ISO 27001 / SOC 2 certification effort.
? What we offer
- Hybrid work model with a flexible schedule
- Based in our Barcelona HQ
- Comprehensive health insurance for your family unit
- Salary range: 55.000€ – 70.000€
- Annual training budget for continuous learning and growth
📌 Senior Site Reliability Engineer (Barcelona)
🏢 TRLLN
📍 Barcelona