Senior Site Reliability Engineer - DevOps (Remote) (Madrid)

Senior Site Reliability Engineer - DevOps (Remote) (Madrid)

31 ago
|
Lodgify
|
Madrid

31 ago

Lodgify

Madrid

SpainPlatform Group – Development /Full-time Contrato fijo /Remote Who we are Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology. Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals.

Role OverviewAre you a systems-minded engineer who cares deeply about reliability, scalability, and production excellence? Join Lodgify as a Senior Site Reliability Engineer and help our engineering teams build and operate services that are reliable, observable, scalable, and resilient by design. Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.

Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures. Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana. Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.

Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks. Implement self-service Internal Developer Platform features via APIs and Kubernetes operators. Improve deployment safety,



rollbackability, and release observability.

Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms. Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability. You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.

SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery. You can write maintainable software to automate operational tasks and reduce manual intervention. You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.

You know how to balance reliability, performance, cost, and delivery speed pragmatically. You collaborate effectively with Engineering, Platform, Security, and Product stakeholders. MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.





Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience. ð Remote Flexibility: The freedom to work from home any day that works for you.ð Time to Recharge: 25 working days of paid vacation and Jornada Intensiva in August..ð Alan Health Insurance: Premium health, dental, and mental health support via Alan. ð Meal Perk: 150/month allowance on your Alan card + 50% off Ametller Origen prepared dishes at the office.ð Tax-Free Savings: Increase your take-home pay by using Flexible Remuneration for extra meal costs (up to 70/mo) and public transport (up to 136/mo).ð️ Home Office Gear: We provide a table, ergonomic chair, and monitor for your home setup.ðað Language Learning: Free Spanish classes.ð Referrals: Cash rewards for bringing in new talent. ð Social Life: Daily office breakfast and monthly team eventsð Dynamic Hub: A high-energy, inclusive environment designed for collaboration and connection with a team that represents over 60 countries.*Benefits offered may differ based on the type of contract that is issuedSo, what are you waiting for? All applications and CVs must be submitted in English ðWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. If you would like more information about how your data is processed, please contact us.

📌 Senior Site Reliability Engineer - DevOps (Remote) (Madrid)
🏢 Lodgify
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer - devops (remote) (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer - devops (remote) (madrid) / madrid