25 ago
|
TK Elevator
|
Madrid
25 ago
TK Elevator
Madrid
To strengthen our global Digital Technology organization, we are looking for an experienced Senior Site Reliability Engineer (d/f/m) who will take ownership of System Health Monitoring across our digital ecosystem and the MAX IoT Platform. In this highly visible role, you will help establish reliability as a core capability across the organization by defining observability standards, driving monitoring excellence, and enabling engineering teams to build resilient, reliable solutions.
Drive Monitoring & Observability Excellence
Own and continuously evolve System Health Monitoring across products and platforms.
Define and govern observability standards, monitoring requirements, health models, and alerting strategies.
Establish a unified view of platform health utilizing Azure observability solutions, Central Log Analytics, Grafana, and DevOps monitoring tools.
Design, implement, and optimize Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budget frameworks.
Promote reliability engineering best practices, including health checks, error‑budget management, and post‑incident learning processes.
Analyze incident trends and identify opportunities to improve monitoring, alerting, synthetic testing, and operational resilience.
Transform monitoring data into actionable insights and predictive analytics that proactively identify risks and performance issues.
Partner closely with Product, Architecture, DevOps, and Incident Operations teams to improve monitoring coverage and alert quality.
Provide technical guidance, coaching,
and mentorship to engineering teams across the organization.
Develop and maintain high‑quality documentation, including runbooks, monitoring standards, and incident‑response procedures.
Contribute actively to the global DevOps community by sharing knowledge, best practices, and lessons learned.
Minimum 5 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or a related discipline.
Strong expertise in Microsoft Azure and cloud‑native technologies.
Azure Monitor, Log Analytics / Kusto Query Language (KQL), Application Insights, Grafana.
Proven experience defining and managing SLIs, SLOs, error budgets, and reliability frameworks for large‑scale distributed systems.
Strong understanding of distributed architectures, cloud platforms, and Azure PaaS services.
Solid scripting and automation skills using technologies such as C#/.NET, PowerShell, or Python.
Fluency in English (written and spoken) is required.
Experience with Power BI, Databricks, or AI‑driven observability and analytics solutions.
Knowledge of Azure DevOps, CI/CD pipelines, Git, and Agile methodologies.
Bachelor's degree in Computer Engineering or a related technical discipline.
Flexibility – We support, for example, through versátil yet regulated working hours and remote working options.
Development – Individual support to help you get started in your new job as well as training and education programs to help you develop professionally and personally.
📌 Senior Site Reliability Engineer (d/f/m) (Madrid)
🏢 TK Elevator
📍 Madrid