Site Reliability Engineer (Madrid)

Site Reliability Engineer (Madrid)

06 ago
|
Remobi
|
Madrid

06 ago

Remobi

Madrid

Site Reliability Engineer at Remobi About the role Remobi is seeking a skilled Site Reliability Engineer to join their team and take ownership of maintaining and enhancing a cloud-native industrial asset management platform. This position focuses on ensuring operational excellence through proactive monitoring, efficient incident response, and continuous platform improvements. You will work alongside operations and DevOps teams to strengthen system reliability, optimize deployment processes, and drive measurable improvements in service level objectives. The role offers an opportunity to shape the operational maturity of a growing platform while working with modern cloud technologies and observability tools.Key facts Location: European Union Engagement: Full-time employmentWhat you'll do Provide hands-on support for daily platform operations, working closely with the operations team to ensure smooth functioning of all cloud-based services and infrastructure componentsLead incident response efforts by quickly identifying root causes, coordinating with relevant stakeholders, and implementing effective resolutions to minimize service disruptions and customer impactDevelop and refine troubleshooting workflows and runbooks to accelerate issue detection and reduce mean time to resolution across all platform componentsOversee deployment activities and coordinate patching schedules across multiple cloud environments, ensuring minimal downtime and adherence to change management protocolsDesign and implement enhanced monitoring solutions using OpenSearch and OpenTelemetry to provide comprehensive visibility into system health and performance metricsConfigure and tune alerting systems to provide early warning indicators of potential issues, enabling proactive intervention before problems escalated to customer-facing incidentsProduce detailed deployment reports and post-incident analyses that document actions taken, lessons learned,



and recommendations for preventing similar issues in the futureChampion initiatives aimed at improving platform stability, increasing release velocity, and advancing the overall operational maturity of the infrastructureManage and optimize Kubernetes clusters and containerized workloads to ensure efficient resource utilization, high availability and seamless scaling capabilitiesPartner with DevOps engineers to identify process bottlenecks and implement platform enhancements that streamline operations and improve developer productivityContribute to the evolution service level objectives and establish meaningful metrics that accurately reflect system reliability and user experienceParticipate in on-call rotations and ensure appropriate coverage for critical systems during off-hours and weekendsRequirements Demonstrated professional experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles with a track record of maintaining production systemsDeep expertise with Kubernetes orchestration, including cluster management, pod scheduling, resource allocation and troubleshooting containerized applicationsProficiency with virtualization technologies and understanding of how virtual infrastructure supports modern cloud deploymentsHands-on experience building and maintaining CI/CD pipelines using GitLab, including pipeline optimization and automated testing integrationStrong working knowledge of Terraform for Infrastructure as Code, including module development, state management, and infrastructure provisioning workflowsPractical experience with Apache Kafka for event streaming, including topic management, consumer group configuration,



and performance tuningFamiliarity with observability platforms, specifically OpenSearch for log aggregation and analysis, and OpenTelemetry for distributed tracing and metrics collectionSolid grasp of incident management frameworks and cloud operations best practices, including change management, capacity planning and disaster recovery proceduresExcellent communication skills with the ability to document technical processes clearly and collaborate effectively with diverse team membersNice to have Previous experience operating large-scale distributed systems or customer-facing cloud platforms where reliability directly impacts end-user satisfactionStrong analytical mindset with proven problem-solving abilities and a systematic approach to diagnosing complex technical issuesDemonstrated success working within cross-functional teams, bridging gaps between development, operations and business stakeholdersExperience with additional cloud providers or multi-cloud architecturesBackground in industrial technology, IoT platforms or asset management systemsFamiliarity with chaos engineering practices and reliability testing methodologiesSkills & tools Kubernetes, Docker, container orchestrationTerraform, Infrastructure as CodeGitLab CI/CD pipelinesApache Kafka, event streamingOpenSearch, OpenTelemetry, observabilityVirtualization technologiesIncident management platformsLinux system administrationCloud platforms and servicesScripting languages for automationPractical notes This position is based within the European Union, and candidates should be eligible to work in EU member states. The role involves collaboration with distributed teams, so strong written and verbal communication skills are essential for effective remote coordination. Expect participation in on-call schedules to support incident response outside regular business hours. Continuous learning is encouraged as the platform evolves and new technologies are adopted.#J-18808-Ljbffr

📌 Site Reliability Engineer (Madrid)
🏢 Remobi
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (madrid) / madrid