Site Reliability Engineer (Madrid)

Site Reliability Engineer (Madrid)

08 ago
|
Remobi
|
Madrid

08 ago

Remobi

Madrid

Site Reliability Engineer at Remobi About the role Remobi is seeking a skilled Site Reliability Engineer to join their team and take ownership of maintaining and enhancing a cloud-native industrial asset management platform. This position focuses on ensuring operational excellence through proactive monitoring, efficient incident response, and continuous platform improvements. You will work alongside operations and DevOps teams to strengthen system reliability, optimize deployment processes, and drive measurable improvements in service level objectives. The role offers an opportunity to shape the operational maturity of a growing platform while working with modern cloud technologies and observability tools.
Key facts Location: European Union Engagement: Full-time employment
What you'll do Provide hands-on support for daily platform operations, working closely with the operations team to ensure smooth functioning of all cloud-based services and infrastructure components
Lead incident response efforts by quickly identifying root causes, coordinating with relevant stakeholders, and implementing effective resolutions to minimize service disruptions and customer impact
Develop and refine troubleshooting workflows and runbooks to accelerate issue detection and reduce mean time to resolution across all platform components
Oversee deployment activities and coordinate patching schedules across multiple cloud environments, ensuring minimal downtime and adherence to change management protocols
Design and implement enhanced monitoring solutions using OpenSearch and OpenTelemetry to provide comprehensive visibility into system health and performance metrics
Configure and tune alerting systems to provide early warning indicators of potential issues, enabling proactive intervention before problems escalated to customer-facing incidents
Produce detailed deployment reports and post-incident analyses that document actions taken, lessons learned, and recommendations for preventing similar issues in the future




Champion initiatives aimed at improving platform stability, increasing release velocity, and advancing the overall operational maturity of the infrastructure
Manage and optimize Kubernetes clusters and containerized workloads to ensure efficient resource utilization, high availability and seamless scaling capabilities
Partner with DevOps engineers to identify process bottlenecks and implement platform enhancements that streamline operations and improve developer productivity
Contribute to the evolution service level objectives and establish meaningful metrics that accurately reflect system reliability and user experience
Participate in on-call rotations and ensure appropriate coverage for critical systems during off-hours and weekends
Requirements Demonstrated professional experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles with a track record of maintaining production systems
Deep expertise with Kubernetes orchestration, including cluster management, pod scheduling, resource allocation and troubleshooting containerized applications
Proficiency with virtualization technologies and understanding of how virtual infrastructure supports modern cloud deployments
Hands-on experience building and maintaining CI/CD pipelines using GitLab, including pipeline optimization and automated testing integration
Strong working knowledge of Terraform for Infrastructure as Code, including module development, state management, and infrastructure provisioning workflows
Practical experience with Apache Kafka for event streaming, including topic management, consumer group configuration, and performance tuning




Familiarity with observability platforms, specifically OpenSearch for log aggregation and analysis, and OpenTelemetry for distributed tracing and metrics collection
Solid grasp of incident management frameworks and cloud operations best practices, including change management, capacity planning and disaster recovery procedures
Excellent communication skills with the ability to document technical processes clearly and collaborate effectively with diverse team members
Nice to have Previous experience operating large-scale distributed systems or customer-facing cloud platforms where reliability directly impacts end-user satisfaction
Strong analytical mindset with proven problem-solving abilities and a systematic approach to diagnosing complex technical issues
Demonstrated success working within cross-functional teams, bridging gaps between development, operations and business stakeholders
Experience with additional cloud providers or multi-cloud architectures
Background in industrial technology, IoT platforms or asset management systems
Familiarity with chaos engineering practices and reliability testing methodologies
Skills & tools Kubernetes, Docker, container orchestration
Terraform, Infrastructure as Code
GitLab CI/CD pipelines
Apache Kafka, event streaming
OpenSearch, OpenTelemetry, observability
Virtualization technologies
Incident management platforms
Linux system administration
Cloud platforms and services
Scripting languages for automation
Practical notes This position is based within the European Union, and candidates should be eligible to work in EU member states. The role involves collaboration with distributed teams, so strong written and verbal communication skills are essential for effective remote coordination. Expect participation in on-call schedules to support incident response outside regular business hours. Continuous learning is encouraged as the platform evolves and new technologies are adopted.
#J-18808-Ljbffr

📌 Site Reliability Engineer (Madrid)
🏢 Remobi
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (madrid) / madrid