Site Reliability Engineer (España)

Site Reliability Engineer (España)

06 ago
|
Remobi
|
España

06 ago

Remobi

España

ph3Site Reliability Engineer at Remobi /h3pAbout the role Remobi is seeking a skilled Site Reliability Engineer to join their team and take ownership of maintaining and enhancing a cloud-native industrial asset management platform. This position focuses on ensuring operational excellence through proactive monitoring, efficient incident response, and continuous platform improvements. You will work alongside operations and DevOps teams to strengthen system reliability, optimize deployment processes, and drive measurable improvements in service level objectives. The role offers an opportunity to shape the operational maturity of a growing platform while working with modern cloud technologies and observability tools. /ph3Key facts /h3pLocation: European Union Engagement: Full-time employment /ph3What you'll do /h3ulliProvide hands-on support for daily platform operations, working closely with the operations team to ensure smooth functioning of all cloud-based services and infrastructure components /liliLead incident response efforts by quickly identifying root causes, coordinating with relevant stakeholders, and implementing effective resolutions to minimize service disruptions and customer impact /liliDevelop and refine troubleshooting workflows and runbooks to accelerate issue detection and reduce mean time to resolution across all platform components /liliOversee deployment activities and coordinate patching schedules across multiple cloud environments, ensuring minimal downtime and adherence to change management protocols /liliDesign and implement enhanced monitoring solutions using OpenSearch and OpenTelemetry to provide comprehensive visibility into system health and performance metrics /liliConfigure and tune alerting systems to provide early warning indicators of potential issues, enabling proactive intervention before problems escalated to customer-facing incidents /liliProduce detailed deployment reports and post-incident analyses that document actions taken, lessons learned, and recommendations for preventing similar issues in the future /liliChampion initiatives aimed at improving platform stability,



increasing release velocity, and advancing the overall operational maturity of the infrastructure /liliManage and optimize Kubernetes clusters and containerized workloads to ensure efficient resource utilization, high availability and seamless scaling capabilities /liliPartner with DevOps engineers to identify process bottlenecks and implement platform enhancements that streamline operations and improve developer productivity /liliContribute to the evolution service level objectives and establish meaningful metrics that accurately reflect system reliability and user experience /liliParticipate in on-call rotations and ensure appropriate coverage for critical systems during off-hours and weekends /li /ulh3Requirements /h3ulliDemonstrated professional experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles with a track record of maintaining production systems /liliDeep expertise with Kubernetes orchestration, including cluster management, pod scheduling, resource allocation and troubleshooting containerized applications /liliProficiency with virtualization technologies and understanding of how virtual infrastructure supports modern cloud deployments /liliHands-on experience building and maintaining CI/CD pipelines using GitLab, including pipeline optimization and automated testing integration /liliStrong working knowledge of Terraform for Infrastructure as Code, including module development, state management, and infrastructure provisioning workflows /liliPractical experience with Apache Kafka for event streaming, including topic management, consumer group configuration, and performance tuning /liliFamiliarity with observability platforms,



specifically OpenSearch for log aggregation and analysis, and OpenTelemetry for distributed tracing and metrics collection /liliSolid grasp of incident management frameworks and cloud operations best practices, including change management, capacity planning and disaster recovery procedures /liliExcellent communication skills with the ability to document technical processes clearly and collaborate effectively with diverse team members /li /ulh3Nice to have /h3ulliPrevious experience operating large-scale distributed systems or customer-facing cloud platforms where reliability directly impacts end-user satisfaction /liliStrong analytical mindset with proven problem-solving abilities and a systematic approach to diagnosing complex technical issues /liliDemonstrated success working within cross-functional teams, bridging gaps between development, operations and business stakeholders /liliExperience with additional cloud providers or multi-cloud architectures /liliBackground in industrial technology, IoT platforms or asset management systems /liliFamiliarity with chaos engineering practices and reliability testing methodologies /li /ulh3Skills tools /h3ulliKubernetes, Docker, container orchestration /liliTerraform, Infrastructure as Code /liliGitLab CI/CD pipelines /liliApache Kafka, event streaming /liliOpenSearch, OpenTelemetry, observability /liliVirtualization technologies /liliIncident management platforms /liliLinux system administration /liliCloud platforms and services /liliScripting languages for automation /li /ulh3Practical notes /h3pThis position is based within the European Union, and candidates should be eligible to work in EU member states. The role involves collaboration with distributed teams, so strong written and verbal communication skills are essential for effective remote coordination. Expect participation in on-call schedules to support incident response outside regular business hours. Continuous learning is encouraged as the platform evolves and new technologies are adopted. /p /p #J-18808-Ljbffr

📌 Site Reliability Engineer (España)
🏢 Remobi
📍 España

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (españa) / españa

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (españa) / españa