Site Reliability Engineer (España)

Site Reliability Engineer (España)

06 ago
|
Remobi
|
España

06 ago

Remobi

España

Site Reliability Engineer at Remobi

About the role Remobi is seeking a skilled Site Reliability Engineer to join their team and take ownership of maintaining and enhancing a cloud-native industrial asset management platform. This position focuses on ensuring operational excellence through proactive monitoring, efficient incident response, and continuous platform improvements. You will work alongside operations and DevOps teams to strengthen system reliability, optimize deployment processes, and drive measurable improvements in service level objectives. The role offers an opportunity to shape the operational maturity of a growing platform while working with modern cloud technologies and observability tools.

Key facts

Location: European Union Engagement: Full-time employment

What you'll do

- Provide hands-on support for daily platform operations, working closely with the operations team to ensure smooth functioning of all cloud-based services and infrastructure components
- Lead incident response efforts by quickly identifying root causes, coordinating with relevant stakeholders, and implementing effective resolutions to minimize service disruptions and customer impact
- Develop and refine troubleshooting workflows and runbooks to accelerate issue detection and reduce mean time to resolution across all platform components
- Oversee deployment activities and coordinate patching schedules across multiple cloud environments, ensuring minimal downtime and adherence to change management protocols
- Design and implement enhanced monitoring solutions using OpenSearch and OpenTelemetry to provide comprehensive visibility into system health and performance metrics
- Configure and tune alerting systems to provide early warning indicators of potential issues, enabling proactive intervention before problems escalated to customer-facing incidents
- Produce detailed deployment reports and post-incident analyses that document actions taken, lessons learned, and recommendations for preventing similar issues in the future




- Champion initiatives aimed at improving platform stability, increasing release velocity, and advancing the overall operational maturity of the infrastructure
- Manage and optimize Kubernetes clusters and containerized workloads to ensure efficient resource utilization, high availability and seamless scaling capabilities
- Partner with DevOps engineers to identify process bottlenecks and implement platform enhancements that streamline operations and improve developer productivity
- Contribute to the evolution service level objectives and establish meaningful metrics that accurately reflect system reliability and user experience
- Participate in on-call rotations and ensure appropriate coverage for critical systems during off-hours and weekends

Requirements

- Demonstrated professional experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles with a track record of maintaining production systems
- Deep expertise with Kubernetes orchestration, including cluster management, pod scheduling, resource allocation and troubleshooting containerized applications
- Proficiency with virtualization technologies and understanding of how virtual infrastructure supports modern cloud deployments
- Hands-on experience building and maintaining CI/CD pipelines using GitLab, including pipeline optimization and automated testing integration
- Strong working knowledge of Terraform for Infrastructure as Code, including module development, state management, and infrastructure provisioning workflows
- Practical experience with Apache Kafka for event streaming, including topic management, consumer group configuration, and performance tuning




- Familiarity with observability platforms, specifically OpenSearch for log aggregation and analysis, and OpenTelemetry for distributed tracing and metrics collection
- Solid grasp of incident management frameworks and cloud operations best practices, including change management, capacity planning and disaster recovery procedures
- Excellent communication skills with the ability to document technical processes clearly and collaborate effectively with diverse team members

Nice to have

- Previous experience operating large-scale distributed systems or customer-facing cloud platforms where reliability directly impacts end-user satisfaction
- Strong analytical mindset with proven problem-solving abilities and a systematic approach to diagnosing complex technical issues
- Demonstrated success working within cross-functional teams, bridging gaps between development, operations and business stakeholders
- Experience with additional cloud providers or multi-cloud architectures
- Background in industrial technology, IoT platforms or asset management systems
- Familiarity with chaos engineering practices and reliability testing methodologies

Skills & tools

- Kubernetes, Docker, container orchestration
- Terraform, Infrastructure as Code
- GitLab CI/CD pipelines
- Apache Kafka, event streaming
- OpenSearch, OpenTelemetry, observability
- Virtualization technologies
- Incident management platforms
- Linux system administration
- Cloud platforms and services
- Scripting languages for automation

Practical notes

This position is based within the European Union, and candidates should be eligible to work in EU member states. The role involves collaboration with distributed teams, so strong written and verbal communication skills are essential for effective remote coordination. Expect participation in on-call schedules to support incident response outside regular business hours. Continuous learning is encouraged as the platform evolves and new technologies are adopted.

#J-18808-Ljbffr

📌 Site Reliability Engineer (España)
🏢 Remobi
📍 España

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (españa) / españa

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (españa) / españa