Requirements
- Strong demonstrable ability to work with Linux systems and cloud platforms (AWS, GCP or Azure)
- Solid Kubernetes knowledge and ability to run production systems
- A clear understanding of observability (monitoring, logging, tracing)
- Capable of designing or operating high-availability, distributed systems
- A mindset focused on automation, scalability, and continuous improvement
- Confidence working in fast-moving environments where reliability really matters
What the job involves
- At Open Cosmos, our Data division transforms satellite data into meaningful insights that drive real-world impact. The team delivers all data products generated by Open Cosmos and its partners, curates and develops DataCosmos (our geospatial data platform) and builds integrations that make satellite imagery easy to access and act on
- We’re now looking for a Site Reliability Engineer to help us ensure our data platform is reliable, scalable, and performing at its best as we grow
- Owning the reliability, performance, and scalability of our data platform and processing pipelines
- Monitoring systems end-to-end, ensuring full visibility across infrastructure and data flows
- Responding to incidents, troubleshooting issues, and driving long-term fixes
- Improving deployments and contributing to CI/CD pipelines for safe, repeatable releases
- Working closely with engineering teams to design resilient, scalable systems
- Automating processes and reducing operational overhead
- Supporting customer-impacting issues alongside Customer Success teams
📌 Site Reliability Engineer (Barcelona)
🏢 EF
📍 Barcelona