11 sep
|
Mirantis
|
Barcelona
11 sep
Mirantis
Barcelona
Who you are
- Proven experience designing and building observability platforms for large-scale, production infrastructure environments (not just consuming an existing setup — actually building one)
- Strong hands-on experience with metrics, logging, and distributed tracing tooling (e.g., Prometheus, Grafana, OpenTelemetry, Loki, Thanos/Cortex/Mimir, Elasticsearch/OpenSearch, Jaeger/Tempo, or equivalents)
- Experience with high-volume telemetry pipelines and the tradeoffs involved (cardinality, retention, cost, query latency)
- Strong software engineering skills in at least one language commonly used in this space (e.g., Go, Python, Rust)
- Experience with Kubernetes and cloud-native infrastructure
- Solid understanding of SLO/SLI/error-budget practices and alerting design that minimizes noise
- Comfortable working in a fast-moving environment where the platform is being built out alongside the infrastructure it's monitoring
- Strong communication skills — able to work directly with operations/service delivery teams to understand real incident-response needs, not just technical specs
- Experience building observability for GPU/HPC infrastructure or other specialized, high-performance compute environments
- Experience with eBPF-based observability tooling
- Familiarity with AIOps/ML-based anomaly detection or automated triage systems
- Experience operating in a managed services or MSP context, where observability directly drives customer-facing SLAs
- Contributions to open-source observability projects
What the job involves
- Mirantis is building out our Neocloud service offering — managing large-scale infrastructure to a high SLA for customers running demanding compute workloads. As that offering scales, so does the volume and complexity of telemetry we need to collect, correlate, and act on
- We're looking for an Observability Platform Engineer to design and build the monitoring, logging, tracing, and alerting platform that our operations teams depend on to detect and resolve incidents fast — at large scale,
across a globally distributed environment
- This is a hands-on, build-it role. You'll be the person who turns "we have no visibility into this" into a platform that surfaces the right signal at the right time, and turns "we found out from the customer" into "we caught it before they noticed."
- Design, build, and operate observability platform components — metrics, logging, distributed tracing, and alerting — for large-scale infrastructure environments
- Build telemetry pipelines capable of handling high cardinality, high volume data from large fleets of infrastructure, with an eye on cost, retention, and query performance
- Define and implement SLO/SLI frameworks and alerting strategies that reduce noise and surface real signal to on-call engineers
- Partner closely with service delivery and operations teams to understand what they need to see during an incident, and build for that — not just for dashboards nobody opens
- Integrate observability tooling with incident management workflows, including root-cause analysis support and post-incident review data
- Continuously improve detection speed and reduce mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR) across the platform
- Contribute to the roadmap for AI-assisted operations tooling (e.g., automated triage, anomaly detection, engineer-assist tooling) as it matures
- Own the reliability, scalability, and security of the observability stack itself — it needs to be up when everything else is on fire
- Document architecture, runbooks, and operational practices so the platform is maintainable beyond you
- What Success Looks Like:
- Operations teams can diagnose incidents faster because the right data is surfaced automatically, not hunted for manually
- Alert volume is high-signal, low-noise — engineers trust what fires
- The observability platform scales cleanly as infrastructure footprint grows, without cost or performance surprises
- Reduced MTTD/MTTR trends, tracked and demonstrable over time
- A platform other engineers actually want to build on, not work around
📌 Observability Platform Engineer (Barcelona)
🏢 Mirantis
📍 Barcelona