19 sep
|
Nexthink
|
Madrid
Experteer Overview As a Senior Site Reliability Engineer at Nexthink, you will strengthen our cloud-native platform and drive reliable, scalable delivery for a multi-tenant SaaS. You will partner with 50+ product teams and security groups to design robust infrastructure, implement automation, and optimize incident response. Your work ensures seamless user experiences for our integral customers and supports Nexthink’s mission of proactive digital employee experience. This role offers hands-on influence across cloud, Kubernetes, IaC, and observability at scale.Compensaciones / Beneficios
- Manage cloud-native systems (AWS) and automation tooling
- Operate and optimize Kubernetes clusters, pipelines, and service meshes
- Build and maintain infrastructure for a multi-tenant SaaS platform with reliability and security
- Define SLOs/SLAs and proactively address availability and performance issues
- Develop infrastructure as code (Terraform or equivalent) for repeatable provisioning
- Create internal platform tools for provisioning, monitoring, and efficiency
- Monitor systems to ensure high-quality user experiences
- Participate in on-call rotation, incident response, and cross-team communication
- Lead incident response as Incident Commander and drive post-incident reviews
- Improve incident metrics (MTTD/MTTR) and reduce outages
- Collaborate with software engineers to embed observability and fault tolerance
- Automate runbooks, health checks, and alerting; support canary deployments and safe rollbacks
- Contribute to security, compliance automation, and cost optimization
- Assist with automated testing and deployment strategiesResponsabilidades
- 5+ years as an SRE or Platform Engineer with strong software development practices
- Hands-on with cloud services (AWS, GCP, Azure) and SaaS support
- Programming/scripting in Python, Go, Bash; IaC experience (Terraform)
- Proficiency in Kubernetes, Docker, and related ecosystems (Helm)
- Experience with multi-tenant microservices and CI/CD (Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane)
- Experience with monitoring (Datadog) and incident management
- Strong Linux, networking, and troubleshooting skills; knowledge of TCP/IP, VPN, VPCs, firewalls, load balancers
- Familiarity with zero-downtime deployments (blue/green, canary) and resilience testing
- Exposure to compliance standards (SOC 2, ISO 27001, HIPAA); FedRAMP a plus
- Excellent English communication and collaborative mindset; able to work in agile environmentsRequisitos principales
- Hybrid work model
- unlimited vacation
- 3 company-paid volunteer days
- fitness centre access
- travel reimbursement
- French-language classes reimbursement
📌 Senior Site Reliability Engineer (Madrid)
🏢 Nexthink
📍 Madrid