12 sep
|
Palo Alto Networks
|
Madrid
12 sep
Palo Alto Networks
Madrid
Experteer Overview In this role, you will join Team Engineering to operate and improve large-scale, multi-cloud production environments, ensuring reliable, secure services for tens of thousands of enterprise customers. You’ll work closely with cross-functional teams to reduce incidents and drive operational excellence in a high-availability, distributed setup. The role offers hands-on engineering in Kubernetes, Terraform, and Python, with a strong emphasis on automation, monitoring, and rapid incident response. This is a remote-friendly, fast-paced position at a market-leading cybersecurity company with a culture that values collaboration and impact.Compensaciones / Beneficios
- Own and operate large-scale, general production environments across multiple cloud providers (GCP, AWS, Azure)
- Monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
- Drive end-to-end troubleshooting across complex, distributed systems with high context switching
- Design, deploy, and improve monitoring and observability systems (Prometheus, Grafana)
- Collaborate with internal teams (CX, CS, Engineering) to ensure reliability and performance
- Work hands-on with DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
- Develop and maintain automation and tooling (primarily in Python)
- Gain deep understanding of system architecture and interconnected services
- Contribute to a culture of operational excellence in a high-scale, high-availability environment
- Champion asynchronous communication, documentation, and tooling standards for distributed teams
- On call responsibilities: Daytime hours 12:00-20:00 CET/CEST; occasional weekends andholidays by rotationResponsabilidades
- 5+ years of experience in SRE roles in production environments at scale
- Strong hands-on experience with Kubernetes and Terraform
- Strong hands-on experience with at least one major cloud platform (GCP or AWS)
- Experience building and configuring monitoring systems (Prometheus, Grafana)
- Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
- Proficiency in Python for scripting and automation
- Proven success in a fully remote or distributed team environment
- Strong troubleshooting and problem-solving skills with incident handling
- Ability to work in fast-paced environments with high context switching
- Ownership-driven, proactive, and collaborative with strong communication skills
- Curious mindset and eagerness to learnRequisitos principales
-
📌 Sr Staff Site Reliability Engineer (Madrid)
🏢 Palo Alto Networks
📍 Madrid