04 ago
|
jobr.pro
|
Madrid
pp#LI-Hybrid /ph3Job Description /h3pAt Nexthink, we empower our customers with industry-leading solutions to enable continuous improvement of employee experience. We deliver unmatched visibility across all environments, so IT teams can consistently see, diagnose, and fix digital workplace issues. As a SaaS provider, our commitment is to deliver a seamless, resilient, and scalable platform around the clock. /ppWe are looking for an experienced, proactive and innovative professional that is keen to join as a Senior Site Reliability Engineer! The mission of Nexthink's SRE team is to strengthen our infrastructure and enhance our ability to deploy, monitor, and scale systems effectively and reliably. They work closely with over 50 Product Engineering teams that develop our products and services, as well as with the Technical Platform Engineering, Security and Architecture teams to understand the reliability requirements, design and implement solutions, and promote them for adoption and usage. /ppJoin our vibrant team of diverse and experienced engineers where cutting-edge technology meets innovation. Be a part of Nexthink's Digital Employee Experience technological revolution, ensuring our general customers enjoy a seamless user experience. Apply now and become a key player in our dynamic SRE organisation. /ppbAs a Senior Site Reliability Engineer, you will: /b /pulliImplement and manage cloud-native systems (AWS) using best-in-class tools and automation. /liliOperate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles. /liliDesign, build, and maintain the infrastructure powering our multi-tenant SaaS platform with reliability, security, and scalability in mind. /liliDefine and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues. /liliDevelop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning. /liliBuild internal platform tools and automation to support provisioning, monitoring, and operational efficiency. /liliMonitor infrastructure and applications ensuring high-quality user experiences. /liliParticipate in a shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution and communication. /liliAct as an Incident Commander during the on-call duty and coordinate cross-team responses effectively to maintain an SLA. /liliDrive and refine incident response processes, reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
/liliDiagnose and resolve complex issues independently, minimizing the need for external escalation. /liliWork closely with software engineers to embed observability, fault tolerance, and reliability principles into service design. /liliAutomate runbooks, health checks, and alerting to support reliable operations with minimal manual intervention. /liliSupport automated testing, canary deployments, and rollback strategies to ensure safe, fast, and reliable releases. /liliContribute to security best practices, compliance automation, and cost optimization. /li /ulh3Qualifications /h3pMinimum Bachelor’s degree in Computer Science or equivalent practical experience. /pulli5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices. /liliStrong hands‑on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS product. /liliStrong programming or scripting skills (e.g., Python, Go, Bash…) and experience with infrastructure-as-code (e.g., Terraform). /liliProficiency with Kubernetes, container‑based deployment (e.g., Docker) and related ecosystems (e.g., Helm). /liliExperience supporting multi‑tenant microservices architectures. /liliExperience with CI/CD pipelines tools (e.g., Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane). /liliExperience with managing monitoring solutions (e.g., Datadog). /liliComfortable participating in a rotating on‑call schedule, managing critical incidents, and leading post‑incident reviews. /liliAt ease with operating and managing production systems, striking the right balance between urgency and methodology. /liliStrong system‑level troubleshooting skills and a proactive mindset toward incident prevention. /liliDeep understanding of Linux systems, networking, and common troubleshooting practices. /liliSolid understanding of the network stack (e.g., TCP/IP, VPN, etc.), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc). /liliKnowledge of zero‑downtime deployment strategies, blue/green and canary releases.
/liliExposure to compliance standards such as SOC 2, ISO 27001, or HIPAA. FedRAMP experience is a big plus. /liliExperience with chaos engineering or resilience testing practices. /liliExcellent problem‑solving skills, collaborative mindset, and a strong grasp of agile, iterative development. /liliSelf‑driven, highly organised, and capable of independently managing priorities. /liliCuriosity to learn new things and discover new technologies. /liliStrong communication, presentation, and team collaboration skills. /liliExcellent written and verbal skills in English. /li /ulpCheck what we offer: /pulli Permanent Contract and a competitive compensation package. /lili? Health insurance through our partnership with ACKO, including OPD coverage for dental, vision, health check‑ups, consultations, and pharmacy expenses. /lili Hybrid work model balancing office and remote work, with a structured approach for new hires to foster connections and onboarding. /lili ️ Flexible Hours and unlimited vacation (employees have unlimited paid time off on top of the 22 days of holidays we offer). Plus, company‑paid bank holidays (12), sick days (10‑30), bereavement leave (5), and 3 days per year for volunteering. /lili Free access to professional training platforms to explore your interests and enhance your skills. /lili?️ Stay covered against accidents, bodily injuries, and disabilities with our personal accident insurance policy, providing assurance with coverage up to three times your annual CTC. /lili New mothers are entitled to up to 26 weeks of maternity leave, with the flexibility to use up to 8 weeks before the expected delivery and the remaining 18 weeks after. Birth fathers can take 6 weeks of paternity leave, while adoptive parents are eligible for 26 weeks of leave for mothers and 6 weeks for fathers. /lili Under the Payment of Gratuity Act, receive gratuity at the rate of 15 days of basic pay for every completed year of service, provided you've been employed by the company for a minimum of 5 years. Gratuity is payable at retirement or resignation based on your last drawn basic pay. /lili Bonuses for referring successful hires after three months of continuous employment. /li /ulpPlease note that not all the benefits listed above are available for temporary, contract, and internship roles. To ensure you have the most up‑to‑date information, we recommend checking with your Recruitment Partner. /p /p #J-18808-Ljbffr
📌 Senior Site Reliability Engineer (Madrid)
🏢 jobr.pro
📍 Madrid