22 sep
|
Jobgether
|
España
Our partner is looking for a Senior Site Reliability Engineer (SRE, Compute Node Team) based in Spain.
This is a senior Site Reliability Engineering role focused on the infrastructure that runs and manages virtual machines across a large-scale cloud platform.
You will work close to the Linux operating system, hypervisor, and node-level services that form the foundation of the compute environment.
The role combines deep Linux systems engineering, virtualization, containerization, observability, and production reliability.
You will investigate complex issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance across user and kernel space.
You will also help shape reliability practices through strong monitoring, incident response, root-cause analysis, and postmortem processes.
This is an opportunity to influence critical compute infrastructure supporting demanding AI and cloud workloads at significant scale.
Ensure the reliability, availability, and performance of compute nodes responsible for running virtual machines.
Analyze and debug complex Linux systems across both user space and kernel space.
Work hands-on with virtualization technologies, primarily QEMU/KVM and Linux-native technologies.
Analyze VM lifecycle behavior, performance characteristics, resource utilization, and failure modes.
Support and improve containerized workloads using Linux-native mechanisms such as namespaces and cgroups.
Conduct structured root-cause analysis and develop corrective actions for recurring or systemic reliability issues.
Lead and contribute to postmortems focused on long-term reliability improvements rather than short-term remediation alone.
Investigate performance issues across multiple layers of the compute stack and develop practical engineering solutions.
Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
Deep expertise in Linux, including strong understanding of both user space and kernel space.
Knowledge of important Linux kernel subsystems, including scheduling, memory management, filesystems, cgroups, and namespaces.
Understanding of virtual machine lifecycles, performance characteristics, resource management, and failure modes.
Practical experience with containers, Linux namespaces, and cgroups.
Strong understanding of resource isolation, allocation, and control in containerized environments.
Strong understanding of the SRE discipline, including the relationship between software engineering, operations, reliability, and system design.
Experience building and operating observability stacks, rather than simply consuming existing monitoring dashboards.
Experience operating production systems and responding effectively to reliability and performance incidents.
Strong communication and collaboration skills when working with multidisciplinary infrastructure and engineering teams.
Experience with Kubernetes internals or node-level components is an advantage.
Hands-on experience with low-level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash dumps is beneficial.
Familiarity with large-scale compute or bare-metal infrastructure is a plus.
Contributions to open-source infrastructure or systems software are advantageous.
Collaborative and innovative engineering environment.
Opportunity to work on impactful AI and cloud infrastructure projects.
Exposure to large-scale compute, Linux systems, virtualization, and distributed infrastructure.
Opportunity to collaborate with highly skilled international engineering teams.
Workplace accommodations available during the application process where required.
📌 Senior Site Reliability Engineer (SRE, Compute Node Team) (España)
🏢 Jobgether
📍 España