04 ago
|
Jobgether
|
España
ppbThis position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer (SRE, Compute Node Team) based in Spain. /b /ppThis role focuses on ensuring the reliability, scalability, and performance of advanced cloud infrastructure supporting AI-driven workloads. /ppYou will work at the intersection of Linux systems engineering, virtualization, and large-scale compute operations. /ppThe position offers the opportunity to design and improve foundational infrastructure used by developers and enterprises worldwide. /ppYou will take ownership of critical reliability challenges, from system debugging and observability to incident response and platform optimization. /ppWorking alongside experienced infrastructure and engineering teams, you will help shape the future of AI cloud platforms. /ppThis is an opportunity for a senior engineer who enjoys solving complex technical problems and building highly reliable systems at scale. /ph3Accountabilities /h3pThe Senior Site Reliability Engineer will be responsible for maintaining and improving the reliability of compute infrastructure, with a strong focus on Linux systems, virtualization, and operational excellence. /pulliEnsure the reliability, availability, and performance of compute nodes running virtual machines across cloud environments. /liliAnalyze and troubleshoot complex Linux systems across both user space and kernel space. /liliInvestigate and resolve production issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance. /liliWork hands‑on with virtualization technologies,
including QEMU/KVM and Linux‑native virtualization solutions. /liliDesign and improve observability capabilities at the infrastructure layer, including metrics, logs, traces, alerts, SLIs, and SLOs. /liliLead incident response activities, perform root‑cause analysis, and drive post‑incident improvements. /liliCollaborate with platform, kernel, hypervisor, GPU, and infrastructure teams to improve system architecture and operational reliability. /liliDevelop solutions that enhance scalability, performance, and maintainability of compute platforms. /liliContribute to engineering practices that promote automation, reliability, and continuous improvement. /li /ulh3Requirements /h3pThe adecuado candidate has extensive experience in site reliability engineering, Linux infrastructure, and distributed systems, with a strong ability to debug and optimize complex production environments. /pulliDeep expertise in Linux systems, including user space, kernel space, and core kernel subsystems such as scheduling, memory management, filesystems, cgroups, and namespaces. /liliStrong understanding of system architecture, boundaries, and performance trade‑offs across different infrastructure layers. /liliHands‑on experience with virtualization technologies, particularly QEMU/KVM,
including VM lifecycle management and performance optimization. /liliPractical experience with container technologies, namespaces, and resource isolation mechanisms. /liliStrong debugging and problem‑solving skills with a structured, hypothesis‑driven approach to incident investigation. /liliSolid understanding of SRE principles, including reliability engineering, system design, and operational ownership. /liliExperience building and operating observability solutions rather than only consuming monitoring tools. /liliAbility to translate system behavior into actionable reliability improvements. /liliExperience with Kubernetes internals, node‑level components, or large‑scale compute platforms is a plus. /liliFamiliarity with low‑level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash analysis is beneficial. /liliExperience with hardware‑level debugging, GPU infrastructure, NVLink, InfiniBand, or open‑source infrastructure projects is considered an advantage. /li /ulh3Benefits /h3ulliCompetitive compensation package. /liliFlexible working environment with ownership over projects and responsibilities. /liliOpportunities for career growth and continuous learning. /liliChance to work on impactful AI infrastructure projects. /liliCollaborative international environment with talented engineering teams. /liliOpportunity to solve challenging technical problems at large scale. /liliCulture focused on innovation, trust, and meaningful impact. /liliPossibility to contribute to the future of cloud infrastructure and AI technology. /li /ul /p #J-18808-Ljbffr
📌 Senior Site Reliability Engineer (SRE, Compute Node Team) (España)
🏢 Jobgether
📍 España