04 ago
|
Mirantis
|
Barcelona
04 ago
Mirantis
Barcelona
ppMirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.br/Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen. Learn more at /ph3Job Description /h3h3Role Overview: /h3We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The idóneo candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure.h3Key Responsibilities: /h3ulliDesign, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.
/liliTroubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance. /liliManage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations. /liliPerform performance tuning, monitoring, and capacity planning for HPC networking systems. /liliImplement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer). /liliDiagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments. /liliCollaborate with compute, storage, and platform teams to support HPC workloads and cluster operations. /liliDevelop and maintain documentation for network architecture, configurations, and operational procedures. /liliParticipate in on-call rotations and provide escalation support for critical incidents. /liliLead or contribute to network upgrades, migrations, and new deployments. /li /ulh3Qualifications /h3h3Required: /h3ulli5+ years of experience in network engineering, with a focus on HPC or data center environments. /liliStrong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA). /liliSolid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
/liliProven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies). /liliExperience with network performance analysis and troubleshooting tools. /liliFamiliarity with Linux systems and scripting for automation (e.g., Bash, Python). /liliStrong analytical and problem-solving skills. /li /ulh3Preferred: /h3ulliExperience with large-scale HPC clusters or AI/ML infrastructure. /liliKnowledge of RDMA, MPI, and low-latency networking concepts. /liliCertifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent. /liliExperience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform). /li /ulh3Soft Skills: /h3ulliStrong communication and collaboration skills. /liliAbility to work independently and handle complex technical challenges. /liliDetail-oriented with a proactive approach to problem-solving. /li /ulh3Additional Information /h3h3We offer: /h3ulliOperate some of the most advanced AI infrastructure environments in production today. /liliWork with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments. /liliHelp define operational standards and reliability practices for next-generation AI infrastructure services. /liliInfluence the adoption of AI-powered operational capabilities through k0rdent AI. /liliWork alongside highly skilled engineers solving complex infrastructure and platform challenges at scale. /liliJoin a growing organisation investing heavily in AI infrastructure, platform services, and operational innovation. /li /ulpWe are a Leader for Container Management in G2 (#2 after AWS)! /p /p #J-18808-Ljbffr
📌 Principal HPC Network Engineer (remote in the EU) (Barcelona)
🏢 Mirantis
📍 Barcelona