ph3Location work modality: /h3 pEurope/ Remote /p h3Start: /h3 pAug 2026 /p h3Type of Contract: /h3 pFull time or Contract /p h3About Radian Arc /h3 pRadian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments. /p h3What Impact You Will Have /h3 pDesign, implement, and operate the network infrastructure powering the GPU cloud platform, including high-performance AI fabrics as well as classical datacenter networking components such as routing, security, and external connectivity. This role spans both high-performance east-west networking for distributed AI workloads and north-south connectivity, security, and inter-datacenter transport. /p pAs the first dedicated networking role in the organization, the Staff Network Engineer combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands‑on execution across design, deployment, troubleshooting, automation, and operational improvement. /p pThe Staff Network Engineer owns the long-term technical direction and operational strategy for Radian Arc’s AI interconnect networks, designing scalable GPU fabrics and ensuring predictable low-latency performance across distributed training and inference workloads. The role includes designing large-scale RoCE and Ethernet fabrics, guiding architecture decisions, and ensuring operational excellence across integral deployments, from hyperscale datacenters to smaller edge locations. /p pYou will collaborate closely with platform, compute, storage, observability, and operations teams to ensure networking is deeply integrated into the overall infrastructure architecture. This role also acts as the senior escalation point for complex networking incidents, driving deep technical investigations and systemic improvements that increase reliability, latency consistency, and operational maturity across the platform. /p pBecause this is currently the primary networking role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of technical direction, standards, cross-team influence, and long-term design, while also directly executing critical networking work that, in a larger organization, would be distributed across multiple engineers. /p h3What You’ll Do /h3 h3AI Fabric HPC Networking /h3 ul liDesign and operate high-performance GPU networking fabrics supporting distributed AI workloads. /li liArchitect large-scale RoCE fabrics optimized for distributed training and inference. /li liOptimize network performance for GPU communication patterns and east-west traffic. /li /ul pDesign fabric topologies such as: /p ul liLeaf-spine /li liFat-Tree /li liRail architectures /li liMulti-plane /li /ul pImplement high-performance networking technologies including: /p ul liRDMA /li liRoCE /li liHigh-bandwidth east-west fabrics /li liSpectrum-X /li /ul ul liCollaborate with compute teams to support distributed training frameworks and GPU communication libraries. /li liDefine reference architectures and design principles for AI fabrics so future deployments follow reusable standards rather than one-off implementations. /li liEvaluate architectural trade-offs across performance, resilience, cost, operability, and deployment speed, and make clear recommendations to stakeholders. /li /ul h3Datacenter Networking /h3 ul liDesign and operate Layer‑2 and Layer‑3 datacenter networks. /li liImplement scalable routing architectures based on BGP and ECMP. /li liDesign tenant network isolation mechanisms across multi‑tenant environments. /li liImplement and maintain: ul liNetwork bridges /li liRouting stacks /li liOverlay networking systems /li /ul /li liMaintain north‑south ingress/egress routing and traffic management. /li liDefine standards and reusable patterns for segmentation, routing, and overlay integration across platform deployments.
/li /ul h3Security Edge Connectivity /h3 ul liDeploy and maintain north‑south security infrastructure /li liImplement WAF and application‑layer protections /li liIntegrate security controls with platform services. /li /ul h3Technologies Include: /h3 ul liCitrix NetScaler / Citrix WAF /li liTLS termination /li liDDoS mitigation /li liAPI and proxy gateway protection /li /ul h3Inter‑Datacenter Networking /h3 ul liDesign and operate private interconnects between datacenters /li liImplement and maintain dark fiber ring architectures /li liOperate high-capacity WAN connectivity between regions /li liIntegrate datacenter fabrics into a global backbone network /li liDefine scalable design principles for backbone evolution, inter‑site routing, redundancy, and failure-domain isolation. /li /ul h3Technologies Include: /h3 ul liDWDM / dark fiber transport /li liBGP inter‑site routing /li liRedundant fiber ring architectures /li li100–400G optical transport /li liSpectrum-XGS /li /ul h3Engineering Execution Delivery /h3 ul liLead end-to-end engineering delivery of networking infrastructure, from design and labvalidation to production deployment /li liValidate network BOMs together with procurement and deployment teams /li liProvide detailed input into datacenter layouts and rack elevations /li liDrive capacity planning, performance modeling, and scaling strategies /li liEnsure network changes are executed safely with minimal customer impact /li liAct as both the architectural owner and the practical execution lead for critical network initiatives during the build‑out phase of the networking function /li liEstablish deployment standards, validation criteria, rollback approaches, and acceptance patterns that future engineers and teams can reuse. /li /ul h3Operational Excellence Reliability /h3 ul liOwn operational performance and reliability of networking infrastructure /li liDrive automation for: ul liProvisioning /li liConfiguration management /li liMonitoring /li liLifecycle management /li /ul /li liImprove day‑2 operations through automation and operational tooling /li liLead incident response and root‑cause analysis for major network events /li liDefine and track SLAs, SLOs, and reliability metrics /li liTranslate major incidents and operational pain points into durable standards, design changes, and long‑term architectural improvements /li liEstablish measurable benchmarks for reliability, latency consistency, operability, and recovery behavior across network deployments. /li /ul h3Cross‑Functional Collaboration /h3 ul liWork closely with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams. /li liProvide technical leadership across infrastructure initiatives. /li liCommunicate architectural decisions, trade‑offs, and risks clearly to stakeholders. /li liInfluence the long‑term platform networking roadmap and architecture. /li liAct as the primary networking design authority across the organization, guiding adjacent teams on how networking constraints and capabilities should shape platform decisions. /li liRaise the technical bar by mentoring engineers in adjacent domains and helping build the future networking function. /li /ul h3Technical Stack /h3 h3Datacenter Networking /h3 ul liBGP /li liEVPN / VXLAN /li liECMP /li liVLAN / VRF /li liOVS / OVN /li liLinux networking /li liBlueField DPU /li /ul h3Routing Control Plane /h3 ul liVyOS /li liBGP-based routing architectures /li liECMP fabrics /li /ul h3Security /h3 ul liCitrix NetScaler / WAF /li liDDoS protection /li /ul h3Transport Backbone /h3 ul liDark fiber /li liMetro fiber rings /li liDWDM transport /li li100–1600G optical networking /li /ul h3AI Networking /h3 ul liRDMA /li liRoCE /li liGPU fabrics /li liLarge-scale east‑west compute networking /li liCongestion control /li /ul h3What You'll Need /h3 h3Core Experience /h3 ul liStrong hands‑on experience designing and operating large-scale datacenter networks /li liExpert knowledge of modern networking protocols including:
ul liBGP /li liOSPF /li liECMP /li liEVPN / VXLAN /li /ul /li liProven experience operating high-speed Ethernet networks in production environments /li liExperience operating NVIDIA / Mellanox networking platforms /li liExperience owning both architecture and direct implementation in lean or fast‑scaling environments is strongly preferred. /li /ul h3Advanced AI Fabric Networking Expertise /h3 pThe candidate should have deep expertise in designing and operating networking fabrics optimized for large-scale GPU clusters and distributed AI workloads. This includes a strong understanding of GPU communication patterns and the networking requirements of distributed training and inference systems. /p ul liDeep understanding of NCCL communication patterns and their impact on network topology and performance. /li liExperience tuning RoCE fabrics for large-scale GPU clusters. /li liStrong knowledge of RDMA transport behavior and failure modes. /li liPractical experience implementing and tuning PFC and ECN for congestion management. /li liUnderstanding of GPU collective communication patterns such as all‑reduce, all‑gather, broadcast, reduce‑scatter, and their impact on east‑west network traffic. /li liExperience designing rail‑optimized GPU networking fabrics for distributed training and inference clusters. /li liFamiliarity with diagnosing performance issues related to: ul liNCCL stalls /li liRDMA congestion /li liFabric hotspots /li liPacket loss impacting distributed training /li /ul /li liUnderstanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow. /li /ul pThe candidate should also be able to collaborate closely with compute platform teams to ensure that networking infrastructure is optimized for distributed training, distributed inference, and high-throughput AI workloads. /p h3Systems Troubleshooting /h3 ul liAbility to debug complex cross-layer issues spanning: ul liHardware /li liFirmware /li liKernel networking /li liDistributed application communication layers /li /ul /li liStrong knowledge of networking hardware, optics, and high-speed interconnects. /li liExperience designing network observability systems. /li liStrong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues. /li /ul h3Automation /h3 ul liStrong automation skills using Python and/or Bash. /li liExperience applying software engineering practices to infrastructure automation. /li liExperience building reusable tooling, standards, or validation approaches that increase leverage across teams. /li /ul h3Leadership /h3 ul liProven ability to lead complex technical initiatives across teams. /li liComfortable collaborating across engineering, operations, and vendors. /li liStrong systems‑level thinking balancing performance, reliability, scalability, and operational cost. /li liDemonstrated ability to set architectural direction and drive adoption of engineering standards across an organization. /li liProven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority. /li liStrong mentoring capability and ability to raise the technical level of adjacent engineering teams. /li liAble to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency. /li /ul h3What We Offer /h3 ul liAttractive compensation package reflecting your expertise and experience. /li liA great work environment characterised by friendliness, international diversity, flexibility, and a hybrid‑friendly approach. /li liYou'll be part of a fast‑growing scale‑up with a mission to make a positive impact, offering an exciting career evolution. /li /ul pOur job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands. /p h3Our inclusive responsibility /h3 pRadian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law. /p /p #J-18808-Ljbffr
📌 Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc (EMEA)
🏢 Submer
📍 España