ph3About Radian Arc /h3 pRadian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments. /p h3What Impact You Will Have /h3 pMission: Design, build, and operate the AI storage layer powering large-scale GPU infrastructure, enabling datasets, model artifacts, checkpoints, and inference state to be delivered to compute clusters with extremely high throughput and predictable latency. /p pYou will play a key role in architecting and evolving the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads including distributed inference, fine‑tuning, and large-scale model training. The role spans multiple storage architectures used across the platform, including hyperconverged storage currently based on StorPool, local NVMe storage for latency-sensitive workloads and edge deployments, and disaggregated AI storage platforms such as VAST Data and Weka. /p pAs the first dedicated storage platform role in the organization, this position combines Staff-level architectural ownership, technical direction, and cross‑functional influence with hands‑on execution across storage design, deployment, performance engineering, troubleshooting, platform integration, and operational improvement. /p pA key responsibility of this role is designing and optimizing the storage architecture underlying distributed inference stacks such as NVIDIA Dynamo, llm-d, or similar inference orchestration frameworks. This includes ensuring that storage systems efficiently support inference workloads through optimized dataset access, model artifact distribution, checkpoint handling, and KV-cache persistence. You will design scalable storage systems capable of feeding thousands of GPUs while balancing throughput, latency, resilience, and cost efficiency, and work closely with compute, networking, and platform engineering teams to ensure seamless integration with the platform orchestration layer. /p pBecause this is currently the primary storage platform role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of long‑term design, standards, cross‑team influence, and platform direction, while also directly executing critical storage work that, in a larger organization, would be distributed across multiple engineers. /p h3What You’ll Do /h3 h3Storage Architecture /h3 ul liDesign scalable AI storage architectures supporting both edge and core deployments. /li liDefine storage strategies for distributed inference, fine‑tuning, and training workloads. /li liArchitect solutions across multiple storage models: ul liHyperconverged infrastructure such as StorPool, /li liLocal NVMe storage, /li liDisaggregated storage systems such as VAST, Weka, and related architectures. /li /ul /li liDefine reference architectures, design principles, and reusable patterns for storage platforms so future deployments follow standards rather than one‑off implementations. /li liEvaluate trade‑offs across throughput, latency, resilience, data locality, cost, and operability, and make clear recommendations to engineering and leadership. /li liInfluence the long‑term storage roadmap, including architecture choices for edge, core, hyperconverged, and disaggregated environments. /li /ul h3AI Workload Optimization /h3 ul liOptimize storage throughput and latency for GPU‑heavy clusters. /li liDesign data locality strategies to minimize dataset movement across the network. /li liBenchmark storage performance under real AI workloads. /li liOptimize I/O patterns for large dataset ingestion, checkpointing, and model artifact distribution. /li liWork directly with compute teams to ensure storage architecture matches the access patterns of distributed training, fine‑tuning, and inference frameworks. /li liEstablish performance baselines and validation methods so storage platforms are tested against realistic AI workload behavior rather than only synthetic benchmarks. /li /ul h3Platform Integration /h3 ul liImplement and maintain CSI drivers. /li liIntegrate storage platforms with Kubernetes and orchestration systems. /li liIntegrate block, object, and shared file storage into the platform. /li liDesign multi‑tenant storage architectures supporting isolated workloads. /li liEnsure storage capabilities are correctly exposed into platform services, workload orchestration, and lifecycle automation.
/li liDefine standards for how storage should be integrated into Kubernetes‑based and platform‑managed environments across different deployment models. /li /ul h3Distributed Storage Systems /h3 ul liContribute to the design of exabyte‑scale storage platforms. /li liSupport S3‑compatible object storage, distributed file systems, and block storage. /li liIntegrate storage clusters into heterogeneous customer environments. /li liDesign storage systems with clear fault domains, lifecycle management approaches, scaling paths, and operational boundaries. /li liDefine reusable operating patterns for multi‑cluster and multi‑site storage environments. /li /ul h3Distributed Inference Storage Architecture /h3 ul liDesign the storage architecture supporting distributed inference platforms such as NVIDIA Dynamo, llm-d, or similar frameworks. /li liOptimize storage performance for large‑scale LLM inference workloads. /li liDesign efficient strategies for KV‑cache persistence and retrieval using distributed storage platforms such as VAST or Weka. /li liOptimize storage access patterns for token generation pipelines and high‑concurrency inference workloads. /li liEnsure inference infrastructure scales efficiently across thousands of GPUs. /li liPartner with platform and inference teams to ensure storage design supports evolving inference architectures and avoids becoming a bottleneck in throughput, latency, or concurrency. /li /ul h3AI Data Path Optimization /h3 ul liDesign high‑performance data paths between GPU clusters and distributed storage. /li liOptimize performance using technologies such as: ul liGPU Direct Storage, /li liRDMA / RoCE, /li liNVMe‑oF. /li /ul /li liEnsure predictable latency for inference serving workloads. /li liDefine architectural approaches for storage‑to‑GPU data movement that balance performance gains with operational complexity and deployment practicality. /li /ul h3Performance Engineering /h3 ul liWork with technologies such as the following, to maximize data throughput to GPU clusters: ul liRDMA, /li liRoCE, /li liGPU Direct Storage, /li liSPDK, /li liNVMe‑oF, /li liLead storage performance investigations across hardware, network, OS, filesystem, and workload interaction points. /li /ul /li liDrive systematic tuning of storage paths for large‑scale GPU environments and define repeatable validation and benchmarking approaches for future deployments. /li /ul h3Reliability Operations /h3 ul liImprove the reliability, durability, and observability of the storage stack. /li liCollaborate with operations teams to monitor storage systems using telemetry and metrics. /li liOptimize performance, latency, and resilience of storage infrastructure. /li liLead incident response and root‑cause analysis for major storage events and chronic performance issues. /li liTranslate operational pain points and incidents into durable design changes, standards, runbooks, and architectural improvements. /li liEstablish measurable benchmarks for storage reliability, performance consistency, recovery behavior, and operability across deployments. /li /ul h3Engineering Execution Delivery /h3 ul liLead end‑to‑end engineering delivery of storage infrastructure from architecture and validation through production rollout. /li liSupport practical implementation of storage platforms in both new deployments and existing environments. /li liValidate storage BOMs and architecture assumptions together with infrastructure, compute, and deployment teams. /li liContribute detailed input into datacenter layouts, node profiles, and storage topology decisions. /li liDrive scaling strategies, capacity planning, and storage lifecycle decisions. /li liEnsure storage changes are executed safely with minimal customer impact. /li liAct as both the architectural owner and the practical execution lead for critical storage initiatives during the build‑out phase of the storage function. /li /ul h3Cross-Team Collaboration /h3 ul liWork closely with compute, networking, platform, DevOps, and operations teams. /li liEnsure storage integrates seamlessly into the AI platform architecture. /li liAct as the primary storage design authority across the organization, guiding adjacent teams on how storage constraints and capabilities should shape platform decisions. /li liCommunicate architectural decisions, trade‑offs, risks, and operational implications clearly to stakeholders. /li liShare knowledge and mentor engineers on high‑performance storage design. /li liRaise the technical bar by helping adjacent teams better understand storage behavior in distributed AI environments.
/li /ul h3Technical Stack /h3 ul liCSI. /li liNVMe / NVMe‑oF. /li liDistributed file systems. /li liObject storage. /li liLinux storage stack. /li liRDMA / RoCE. /li liGPU Direct Storage. /li liSPDK. /li liStorPool. /li liVAST Data. /li liWeka. /li liMinioFS. /li liRook Ceph. /li /ul h3What You’ll Need /h3 h3Core Experience /h3 ul liStrong hands‑on experience designing and operating distributed storage systems for high‑performance compute environments. /li liProven experience designing storage architectures for large‑scale AI inference or training platforms, including dataset distribution, checkpointing, and KV‑cache storage patterns. /li liDeep knowledge of the Linux storage and I/O stack. /li liStrong understanding of AI workload data access patterns. /li liExperience optimizing storage for GPU‑accelerated workloads. /li liFamiliarity with Kubernetes storage integrations such as CSI. /li liExperience operating large‑scale storage clusters. /li liExperience owning both architecture and direct implementation in lean or fast‑scaling environments is strongly preferred. /li /ul h3Advanced AI Storage Expertise /h3 pThe candidate should have deep expertise in designing and operating storage platforms optimized for GPU‑heavy environments and distributed AI workloads. /p pThis includes a strong understanding of how training, fine‑tuning, and inference systems interact with storage, and how storage architecture affects throughput, latency, concurrency, checkpoint recovery, dataset distribution, and serving performance. /p h3Relevant Expertise Includes /h3 ul liStrong understanding of storage access patterns for distributed inference and training. /li liExperience designing storage platforms that support large dataset ingestion and model artifact distribution at scale. /li liPractical experience tuning storage architectures for checkpointing, distributed file access, object access, and high‑concurrency inference. /li liFamiliarity with storage patterns for KV‑cache persistence and retrieval. /li liExperience optimizing data locality and reducing unnecessary network movement between storage and compute. /li liUnderstanding of how storage performance affects large‑scale AI frameworks, model‑serving systems, and inference orchestration layers. /li /ul h3Systems Troubleshooting /h3 ul liAbility to debug complex cross‑layer issues spanning: /li ul liStorage hardware, /li liNetworking, /li liLinux kernel and I/O paths, /li liFilesystems, /li liObject and block storage layers, /li liKubernetes integrations, /li liDistributed workload behavior. /li /ul liStrong knowledge of storage hardware, NVMe devices, storage fabrics, and high‑performance data paths. /li liExperience designing storage observability systems. /li liStrong ability to act as the senior escalation point for ambiguous, high‑impact, and multi‑domain technical issues. /li /ul h3Automation /h3 ul liStrong automation skills using Python and/or Bash. /li liExperience applying software engineering practices to storage automation and operational tooling. /li liExperience building reusable tooling, standards, validation patterns, or lifecycle automation that increase leverage across teams. /li /ul h3Leadership /h3 ul liProven ability to lead complex technical initiatives across teams. /li liComfortable collaborating across engineering, operations, deployment teams, vendors, and platform stakeholders. /li liStrong systems‑level thinking balancing performance, reliability, scalability, operability, and cost efficiency. /li liDemonstrated ability to set architectural direction and drive adoption of engineering standards across an organization. /li liProven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority. /li liStrong mentoring capability and ability to raise the technical level of adjacent engineering teams. /li liAble to balance short‑term execution needs with long‑term platform design, operational sustainability, and cost efficiency. /li /ul h3What We Offer /h3 ul liAttractive compensation package reflecting your expertise and experience. /li liA great work environment characterised by friendliness, international diversity, flexibility, and a hybrid‑friendly approach. /li liYou’ll be part of a fast‑growing scale‑up with a mission to make a positive impact, offering an exciting career evolution. /li /ul h3Our inclusive responsibility /h3 pRadian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law. /p /p #J-18808-Ljbffr
📌 Staff Storage Platform Engineer (AI Storage) - Radian Arc (España)
🏢 Submer
📍 España