04 ago
|
Pragmatike
|
Madrid
ph3Overview /h3 pbLocation: /b Fully remote (EMEA timezone) /p pbStart date: /b ASAP /p pbLanguages: /b Fluent English required /p pbIndustry: /b Cloud Computing / AI / European Deep-Tech SaaS /p h3About The Role /h3 pPragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers. /p pWe are seeking a bAI Infrastructure Engineer /b with strong experience in production-grade model serving and infrastructure for AI systems. This is a highly technical, hands-on role focused on building scalable, reliable, and efficient ML inference platforms powering real-time AI applications. /p pYou will be responsible for designing and operating the core infrastructure that serves machine learning models at scale. You will work closely with infrastructure, platform, and applied AI teams to ensure high availability, low latency, and cost-efficient inference systems. Strong ownership, production mindset, and experience with distributed GPU systems are essential. /p h3Your Responsibilities /h3 ul liBuild and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent /li liDesign and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models /li liDevelop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers /li liOptimize GPU utilization,
memory efficiency, network throughput, and model artifact storage performance /li liDesign observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health /li liManage model registries and CI/CD pipelines enabling automated and reproducible model deployments /li liOwn the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities /li liDefine engineering best practices and contribute to platform scalability in a fast-moving startup environment /li /ul h3Required Qualifications /h3 ul li4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems /li liHands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent /li liStrong background in container orchestration and operating GPU-based workloads in production /li liExperience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines /li liProficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar) /li liStrong understanding of distributed systems, performance tuning,
and production reliability engineering /li liAbility to effectively use AI coding assistants to accelerate development and debugging workflows /li liOwnership mindset with the ability to operate independently in a remote-first environment /li /ul h3Preferred Qualifications /h3 ul liExperience with ML platforms such as Kubeflow, MLflow, or KubeAI /li liKnowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems /li liExperience with cost optimization across different GPU types and inference workloads /li liBackground in early-stage startups or greenfield infrastructure projects /li liProven experience building production systems from scratch rather than maintaining legacy platforms /li /ul h3Why Join Us /h3 ul liTake ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform /li liBuild foundational ML inference systems from the ground up in a high-growth, well-funded startup /li liWork at the intersection of distributed systems, GPU computing, and sustainable cloud architecture /li liGain deep expertise in next-generation AI infrastructure and large-scale model serving systems /li liInfluence core engineering decisions and define best practices that will scale with the company. /li /ul pPragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation. /p pIn accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration. /p /p #J-18808-Ljbffr
📌 AI Infrastructure Engineer (GPU) - Remote EMEA (Madrid)
🏢 Pragmatike
📍 Madrid