Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna LabsThe Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. The Acceleration Kernel Library team focuses on maximizing performance for AWS's custom ML accelerators, crafting high-performance kernels for ML functions at the hardware‑software boundary.Key Job ResponsibilitiesDesign and implement high-performance compute kernels for ML operations leveraging the Neuron architecture and programming models.Analyze and optimize kernel-level performance across multiple generations of Neuron hardware.Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.Implement compiler optimizations such as fusion, sharding, tiling, and scheduling.Work directly with customers to enable and optimize their ML models on AWS accelerators.Collaborate across teams to develop innovative kernel optimization techniques.A Day in the LifeBuild high-impact solutions for a general customer base.Participate in design discussions, code reviews, and communicate with internal and external stakeholders.Work cross‑functionally to help drive business decisions with technical input.Operate in a startup‑like development environment, solving the most important problems.Qualifications5+ years of professional software development experience (non‑internship).5+ years of programming in at least one software programming language.5+ years of leading design or architecture of scalable,
reliable systems.Experience as a mentor, technical lead, or engineering team leader.Full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.Bachelor’s degree in computer science or equivalent (preferred).Expertise in accelerator architectures for ML or HPC such as GPUs, CPUs, FPGAs, or custom architectures.Experience with GPU kernel optimization and GPGPU computing (CUDA, NKI, Triton, OpenCL, SYCL, ROCm).Proficiency in low‑level performance optimization and memory hierarchy understanding.Experience developing high‑performance libraries for HPC applications.Knowledge of ML frameworks (PyTorch, TensorFlow) and their GPU backends.Experience with LLVM/MLIR backend development for GPUs.Experience with parallel programming and optimization techniques.Understanding of GPU memory hierarchies and optimization strategies.Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Base salary range: CAN, ON, Toronto – 150,700.00 – 251,700.00 CAD annually. Compensation may include sign‑on payments, restricted stock units (RSUs), and other elements. Amazon offers comprehensive benefits including health insurance, a Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and resources to improve health and well‑being.Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.#J-18808-Ljbffr
📌 Sr. Ml Kernel Performance Engineer, Aws Neuron, Annapurna Labs - C$150,700 - C$251,700 A Year (Madrid)
🏢 Amazon
📍 Madrid