OverviewIn this role you enable data-science models to run reliably in production by owning the ML infra and operations on AWS and Kubernetes. You will design, deploy, monitor, and scale production ML services, collaborating with Data Scientists who retain model development. You’ll build CI/CD pipelines, IaC workflows, and robust APIs while enforcing security, observability, and cost-efficient operation. This position supports both existing products and new development, offering impact at scale and opportunities to mentor teammates.Compensaciones / Incentivos competitive benefits and salariespersonal and professional development opportunitiesflexibilitygrowth opportunitiescollaborative cultureResponsabilidades Lead complex MLOps projects with autonomy and judgmentDesign, develop, test, and debug software for customer-facing and internal appsOperate production ML services including deployment, scaling, and monitoringMaintain Kubernetes deployments and AWS infrastructure for model-serving workloadsDevelop and maintain CI/CD pipelines and infrastructure-as-code for model-serving systemsCreate and maintain APIs, orchestration layers, and data interfaces for production modelsCollaborate with Data Scientists to productionize models and resolve integration issuesBenchmark and stress-test ML/LLM services; ensure reliability and performanceEstablish logging, metrics, alerting, dashboards, incident response,
and cost-effective operationOperate MLflow and Kubeflow for lifecycle, pipelines, and workflowsEvaluate new technologies to strengthen systems and enforce engineering standardsMentor junior engineers and contribute to roadmaps and cross-team initiativesShare MLOps expertise and stay current with industry developmentsDevelop domain expertise in at least one cybersecurity application areaDocument technical approaches and decisionsRequisitos principales 7+ years in engineering or related roles with hands-on production operationsStrong cloud and software engineering foundations; MLOps experienceProduction architecture and API design experience in Python; Java or C++ a plusComprehensive MLOps background: containerized model serving, CI/CD, IaC, observability, release managementHands-on AWS and Kubernetes operations including Docker, IAM, networking, monitoring, and incident troubleshootingCI/CD ownership; Jenkins and ArgoCD are strong plusesExperience deploying ML services in production and performing benchmarking/load testingMLflow & Kubeflow experience (learning ability okay if not yet proficient)Knowledge of PyTorch, TensorFlow, scikit-learn for integrating model workCross-functional collaboration and project leadership experienceMentorship and strong communication skillsProblem-solving and risk management abilitiesCybersecurity domain interestcollaborationclear communicationmentorshipMLOpsKubernetesAWS
📌 Sr. Machine Learning Engineer (Madrid)
🏢 Fortra
📍 Madrid