04 ago
|
Tamarind Intelligence
|
Alcobendas
04 ago
Tamarind Intelligence
Alcobendas
ph3Tasks /h3pYou will have bFull Ownership of platform operations /b, serving as the operational backbone that keeps our Kubernetes-based ecosystem running smoothly and troubleshooting issues observed by platform users. /ppYour main tasks will include: /pullibPlatform Operations Stability: /b Take full responsibility for running, maintaining, and maintaining the operational health of our Kubernetes-based cloud Big Data platform. /lilibTroubleshooting L2/L3 Incident Response: /b Act as the primary point of contact for platform users (data engineers, analysts, internal teams). You will dive deep into application logs, investigate container failures, analyze pod behavior, and diagnose performance bottlenecks or process crashes. /lilibUser Application Operations: /b Support and maintain internal user applications running on top of Docker and Kubernetes. Ensure deployed services are healthy, scalable, and resilient. /lilibRoot Cause Analysis (RCA): /b Apply an investigation-first mindset to identify why platform or application failures occur, preventing recurring incidents rather than just applying temporary fixes. /lilibAgile Collaboration Support: /b Work in 2-week sprints, addressing operational tickets, incidents, and platform health tasks while keeping team members and internal stakeholders aligned. /lilibContinuous Operational Improvement: /b Continuously refine operational runbooks, monitoring setups, and incident response definitions to improve mean time to resolution (MTTR). /li /ulh3Qualifications /h3ullibExperience: /b Minimum 5 years of proven experience in Systems Operations, L2/L3 Application Support, Site Reliability Engineering (SRE), or Infrastructure Operations. /lilibKubernetes Docker (Core): /b Deep hands-on experience operating, debugging, and troubleshooting containerized workloads in Kubernetes and Docker environments (e.g., pod lifecycle, ingress/service issues, resource limits, volume mounts). /lilibTroubleshooting Mindset: /b Exceptional diagnostic skills. You enjoy reading logs,
analyzing metrics, tracing application failure modes, and asking the right questions to solve user-reported problems. /lilibLinux Fundamentals: /b Confident command-line usage, high-level handling of Linux systems, process management, log analysis, and system recovery. /lilibOperational DevOps Tooling: /b Practical experience managing deployments and platform states using tools like ArgoCD, Helm, or Terraform from an operations perspective. /lilibObservability Open Source Stack: /b Hands-on experience using monitoring and logging stacks (e.g., Prometheus, Grafana) to diagnose issues. Familiarity with applications like Airflow, Trino, Hive, or OpenShift is a strong plus. /lilibNetworking Basics: /b Practical understanding of network concepts (DNS, TCP/IP, load balancers, network policies) to troubleshoot connection or traffic routing issues. /lilibMindset: /b Strong user-empathy, autonomy to take ownership of production issues, and a supportive team player attitude focused on knowledge sharing. /lilibCommunication: /b Good English skills (written and spoken) to collaborate effectively with platform users and international teams. /li /ulh3What We Offer /h3ullibHybrid work model /b (in any of our offices in Alcobendas, Valencia, Barcelona, Sevilla, or Logroño) /lilibFlexible working hours /b /lilibFlexible compensation package /b /lilibDiscounts on Arsys products /b /lilibDiscounts on technology brands /b /lilibChallenging technical environment /b with real high-performance system problems /lilibAccess to advanced AI tools /b with no usage restrictions /lilibClose collaboration with senior profiles /b and a strong culture of continuous improvement /lilibPersonalized training and development plans /b /lilibBiannual team events /b /li /ulpWe value diversity and welcome all applications - regardless of, for example, gender, nationality, ethnic or social origin, religion, disability, age, as well as sexual orientation and identity, physical characteristics, marital status, or any other irrelevant factor subject to applicable law. /p /p #J-18808-Ljbffr
📌 Senior Platform Operations Engineer (DevOps - Ops Focus) (Alcobendas)
🏢 Tamarind Intelligence
📍 Alcobendas