15 sep
|
Enfint
|
Barcelona
Описание: Codeway is a general consumer tech company that develops and scales mobile apps across creativity, productivity, wellness, language learning, and entertainment. Its products serve more than 400 million users worldwide.
Задачи Define, instrument, and report on SLIs, SLOs, and error budgets across critical services
Own observability end-to-end, including metrics, logs, traces, dashboards, and alerting
Reduce alert noise and false positives
Run reliability reviews and an error-budget policy
Operate, scale, and upgrade the multi-cluster Kubernetes environment
Act as the escalation point for cluster and platform issues
Own capacity planning, performance, and cloud cost efficiency
Build self-service platform tooling for product teams
Embed security through RBAC, least privilege, secrets management, scanning, network policies, and patching
Partner with the security function on vulnerability remediation, audit readiness, and secure‐by‐default infrastructure
Own disaster recovery and validate RTO/RPO targets through drills and failure testing
Contribute to architecture and production‐readiness reviews
Lead the on‐call rotation and act as incident commander during production incidents
Run blameless postmortems and track corrective actions through closure
Build and maintain Infrastructure as Code with Terraform and CI/CD pipelines
Enforce GitOps and progressive delivery with automated rollbacks
Identify, measure, and eliminate operational toil through automation
Establish and mature reliability capabilities during the first 12 months
Establish shared security measurement and a common view of operational health across engineering and security
Connect platform performance, reliability, delivery speed, security, and infrastructure cost to business operations
Provide metrics that support company investment decisions
Требования 5–8 Years of experience in SRE, Platform,
or DevOps roles
Experience operating high‐traffic, always‐on production systems at meaningful scale
Hands‐on production Kubernetes experience, including upgrades, autoscaling, and troubleshooting under load
Strong cloud engineering background
Solid Linux and networking fundamentals
Experience defining and operating with SLOs and error budgets
Experience with Infrastructure as Code and CI/CD pipeline design
Depth in observability tooling, instrumentation, dashboarding, and alert design
Security‐first mindset with experience in least privilege, secrets hygiene, and vulnerability management
Fluency in scripting and automation in at least one language
Incident‐command experience, including on‐call ownership and blameless postmortems
Ability to communicate clearly with engineers and leadership under pressure
Nice to have: high‐scale consumer or mobile app backends, AI/ML inference workloads, GitOps, progressive delivery, service mesh, API gateways, multi‐region or multi‐cluster topologies, cloud cost optimization, FinOps, SOC 2, ISO 27001, GDPR, DevSecOps, chaos engineering, resilience testing, Kubernetes/cloud/DevOps certifications, supporting many independent services and teams
Условия Competitive compensation package
Meal compensation
Unlimited private health insurance and HPV vaccine coverage
Pet adoption support covering primary healthcare expenses, parasite vaccinations, and microchip costs during the first year
MacBook, iPhone 15 Pro, Magic Mouse, Magic Keyboard, adjustable desk, 4K screen, and other required gadgets
Gym membership support
Flexible schedule
English course support
Office located in central Barcelona at Edifici Estel
Coffee shop, healthy snacks, free breakfast, and free lunch
No dress code
Gaming area with PS5
Software subscription support
Additional monthly compensation for public transportation
Case study may be included in the recruiting process
📌 sre engineer for high-availability services (Barcelona)
🏢 Enfint
📍 Barcelona