Site Reliability Engineer (Madrid)

Site Reliability Engineer (Madrid)

09 ago
|
Talent
|
Madrid

09 ago

Talent

Madrid

Overview¿Listo para inscribirse? Antes de hacerlo, asegúrese de leer todos los detalles pertenecientes a este trabajo en la descripción a continuación.At Tinybird, we help developers and data teams take flight by unlocking the power of real-time data to quickly build data pipelines and innovative data products. With Tinybird, you can ingest multiple data sources at scale, query and shape it using SQL, and publish results as low-latency, high-concurrency APIs for applications. Developers can create fast APIs quickly, enabling innovation and efficiency.About Tinybird: At Tinybird, we help developers and data teams take flight by unlocking the power of real-time data to quickly build data pipelines and innovative data products. With Tinybird, you can ingest multiple data sources at scale, query and shape it using SQL, and publish results as low-latency, high-concurrency APIs for applications.What you will be doingWe are looking for someone to help us scale and to keep our software and infrastructure reliable and elastic as we scale. You will participate as part of the on-call team, to understand not only our product, but also the issues our clients face.We run our stack in Linux. Technologies we use:OpenResty: SSL termination and load balancingVarnish: load balancing and cachingRedis: metadata storePython: most backend uses Python with some C++ for hot pathsClickHouse: main data storeZookeeper: replication coordination for ClickHouseGrafana, Loki and Mimir for monitoring and alertingTerraform: cloud provisioning (VMs, networks, Kubernetes clusters)Ansible: deploys software and configurationKubernetes: base of infrastructure with autoscalingWe operate a large-scale distributed system focusing on efficiency, building a self-service platform that adapts to workload changes and autoscales itselfYou’ll work with product and backend teams to design system architecture, optimize resource usage,



and improve elasticity and autonomyYou’ll need to understand how ClickHouse works to extract the best performanceSome challenges and things we want to improve:High availability and elasticity: the platform should scale automatically and efficiently without manual intervention, making capacity decisions transparently and safelyObservability: good understanding of storage, networking, and compute; monitoring of resources and service metricsDisaster recovery: better tooling, incident discovery, and on-call experienceWhat you bringExperience designing, building and running distributed cloud architectures and large-scale web applicationsProgramming skills and willingness to dive into our codebase, including ClickHouse source code; we work mainly with Python and C++Accountable and enthusiastic about owning and managing the platform, proactive about fixing issuesBias for action, iteration, and delivery; comfortable with quick reversals when neededSystems thinking with attention to edge cases, failure modes, behaviors, and implementationsComfortable collaborating asynchronously; expects direct daily team communicationBuild software with empathy, intuitive and maintainable; document key insights and solutions for easy understandingExperience with OpenResty, Varnish, Redis, Terraform or Ansible is helpful, but we expect you to recommend the right technology for each challengeExperience with ClickHouse or rolling out databases at scale is a plusDeep expertise in Kubernetes: designing and operating production-grade clusters, writing custom controllers, and tuning autoscaling (KEDA, Karpenter, etc.). Understand networking, storage, scheduling, and resource management, and reason about performance and failures at scaleProficiency with AWS and GCP cloud providersHow We WorkWe’re a fully remote company, committed to a remote-first culture. xcskxlj We have offices in Madrid and New York City; you can visit as it suits you.As we’re in the early stages, your contributions will have a significant impact on everything we do.We believe in transparency, so you’ll always be in the loop about what’s happening.Check out our blog or follow us on LinkedIn to learn more about what’s important to us.#J-18808-Ljbffr

📌 Site Reliability Engineer (Madrid)
🏢 Talent
📍 Madrid

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (madrid) / madrid

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: site reliability engineer (madrid) / madrid