
Posted 21 days ago
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
BerlinHybridFull-time
AI Summary
Senior SRE who will define reliability practices for a private Kubernetes platform running on own hardware, owning SLOs, incident response, automation, and observability while enabling Berlin development teams.
About this role
Introduction
At a glance
- Location & work model: Berlin, hybrid
- Tech stack: Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph
- Team: A growing SRE team – you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardware
- Process: Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team
- Languages: Fluent English required; German is a plus, not a must
Why this role is special
Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.
Your first 90 days
You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.Your mission
- Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions
- Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case
- Eliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks
- Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hard
- Plan capacity, performance and cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations
Your profile
Must-haves:
- Kubernetes in production – built, not just used: you've set up and maintained clusters on your own servers (e.g. kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experience isn't enough for this role
- Lived SRE practice: SLOs, error budgets, incident management, on-call
- Hands-on experience with GitOpsor comparable infrastructure/deployment automation – experience with Argo CD or Flux is a strong plus
- Solid observability skills – metrics, logs, traces, alerting that people trust
- A strong automation instinct – you'd rather fix a problem's cause than repeat its workaround
- A collaborative, enabling mindset – you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solution
- Harvester, KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms
- Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
- Auto-scaling (HPA, VPA, KEDA, cluster auto scaler) and capacity/cost planning
- Experience building Kubernetes operators/CRDs
- German language skills
You don't tick every box – or your title was never “SRE”? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you – production experience and engineering mindset matter more to us than titles or buzzwords.
Bonus Points
THE JOY OF WORKING WITH US
- Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
- Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.
- AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
- Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
- Flexible work: Hybrid work model three office days per week with a focus on outcomes.
- Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.
Job Location
Berlin, Munich, Pforzheim or Stockholm (all Hybrid)
Benefits
Skills
Argo CDCapacity PlanningCephCluster AutoscalerCRDError BudgetESXiFluxGitOpsGrafanaHarvesterHPAIncident ManagementKedaKubernetesKubernetes OperatorKubeVirtLonghornObservabilityOn-callOpenStackPrometheusSLISLOVPAVSphere
Explore related jobs
More jobs at Jobs at FactFinder
Similar Argo CD jobs
Jobs in Berlin
(Mini-/Studentenjob) Stewarding für Events, Gastro, Stadien und Hochzeiten (gn)Young Talents GmbH · Berlin, Berlin
(Mini-/Studentenjob) in der Logistik: Event-, Veranstaltungs-, und Messelogistik (gn)Young Talents GmbH · Berlin, Berlin
Aushilfe (gn) in der Logistik: Event-, Messe- und Veranstaltungsbereich (Mini-/Studentenjob)Young Talents GmbH · Berlin, Berlin
Senior Account Executive - Enterprise Sales (f/m/x)Netconomy GmbH · Berlin, Dortmund
Verkaufstalent Shop in Shop KaDeWe BerlinMy Jewellery · Berlin, Berlijn
Minijobber Shop in Shop KaDaWe BerlinMy Jewellery · Berlin, Berlijn
Browse these categories
Market data for devops / sre / platform engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.