Site Reliability Engineer (m/f/d) at Codesphere
Codesphere is hiring a Site Reliability Engineer (m/f/d) in Munich, Germany +1 more. Remote.
About Codesphere
Codesphere is a virtual cloud provider based in Germany. It offers a sovereign cloud platform that aims to free organizations from dependence on cloud providers and reduce operational complexity. The platform serves enterprises, governments, and regulated industries. Codesphere has partnered with Celonis for sovereign deployments and is involved in building an AI platform for the German federal digital ministry and the core of government tech marketplaces such as Deutschlandplattform and MEDI:CUS. Founded in Karlsruhe in 2020, the company has an international team of over 60 experts and offices in Karlsruhe and Munich.
Site Reliability Engineer (m/f/d) job description
About Codesphere
Codesphere is a Virtual Cloud Provider from Germany building the future of sovereign cloud infrastructure. Our platform gives enterprises and governments full sovereignty without giving up modern cloud capability β a vision recently validated by a series of multi-million European government tenders.
Since our founding in Karlsruhe in 2020, weβve expanded into an international team of 60+ experts. Based in Karlsruhe and Munich and backed by top-tier investors, we are chasing a bold vision.
Weβre scaling fast and would love for you to join us and grow alongside us π
What you'll drive
You define and enforce SLOs, SLIs, and SLAs across production
You monitor system health, plan capacity, and automate deployments, patching, and infrastructure provisioning
You diagnose and resolve production incidents fast β including 24/7 on-call participation
You lead post-mortems and turn findings into prevention; maintain runbooks and escalation procedures
You manage cloud infrastructure via IaC and own CI/CD pipeline design and maintenance
You drive scalability, fault tolerance, disaster recovery, and security compliance
You partner with Dev teams on production readiness, Shift Left practices, and error budget management
What makes you a great fit
Proven experience in an SRE, DevOps, or platform engineering role with hands-on production ownership
Strong knowledge of Kubernetes, Terraform, and Ansible
Familiarity with Ceph or comparable distributed storage systems
Experience with SLOs, SLIs, error budgets, and CI/CD pipeline design
Degree in a relevant field or comparable qualification
Calm, structured, and fast under pressure β strong debugging and incident response skills
Good communicator, able to translate operational concerns into guidance for Dev teams
Go development experience is a plus
What's in it for you
32 days of paid time off β 30 regular vacation days plus Christmas Eve and New Year's Eve off
Meal allowance β up to 15 digital vouchers per month, adding up to over β¬100 net for you
Flexibility β hybrid work setup with mobile work options and flexibility around core hours
Steep learning curve β fast-moving environment, real ownership, and a front-row seat to scaling a company
Job-Rad β lease a bike through us, tax-free
Gym access β stay active on site (Karlsruhe office only)
Employee events β from team offsites to regular get-togethers
Company pension scheme β company-supported pension to set you up for later
Great public transport links β both offices are within walking distance of tram and metro stops
Apply now
Applications go straight to Codesphere. We never sit between you and the employer.
Apply now β

