Founded in Switzerland in 1968, Zühlke is a team of colleagues across Europe and Asia, empowering ideas and creating new business models by developing services and products based on new technologies. While we work with the latest technologies on complex business challenges globally, our priority is to nurture what sets us apart: our people.
Working with us, you’ll be part of agile, collaborative teams with the opportunity to deliver transformational impact through technology, engineering excellence, and meaningful client outcomes.
The role
- Responsibilities include platform monitoring, incident troubleshooting and resolution, root-cause analysis, service reliability improvements, operational automation, runbook maintenance, and on-call support where required.
- The role will help maintain platform availability, stability, and performance while reducing manual operational effort.
What’s important to us
- Possess a degree in Computer Science/Information Technology or related fields.
- 5 to 7 years of strong software engineering experience with proficiency in at least one programming language, i.e. JavaScript, Java, Python, or NET. 2 to 4 years of hands-on experience supporting the reliability and availability of production systems in an SRE or production support environment.
- At least 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
- At least 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous.
- Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
- Proficiency with CI/CD workflows and GitHub Actions.
- Knowledge of artifact repository management systems such as Frog.
- Strong Linux administration skills and Shell scripting expertise.
- Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
- Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root-cause analysis, and post-incident follow-up.
- Experience in diagnosing and resolving complex system issues across multiple technology layers.
- Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
- Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
- Willingness to participate in an on-call support rotation for critical production services, where required.
- Able to optimize developer workflows and enhance developer experience.
- Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
- Excellent communication skills to work effectively with diverse engineering teams.
- Strong team-player mindset, focused on leveraging experience to help the team succeed.
- Possess positive learning and collaborative mindset.
- Strong analytical, problem-solving and troubleshooting skills.
- Good written and verbal communication skills.
- Agile, fast learner and able to adapt to changes.