- Deploy, operate, monitor and troubleshoot Kubernetes (K8s) clusters and containerised workloads to ensure stable production environments.
- Develop automation scripts and internal tools using Python, Java or Go to reduce manual workload.
- Manage and maintain databases, including routine maintenance, performance tuning, backup and recovery, and fault resolution.
- Build and maintain observability systems, including log collection, metrics monitoring and performance tracking.
- Collaborate with development teams to optimise CI/CD workflows and improve delivery efficiency.
- Perform daily system checks, incident handling, root cause analysis, and implement optimisation plans.
Required Qualifications & Core Skills
Minimum 5 years of experience in DevOps, system operations, or cloud-native environments.
Hands-on experience deploying, managing and troubleshooting Kubernetes clusters.
Proficiency in at least one programming language: Python, Java or Go.
Experience operating or maintaining AI/ML systems and related infrastructure.
Fluent Mandarin for daily communication and strong English for documentation.
Good understanding of Linux, networking and cloud-native architecture.
Strong problem-solving and troubleshooting skills.