We are looking for a Cloud Operations & DevOps Engineer to maintain overseas cloud platforms, ensure service reliability, resolve production incidents, and improve operational efficiency through automation.
Key Responsibilities:
- Operate and maintain cloud platforms and core systems.
- Monitor system health and address performance bottlenecks.
- Deploy platform updates and manage operational changes.
- Respond to production incidents, restore services, and conduct root-cause analysis.
- Develop automation tools and improve operational processes.
- Maintain technical documentation, SOPs, and incident records.
- Participate in a 7×24 On-Call rotation.
Requirements:
- Bachelor’s degree in computer science, IT, or a related field.
- At least two years of relevant production operations experience.
- Strong hands-on Linux administration and troubleshooting skills.
- Experience with Docker, Kubernetes, and containerized environments.
- Production experience with AWS or Azure.
- Proficiency in Python or Shell scripting.
- Experience with system deployment, monitoring, incident response, and root-cause analysis.
- Working proficiency in Mandarin and English.
- Strong communication, service mindset, and willingness to learn.
Experience with Ansible, Terraform, CI/CD tools, Go, C/C++, Prometheus, Grafana, ELK, Splunk, or Datadog will be advantageous.