- Kuala Lumpur Federal Territory Malaysia
Working Location
Job Description
Responsibilities
Responsibilities:
• Responsible for daily SRE operations, including CI/CD, capacity planning, system monitoring, alerting, incident response, and troubleshooting.
• Ensure high system availability, stability, performance, and security while continuously identifying opportunities to optimize infrastructure and reduce computing costs.
• Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.
• Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.
• Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams.
• Collaborate closely with development and infrastructure teams to identify potential reliability risks and implement proactive solutions.
• Participate in on-call rotations and respond to critical incidents when required to maintain system availability and business continuity.
Requirements:
• Bachelor’s degree or above in Computer Science, Information Technology, or a related field.
• At least 2 years of relevant experience in SRE, DevOps, Cloud Infrastructure, System Operations, or a related field.
• Strong understanding of Linux, TCP/IP, Kubernetes, databases, SQL, Shell scripting, and Python.
• Hands-on experience with at least one major cloud platform, such as AWS, Tencent Cloud, or Microsoft Azure.
• Experience with CI/CD pipelines, system monitoring, alerting, troubleshooting, and infrastructure automation.
• Familiarity with AI-assisted development and troubleshooting tools, such as Claude Code and Codex, with the ability to leverage AI tools to improve development and operational eAiciency.
• Good command with the ability to communicate eAectively with cross-functional and technical teams.
• Strong sense of ownership, problem-solving ability, and initiative, with a proactive mindset toward improving system reliability and operational eAiciency.
• Willingness and ability to participate in 24/7 on-call support for critical incidents when required.
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.