- Jalan Sarawak Bandar Kuala Lumpur WP Kuala Lumpur Malaysia 55100
Working Location
Job Description
Requirements
5+ years of operations experience; 2+ years leading a team
Expert in Linux and Shell/Python/Ansible automation
Familiar with multi-stack deployment, CI/CD, and containerization (Docker/K8s)
Familiar with mainstream databases and monitoring systems; capable of security hardening and troubleshooting
Strong communication, coordination, and project-driving skills
Responsibilities
Key Responsibilities
• Infrastructure Automation: design and standardize the team's automation scripts (Bash/Python) and Ansible Playbook standards; drive IaC adoption; review automation proposals; significantly reduce manual-ops workload
• Multi-Stack Environment Management: plan and maintain Web services and application architectures; support deployment and tuning across Java, Python, Go, PHP, Node.js (Nginx, Tomcat, Docker, K8s)
• Cloud Resources & CDN Management: manage multi-cloud environments (Alibaba, Tencent, Huawei, AWS) with cost and capacity control; lead CDN configuration for Web traffic security, acceleration, and HA
• Domain & Certificate Systems: manage all DNS resolution and filings; build a unified SSL certificate issuance, deployment, monitoring, and auto-renewal mechanism to prevent expiry incidents
• Cloud Security: lead server and business security-protection strategies (DDoS, brute-force, WAF, intrusion detection); coordinate security-incident response
• DevOps Modernization: design and implement containerized CI/CD pipelines (Docker, Jenkins, GitLab CI); define build/test/deploy workflow standards for release consistency and efficiency
• Code & Repository Management: manage GitLab repositories, permissions, and branch strategies; standardize team development and release processes
• Network Security: define and implement hardening for firewalls, security groups, proxies, and multi-protocol VPN (OpenVPN, WireGuard) to keep remote access secure and reliable
• Monitoring & Alerting: build and optimize monitoring/alerting platforms (Zabbix, Prometheus, Grafana, ELK); refine metrics, alert rules, and incident-response mechanisms
• Database Operations: manage deployment, backup, replication, performance tuning, and HA of mainstream databases (MySQL, Redis, MongoDB, PostgreSQL)
• Disaster Recovery: design backup and recovery strategies; run DR drills to ensure fast recovery and minimal downtime for critical systems
• Team Management: assign tasks, mentor junior engineers, own the on-call and incident-response process, and report ops status and improvement plans upward
Benefits
Skills
LRT - MASJID JAMEK
0.5 km
LRT - PLAZA RAKYAT
0.7 km
LRT - DANG WANGI
0.8 km
MRL - BUKIT NANAS
0.8 km
MRT - MERDEKA
0.9 km
LRT - BANDARAYA
0.9 km
KTM - BANK NEGARA
1.0 km
MRL - MEDAN TUANKU
1.0 km
MRT - PASAR SENI
1.1 km
LRT - PASAR SENI
1.1 km
MRL - RAJA CHULAN
1.1 km
MRT - BUKIT BINTANG
1.2 km
MRL - BUKIT BINTANG
1.2 km
MRL - HANG TUAH
1.2 km
MRL - IMBI
1.3 km
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.