- Kuala Lumpur Federal Territory Malaysia
Working Location
Job Description
Responsibilities
Job Summary:
We are looking for an AI Infra Engineer to support the development and operation of AI infrastructure platforms. The ideal candidate will have a strong background in Backend Engineering, Machine Learning Engineering, or MLOps, with hands-on experience in AI model training/inference deployment, cloud infrastructure, and GPU computing environments.
This role will work closely with algorithm teams to enable efficient utilization of AI computing resources, optimize AI workloads, and ensure reliable deployment of AI training and inference services.
Key Responsibilities:
• Operate and maintain GPU-based AI infrastructure, including AWS cloud and bare-metal GPU Kubernetes (K8s) clusters.
• Support GPU cluster setup, monitoring, troubleshooting, and incident resolution.
• Develop and maintain tools related to AI computing resource management and scheduling.
• Support algorithm teams in deploying, running, and optimizing AI model training and inference services.
• Improve AI workload efficiency through infrastructure optimization and automation.
• Build and maintain ML infrastructure pipelines to support AI development and production environments.
• Collaborate with engineering and algorithm teams to improve AI platform reliability and performance.
Requirements:
• Bachelor’s degree in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
• 1+ years of relevant working experience in Backend Engineering, Machine Learning Engineering, MLOps, Cloud Engineering, or related fields.
• Strong programming experience in Python or other backend development languages.
• Experience with Linux systems, Docker, and Kubernetes (K8s).
• Hands-on experience with cloud platforms, preferably AWS.
• Experience with AI model training, inference deployment, or ML production environments is highly preferred.
• Familiarity with AI/ML frameworks such as:
o PyTorch
o DeepSpeed
o Megatron
o vLLM
• Understanding of Large Language Model (LLM) training/inference workflows and distributed training concepts is a plus.
• Experience with GPU computing environments, resource scheduling, or AI infrastructure operation is highly preferred.
Preferred Qualifications:
• Experience in MLOps or ML Platform Engineering.
• Experience supporting GPU clusters or AI computing platforms.
• Understanding of distributed training strategies, including data parallelism, tensor parallelism, and pipeline parallelism.
• Strong problem-solving skills and ability to troubleshoot production issues.
• Passionate about AI technologies and capable of leveraging AI tools to improve engineering efficiency.
Important Information
Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.