- Kuala Lumpur Federal Territory Malaysia
工作地点
职位描述
岗位职责
Job Summary:
We are looking for an AI Infra Engineer to support the development and operation of AI infrastructure platforms. The ideal candidate will have a strong background in Backend Engineering, Machine Learning Engineering, or MLOps, with hands-on experience in AI model training/inference deployment, cloud infrastructure, and GPU computing environments.
This role will work closely with algorithm teams to enable efficient utilization of AI computing resources, optimize AI workloads, and ensure reliable deployment of AI training and inference services.
Key Responsibilities:
• Operate and maintain GPU-based AI infrastructure, including AWS cloud and bare-metal GPU Kubernetes (K8s) clusters.
• Support GPU cluster setup, monitoring, troubleshooting, and incident resolution.
• Develop and maintain tools related to AI computing resource management and scheduling.
• Support algorithm teams in deploying, running, and optimizing AI model training and inference services.
• Improve AI workload efficiency through infrastructure optimization and automation.
• Build and maintain ML infrastructure pipelines to support AI development and production environments.
• Collaborate with engineering and algorithm teams to improve AI platform reliability and performance.
Requirements:
• Bachelor’s degree in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
• 1+ years of relevant working experience in Backend Engineering, Machine Learning Engineering, MLOps, Cloud Engineering, or related fields.
• Strong programming experience in Python or other backend development languages.
• Experience with Linux systems, Docker, and Kubernetes (K8s).
• Hands-on experience with cloud platforms, preferably AWS.
• Experience with AI model training, inference deployment, or ML production environments is highly preferred.
• Familiarity with AI/ML frameworks such as:
o PyTorch
o DeepSpeed
o Megatron
o vLLM
• Understanding of Large Language Model (LLM) training/inference workflows and distributed training concepts is a plus.
• Experience with GPU computing environments, resource scheduling, or AI infrastructure operation is highly preferred.
Preferred Qualifications:
• Experience in MLOps or ML Platform Engineering.
• Experience supporting GPU clusters or AI computing platforms.
• Understanding of distributed training strategies, including data parallelism, tensor parallelism, and pipeline parallelism.
• Strong problem-solving skills and ability to troubleshoot production issues.
• Passionate about AI technologies and capable of leveraging AI tools to improve engineering efficiency.
重要安全守则
申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。