jobs in ASUS GLOBAL PTE. LTD.

全职 ASUS AICS SG - Machine Learning Engineer 工作, 薪水 up to SGD 13,000, ASUS GLOBAL PTE. LTD. East Region (Singapore) 公司招聘中 - Ricebowl

ASUS AICS SG - Machine Learning Engineer

ASUS GLOBAL PTE. LTD.

Bedok, East Region (Singapore)

分享
保存

工作地点

  • Bedok East Region (Singapore) Singapore

职位描述

岗位职责

About ASUS

AICS is part of ASUS, a multinational company known for the world’s best motherboards, PCs, monitors, graphics cards and routers. Along with an expanding range of superior gaming, content-creation and AIoT solutions, ASUS leads the industry through cutting-edge design and innovations made to create the most ubiquitous, intelligent, heartfelt and joyful smart life for everyone. With a global workforce that includes more than 5,000 R&D professionals, ASUS is driven to become the world’s most admired innovative leading technology enterprise.

About AICS

The mission of ASUS Intelligent Cloud Services (AICS) is to build revolutionary healthcare solutions with natural language processing, computer vision, and big data analytics. We provide Software as a Service (SaaS) applications to accelerate the effective use of medical data and improve the efficiency of hospital operations, unleashing the power of data for precision healthcare and bringing transformative impact to the industry.

Job Overview

We are looking for an experienced

Machine Learning Engineer (ML Ops Engineer)

to build and operate the infrastructure that enables reliable, scalable, and efficient AI model deployment.

You will work closely with ML Engineers, AI Researchers, Software Engineers, and Product Teams to manage the production lifecycle of AI models, including model evaluation, release, deployment, monitoring, and updates.

The role focuses on GPU infrastructure, model serving, model lifecycle management, and AI platform reliability.

Responsibilities

GPU & Compute Resource Management

Design and operate infrastructure for efficient GPU resource allocation and utilization across AI workloads.

Manage GPU workloads in Kubernetes and containerized environments.

Implement resource scheduling, quotas, priorities, and workload isolation for multiple AI workloads.

Monitor GPU utilization, capacity, performance, and resource consumption.

Optimize GPU utilization and inference efficiency as workloads scale.

Troubleshoot GPU, container, networking, and infrastructure issues in production.

Model Lifecycle & Evaluation

Build and maintain processes for model versioning, evaluation, release and rollback.

Establish automated workflows to evaluate new model versions against defined quality, performance, and reliability criteria.

Design model release gates to ensure new models meet predefined requirements before production deployment.

Compare model versions across metrics such as model quality, latency, throughput, and resource consumption.

Support controlled model rollout, including canary deployment, A/B testing, and rollback.

Maintain model metadata, evaluation results, deployment history, and release status for traceability.

Model Serving &Deployment

Build and operate reliable infrastructure for self-hosted AI model inference and serving.

Deploy and optimize LLM inference services using vLLM or similar inference engines.

Optimize model serving for latency, throughput, GPU utilization, and reliability.

Automate model deployment and configuration changes across environments.

Implement deployment strategies that minimize service disruption during model updates.

Monitor and troubleshoot production inference workloads.

Model Gateway & AI Platform

Build and maintain model gateway / model routing infrastructure that provides a unified interface to multiple AI models and inference backends.

Support model routing, traffic management, authentication, rate limiting, and observability.

Enable applications and AI agents to consume models through a consistent and reliable interface.

Integrate different model providers and self-hosted inference services into a unified platform.

Platform Reliability &Observability

Build monitoring and observability for AI workloads, including:

GPU utilization and health

Inference latency and throughput

Model quality metrics

Service availability

Model version and deployment status

Resource consumption

Establish logging, metrics, tracing, alerting, and operational dashboards.

Investigate production incidents and perform root-cause analysis.

Continuously improve system reliability, scalability, and operational efficiency.

Requirements

4+ years of experience in software engineering, MLOps, ML infrastructure, DevOps, or a related field.

Strong programming skills in Python and/or Go.

Hands-on experience operating production AI/ML infrastructure.

Strong experience with Docker and Kubernetes.

Solid understanding of GPU infrastructure and resource management.

Experience with GPU scheduling, resource allocation, monitoring, or capacity planning.

Experience operating model serving or inference infrastructure in production.

Familiarity with LLM inference and serving, preferably with hands-on experience using vLLM.

Experience designing or operating model evaluation and release processes.

Experience with CI/CD and infrastructure automation.

Strong understanding of Linux, networking, distributed systems, and cloud-native infrastructure.

Experience with monitoring and observability.

Strong troubleshooting and problem-solving skills.

Ability to collaborate effectively with ML researchers, ML engineers, and software engineers.

Nice to Have

Experience building or operating a Model Gateway / AI Gateway.

Experience with model routing, traffic management, rate limiting, and multi-model serving.

Experience with vLLM internals and performance tuning, such as batching, KV cache, GPU memory utilization, and concurrency.

Experience with Kubernetes GPU scheduling and NVIDIA GPU infrastructure.

Experience operating on-premises GPU clusters.

Experience with LLM / Generative AI / Agentic AI infrastructure.

Experience implementing canary deployment, A/B testing, or automated model rollback.

Experience building internal AI / ML platforms used by multiple teams.

Experience with Infrastructure as Code such as Terraform.

Experience working in healthcare or other regulated environments is a plus.

ASUS is a global technology leader delivering incredible experiences that enhance the lives of people everywhere. World renowned for continuously reimagining today’s technologies for tomorrow, ASUS puts users first In Search of Incredible to provide the world’s most innovative and intuitive devices, components, and solutions. Today’s ASUS is more ambitious than ever, unleashing remarkable gaming, content-creation, AIoT, and cloud solutions that solve user needs and infuse delight.

ASUS is home to industry-leading experts who are encouraged to pursue their passion for innovation and entrepreneurial spirit to deliver the future of technology to the world.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多