Our client is a technology-driven organization based in Singapore, scaling its AI capabilities across core products. They are investing in a robust ML Platform to move from experimentation to reliable, cost-efficient production.
The Opportunity
Our client is looking for a Senior MLOps Engineer to be the key engineer responsible for productionizing ML and GenAI workloads. This is a senior Individual Contributor role - high ownership, deep hands-on, working directly with Data Science, Data Engineering, and Platform teams.
You will be the go-to expert for "how we ship models to production reliably."
Key Responsibilities
1. ML Lifecycle & Productionization
- Build and maintain CI/CD pipelines for ML: training, validation, versioning, deployment, and automated rollback
- Productionize batch and real-time models, including LLM / RAG based services
- Manage Model Registry, Experiment Tracking (MLflow / W&B), and Feature Store
2. ML Platform & Infrastructure
- Build scalable ML infrastructure on AWS / GCP using Docker, Kubernetes (EKS / GKE), Terraform
- Implement Infrastructure as Code (Terraform) and GitOps best practices
- Optimize compute - GPU/TPU, spot instances, inference cost and latency
3. Reliability, Observability & Governance
- Implement monitoring for data drift, concept drift, model performance, and pipeline health using Prometheus / Grafana / CloudWatch
- Ensure reproducibility, data/model versioning, security, and auditability
- Partner with Security & Compliance on model governance and Responsible AI practices (PDPA)
4. Collaboration
- Work closely with Data Scientists to optimize training jobs and inference pipelines
- Work with Data Engineers on feature pipelines and data validation
- Define and document MLOps best practices, templates, and runbooks
Requirements
Must-Have:
- Experience in Software Engineering / DevOps / MLOps, with at least 2+ years focused on MLOps
- Strong Python, solid software engineering (Git, CI/CD, unit testing, code reviews)
- Hands-on experience taking ML models to production in a commercial setting
- Strong Cloud experience (AWS preferred: SageMaker, S3, EKS, Lambda) + Kubernetes + Docker
- Experience with orchestration: Airflow / Prefect / Kubeflow
- Experience with MLflow, model packaging, and model serving (REST/gRPC, Seldon / BentoML / FastAPI)
Good to Have:
- Experience with Feature Store (Feast) and Vector DBs (Pinecone, Weaviate, pgvector) for RAG
- LLMOps: deployment and evaluation of LLMs, prompt versioning, guardrails
- Streaming: Kafka / Kinesis
- Spark / Data processing at scale
What Success Looks Like in 6 Months
- ML training and deployment pipelines are automated and standardized
- Model monitoring and drift detection is live for all production models
- Inference cost/latency is optimized and documented