Own the design, development, maintenance, and evolution of the in-house AIOps / ML / LLM platform, including related cloud and on-premise Kubernetes solutions.
Translate client, security, compliance, and internal requirements into practical platform designs with cross-functional teams.
Build and operate production ML / LLM workflows, including retraining, deployment, inference serving, monitoring, rollback, and optimisation.
Troubleshoot production issues across application, infrastructure, networking, Linux, Kubernetes, and ML serving layers.
Qualifications / Requirements
Strong software/platform engineering fundamentals, including system design, API design, distributed systems, scalability, reliability, observability, authentication/authorization, testing, and maintainable code design.
Practical understanding of the ML / LLM lifecycle, including data pipelines, model training/retraining, evaluation, experiment tracking, deployment, monitoring, and production feedback loops.
Strong development experience in Python, with working proficiency in Go and C++ for reading, debugging, maintaining, and extending existing production codebases.
Strong Linux, networking, and Kubernetes fundamentals, including production troubleshooting, service connectivity, ingress, resource limits, workload debugging, and deployment operations.
Experience designing, deploying, and operating production platforms on AWS, Azure, GCP, or on-premise environments.
Experience building CI/CD, automation, and MLOps / LLMOps workflows for production ML / LLM systems.
Strong communication skills and ability to work with AI, deployment, infrastructure, and security teams.
Good to Have
Deep experience operating Kubernetes in bare-metal, air-gapped, or restricted on-premise environments.
Experience with MLflow, Kubeflow, vLLM, TensorRT, TGI, or similar ML / LLM platform tools.
Exposure to TypeScript / React or Java-based services.