- Develop and maintain AI evaluation frameworks, regression test suites and automated evaluation pipelines.
- Build datasets, quality dashboards and evaluation tools to improve the performance of AI agents and LLM-powered applications.
- Design evaluation methods such as LLM-as-a-Judge, multi-turn conversations, tool evaluation and agent workflow analysis.
- Integrate AI evaluation into CI/CD and model release processes.
- Analyse production issues and improve datasets, prompts, evaluation logic and AI workflows.
- Conduct experiments to optimise model accuracy, retrieval performance, latency and cost.
- Build reusable platform capabilities and developer tools to support multiple AI products.
- Prepare technical documentation and collaborate with cross-functional engineering teams.
Requirements
- Bachelor's Degree in Computer Science, Software Engineering or a related discipline.
- Minimum 2 years of relevant experience in AI Platform Engineering, LLMOps, Machine Learning Engineering or Software Engineering.
- Experience with most of the following technologies: LangSmith, LangGraph / LangChain, FastAPI, Temporal, Grafana, Redis, React / TypeScript, EvalsHub or similar AI evaluation tools
- Experience with LLM evaluation, Retrieval-Augmented Generation (RAG), semantic search, embeddings or model benchmarking is preferred.
- Strong analytical and problem-solving skills.
- Good communication and stakeholder management skills.
- Able to work independently and collaboratively in a fast-paced environment.
Interested candidates are encouraged to submit their resumes outlining their relevant experience and achievements to sherriyang(@)************* or click apply!
**We regret to inform that only shortlisted candidates would be notified**
EA License No: 04C3537
EA Personnel No: R22106683
EA Personnel Name: Yang Hui Shan, Sherri