ABOUT THE ROLE
We are looking for an experienced AI Engineer to design, build, deploy, and operate production-grade AI applications powered by Large Language Models (LLMs). The ideal candidate has hands-on experience taking AI solutions from prototype to production, with a strong understanding of modern AI infrastructure, Retrieval-Augmented Generation (RAG), evaluation frameworks, observability, and cloud AI platforms.
You will work closely with product managers, software engineers, and platform teams to build scalable, reliable, and secure AI-powered products.
KEY RESPONSIBILITIES:
AI Application Development
- Design, develop, and deploy production-ready AI applications using Large Language Models (LLMs).
- Own the complete AI application lifecycle—from experimentation and development to deployment, monitoring, and continuous improvement.
- Build scalable AI systems capable of handling real-world production traffic with high reliability and low latency.
- Collaborate with cross-functional teams to convert business requirements into AI-powered solutions.
Production AI & LLM Engineering
- Build and optimize Retrieval-Augmented Generation (RAG) pipelines for enterprise use cases.
- Develop robust document ingestion pipelines supporting structured and unstructured data sources.
- Implement embedding pipelines, vector databases, chunking strategies, metadata enrichment, and retrieval optimization.
- Design prompt engineering strategies and optimize prompt performance across different LLMs.
- Fine-tune open-source or proprietary LLMs where appropriate to improve domain-specific performance.
- Optimize inference cost, latency, throughput, and response quality.
AI Infrastructure & Platform Engineering
- Work with cloud AI platforms such as AWS AI/Bedrock, Azure AI Foundry, Google Vertex AI
- Integrate managed AI services with enterprise applications.
- Build scalable inference pipelines and model-serving infrastructure.
- Configure secure access to AI services, APIs, and enterprise data sources.
AI Evaluation & Quality Assurance
- Design and implement automated evaluation frameworks for LLM applications.
- Define and monitor AI quality metrics including Faithfulness, Relevance, Groundedness, Hallucination Detection, Answer Correctness, Latency and Cost Per Request.
- Build offline and online evaluation pipelines.
- Perform A/B testing and continuously improve AI application performance using evaluation insights.
AI Observability & Monitoring
- Implement end-to-end observability for AI systems.
- Monitor Model Performance, Prompt Quality, Retrieval Effectiveness, Token Usage, Latency, Error Rates, Cost, User Feedback.
- Build dashboards, alerts, and monitoring pipelines for production AI applications.
- Identify and troubleshoot production issues affecting AI application quality.
AI Gateway & Optimization
- Implement AI gateways for routing requests across multiple LLM providers.
- Build intelligent fallback mechanisms and model routing strategies.
- Implement semantic caching and response caching to reduce latency and inference costs.
- Optimize token utilization and API usage.
Data Ingestion & Knowledge Systems
- Design scalable ingestion pipelines for enterprise knowledge bases.
- Process PDFs, Office documents, web content, databases, APIs, and other structured/unstructured sources.
- Build data transformation, indexing, and synchronization pipelines.
- Ensure data quality, freshness, and governance within AI systems.
Software Engineering & APIs
- Develop RESTful APIs to expose AI capabilities.
- Integrate AI services with existing enterprise applications.
- Collaborate with frontend developers to build intuitive AI-powered user experiences.
- Follow software engineering best practices, including testing, version control, CI/CD, and code reviews.
DevOps & Cloud Deployment
- Deploy AI applications using modern cloud-native practices.
- Build CI/CD pipelines for AI workloads.
- Manage infrastructure, scalability, security, and production releases.
- Collaborate with DevOps teams to ensure high availability and operational excellence.
MANDATORY QUALIFICATIONS
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, or a related field.
- Proven experience building, deploying, and operating at least one production AI application.
- Demonstrated experience monitoring and supporting AI applications in production environments.
- Hands-on experience with AI data ingestion pipelines, AI evaluation frameworks, AI observability and monitoring.
- Experience with at least one major AI cloud platform such as AWS (including Bedrock or SageMaker), Azure AI Foundry, Google Vertex AI.
- Strong understanding of LLM fine-tuning techniques, Production-grade Retrieval-Augmented Generation (RAG), AI caching strategies, AI gateways and model routing, Evaluation methodologies, Observability and monitoring.
- Experience working with vector databases, embeddings, and semantic search.
- Strong software engineering fundamentals with proficiency in Python.
- Experience building and consuming REST APIs.
- Familiarity with Git, CI/CD, containerization, and cloud-native development.
GOOD TO HAVE
- AWS Certified Developer – Associate or higher (or equivalent Azure/GCP certification).
- Experience with cloud infrastructure, Kubernetes, Docker, and infrastructure-as-code.
- Experience deploying AI applications at scale.
- Knowledge of frontend technologies (React, Angular, Vue.js, or similar).
- Experience integrating AI applications with enterprise systems.
- Familiarity with agentic AI frameworks such as LangGraph, AutoGen, CrewAI, Semantic Kernel, or similar.
- Experience with LLMOps tooling such as LangSmith, Arize Phoenix, MLflow, Weights & Biases, Promptfoo, or similar.
- Knowledge of AI security, guardrails, responsible AI, and governance practices.
- Experience optimizing AI systems for cost, latency, scalability, and reliability.
PREFERRED EXPERIENCE
- 2–4+ years of software engineering experience with at least 1 years focused on production AI/LLM applications.
- Experience delivering enterprise-grade AI solutions from proof of concept through production.
- Strong understanding of distributed systems, cloud architecture, and scalable application design.
- Ability to balance AI quality, operational cost, performance, and user experience when designing solutions.