Home/Jobs/AI Infrastructure Engineer
AI Infrastructure Engineer
Pokee AI
Singapore
3+ years
Today
$130K–150K/yr
Full-time
Remote
Skills Required
vLLM
TensorRT-LLM
Python
Rust
C++
Go
TensorRT
Triton
ONNX Runtime
Kubernetes
Docker
AWS
GCP
GPU computing
distributed systems
Description
Pokee AI is hiring an AI Infrastructure Engineer to build and optimize the systems behind RL-trained AI agents. The role focuses on scalable training, high-performance inference serving, and production infrastructure for enterprise use.
Company: Pokee AI
Role: AI Infrastructure Engineer
Location: Remote (US/Singapore Preferred)
Experience
- 3+ years of experience in ML infrastructure, ML platform engineering, or a related systems role
- Strong proficiency in Python and systems-level languages such as Rust, C++, or Go
- Hands-on experience with ML serving frameworks such as vLLM, TensorRT, Triton, or ONNX Runtime
- Experience with container orchestration and cloud infrastructure such as Kubernetes, Docker, AWS, or GCP
- Solid understanding of GPU computing, distributed systems, and performance profiling
- Familiarity with ML experiment tracking and pipeline orchestration tools such as MLflow, Weights & Biases, or Airflow
Responsibilities
- Build and optimize scalable training and inference infrastructure for RL-based AI agent models
- Optimize model serving for latency, throughput, and cost across cloud and on-device deployments
- Develop and manage CI/CD pipelines, experiment tracking, and model versioning systems
- Implement data pipelines for training data collection, preprocessing, and reward signal computation
- Collaborate with research scientists to productionize new algorithms and model architectures
- Ensure infrastructure meets enterprise requirements for reliability, security, and compliance
Additional Responsibilities
- Support cloud, on-premise, and on-device deployments
- Align infrastructure with enterprise reliability, security, and compliance needs
Nice To Have
- Experience with on-device or edge inference optimization such as GGUF quantization, TensorRT-LLM, CoreML, or QNN
- Familiarity with on-premise GPU deployments such as NVIDIA DGX, Dell PowerEdge, or Lenovo ThinkStation
- Experience supporting RL training loops or online learning systems in production
- Background in enterprise software with knowledge of security and compliance frameworks
- Contributions to open-source ML infrastructure projects
More Skills
performance profiling, MLflow, Weights & Biases, Airflow, CI/CD pipelines, experiment tracking, model versioning, data pipelines, reward signal computation, RL-based AI agent models, SOC 2, data residency
Prepare for this role
Recommended resources to build the skills for this position. Sponsored.
Top 15 Together AI Interview Questions
Zenaique
A Together AI-focused interview question set from Zenaique.
Top 25 xAI Interview Questions
Zenaique
An xAI-focused interview question set from Zenaique.
Top 25 LLM Evaluation Interview Questions
Zenaique
Curated LLM evaluation questions covering benchmarks, evals, and human review.
More vLLM jobs
AI Engineer
Ditto
San Francisco
4 days ago
Lead AI Engineer
Bridge-it
Delhi
6 days ago
Software AI Engineer
Tredence
Bengaluru
11 days ago
Staff Software AI Engineer
Equinix
Bengaluru
12 days ago
Senior AI/ML Engineer
AlifCloud IT Consulting Pvt. Ltd.
Pune
16 days ago
Senior System Software Engineer - Local AI
Nvidia
Pune
17 days ago