About the company:
We are in the tech industry, specializing in cutting-edge AI solutions and advanced platform development. With a mission to empower businesses through intelligent automation and data-driven insights, we foster a collaborative and forward-thinking environment where creativity and technical excellence thrive.
Job responsibilities:
Own the end-to-end build of an AI model service platform, starting from zero, engineered to grow into a system that can reliably serve user bases in the hundreds of thousands to millions.
Architect and continuously scale a unified layer of model API interfaces, overseeing how requests are routed across models, how load gets balanced, and how the service expands over time — all while keeping AI-driven services running reliably for connected hardware.
Deep working knowledge of today's leading AI models and their APIs
Solid backend development skills, with experience in cloud deployment, microservices, and inference frameworks, and the ability to iteratively scale and improve systems in production.
Business-savvy and outcome-driven, comfortable partnering across Agent teams, hardware teams, and other functions to keep the platform's direction tied to broader company goals.
Requirements:
Bachelor's degree in Computer Science, Software Engineering, or a related field
Strong backend development experience (Java, Go, Python, or similar)
Hands-on experience with cloud platforms (AWS, Azure, GCP, or Alibaba Cloud)
Experience with microservices architecture and API design
Familiarity with AI/LLM APIs (OpenAI, Anthropic, Qwen, etc.) and inference frameworks (e.g., vLLM, Triton, TGI)
Experience building systems that scale to high user volumes
Understanding of load balancing, service routing, and distributed systems
Good communication skills; able to work with cross-functional teams (Agent, hardware, product)