Job Summary
You will lead the product and project lifecycle of smart computing and AI cluster offerings, coordinating cross-functional teams to deliver operational excellence and continuous improvements that drive business success.
Responsibilities
Own the end-to-end product and project lifecycle of smart computing and AI cluster offerings, ensuring successful launch and ongoing operations while building a reusable knowledge base
Align stakeholder expectations across business, product, engineering, and operations teams to coordinate execution and deliver post-launch support
Translate business and technical requirements into clear, actionable operational plans that meet compute, data, and network milestones on time and with quality
Drive the product roadmap by setting plans, tracking progress, and ensuring key milestones are met on schedule
Proactively identify issues and risks, implementing solutions and mitigations before escalation
Continuously enhance product operations and cross-team workflows by establishing more efficient processes to close operational gaps
Preferred competencies and qualifications
Bachelor’s degree or higher in engineering, computer science, or a related field preferred
5-8 years of experience in product operations, program management, or related cross-functional roles
Experience with AI compute, high-performance computing (HPC), big data, or cloud platforms is a plus
Familiarity with cloud product operations (IaaS, PaaS, MaaS) and workflows
Experience with GPU cluster operations or AI platform operations preferred
Strong logical communication and stakeholder management skills, with a dependable, detail-oriented, and accountable approach under pressure
ScitiX – Delivering Intelligent Computing for the Future of AI.
ScitiX is a next-generation provider of intelligent computing infrastructure, dedicated to empowering artificial intelligence innovation through high-performance, scalable solutions. Built on deep technical expertise and large-scale cluster operation experience, ScitiX offers a full-stack platform spanning IaaS, PaaS, and MaaS, delivering end-to-end support for AI development workflows.
At the core of ScitiX’s offering are high-performance GPU clusters capable of supporting large-scale model training with ultra-low latency interconnects and optimized storage bandwidth. Leveraging multi-cluster management capabilities, ScitiX delivers serverless GPU cloud services with minute-level billing, enabling flexible and efficient resource utilization tailored to diverse AI workloads.
The company’s all-in-one AI platform, SiFlow , integrates data ingestion, preprocessing, distributed training, fine-tuning, and inference into a unified system. Designed for both researchers and enterprise users, SiFlow significantly reduces infrastructure complexity, allowing algorithm engineers to focus entirely on model innovation.
ScitiX also delivers Model as a Service (MaaS), providing access to a wide range of pre-trained models and tools for evaluation, inference, and deployment. This approach accelerates the development cycle, lowers entry barriers for AI adoption, and supports rapid iteration in complex application scenarios.
With a proven track record in managing large-scale GPU clusters at scale, ScitiX delivers reliable, secure, and high-availability services backed by automated monitoring, fault detection, and intelligent scheduling systems. The platform is deployed across Tier-3+ data centers globally, ensuring consistent performance and business continuity for mission-critical AI applications.