Technical Strategy: Define the AIGC roadmap across pre-training and post-training workstreams — from distributed training infrastructure to alignment, evaluation, and inference deployment.
Pre-training & Infrastructure: Oversee the design of distributed training toolchains for ultra-large-scale AIGC models. Drive system-level optimization across computation, communication, and storage layers. Ensure training stability and efficiency at scale.
Post-training & Alignment: Guide architecture design for video generation post-training — including high-quality instruction data curation, preference alignment (RLHF, DPO, GRPO, PPO), and video quality enhancement pipelines.
Capability Expansion: Push the frontier on long-video modeling, storyline consistency, precise camera control, and multi-modal generation.
Evaluation & Quality: Establish video quality evaluation frameworks and multi-dimensional Reward Models to systematically measure and improve output quality.
Inference & Deployment: Drive model distillation, quantization, and inference acceleration to bring research models into production.
Requirements
Masters and above in Computer Science or any related field.
At least 5 years of relevant experience in AI/ML, with a strong focus on generative models.
Leadership: Demonstrated experience managing and growing a technical team with direct reports. Track record of hiring, mentoring, and retaining top talent.
AIGC Depth: Hands-on experience in AIGC pre-training OR post-training (or both). Deep familiarity with Transformer architectures and Diffusion models (e.g., Stable Diffusion, Flux, DiT).
Distributed Systems: Strong understanding of distributed training principles (Data/Pipeline/Tensor/Expert Parallelism) and frameworks such as PyTorch, DeepSpeed, and Megatron-LM, OR
Post-training Expertise: Solid grasp of preference alignment methods (RLHF/DPO/GRPO/PPO), fine-tuning techniques (LoRA/QLoRA/DoRA), and distillation approaches (Consistency Models, Flow Matching).
Communication: Excellent cross-functional communication skills. Comfortable presenting to senior leadership and collaborating across engineering, product, and research teams.
Plus Points
Experience building a team or function from zero.
End-to-end ownership of the full lifecycle of a video generation model, from data to deployment.
Research leadership in physical simulation, world consistency, temporal consistency, or causal reasoning.
Expertise in high-quality video evaluation and human preference alignment at scale.
Publications at top-tier venues (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP).
Familiarity with GPU hardware architecture, CUDA programming, NCCL, and cuDNN.
Experience with extreme efficiency optimization such as inference acceleration, VRAM compression, quantization.