164 KALLANG WAY Central Region (Singapore) Singapore
职位描述
岗位职责
Responsibilities
Design pre-training technical strategies, including data composition, training hyperparameters (LR schedule, batch size ramp-up, progressive sequence length strategies), etc.
Lead pre-training data engineering, including data cleaning (deduplication, PII filtering, quality classification), tokenization, data mixing strategies, and curriculum learning.
Configure and optimize large-scale distributed training, including 3D parallelism, GPU memory optimization, and communication optimization.
Develop and implement long-context training solutions, including RoPE scaling, NTK-aware interpolation, and YaRN.
Monitor and optimize training performance, including loss curve analysis, gradient anomaly detection, training stability, and checkpoint management.
Conduct regular intermediate evaluations and adjust data composition and training strategies based on evaluation results.
Responsibilities
Bachelor’s degree or above in AI, NLP, Computer Science, Systems, or a related field.
Proven end-to-end experience in large language model (LLM) pre-training projects, with hands-on experience pre-training models of 7B+ parameters.
Strong expertise in large-scale distributed training frameworks, such as Megatron-LM, DeepSpeed, or FSDP, with hands-on experience training on 64+ GPUs.
Strong understanding of pre-training data engineering, including data cleaning, deduplication, quality filtering, tokenization, and data mixture optimization.
Familiarity with training monitoring and debugging, including loss analysis, gradient monitoring, and training stability management.
Familiarity with long-context training techniques, including RoPE scaling, NTK-aware interpolation, and YaRN.