Bachelor's degree or above in Computer Science, Artificial Intelligence, Natural Language Processing (NLP), Machine Learning, Distributed Systems, or a related technical field.
Hands-on experience delivering or contributing to end-to-end LLM pre-training projects.
Proven experience participating in the pre-training of models with 7B+ parameters.
Strong hands-on expertise with distributed training frameworks such as Megatron-LM, DeepSpeed, or FSDP.
Practical experience working with large-scale training environments involving 64+ GPUs.
Strong understanding of large-scale pre-training data pipelines, including data cleaning, deduplication, quality filtering, tokenization, data mixing, and data quality optimization.
Experience designing or optimizing data pipelines for large-scale LLM training.
Strong understanding of large-scale pre-training data pipelines, including data cleaning, deduplication, quality filtering, tokenization, data mixing, and data quality optimization.
Experience designing or optimizing data pipelines for large-scale LLM training.
Familiarity with long-context training and context-extension techniques, including RoPE scaling, NTK-aware interpolation, and YaRN.
Preferred Qualifications
Experience with 70B+ parameter model pre-training.
Experience with Mixture-of-Experts (MoE) model pre-training.
Experience optimizing large-scale GPU clusters, distributed training systems, and AI training infrastructure.
Publications in top-tier AI/ML conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP, particularly in areas related to LLM pre-training, model architecture, or training optimization.