Dive deep into machine learning methods, including large language models (LLMs), multimodal models, and image/video generation models, adapting these methods to TikTok Live scenarios and serving as foundational models and features.
Currently pursuing a Master's degree in Computer Science.
Familiarity with large language models and its applications
...
Leverage AWS services (SageMaker, ECR, EKS/ECS, Lambda, Step Functions, S3, CloudWatch, CloudFormation, Terraform, etc.) to host, scale, and manage model training and inference pipelines.
Develop monitoring and alerting solutions for model latency, accuracy, data drift, and infrastructure health; integrate with Prometheus, Grafana, CloudWatch, or similar tools.
Automate model versioning, artifact storage, and metadata tracking using Mlflow or SageMaker model registry.
...
Architect and implement solutions leveraging AWS services (SageMaker, ECR, EKS/ECS, Lambda, S3, CloudWatch, CloudFormation, Terraform, etc.) to host, scale, and manage model training and inference pipelines.
Develop comprehensive monitoring and alerting solutions for model latency, accuracy, data drift, and infrastructure health; integrate with Prometheus, Grafana, CloudWatch, or similar tools.
Oversee the automation of model versioning, artifact storage, and metadata tracking using Mlflow or SageMaker model registry.
...
Collaborate closely with IT operations and DevOps teams to ensure the smooth integration of infrastructure platforms with other applications and processes.
Establish and maintain systems for monitoring machine learning models in production. Oversee the development of machine learning models at MoneyLion to ensure proper model governance.
Effectively manage cloud infrastructure costs by monitoring and optimizing spending, and provide transparency and accountability in cost-related matters.
...
You will also be responsible for training stability and reliability. This includes identifying the root causes of loss spikes, divergence, slow nodes, communication bottlenecks, checkpoint failures, and data-related instability, as well as designing mechanisms for fast checkpoint recovery and automatic exclusion of problematic nodes.
The ideal candidate has strong hands-on experience with PyTorch distributed training and a solid understanding of CUDA architecture, GPU memory hierarchy, NCCL communication, and performance profiling.
You should have source-level familiarity with at least one major large-scale training framework, such as Megatron-LM, DeepSpeed, PyTorch FSDP, or TorchTitan, and be comfortable reading, modifying, and debugging framework internals.
...