jobs in NEXadept

全职 LLM Inference Performance Engineer 工作, 薪水, NEXadept 公司招聘中 - Ricebowl

LLM Inference Performance Engineer

NEXadept

Undisclosed

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

About the Role

We are looking for an LLM Inference Performance Engineer to optimize large-scale LLM inference performance across TPU/GPU and other AI accelerators.

You will work across LLM inference, kernels, compilers, and runtime systems, improving latency, throughput, scalability, and overall inference efficiency.


Responsibilities

  • Optimize LLM inference performance on TPU/GPU and other AI accelerators
  • Develop and optimize inference backends, kernels, and runtime components
  • Optimize Attention, GEMM, KV Cache, Sampling, fused kernels and other performance-critical workloads
  • Work with technologies such as JAX, XLA, Pallas, CUDA, Triton or related frameworks
  • Build benchmarking and profiling tools to identify performance bottlenecks
  • Collaborate with model, inference, compiler, and hardware teams to improve production performance


Requirements

  • Bachelor’s degree or equivalent experience in CS, Engineering, ML, Systems, or related fields
  • Experience in LLM inference, ML systems, hardware acceleration, or performance optimization
  • Strong programming skills in C++ and/or Python
  • Understanding of GPU/TPU architecture, memory behavior, kernels, or ML workload performance
  • Experience with performance profiling and benchmarking

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多