Technical Lead, Machine Learning
The role
We need someone who can take research ideas and turn them into ML systems that actually work in production — reliably, at scale, day after day. You'll sit right at the crossroads of research, infrastructure, and product, making sure models don't just perform well in a notebook but hold up under real users, real load, and real constraints.
What you'll be doing
- Owning ML systems end-to-end — data pipelines, training workflows, evals, inference, deployment. The whole pipe, not just one slice of it.
- Fine-tuning and adapting models with modern techniques (LoRA, QLoRA, SFT, DPO, distillation) — whatever gets the best result for the constraints we're working with.
- Building inference systems that don't force you to choose between fast, cheap, and reliable — or at least finding the right trade-off when you do.
- Keeping our training data (synthetic and real) clean, well-structured, and useful.
- Building evaluation pipelines that check not just "does it perform well" but "is it safe, robust, and fair" — working closely with research leadership on this.
- Getting models into production and keeping them there: GPU optimization, memory efficiency, latency, scaling — the unglamorous work that makes everything else possible.
- Working shoulder-to-shoulder with application engineers so ML doesn't feel bolted on to our backend, mobile, and desktop products — it feels native.
- Making smart, pragmatic calls and shipping fast, then learning from how things actually get used in the wild.
What success looks like
- Research turns into shipped, production-ready work — with real targets, not vague hopes.
- Our pipelines and inference systems are stable and easy to maintain — not held together with duct tape.
- When something breaks in production, it gets caught and fixed fast, before users really feel it.
- The people around you feel supported and unblocked, not stuck waiting on you.
- Every iteration on our models makes the product measurably better and safer over time.
Stack
Python, PyTorch/JAX, GPU-based training and inference systems.
Who you probably are
- You've shipped ML systems real people use — not just impressive demos.
- You know large models well enough to know how and where they break.
- You write code you're proud of, and you care about correctness, not just "it works on my machine."
- You're self-directed and pragmatic — you take ownership of outcomes, not just tasks.
- You communicate clearly and work well in a small team where trust matters more than process.