Obsessed about training data? We're building the data infrastructure layer for AI training.
The role
As our founding research engineer, you'll own data curation as a measured discipline across every stage of training, from pre-training to post-training. Working alongside the founders, you'll turn what works into product features — and as one of the first researchers, the methods and data infrastructure you build become the foundation the company runs on.
Examples of work you might do
Measure what's in a dataset — quality, provenance, safety, and coverage metrics that predict downstream performance
Diagnose where data will hurt a model — decontamination, difficulty annotation, multilingual asymmetries, long-tail gaps
Design data curation interventions — pruning, filtering, synthetic augmentation, relabeling — and attribute the gains to specific changes
Turn the latest research into running experiments — training and evaluation loops that prove a data hypothesis
Ship the methods that work as product features, and share findings as technical reports or papers
What we’re looking for
Data obsession — you want to know exactly which properties of a dataset move a model, and you won't trust a result you can't measure
Bias toward iteration — you design the smallest experiment that answers the question and move fast on the answer
Research that ships — you optimize for models that measurably improve, not for leaderboard wins
Product engineer — you own problems end to end and turn research into features that ship, not just results in a notebook
Drawn to data-centric AI — you're excited by data-centric methods and the move toward autonomous AI research
Required qualifications
Strong machine learning and deep-learning fundamentals
Enough software engineering and PyTorch / Jax experience (or willingness to learn) to run ML experiments and build production prototypes
Hands-on experience in one or more stages of training and evaluating LLMs/vLLMs
Industry or research experience in one or more of: data curation, data pruning & selection, synthetic data generation, curriculum learning, dataset distillation, large-scale language/multimodal training
Comfortable reading ML research — sourcing, vetting, and implementing promising ideas from the literature
Able to drive applied research independently in a fast, ambiguous, early-stage environment
Nice to have
Post-training experience — SFT, preference optimization (DPO, RLVR), or reward modeling
Multilingual or multimodal data work
Multi-GPU / distributed training experience
Open-source, HuggingFace contributions
Public technical writing, published research papers