jobs in Talotrace

Talotrace Hiring! Full Time Founding AI Engineer in - Ricebowl

Founding AI Engineer

Talotrace

Undisclosed

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

Founding AI Engineer

Location: Singapore. Hybrid.


About the job

TaloTrace builds AI agents that test software on real devices. Point us at a web app or a mobile build and our agents explore it, work out what it should do, find the bugs, and file them. Web, Android and iOS, on emulators and physical devices.

We are a small team of ambitious people who work at high iteration speed.

We are looking for a Founding AI Engineer to own the model layer behind that. Our agents look at a screen, decide what to do, act, and judge what happened. Each of those steps is a model, and each can be smaller, faster, cheaper and more accurate than it is today.

You will be the first person here whose whole job is the models, so how you set that work up is how it will run for a long time.


Core Responsibilities

  • Own the model layer end to end: data, training, evaluation, deployment, and what it costs to run
  • Turn research results into production systems, and drop the ones that don’t hold up
  • Build evaluation that tells us the truth, including when the answer is that something didn’t work
  • Make cost, quality and latency tradeoffs and back them with numbers
  • Set the standards for model work here: what gets measured, what counts as a real improvement, and what must be true before a model ships


What You Could Own

You will not build all of these at once. Expect to pick up two or three in your first year.

  • Small models replacing large ones. Distillation, fine-tuning and quantization so the pipeline runs on a fraction of the compute
  • Grounding and perception. Reliable element localisation across web, Android and iOS, including after a visual redesign
  • The judgment layer. Deciding whether a finding is real, how serious it is, and when to say the answer is unclear
  • Exploration policies. Models that cover an application the way real users move through it, rather than only the happy path
  • State prediction. Anticipating what a screen becomes after an action, which helps both navigation and catching unexpected behavior
  • Evaluation infrastructure. Offline benchmarks, regression gates on model changes, and metrics that measure what they claim to
  • Inference optimisation. Making the models themselves cheaper to serve quantization, batching, KV-cache work, on-device deployment. You set what each tier costs and how good it is; the platform side decides when to call which
  • Training data pipelines. Collecting, cleaning, labelling and versioning the interaction data everything above runs on


What We Are Looking For

  • 4+ years in ML/AI engineering, or strong software engineering with substantial hands-on model work. The floor matters less than whether you have owned a model in production and can show it
  • Strong Python and practical PyTorch or equivalent. You can read a paper’s repo and get it running
  • You have trained or improved models yourself: supervised fine-tuning, LoRA/PEFT, RL, behaviour cloning or distillation
  • You have built evaluation harnesses and can tell when a metric is measuring the wrong thing
  • Production experience with LLM systems: structured output, tool use, failure modes, latency and cost
  • Comfort with data pipelines. Collecting, cleaning and versioning training data is most of this job
  • You check your results before you report them, and you say when something didn’t work
  • You build quick prototypes to answer a question and throw them away afterwards
  • Clear writing. You can write up an experiment so someone else can act on it
  • You get a lot out of AI coding tools and your output holds up to review


You Will Stand Out If You Have

  • Been an early engineer at a startup. You know what a first version looks like
  • Worked with vision-language models or multimodal training
  • Worked on GUI agents, computer-use models, or robotic and embodied policies
  • Done imitation learning or behaviour cloning from human demonstration data
  • Optimised inference: FP8 or INT4 quantization, vLLM or TensorRT, KV-cache work, on-device deployment
  • Used reinforcement learning, especially RLHF/GRPO or verifiable-reward setups
  • Modelled human behaviour in any domain, whether game AI, recommender systems, user simulation or behavioural biometrics
  • Published, open-sourced or reproduced work in an adjacent area



Process

Intro call, a take-home challenge scoped to about 2 to 3 hours, then a live technical interview.


Why TaloTrace?

Our motto is leave quality to us.

We think the next era of builders and changemakers should be able to reach far more people than the last one, and software is how they will do it. Our mission is to let them get from a problem statement to delivered software without stopping worrying about whether it works.

Model changes land against real customer runs, so you find out whether something worked in days rather than quarters. Building the measurement that makes that trustworthy is part of the job.

We have hired evidence of what you have built. We welcome applications from everyone and will accommodate whatever you need during the process, so just ask.


How We Work

Small team, fast pace, not much process. Where an experiment can settle a question we run it instead of debating it. Where it can’t, on direction, priorities, and what “good” means, we argue it out in the open, and we expect you to push back when you disagree. People here regularly work outside their main area.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More