jobs in Nucleus AI

Nucleus AI Hiring! Full Time Tokenisation Engineer in - Ricebowl

Tokenisation Engineer

Nucleus AI

Undisclosed

Singapore

Share
Save

Working Location

  • Singapore

Job Description

Responsibilities

Company Description Nucleus AI builds and operates frontier-level enterprise large language models (LLMs) on organizations’ proprietary data, enabling secure, high-performance AI tailored to specific workflows. The company helps teams move beyond general-purpose AI tools by providing privately deployed models that remain under the organization’s control. Its platform manages end-to-end AI operations, including training, secure inference, deployment, monitoring, automation, and ongoing maintenance, so enterprises can scale from experimentation to production without running their own infrastructure. Based in Singapore, Nucleus AI partners with enterprise and public-sector clients to integrate purpose-built AI systems with existing data and processes. More information is available at *************
Role Description This is a full-time Tokenisation Engineer role based in Singapore with a hybrid work arrangement, combining on-site collaboration and some work-from-home flexibility. The Tokenisation Engineer will design, implement, and optimize tokenization pipelines for multilingual and domain-specific text used in Nucleus AI’s LLMs. Day-to-day responsibilities include analyzing data characteristics, defining vocabulary and tokenization strategies, evaluating trade-offs between performance, efficiency, and accuracy, and integrating tokenizers with training and inference systems. The role involves collaborating closely with ML engineers and researchers to test new tokenization approaches, running experiments, and interpreting metrics to improve model quality. The Tokenisation Engineer will also contribute to tooling, documentation, and best practices that make tokenization workflows repeatable, robust, and secure in enterprise environments.
Qualifications
  • Strong foundations in computer science or a related discipline, with experience in natural language processing, language modeling, or text preprocessing.
  • Hands-on skills in tokenization and data processing, including implementing and tuning subword or byte-level tokenizers, designing vocabularies, and handling multilingual or domain-specific corpora.
  • Proficiency in programming for ML and data engineering (e.g., Python, relevant NLP/ML libraries, scripting and automation for data pipelines).
  • Experience working with large-scale datasets, performance optimization, and model training/inference workflows in production or research environments.
  • Ability to collaborate with cross-functional teams, clearly communicate technical decisions, and document methodologies and results for reproducibility.
  • Comfort working in a hybrid setup in Singapore and engaging with enterprise or public-sector stakeholders on secure, data-centric AI solutions.
  • Bachelor’s or higher degree in Computer Science, Mathematics, Engineering, or a related field, or equivalent practical experience; prior work with LLMs or enterprise AI platforms is an advantage.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More