jobs in Nucleus AI

全职 Tokenisation Engineer 工作, 薪水, Nucleus AI 公司招聘中 - Ricebowl

Tokenisation Engineer

Nucleus AI

Undisclosed

Singapore

分享
保存

工作地点

  • Singapore

职位描述

岗位职责

Company Description Nucleus AI builds and operates frontier-level enterprise large language models (LLMs) on organizations’ proprietary data, enabling secure, high-performance AI tailored to specific workflows. The company helps teams move beyond general-purpose AI tools by providing privately deployed models that remain under the organization’s control. Its platform manages end-to-end AI operations, including training, secure inference, deployment, monitoring, automation, and ongoing maintenance, so enterprises can scale from experimentation to production without running their own infrastructure. Based in Singapore, Nucleus AI partners with enterprise and public-sector clients to integrate purpose-built AI systems with existing data and processes. More information is available at *************
Role Description This is a full-time Tokenisation Engineer role based in Singapore with a hybrid work arrangement, combining on-site collaboration and some work-from-home flexibility. The Tokenisation Engineer will design, implement, and optimize tokenization pipelines for multilingual and domain-specific text used in Nucleus AI’s LLMs. Day-to-day responsibilities include analyzing data characteristics, defining vocabulary and tokenization strategies, evaluating trade-offs between performance, efficiency, and accuracy, and integrating tokenizers with training and inference systems. The role involves collaborating closely with ML engineers and researchers to test new tokenization approaches, running experiments, and interpreting metrics to improve model quality. The Tokenisation Engineer will also contribute to tooling, documentation, and best practices that make tokenization workflows repeatable, robust, and secure in enterprise environments.
Qualifications
  • Strong foundations in computer science or a related discipline, with experience in natural language processing, language modeling, or text preprocessing.
  • Hands-on skills in tokenization and data processing, including implementing and tuning subword or byte-level tokenizers, designing vocabularies, and handling multilingual or domain-specific corpora.
  • Proficiency in programming for ML and data engineering (e.g., Python, relevant NLP/ML libraries, scripting and automation for data pipelines).
  • Experience working with large-scale datasets, performance optimization, and model training/inference workflows in production or research environments.
  • Ability to collaborate with cross-functional teams, clearly communicate technical decisions, and document methodologies and results for reproducibility.
  • Comfort working in a hybrid setup in Singapore and engaging with enterprise or public-sector stakeholders on secure, data-centric AI solutions.
  • Bachelor’s or higher degree in Computer Science, Mathematics, Engineering, or a related field, or equivalent practical experience; prior work with LLMs or enterprise AI platforms is an advantage.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多