jobs in ANCHOR SEARCH GROUP PTE. LTD.

全职 Senior Data Engineer – AI 工作, 薪水 up to SGD 7,000, ANCHOR SEARCH GROUP PTE. LTD. Islandwide (Singapore) 公司招聘中 - Ricebowl

Senior Data Engineer – AI

ANCHOR SEARCH GROUP PTE. LTD.

Islandwide (Singapore)

分享
保存

工作地点

  • Islandwide (Singapore) Singapore

职位描述

岗位职责

You will operate across both fast-moving Forward Deployed Engineering (FDE) engagements (POC/POV, pilot deployments for strategic and lighthouse clients) and steady-state system development and maintenance work — bringing the same rigor and a reusable, asset-fed approach to both.

Responsibilities:

Data Pipeline Engineering & AI-Readiness

•    Design and build ingestion, cleaning, and transformation pipelines that turn messy, real-world client data into AI-ready datasets.

•    Build batch and streaming pipelines (Airflow/Prefect/Kafka) that keep data flowing reliably into AI systems without manual intervention.

•    Own data quality — deduplication, schema validation, completeness checks — upstream of any model or RAG pipeline.

•    Proactively flag data gaps or quality issues that would degrade model/RAG performance downstream, before they surface as an AI Engineer's problem in testing.

RAG & Vector Store Architecture

•    Architect document/data ingestion and indexing pipelines for Retrieval-Augmented Generation (RAG) systems — chunking strategy, embeddings, hybrid/vector search.

•    Design and operate vector database and search infrastructure (pgvector/Pinecone/OpenSearch) at production scale and query volume.

Data Governance &Compliance

•    Implement PII redaction, data residency, and access-control patterns aligned to PDPA and sector-specific requirements (Healthcare, Government, Transport).

•    Maintain clear data lineage and metadata governance so engagement teams and auditors can trace how client data flows into AI outputs.

FDE &Development/Maintenance Coverage

•    During FDE engagements: rapidly assess and prepare a client's data landscape during Discover/POC, identifying data-readiness gaps early.

•    During system development & maintenance engagements: build and operate production-scale data pipelines handling the full volume and complexity of live client systems (e.g., Healthcare or Transport data at scale).

•    Contribute reusable ingestion/indexing patterns back into the shared internal asset library to accelerate future engagements.

Collaboration &Leadership

•    Partner closely and continuously with AI Engineers and AI Architects — understanding what a given model, RAG pipeline, or agent actually needs from the data layer and translating that into concrete pipeline and schema design decisions.

•    Own the definition of "AI-ready" data for each engagement jointly with AI Engineers — agreeing on chunking strategy, metadata, freshness, and quality thresholds before pipelines are built, not after retrieval quality suffers.

•    Sit in solution design conversations alongside AI Engineers and AI Architects, so data architecture and model/RAG architecture are designed together rather than data being treated as a downstream dependency.

•    Mentor junior data engineers and set data engineering standards across engagements.

Requirements:

•    10+years in data engineering, including production-scale pipeline design (not just analytics/reporting pipelines).

•    Strong SQL and at least one systems language (Python/Scala/Java); hands-on with batch and streaming frameworks (Airflow, Spark, Kafka).

•    Experience building data pipelines for AI/ML or RAG use cases — embeddings, vector indexing, hybrid search.

•    Solid understanding of data governance, PII handling, and access-control patterns in regulated environments.

•    Comfortable moving between fast, exploratory data assessment (FDE/POC) and disciplined, high-volume production pipeline engineering (system development &maintenance).

•    Working understanding of core AI/LLM concepts — tokenization, embeddings, chunking strategy, context windows, RAG, and agentic workflows — sufficient to hold areal technical conversation with AI Engineers and AI Architects about what "AI-ready" data means for a given use case, not just how to move and clean it.

Preferred Qualifications

•    Experience with vector databases (pgvector, Pinecone, Weaviate) and search platforms(OpenSearch/Azure AI Search).

•    Exposure to Singapore Government data environments (GCC/HCC) and compliance regimes(IM8, PDPA).

•    Experience with sector-specific data complexity — Healthcare (clinical data governance) or Transport/Aviation systems.

•    Familiarity with data cataloguing and lineage tooling.

•    Prior experience embedded within an AI/ML delivery team (not just a data platform team) — i.e., has sat alongside AI Engineers day-to-day and adjusted pipeline/schema design based on model or RAG performance feedback.

Tech Stack(Illustrative)

•    Languages: Python, SQL (Scala/Java a plus)

•    Pipelines: Airflow/Prefect, Spark, Kafka/Debezium

•    Storage/Search: Postgres, S3/Blob, pgvector/Pinecone/Weaviate, OpenSearch/Azure AI Search

•     Governance: Presidio (PII redaction), data catalogue/lineage tooling

•    Cloud: AWS/Azure/GCP; GCC/HCC exposure a plus

Interested candidates may send their CV to MAC (Reg No. R1221300) ************* quoting the job title in the Subject line. We regret that only shortlisted candidates will be notified.

重要安全守则

申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。

了解更多