About the Role
We are seeking a Data Engineer to join our AI Business Development Team and support the development of our AI-powered platforms and solutions across multiple project domains. The role’s primary focus is the platform, which is an AI-driven system that matches candidate CVs to job descriptions using semantic, embedding-based techniques, with involvement in other AI initiatives as required.
Responsibilities
Document Ingestion & Data Pipelines
- Design, build, and maintain OCR-based ingestion pipelines that convert CVs and job descriptions into structured, searchable data , operating on cloud-based infrastructure with secure storage and point-in-time data recovery, to support the AI matching workload (primarily batch processing).
- Develop and maintain scalable ETL/ELT data pipelines (e.g. on Apache Airflow) that extract, normalise, and feature-engineer CV and JD data for the matching engine.
- Integrate data from diverse sources, including internal databases, document files, and APIs.
- Design and maintain Entity Relationship Diagrams (ERD) and data schemas; produce data quality reports and implement data-validation frameworks.
- Ensure data quality, consistency, and reliability across the data lifecycle.
Semantic Matching & Embeddings
- Build and maintain the components that generate vector embeddings of CV and JD content (e.g. BERT) into the matching pipeline, and maintain vector storage and similarity search (e.g. PostgreSQL with pgvector).
- Support and extend the LLM-assisted domain dictionary used to map related skills and terminology (so that, for example, a requirement for one cloud platform recognises an equivalent platform).
- Tune and evaluate matching quality the relevance and ranking of CV–JD matches and iterate on the pipeline to improve accuracy.
- Support deployment of matching/embedding workloads into production and monitor their performance.
Cloud Infrastructure
- Implement and manage cloud-based data solutions (e.g. AWS S3 for storage, EC2/SageMaker for batch model compute, Lambda for orchestration where appropriate).
- Optimise data storage and processing for cost-efficiency and performance, particularly for batch embedding workloads.
- Support data-store and pipeline architecture design and maintenance.
Data Quality, Governance & Privacy
- Implement data validation, monitoring, and quality-assurance processes.
- Maintain documentation for data models, pipelines, and system architecture.
- Handle sensitive personal data (CV information) in compliance with the Personal Data (Privacy) Ordinance (PDPO) and internal data-protection and egress controls.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Data Engineering, or a related field.
- Hands-on experience in data engineering or a related role, with strong Python and SQL for data processing. (Seniority open, see note below.)
- Experience with OCR-based document ingestion/parsing (e.g. extracting structured data from CVs, forms, or similar documents).
- Experience with NLP/text-processing techniques for document parsing and feature extraction.
- Experience with embeddings and vector search/vector databases (e.g. pgvector or equivalent).
- Experience using LLMs for data tasks, such as skill/terminology extraction, tagging, or generating a domain dictionary from text data.
- Experience with ETL tools and frameworks (e.g. Apache Airflow, dbt), and with relational databases (PostgreSQL essential; MongoDB or similar an advantage).
- Understanding of information-retrieval / matching concepts (ranking, relevance, similarity).
- Familiarity with cloud platforms, preferably AWS (S3, EC2, Lambda; SageMaker for model compute).
- A relevant degree (Computer Science, IT, Data Engineering, or related), strong problem-solving skills, and attention to detail.
Preferred (Bonus Skills)
- Experience with ML frameworks (e.g. scikit-learn, PyTorch, TensorFlow) or managed ML pipelines (e.g. SageMaker).
- Understanding of REST API development for data-service integration.
- Understanding of CI/CD practices and version control (Git).
- Experience working in Agile/Scrum environments with sprint-based delivery.
- Relevant cloud or data-engineering certification (e.g. AWS Certified Data Engineer Associate) is advantageous.
- Familiarity with data-visualisation / BI tools (e.g. Tableau, Power BI, QuickSight).
Salary:
Salary will be commensurate with qualifications and experience.
Application:
Please submit your resume along with your current and expected salary. Personal data collected will be treated confidentially for recruitment purposes. Candidates not invited for an interview within six weeks may consider their applications unsuccessful.
Equal Opportunity Employer:
We are an equal opportunity employer and welcome applications from all qualified candidates. Personal data collected will be handled confidentially by authorized personnel for recruitment-related purposes.
Interested parties please send detailed resume with present/expected salary to HR Department.
Address:
Room 1202, 12/F, Harcourt House, 39 Gloucester Road, Wanchai, Hong Kong.
Website: *************
Full-time