- Islandwide (Singapore) Singapore
工作地点
职位描述
岗位职责
Responsibilities
• Design, build, optimise, and maintain batch and streaming data ingestion pipelines using platforms such as Databricks and Kafka, ensuring scalability, reliability, observability, and alignment with enterprise data architecture standards.
• Perform data transformation and cleansing using PySpark or SQL based on business and technical requirements
• Monitor and troubleshoot data workflows to ensure data quality and pipeline reliability
• Provide technical guidance to engineers and delivery partners on data platform patterns, reusable components, code quality, deployment readiness, and production support practices.
• Lead integration of data from diverse source systems including files, APIs, databases, and streaming platforms, working with source-system owners and consuming teams to define fit-for-purpose ingestion patterns and delivery timelines.
• Help maintain metadata and pipeline documentation for transparency and traceability
• Own production readiness for assigned data and AI platform components, including observability, incident triage, root-cause analysis, release coordination, and continuous improvement of operational runbooks.
• Participate in integrating pipelines with tools such as Microsoft Fabric, Databricks, Delta Lake, and other platform components
• Build and maintain knowledge base and RAG solution on variety of hosting platforms
• Implement and operate knowledge base storage, lifecycle management and embedding/vectorization
• Contribute to automation efforts using version control and CI/CD workflows
• Apply data governance, security, access control, and operational risk policies during solution design and implementation, ensuring pipelines and knowledge platforms meet enterprise compliance requirements.
Requirements
• Bachelor’s degree in Computer Science, Engineering, or a related field
• 5–8 years of experience in data engineering, data platform engineering, or cloud-scale analytics solution delivery, with demonstrated ownership of production pipelines and platform components.
• Proven ability to independently design, build, optimise, and operate production-grade batch or streaming data pipelines, including orchestration, observability, error handling, performance tuning, and operational support.
• Hands-on experience with Python and SQL for data transformation and validation
• Familiarity with Apache Spark (especially PySpark) and large-scale data processing concepts
• Experience with implementing knowledge base and RAG solutions for agentic AI use cases
• Self-starter with strong problem-solving skills and a keen attention to detail
• Able to work independently and lead technical discussions with engineers, architects, product owners, source-system teams, and business stakeholders to translate requirements into secure and maintainable platform solutions.
• Strong documentation and communication skills
• Strong understanding of enterprise data architecture, cloud security, access control, CI/CD, release management, and production operations for data and AI platform solutions.
Licence no: 12C6060
重要安全守则
申请工作时,切勿提供您的银行或信用卡详细资料。不要转账或完成无关的在线调查问卷。如果您发现可疑内容,请举报此招聘广告。