About the Role
Rhino Partners is looking for an experienced Data Engineer to design, build, and deploy a scalable and reliable data lakehouse platform on AWS.
The successful candidate will be responsible for developing end-to-end data pipelines, implementing data quality and governance frameworks, and delivering production-ready data solutions. This role requires strong hands-on AWS expertise, experience with modern lakehouse architectures, and the ability to build robust, cost-efficient data infrastructure.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using AWS services, including AWS Glue, Step Functions, Lambda, and S3.
- Architect and implement data lakehouse solutions using Apache Iceberg or Amazon S3 Tables, incorporating schema evolution, partition evolution, and ACID transactions.
- Optimise data pipelines and storage for performance, scalability, reliability, and cost efficiency.
- Define and implement automated data quality validation frameworks to ensure data accuracy, completeness, and consistency.
- Establish data quality metrics, monitoring, and alerting mechanisms to identify and resolve data issues.
- Implement and maintain data governance standards and ensure compliance with organisational requirements.
- Develop production-quality code and deploy data engineering solutions on AWS cloud infrastructure.
- Implement CI/CD pipelines and infrastructure-as-code using Terraform or AWS CloudFormation to support repeatable and auditable deployments.
- Collaborate with cross-functional teams to translate business and technical requirements into effective data solutions.
- Produce comprehensive technical documentation and facilitate knowledge transfer and handover to Day 2 operations teams.
Requirements
- Bachelor's degree in Computer Science, Data Engineering, Information Systems, or a related discipline.
- Minimum 3–5 years of hands-on experience in data engineering, ETL/ELT development, or data platform engineering.
- Strong experience with AWS data engineering services, particularly AWS Glue, Step Functions, Lambda, and S3.
- Hands-on experience designing and implementing data lakehouse architectures using Apache Iceberg or similar open table formats.
- Strong understanding of Apache Iceberg capabilities, including schema evolution, partition evolution, ACID transactions, and table optimisation.
- Experience implementing automated data quality checks, validation frameworks, and data governance practices.
- Proficiency in writing production-quality code and optimising data pipelines for performance and reliability.
- Experience with CI/CD practices and infrastructure-as-code tools such as Terraform or AWS CloudFormation.
- Strong problem-solving skills and the ability to communicate technical concepts effectively to both technical and non-technical stakeholders.
- Ability to work independently and collaboratively in a cross-functional environment.
Good to Have
- Experience working within a Government Commercial Cloud (GCC) environment.
- AWS certifications, such as AWS Certified Data Analytics or AWS Certified Solutions Architect.
- Experience supporting production data platforms, including operational handover and ongoing maintenance.