Key Responsibilities
Design, develop and maintain scalable
data pipelines and data ingestion frameworks
for large-volume datasets.
Develop data transformation and processing applications using
Apache Spark, PySpark, Scala and Python
.
Build and optimize data pipelines using
Azure Databricks, Azure Data Factory, AWS EMR
and related cloud services.
Work with
Hadoop, HDFS, Hive, Snowflake, Teradata and Data Lake
environments.
Develop batch and real-time data processing solutions using
Spark Structured Streaming and Kafka
.
Perform data extraction, transformation and loading across heterogeneous source and target systems.
Develop and optimize
Spark SQL, HiveQL and SQL
queries for performance and cost efficiency.
Design data models, partitioning strategies and scalable data storage architectures.
Build and manage workflow orchestration using
Apache Airflow
.
Implement CI/CD pipelines and automated testing using tools such as
Jenkins, Docker, GitHub Actions and pytest
.
Troubleshoot data pipeline, performance and production issues and implement sustainable solutions.
Collaborate with business stakeholders, architects and technology teams to understand requirements and deliver data engineering solutions.
Ensure data quality, reliability, security and operational stability across enterprise data platforms.
Required Skills
6+ years of experience in
Data Engineering / Big Data Engineering
.
Strong hands-on experience with
Apache Spark / PySpark
.
Strong programming skills in
Python and/or Scala
.
Good experience with
Hadoop, HDFS and Hive
.
Experience developing
ETL/ELT and data ingestion pipelines
.
Strong SQL and data processing skills.
Experience with
Azure Databricks, Azure Data Factory, AWS EMR
or equivalent cloud data platforms.
Experience with
Kafka / real-time streaming
is an advantage.
Hands-on experience with
Airflow
and data pipeline orchestration.
Experience with
Snowflake, Teradata, SQL Server or other enterprise databases
.
Good understanding of
Data Lake, Delta Lake, Data Warehousing and Data Modelling
.
Experience with
Git, CI/CD, Docker and automated testing
.