️ Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT processes.
- Build and optimize enterprise data warehouses and data lake solutions.
- Develop high-performance data processing solutions using Spark/PySpark.
- Integrate data from multiple internal and external sources.
- Optimize SQL queries and improve data pipeline performance.
- Implement batch and real-time data processing solutions.
- Work with Data Scientists, BI Developers, and Software Engineers to support analytics initiatives.
- Ensure data quality, governance, security, and compliance.
- Deploy and maintain cloud-based data solutions.
- Perform code reviews and mentor junior team members.
- Troubleshoot production issues and implement performance improvements.
- Follow Agile development methodologies and DevOps best practices.
Required Skills
Category
Skills
Experience
12+ Years in Data Engineering
Programming
Python, SQL, PySpark, Scala (Preferred)
Big Data
Apache Spark, Hadoop, Hive
Data Integration
ETL/ELT, Apache Airflow
Streaming
Apache Kafka
Cloud
AWS / Azure / GCP
Data Warehouse
Snowflake, Redshift, BigQuery, Azure Synapse
Data Lake
Delta Lake, Apache Iceberg
Databases
PostgreSQL, MySQL, SQL Server, MongoDB
DevOps
Git, Docker, Kubernetes, Jenkins
Operating System
Linux / Unix
APIs
REST APIs
Data Modeling
Star Schema, Snowflake Schema, Dimensional Modeling
Performance
SQL Optimization, Spark Performance Tuning
Methodology
Agile / Scrum
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Relevant cloud or data engineering certifications are an advantage.
Preferred Skills
- Experience with Databricks.
- Experience with dbt and Terraform.
- Knowledge of CI/CD pipelines.
- Experience with real-time streaming architectures.
- Exposure to machine learning data pipelines.