We are seeking an experienced Data Engineer (PySpark) to design, build, optimize, and maintain scalable data pipelines for production environments. The role requires strong hands-on experience in big data processing, pipeline optimization, and deployment using modern data engineering tools and frameworks.Key ResponsibilitiesDesign, develop, and maintain robust, scalable data pipelines using Python and PySparkPerform data ingestion, transformation, cleansing, and validation across structured and unstructured datasetsConduct Exploratory Data Analysis (EDA) to identify data patterns, anomalies, and quality issuesApply data imputation techniques, data linking, and cleansing to ensure high data qualityImplement feature engineering pipelines to support analytics and downstream use casesOptimize Spark jobs for performance, scalability, and cost efficiencyDeploy and tune production-grade data pipelines, ensuring reliability and performanceAutomate workflows using Apache Airflow and/or JenkinsCollaborate with cross-functional teams to integrate data solutions into production systemsWrite and maintain unit tests to ensure code quality and reliabilityManage source code, CI/CD, and deployments using Git, GitHub, and GitHub ActionsRequirementsTo be considered for this role, you need to meet the following criteria:Required Technical SkillsStrong proficiency in PythonExtensive hands-on experience with Apache Spark (PySpark)Experience working with Jupyter NotebooksStrong knowledge of SQL and NoSQL databasesProven experience with Git for version control and CI/CDHands-on experience with Apache Airflow and/or Jenkins for scheduling and automationSolid understanding of data engineering best practices in production environmentsDemonstrated experience in Spark performance tuning and optimizationAbility to write clean, testable, and maintainable Python codeMandatory RequirementPrevious production experience is a MUST, specifically in deploying, tuning, and maintaining data pipelines in production environmentsPreferred QualificationsExperience working in high-volume or big data environmentsStrong problem-solving and analytical skillsAbility to work independently in a fast-paced environmentWhy Join?Competitive salary packageOpportunity to work on production-scale data platformsExposure to modern data engineering tools and practicesDubai-based role with a dynamic and collaborative work environmentTo view other requirements we have, please visit our website - www.blackpearlconsult.comWe are seeking an experienced Data Engineer (PySpark) to design, build, optimize, and maintain scalable data pipelines for production environments. The role requires strong hands-on experience in big data processing, pipeline optimization, and deployment using modern data engineering tools and frameworks.Key ResponsibilitiesDesign, develop, and maintain robust, scalable data pipelines using Python and PySparkPerform data ingestion, transformation, cleansing, and validation across structured and unstructured datasetsConduct Exploratory Data Analysis (EDA) to identify data patterns, anomalies, and quality issuesApply data imputation techniques, data linking, and cleansing to ensure high data qualityImplement feature engineering pipelines to support analytics and downstream use casesOptimize Spark jobs for performance, scalability, and cost efficiencyDeploy and tune production-grade data pipelines, ensuring reliability and performanceAutomate workflows using Apache Airflow and/or JenkinsCollaborate with cross-functional teams to integrate data solutions into production systemsWrite and maintain unit tests to ensure code quality and reliabilityManage source code, CI/CD, and deployments using Git, GitHub, and GitHub Actions
We are seeking an experienced Data Engineer (PySpark) to design, build, optimize, and maintain scalable data pipelines for production environments. The role requires strong hands-on experience in big data processing, pipeline optimization, and deployment using modern data engineering tools and frameworks.
Key Responsibilities
- Design, develop, and maintain robust, scalable data pipelines using Python and PySpark
Design, develop, and maintain robust, scalable data pipelines using Python and PySpark
- Perform data ingestion, transformation, cleansing, and validation across structured and unstructured datasets
Perform data ingestion, transformation, cleansing, and validation across structured and unstructured datasets
- Conduct Exploratory Data Analysis (EDA) to identify data patterns, anomalies, and quality issues
Conduct Exploratory Data Analysis (EDA) to identify data patterns, anomalies, and quality issues
- Apply data imputation techniques, data linking, and cleansing to ensure high data quality
Apply data imputation techniques, data linking, and cleansing to ensure high data quality
- Implement feature engineering pipelines to support analytics and downstream use cases
Implement feature engineering pipelines to support analytics and downstream use cases
- Optimize Spark jobs for performance, scalability, and cost efficiency
Optimize Spark jobs for performance, scalability, and cost efficiency
- Deploy and tune production-grade data pipelines, ensuring reliability and performance
Deploy and tune production-grade data pipelines, ensuring reliability and performance
- Automate workflows using Apache Airflow and/or Jenkins
Automate workflows using Apache Airflow and/or Jenkins
- Collaborate with cross-functional teams to integrate data solutions into production systems
Collaborate with cross-functional teams to integrate data solutions into production systems
- Write and maintain unit tests to ensure code quality and reliability
Write and maintain unit tests to ensure code quality and reliability
- Manage source code, CI/CD, and deployments using Git, GitHub, and GitHub Actions
Manage source code, CI/CD, and deployments using Git, GitHub, and GitHub Actions