Responsibilities
- Design and develop high-quality, scalable ETL/ELT pipelines for processing big data using AWS analytical services.
- Leverage no-code tools and reusable Python libraries to ensure efficiency, maintainability, and reusability.
- Ensure data pipelines and platform components are scalable, performant, and cost-efficient.
- Build and enhance reusable platform capabilities across ingestion, transformation, orchestration, data quality, cataloguing, and monitoring.
- Develop scalable and reliable AWS data pipelines and data products using standardised data architecture patterns such as medallion architecture, lakehouse, and other modern design paradigms.
- Drive configuration-driven, declarative, and automated engineering approaches that reduce bespoke development and improve productivity.
- Champion automation across the data engineering lifecycle to minimize manual intervention and accelerate delivery.
- Work closely with cross-functional teams, including Tech Leads, Engineering Managers, and Business Analysts, to understand project objectives and deliver robust data solutions.
- Follow Agile/Scrum principles to drive consistent progress and iterative delivery.
- Perform data discovery and analysis to uncover data anomalies.
- Identify and resolve data quality issues through root cause analysis, and provide informed recommendations for data quality improvement and remediation.
- Champion the integration of Claude Code and other LLM tools into our software development lifecycle (SDLC).
- Lead the team in using AI to accelerate coding, debugging, automated testing, and documentation.
- Manage the automated deployment of code and ETL workflows within cloud infrastructure (AWS preferred) using tools such as GitHub Actions, AWS CodePipeline, or other modern CI/CD systems.
- Implement Infrastructure as Code (IaC), automated testing frameworks, observability solutions, security best practices, and operational reliability measures to ensure robust and resilient data platform operations.
- Demonstrate strong organizational and time management skills.
- Prioritize tasks effectively and ensure the timely delivery of key project milestones.
- Set the gold standard for the team by leading code reviews, defining CI/CD patterns, and enforcing data governance standards via AWS Lake Formation.
- Architect the infrastructure for our AI initiatives.
- Implement our initial AWS SageMaker footprint (Data Wrangler, Feature Store) and manage AWS Bedrock integrations (Knowledge Bases, RAG pipelines) to support downstream Data Science and GenAI initiatives.
- Develop and maintain comprehensive data catalogs, including data mapping and documentation, to ensure data governance, transparency, and accessibility for all stakeholders.
- Continuously improve your skills by learning and implementing data engineering best practices.
- Stay updated on industry trends and contribute to team knowledge-sharing and codebase optimization.
Requirements
- Strong hands-on experience with diverse data management methodologies.
- Proven track record of implementing data management methodologies across a wide range of data projects to generate highly relevant data products.
- Experience building scalable ETL/ELT pipelines for processing big data using AWS analytical services.
- Experience leveraging no-code tools and reusable Python libraries.
- Ability to ensure data pipelines and platform components are scalable, performant, and cost-efficient.
- Experience building and enhancing reusable platform capabilities across ingestion, transformation, orchestration, data quality, cataloguing, and monitoring.
- Experience developing scalable and reliable AWS data pipelines and data products using standardised data architecture patterns such as medallion architecture, lakehouse, and other modern design paradigms.
- Experience driving configuration-driven, declarative, and automated engineering approaches.
- Proven ability to champion automation across the data engineering lifecycle.
- Experience working closely with cross-functional teams including Tech Leads, Engineering Managers, and Business Analysts.
- Familiarity with Agile/Scrum principles and iterative delivery.
- Experience performing data discovery and analysis to uncover data anomalies.
- Ability to identify and resolve data quality issues through root cause analysis.
- Experience managing automated deployment of code and ETL workflows within cloud infrastructure (AWS preferred) using tools such as GitHub Actions, AWS CodePipeline, or other modern CI/CD systems.
- Experience implementing Infrastructure as Code (IaC), automated testing frameworks, observability solutions, security best practices, and operational reliability measures.
- Demonstrated strong organizational and time management skills.
- Experience leading code reviews and defining CI/CD patterns.
- Experience enforcing data governance standards via AWS Lake Formation.
- Ability to architect infrastructure for AI initiatives.
- Experience implementing AWS SageMaker (Data Wrangler, Feature Store) and managing AWS Bedrock integrations (Knowledge Bases, RAG pipelines).
- Experience developing and maintaining comprehensive data catalogs, data mapping, and documentation.
- Commitment to continuous learning and implementation of data engineering best practices.