Job Title
Senior Data Engineer
Description
Location: Cape Town — Mon–Thu work from office, Friday remote. Rate: R750/hr. Duration: 12 months (renewal possible).
We are seeking an experienced Senior Data Engineer to lead architecture and delivery of enterprise-scale EDL/ETL pipelines, streaming platforms, and AI-ready data ecosystems across AWS and Azure. You will define end-to-end solutions, drive adoption of cloud-native services, ensure data quality and governance, and enable ML/AI workflows at scale.
Responsibilities
- Architectural leadership: define and own end-to-end architecture for EDL/ETL pipelines, streaming platforms, and AI-ready data ecosystems.
- AWS data engineering: lead large-scale implementations using AWS services including S3, Glue, EMR, Lambda, Kinesis, and Athena.
- Azure data integration: drive adoption and implementation of Azure Data Factory, Synapse, Databricks, and Event Hubs for enterprise workloads.
- Streaming & real-time processing: design resilient, high-throughput streaming pipelines leveraging Kafka, Kinesis, and Event Hubs.
- Kafka topic consumption: design and implement consumer applications using Kafka Streams, Spark Structured Streaming, and Flink; ensure exactly-once semantics, offset handling, consumer lag monitoring, and fault-tolerant replay.
- AI-ready platforms: build feature stores, ML-ready datasets, and automated retraining pipelines to accelerate ML adoption and productionization.
- Data quality & governance: establish enterprise-wide frameworks for validation, reconciliation, metadata management, and lineage.
- Performance & scalability: optimize pipelines for petabyte-scale datasets while managing cost efficiency and high availability.
- Mentorship & collaboration: mentor junior engineers, collaborate with data scientists and stakeholders, and align delivery with business priorities.
- Infrastructure automation: implement CI/CD, Infrastructure as Code, and platform engineering practices to standardize deployments and reduce risk.
Qualifications
- 10+ years of hands-on data engineering experience with strong emphasis on AWS services (S3, Glue, EMR, Kinesis, Lambda, Athena).
- Proven expertise with Azure data services (Azure Data Factory, Synapse, Azure Databricks, Azure Event Hubs).
- Advanced proficiency in Python, PySpark, and SQL for large-scale data workloads.
- Deep knowledge of streaming architectures and technologies (Apache Kafka, Kinesis, Event Hubs) and exactly-once processing patterns.
- Proven experience designing and operating Kafka consumers, including consumer group management, offset handling, and schema evolution strategies.
- Experience building AI/ML workflows including feature engineering, feature stores, model training pipelines, deployment, and MLOps practices.
- Familiarity with Data Mesh, Data Fabric, and data lakehouse paradigms; strong metadata modeling and governance experience.
- Experience with CI/CD pipelines and Infrastructure-as-Code using Terraform or CloudFormation.
- Track record of leading teams, delivering enterprise-scale solutions, and influencing platform strategy across multi-cloud environments.
- Nice-to-have: certifications such as AWS Certified Solutions Architect – Professional or Azure Data Engineer Expert; experience with Feast, Databricks Feature Store, or ML frameworks such as TensorFlow, PyTorch, or Scikit-learn.
Skills
The ideal candidate will demonstrate practical expertise across cloud platforms, streaming systems, data platform design, and ML operationalization.