Skip to content
Selected Work

AWS Data Engineering Case Studies

20 engineering case studies across data, AI, automation and analytics, each with the problem, the architecture, and the numbers that moved.

Data EngineeringDelivered

BI Dashboard Migration & Data-Mart Architecture

Securing a university master database while modernising reporting

Moved a South-African university’s reporting stack off QlikView querying a live Oracle master DB onto Power BI fed by purpose-built data marts and an…

+80%master DB protected
SQLMeshDagsterMySQLOracle
View Case Study
Data EngineeringDelivered

Automated Distributor ETL for Unilever

Hands-off SFTP-to-analytics for daily sales & stock data

An automated ELT pipeline that detects distributor files landing on SFTP, validates and homologates them, and lands analytics-ready data in S3 for rea…

ZeroManual intervention
DagsterdbtPostgreSQLAWS S3
View Case Study
Data EngineeringDelivered

Enterprise Data Warehouse Migration — Sybase to Amazon Redshift

Re-platforming a tier-1 banking EDW onto a cloud-native AWS warehouse

Led the migration of a major retail bank’s on-premises Enterprise Data Warehouse from Sybase to Amazon Redshift on AWS, re-engineering 800+ stored pro…

+40%Redshift compute
PythonSQLPL-SQLAmazon Redshift
View Case Study
Data EngineeringDelivered

Azure Lakehouse Modernisation for a National Utility

Cutting data operating costs in half with Azure Data Factory & Databricks

Re-architected the data infrastructure of a national water & wastewater utility onto Azure, using Azure Data Factory, Azure SQL and Databricks with Py…

−50%cloud re-architecture
PythonPySparkAzureAzure Data Factory
View Case Study
Data EngineeringDelivered

Cloud-Native ETL for Environmental Analytics

A 3-tier AWS Glue pipeline for assessment, measurement & analysis data

Built a 3-tier data architecture on AWS for an environmental-solutions provider, using AWS Glue, S3, Redshift, Lambda and PySpark to lift data-process…

+30%AWS Glue
PythonAWS GlueAWS S3Amazon Redshift
View Case Study
Data EngineeringDelivered

Cloud Data Warehouse Migration on AWS

Re-architecting on-prem ETL into AWS for 80% better resource utilisation

Led a data-warehouse migration to AWS for a pan-African financial-services group, moving 20–30 ETL scripts and stored procedures onto Glue, Redshift a…

+80%job re-architecting
AWSAWS GluePySparkSQL
View Case Study
Data EngineeringDelivered

Finance Cockpit & Intelligent Collections on Cloudera

A big-data platform that cut financial-reporting errors by 50%

Operationalised a Cloudera-based big-data platform for a telecom operator’s finance function: 10 workflows and ~50 scripts across Spark, Impala, Hive…

+30%optimised workflows
PythonPySparkSQLApache Spark
View Case Study
Data EngineeringDelivered

Dynamic, Fully-Automated ETL Pipeline on AWS

100% API automation processing millions of records daily

Built a dynamic ETL pipeline on AWS for a sales-engagement SaaS platform, using Glue, Redshift, Apache Hudi and Athena to automate extraction of milli…

−40%automation
PythonPySparkAWSAWS Glue
View Case Study
Data EngineeringDelivered

Data & Machine-Learning Platform for a Payments Leader

A delta-lake platform that drove $2M+ in new revenue

Designed and built a data & machine-learning platform for a publicly-listed payments processor: 20 Azure pipelines on a delta-lake architecture, proce…

$2M+scalable analytics
PythonPySparkAzureAzure Data Factory
View Case Study
Data EngineeringDelivered

Azure Lake House Data Platform for Retail Banking

Unifying 10+ banking sources for 1,000+ business users

Built a centralised Lake House data platform on Azure for a leading commercial bank: 15 Azure Data Factory & Synapse pipelines integrating 10+ sources…

10+unified platform
PythonPySparkAzureAzure Data Factory
View Case Study
Industries

Which sectors I build for

The same AWS data-engineering playbook, applied to the realities of each industry and backed by the case studies above.

Banking & Financial Services

Data warehouse migrations, regulatory reporting and risk analytics for retail banks, payments processors and financial-services groups. Recent work includes a tier-1 bank Sybase-to-Redshift migration and a delta-lake platform that unlocked $2M+ in revenue.

Utilities & Energy

Lakehouse modernisation and operational analytics for national water, wastewater and energy utilities, cutting data-platform operating cost in half while letting analytics scale to demand.

Aviation & Aerospace

Pipeline and RFQ-operations automation for aerospace-logistics and parts suppliers, turning a five-hour manual quoting backlog into sub-15-minute automated quotes on AWS.

Consumer Goods & Retail

Distributor sales-and-stock ETL, data homologation and real-time analytics for FMCG and retail brands, including hands-off SFTP-to-analytics pipelines for Unilever distributors.

Telecom

Big-data finance platforms and reporting for telecom operators: Cloudera, Spark and Hive workflows that cut financial-reporting errors by 50%.

Healthcare, Public Sector & Research

Document intelligence, multilingual OCR and NLP analytics for healthcare, government digitisation and research, from automated form extraction to mining 50M+ posts for public-health signal.

Taking on new projects · Outside IR35

Have a data pipeline or warehouse problem worth solving?

From messy source data to analytics-ready warehouses that cut cost. Let's scope it. I reply within one business day.