Skip to content
All Projects
Data EngineeringDelivered

Azure Lake House Data Platform for Retail Banking

Unifying 10+ banking sources for 1,000+ business users

Built a centralised Lake House data platform on Azure for a leading commercial bank: 15 Azure Data Factory & Synapse pipelines integrating 10+ sources and processing 10GB daily to serve 1,000+ users with reliable, governed data.

PythonPySparkAzureAzure Data FactoryAzure SynapseAzure Data LakeARM TemplatesSQLPL-SQL
Problem Statement

Fragmented data sources, slow processing and limited scalability were holding a major bank back. Pulling and transforming data from core-banking, CRM, transaction systems and external APIs was slow, and with no central platform, 1,000+ users had no single source of truth to rely on.

  • Fragmented data across core-banking, CRM, transaction systems and external APIs.
  • Inefficient processing and limited scalability hindered decision-making.
  • No centralised platform to serve 1,000+ business users reliably.
Headline Outcomes
10+unified platform

Data sources integrated

10GB15 pipelines

Daily data processed

90%+cleansed & enriched

Data coverage

The Solution

A Lake House on Microsoft Azure, built on Azure Data Lake and Synapse, with 15 Azure Data Factory pipelines integrating diverse banking sources. PySpark cleanses, validates and enriches over 90% of the data for reliable, enterprise-wide analytics.

Lake House architecture on Azure Data Lake and Azure Synapse.

15 Azure Data Factory pipelines integrate 10+ sources, processing 10GB of data daily.

Python and PySpark cleanse, validate and enrich 90%+ of the data for reliability.

Deployed and governed with Azure DevOps, ARM Templates, VNets, IAM and monitoring.

System Architecture

How the data flows

01

10+ Sources

Core-banking · CRM · APIs

02

Azure Data Factory

15 pipelines

03

Lake House

Data Lake + Synapse

04

PySpark Quality

90%+ data enriched

05

1,000+ Users

Governed analytics

Result 01

Gave 1,000+ users a reliable, centralised single source of truth.

Result 02

Improved processing efficiency and accuracy across the bank.

Result 03

Delivered a scalable, governed platform built for future growth.

Further reading

From the blog

Security

How to Mask PII on Ingest Into an S3 Data Lake

Mask PII in an S3 data lake at the ingest boundary: deterministic hashing that keeps joins working, Glue detection, Lake Formation filters, KMS and erasure.

PII MaskingData GovernanceAWS Lake FormationAWS Glue
Taking on new projects · Outside IR35

Have a data pipeline or warehouse problem worth solving?

From messy source data to analytics-ready warehouses that cut cost. Let's scope it. I reply within one business day.