Skip to content
All Projects
Data EngineeringDelivered

Azure Lakehouse Modernisation for a National Utility

Cutting data operating costs in half with Azure Data Factory & Databricks

Re-architected the data infrastructure of a national water & wastewater utility onto Azure, using Azure Data Factory, Azure SQL and Databricks with PySpark to cut operational costs 50% and lift system performance 20%.

PythonPySparkAzureAzure Data FactoryAzure SQLAzure Data LakeDatabricksARM TemplatesSQL
Problem Statement

A national utility was paying too much for an ageing data estate: rising operational costs, slow queries, and an architecture that could not stretch to meet growing data-processing and analytics demand.

  • Escalating operational costs from an inefficient, hard-to-scale data platform.
  • Performance bottlenecks left analytics and reporting workloads slow and unreliable.
  • The legacy system could not scale to growing data-processing and analytics demand.
Headline Outcomes
−50%cloud re-architecture

Operational cost

+20%query optimisation

System performance

CI/CDfully automated

Deployment

The Solution

An ACID-compliant lakehouse on Microsoft Azure: orchestrated with Azure Data Factory, transformed at scale with Databricks and PySpark, and shipped through DevOps CI/CD pipelines so every release was repeatable and observable.

Designed an Azure-native architecture on Azure Data Factory, Azure SQL and Databricks.

PySpark and SQL transformations handled data processing and analytics workloads end-to-end.

ACID guarantees and query optimisation improved data integrity and response time.

DevOps CI/CD pipelines automated deployment for scalability and operational agility.

System Architecture

How the data flows

01

Source Systems

Operational data

02

Azure Data Factory

Ingestion & orchestration

03

Databricks + PySpark

Scaled transforms

04

Azure SQL

ACID serving layer

05

CI/CD Deploy

Azure DevOps

Result 01

Halved the cost of running a national-scale data platform.

Result 02

Delivered a responsive, ACID-compliant foundation for analytics at scale.

Result 03

Made releases fast and low-risk through automated CI/CD pipelines.

Further reading

From the blog

Data Engineering

Cloud Data Warehouse Migration: Snowflake vs Redshift vs BigQuery

A cloud data warehouse migration guide to Snowflake vs Redshift vs BigQuery vs Databricks: how to choose on cost, lock-in and performance, and how to de-risk the move.

Data WarehouseSnowflakeRedshiftBigQuery
Taking on new projects · Outside IR35

Have a data pipeline or warehouse problem worth solving?

From messy source data to analytics-ready warehouses that cut cost. Let's scope it. I reply within one business day.