All projects

The warehouse
you already have

The most valuable analytical dataset in a system is often something it is already producing and throwing away. Before buying a warehouse, look at what the pipeline emits as a side effect and ask what it would cost to simply keep it.

System
Passenger profiling platform
Scale
Full dataset, permanently retained
Role
Sole architect and engineer
Stack
Parquet / DuckDB / S3
01

The challenge

Analytical queries were competing with live traffic on the same database, and the standard remedy is a warehouse: a new system, a new bill, a new pipeline.

02

The approach

Retain what the pipeline was already producing, lay it out so it can be skipped cheaply, and query it with the same code that queries the database.

03

The result

Analytical work moved off the live database at storage cost, with nothing new to operate and a fifty-fold faster path to get the data there.

The columnar files the pipeline already produces are retained permanently and laid out so that engines can skip data by path; a single query builder then drives both the transactional and the columnar engine. 01 Byproduct already being written 02 Retained zone kept, not expired 03 One query builder written once 04 Transactional engine the live database 05 Columnar engine reads files directly One definition, two places to run it
The columnar files the pipeline already produces are retained permanently and laid out so that engines can skip data by path; a single query builder then drives both the transactional and the columnar engine. 01 Byproduct already being written 02 Retained zone kept, not expired 03 One query builder written once 04 Transactional engine the live database 05 Columnar engine reads files directly
Fig. 04 — One definition of the query, two engines to run it on

Hover any step to read what it doesTap any step to read what it does