All projects

Decoupled
by design

Two halves of the same job can have completely different limits. Building the data needs processing power; saving it is limited by how fast a single database will accept writes. Join them together and the slower half sets the pace for everything.

System
Passenger profiling platform
Scale
535M rows, written continuously
Role
Sole architect and engineer
Stack
Python / Parquet / S3 / PostgreSQL
01

The challenge

Profile generation and database writes were bound together, so the fleet spent most of its time waiting on write locks instead of working.

02

The approach

Put a durable buffer between them and make the handoff a single atomic write, so neither side has to trust the other to be available.

03

The result

Compute and storage now scale independently, a failure is contained to one file, and the buffer became the foundation of the analytical layer.

Generation writes columnar part files to object storage and then publishes a single manifest, which the consumer claims under a lease before merging the data and recording the receipt in one transaction. Producer — scales with compute Consumer — bounded by the database 01 Generate compute-bound 02 Part files columnar, compressed 03 Manifest commit token 04 Claim lease, idempotent 05 Merge + receipt one transaction Decoupling boundary
Generation writes columnar part files to object storage and then publishes a single manifest, which the consumer claims under a lease before merging the data and recording the receipt in one transaction. Producer — scales with compute 01 Generate compute-bound 02 Part files columnar, compressed 03 Manifest commit token Consumer — bounded by the database 04 Claim lease, idempotent 05 Merge + receipt one transaction Decoupling boundary
Fig. 01 — The manifest is the only atomic write in the handoff

Hover any step to read what it doesTap any step to read what it does