
Test data for a streaming pipeline, from a settings file rather than a program: a plain YAML file describes a simulation, and one command runs it and sends its events to Kafka, PostgreSQL, Parquet or Iceberg.

Test data for a streaming pipeline, from a settings file rather than a program: a plain YAML file describes a simulation, and one command runs it and sends its events to Kafka, PostgreSQL, Parquet or Iceberg.

Keeping top 10 leaderboards up to date as game scores arrive: four Flink SQL jobs read the scores from Kafka and keep the rankings in PostgreSQL, and a web dashboard shows them as they change.

Capturing every insert and update in PostgreSQL without changing the application: Debezium streams the changes to Kafka, and an S3 sink connector saves them as files.

Small data engineering, stream processing, machine learning and MLOps projects that run on a laptop from a fresh clone, each explaining the system it builds.

Learning MLOps hands-on: the three projects from Jim Dowling's feature store book, rebuilt with open source tools that run locally.

Stopping a language model from inventing SQL over a lakehouse: a semantic layer between the model and Iceberg tables, built with Strands, WrenAI and Trino.

Generating test data that changes while it runs: SimPy simulations written with a declarative API that read settings from and write events to PostgreSQL and Redis.

Running Kafka, Flink, Spark, Trino, Iceberg and Airflow together on a laptop: one CLI starts them as a single local stack, with MLOps and observability tools.

One simulation that writes Parquet for model training and streams Kafka events for live inference, so training and serving use the same simulation code.

Tracking data across Kafka, Flink and Spark pipelines with OpenLineage, to show where each dataset came from and what reads it. Slides from my Current London 2026 session.