
Test data for a streaming pipeline, from a settings file rather than a program: a plain YAML file describes a simulation, and one command runs it and sends its events to Kafka, PostgreSQL, Parquet or Iceberg.

Test data for a streaming pipeline, from a settings file rather than a program: a plain YAML file describes a simulation, and one command runs it and sends its events to Kafka, PostgreSQL, Parquet or Iceberg.

Learning MLOps hands-on: the three projects from Jim Dowling's feature store book, rebuilt with open source tools that run locally.

Stopping a language model from inventing SQL over a lakehouse: a semantic layer between the model and Iceberg tables, built with Strands, WrenAI and Trino.

Running Kafka, Flink, Spark, Trino, Iceberg and Airflow together on a laptop: one CLI starts them as a single local stack, with MLOps and observability tools.

Apache Paimon, Fluss and Apache Iceberg compared as table layers for streaming and batch, then combined into one architecture with Flink and Spark.

Run a data warehousing ETL job with Apache Iceberg for storage and PySpark for processing in an EMR local environment, then verify results in Athena.