Skip to content

Open Data Stack (odctl)

odctl is a command-line tool that runs an open data stack on your machine with Docker Compose. You name the profiles you want, such as kafka-lite or airflow, and odctl starts them together with the services they depend on.

odctl architecture

What it runs

Area Technologies
Messaging Kafka (KRaft), Karapace schema registry, Kafka Connect, Kafka UI (kafbat)
Stream and batch processing Apache Flink, Apache Spark
Analytics ClickHouse, Trino, Metabase
Orchestration Apache Airflow, Temporal
MLOps MLflow, Feast, Evidently
Metadata and lineage OpenMetadata, Marquez
Observability grafana/otel-lgtm (OpenTelemetry Collector, Prometheus, Tempo, Loki, Grafana)
Storage and catalog PostgreSQL 18 with pgvector, pg_textsearch and PostGIS, SeaweedFS (S3), Iceberg REST catalog, Valkey Bundle, Apache Fluss

Kafka Connect comes with a set of connectors, such as the Iceberg sink and the Debezium PostgreSQL source. Kafka Connect lists them all.

Where to go next

  • Getting started installs odctl and starts a first profile.
  • Profiles lists every profile with its ports, images and memory limits.
  • Guides show how to use the services together.

Projects that use odctl

  • benchtop: small data engineering and machine learning projects, each run from a cold clone on odctl.
  • dynamic-des: a SimPy library for simulations that stream to Kafka, PostgreSQL, Redis and Iceberg.
  • agentic-analytics-system: conversational analytics over an Iceberg lakehouse.

Blog posts about odctl and these projects are tagged odctl on jaehyeon.me.