Open Data Stack (odctl)
odctl is a command-line tool that runs an open data stack on your machine with Docker Compose. You name the profiles you want, such as kafka-lite or airflow, and odctl starts them together with the services they depend on.

What it runs
| Area | Technologies |
|---|---|
| Messaging | Kafka (KRaft), Karapace schema registry, Kafka Connect, Kafka UI (kafbat) |
| Stream and batch processing | Apache Flink, Apache Spark |
| Analytics | ClickHouse, Trino, Metabase |
| Orchestration | Apache Airflow, Temporal |
| MLOps | MLflow, Feast, Evidently |
| Metadata and lineage | OpenMetadata, Marquez |
| Observability | grafana/otel-lgtm (OpenTelemetry Collector, Prometheus, Tempo, Loki, Grafana) |
| Storage and catalog | PostgreSQL 18 with pgvector, pg_textsearch and PostGIS, SeaweedFS (S3), Iceberg REST catalog, Valkey Bundle, Apache Fluss |
Kafka Connect comes with a set of connectors, such as the Iceberg sink and the Debezium PostgreSQL source. Kafka Connect lists them all.
Where to go next
- Getting started installs odctl and starts a first profile.
- Profiles lists every profile with its ports, images and memory limits.
- Guides show how to use the services together.
Projects that use odctl
- benchtop: small data engineering and machine learning projects, each run from a cold clone on odctl.
- dynamic-des: a SimPy library for simulations that stream to Kafka, PostgreSQL, Redis and Iceberg.
- agentic-analytics-system: conversational analytics over an Iceberg lakehouse.
Blog posts about odctl and these projects are tagged odctl on jaehyeon.me.