Jaehyeon Kim
Jaehyeon Kim

  • Blog
    • Categories

      List of categories.

    • Tags

      List of tags.

    • Series

      List of series.

    • Archives

  • Projects
  • Slides

/

  • Github Linkedin RSS

  • Font Size
  • Palette
  • Mode
  1. Home
  2. Tags
  3. Apache Spark

Introducing odctl: One CLI for a Local Open Data Stack

Introducing odctl: One CLI for a Local Open Data Stack
July 16, 20264 min read Data-Engineering Open SourceApache FlinkApache IcebergApache KafkaApache SparkOdctlOpen Data StackPythonTrino

odctl is a CLI on PyPI that launches Kafka, Flink, Spark, Trino, Iceberg, Airflow and an MLOps and observability suite as one local stack.

Read More: Introducing odctl: One CLI for a Local Open Data Stack

Current London 2026: Building End-to-End Data Lineage

Current London 2026: Building End-to-End Data Lineage
May 22, 20263 min read Data-EngineeringApache FlinkApache KafkaApache SparkData LineageOpenLineage

A comprehensive walkthrough from my session at Current London 2026 on capturing and visualizing data lineage across a production-style data stack.

Read More: Current London 2026: Building End-to-End Data Lineage

Self-service Data Platform via a Multi-tenant SQL Gateway

Self-service Data Platform via a Multi-tenant SQL Gateway
July 17, 202515 min read Big Data Data Architecture Data-Engineering Data-StreamingApache FlinkApache KyuubiApache RangerApache SparkMarquezOpenLineageSQL GatewayTrino

Apache Kyuubi as a multi-tenant SQL gateway that provisions on-demand Spark, Flink and Trino engines, giving self-service analytics with central governance.

Read More: Self-service Data Platform via a Multi-tenant SQL Gateway

Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images

Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images
December 7, 202317 min read Data-Engineering Data-Streaming DevelopmentAmazon EMRApache FlinkApache KafkaApache SparkAWSDocker

A local Flink and Spark environment built from EMR container images, where Flink ingests data in real time and Spark queries it via the Glue Data Catalog.

Read More: Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images

EMR on EKS - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 4

EMR on EKS - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 4
November 1, 202220 min read Data-Engineering Dbt for Effective Data Transformation on AWSAmazon EKSAmazon EMRApache SparkAWSDbtEMR on EKS

Amazon EMR on EKS data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

Read More: EMR on EKS - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 4

EMR on EC2 - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 3

EMR on EC2 - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 3
October 19, 202219 min read Data-Engineering Dbt for Effective Data Transformation on AWSAmazon EMRAmazon QuickSightApache SparkAWSDbt

Amazon EMR on EC2 data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

Read More: EMR on EC2 - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 3

Data Build Tool (dbt) for Effective Data Transformation on AWS - Part 2 Glue

Data Build Tool (dbt) for Effective Data Transformation on AWS - Part 2 Glue
October 9, 202218 min read Data-Engineering Dbt for Effective Data Transformation on AWSAmazon QuickSightApache SparkAWSAWS GlueDbt

AWS Glue data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

Read More: Data Build Tool (dbt) for Effective Data Transformation on AWS - Part 2 Glue

Develop and Test Apache Spark Apps for EMR Remotely Using Visual Studio Code

Develop and Test Apache Spark Apps for EMR Remotely Using Visual Studio Code
September 7, 202216 min read Data-EngineeringAmazon EMRApache SparkAWSPySpark

Develop Spark apps on an EMR cluster in a private subnet over VPN and the VS Code remote SSH extension, with the cluster shared by several users.

Read More: Develop and Test Apache Spark Apps for EMR Remotely Using Visual Studio Code

Manage EMR on EKS with Terraform

Manage EMR on EKS with Terraform
August 26, 202212 min read Data-EngineeringAmazon EKSAmazon EMRApache SparkAWSEMR on EKSKarpenterTerraform

Provision EMR on EKS with Terraform and EKS Blueprints, then compare two Spark jobs with and without Dynamic Resource Allocation under Karpenter.

Read More: Manage EMR on EKS with Terraform

Data Warehousing ETL Demo with Apache Iceberg on EMR Local Environment

Data Warehousing ETL Demo with Apache Iceberg on EMR Local Environment
June 26, 202212 min read Data-EngineeringAmazon EMRApache IcebergApache SparkAWSPySparkPython

Run a data warehousing ETL job with Apache Iceberg for storage and PySpark for processing in an EMR local environment, then verify results in Athena.

Read More: Data Warehousing ETL Demo with Apache Iceberg on EMR Local Environment
  • ««
  • «
  • 1
  • 2
  • »
  • »»
Profile
Jaehyeon Kim
Jaehyeon Kim
Data Engineer | Data Streaming | Powering ML & AI in Real Time
Taxonomies
Data Streaming 70 Data Engineering 41 Development 28 Data Analysis 17 Data Integration 12 Open Source 11 Machine Learning 7 Kubernetes 6 Security 5 Data Architecture 3 Data Processing 3 Big Data 2 System Architecture 2 Web Development 2
Python 75 Apache Kafka 71 AWS 50 Docker 50 Apache Flink 36 Apache Beam 17 Apache Spark 16 Kafka Connect 16 Amazon MSK 14 AWS Lambda 14 dbt 13 Amazon EMR 11 odctl 11 dynamic-des 10 PostgreSQL 9 Kubernetes 8 PyFlink 8 Debezium 7 Kotlin 7 Apache Airflow 6 Change Data Capture (CDC) 6 Kafka UI 6 Amazon DynamoDB 5 Apache Iceberg 5 Discrete Event Simulation 5 PySpark 5 Amazon API Gateway 4 Amazon Athena 4 AWS Glue 4 AWS Glue Schema Registry 4 BigQuery 4 Digital Twin 4 FastAPI 4 Minikube 4 SimPy 4 Amazon EKS 3 Amazon QuickSight 3 Amazon S3 3 Apache Hudi 3 EMR on EKS 3 ALL 142
Kafka Development with Docker 11 Apache Beam Python Examples 10 Real Time Streaming with Kafka and Flink 7 dbt Pizza Shop Demo 6 Apache Beam Local Development with Python 5 Building Real-Time Digital Twins with dynamic-des 5 dbt for Effective Data Transformation on AWS 5 Getting Started with Real-Time Streaming in Kotlin 5 Kafka Connect for AWS Services Integration 5 Serverless Data Product 4 Data Lake Demo Using Change Data Capture 3 Getting Started with PyFlink on AWS 3 Kafka Development on Kubernetes 3 Parallel processing on single machine 3 Realtime Dashboard with FastAPI, Streamlit and Next.js 3 API development with R 2 dbt Guide for Production 2 Deploy Python Stream Processing App on Kubernetes 2 Download Stock Data 2 From Prototype to Production: Real-Time Product Recommendation with Contextual Bandits 2 ALL 25
2026 15 2025 13 2024 29 2023 39 2022 15 2021 7 2020 1 2019 5 2018 2 2017 6 2016 6 2015 15 2014 5
Posts
  • Live Game Leaderboards with Kafka, Flink SQL and a Discrete-Event Simulation
    Live Game Leaderboards with Kafka, Flink SQL and a Discrete-Event Simulation
    October 2, 2026
  • Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    October 1, 2026
  • Introducing Benchtop: Data Streaming and Machine Learning Projects You Can Run Locally
    Introducing Benchtop: Data Streaming and Machine Learning Projects You Can Run Locally
    September 30, 2026
  • Learning MLOps with a Feature Store: A New Series
    Learning MLOps with a Feature Store: A New Series
    September 28, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • Introducing odctl: One CLI for a Local Open Data Stack
    Introducing odctl: One CLI for a Local Open Data Stack
    July 16, 2026
  • One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    May 25, 2026
  • Building an Event-Driven Hybrid Digital Twin with dynamic-des
    Building an Event-Driven Hybrid Digital Twin with dynamic-des
    April 29, 2026
  • Why Digital Twins Are Rewiring Industry 4.0
    Why Digital Twins Are Rewiring Industry 4.0
    April 22, 2026
  • Live Game Leaderboards with Kafka, Flink SQL and a Discrete-Event Simulation
    Live Game Leaderboards with Kafka, Flink SQL and a Discrete-Event Simulation
    October 2, 2026
  • Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    October 1, 2026
  • Introducing Benchtop: Data Streaming and Machine Learning Projects You Can Run Locally
    Introducing Benchtop: Data Streaming and Machine Learning Projects You Can Run Locally
    September 30, 2026
  • Learning MLOps with a Feature Store: A New Series
    Learning MLOps with a Feature Store: A New Series
    September 28, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • Introducing odctl: One CLI for a Local Open Data Stack
    Introducing odctl: One CLI for a Local Open Data Stack
    July 16, 2026
  • One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    May 25, 2026
  • Current London 2026: Building End-to-End Data Lineage
    Current London 2026: Building End-to-End Data Lineage
    May 22, 2026
  • Building an Event-Driven Hybrid Digital Twin with dynamic-des
    Building an Event-Driven Hybrid Digital Twin with dynamic-des
    April 29, 2026
Actions
Go back Reload Copy URL

Jaehyeon Kim

Data Engineer | Data Streaming | Powering ML & AI in Real Time

Copyright © 2023-2026 Jaehyeon Kim. All Rights Reserved.