Jaehyeon Kim
Jaehyeon Kim

  • Blog
    • Categories

      List of categories.

    • Tags

      List of tags.

    • Series

      List of series.

    • Archives

  • Projects
  • Slides

/

  • Github Linkedin RSS

  • Font Size
  • Palette
  • Mode
  1. Home
  2. Categories
  3. Data Engineering

Data Lake Demo using Change Data Capture (CDC) on AWS - Part 2 Implement CDC

Data Lake Demo using Change Data Capture (CDC) on AWS - Part 2 Implement CDC
December 12, 202117 min read Data-Engineering Data Integration Data-Streaming Data Lake Demo Using Change Data CaptureAmazon EMRAmazon MSKApache HudiApache KafkaAWSChange Data Capture (CDC)DebeziumKafka Connect

Build change data capture on AWS with Amazon MSK and MSK Connect, streaming PostgreSQL row changes into the topics that feed the data lake.

Read More: Data Lake Demo using Change Data Capture (CDC) on AWS - Part 2 Implement CDC

Data Lake Demo using Change Data Capture (CDC) on AWS - Part 1 Local Development

Data Lake Demo using Change Data Capture (CDC) on AWS - Part 1 Local Development
December 5, 202118 min read Data-Engineering Data Integration Data-Streaming Data Lake Demo Using Change Data CaptureAmazon EMRAmazon MSKApache HudiApache KafkaAWSChange Data Capture (CDC)DebeziumKafka Connect

Set up the source PostgreSQL database with an outbox table, then run Debezium and an S3 sink connector locally with Docker Compose.

Read More: Data Lake Demo using Change Data Capture (CDC) on AWS - Part 1 Local Development

Local Development of AWS Glue 3.0 and Later

Local Development of AWS Glue 3.0 and Later
November 14, 20218 min read Data-EngineeringAWSAWS GlueDockerPySparkPython

Create a development environment for AWS Glue 3.0 and later by building a custom Docker image, because AWS publishes no image for those versions.

Read More: Local Development of AWS Glue 3.0 and Later

AWS Glue Local Development with Docker and Visual Studio Code

AWS Glue Local Development with Docker and Visual Studio Code
August 20, 20219 min read Data-EngineeringApache SparkAWSAWS GlueDockerPySparkPython

Build development environments for AWS Glue 1.0 and 2.0 with the published Docker image and the Visual Studio Code Remote Containers extension.

Read More: AWS Glue Local Development with Docker and Visual Studio Code

Thoughts on Apache Airflow AWS Lambda Operator

Thoughts on Apache Airflow AWS Lambda Operator
April 13, 202010 min read Data-EngineeringApache AirflowAWSAWS LambdaDockerPython

In this post, it is demonstrated how AWS Lambda can be integrated with Apache Airflow using a custom operator inspired by the ECS Operator.

Read More: Thoughts on Apache Airflow AWS Lambda Operator

Boost SparkR with Hive

April 30, 20168 min read Data-EngineeringApache HiveApache SparkHiveQLRSparkR

Run SparkR in Hive Context to reach the Hive UDFs and window functions the SQL Context lacks, compared against dplyr, plus building Spark with Hive.

Read More: Boost SparkR with Hive

Quick Start SparkR in Local and Cluster Mode

March 2, 20168 min read Data-EngineeringApache SparkRSparkR

Run SparkR in local and cluster mode from an R project, setting environment variables and the library path, then reading JSON and CSV files.

Read More: Quick Start SparkR in Local and Cluster Mode

Spark Cluster Setup on VirtualBox

February 22, 20164 min read Data-EngineeringApache SparkRSparkR

Set up a two node Spark standalone cluster on VirtualBox Ubuntu guests, covering machine preparation, copying the VDI image and password-less SSH.

Read More: Spark Cluster Setup on VirtualBox
  • ««
  • «
  • 1
  • 2
  • 3
  • 4
  • »
  • »»
Profile
Jaehyeon Kim
Jaehyeon Kim
Data Engineer | Data Streaming | Powering ML & AI in Real Time
Taxonomies
Data Streaming 70 Data Engineering 38 Development 28 Data Analysis 17 Data Integration 12 Open Source 8 Kubernetes 6 Machine Learning 6 Security 5 Data Architecture 3 Data Processing 3 Big Data 2 System Architecture 2 Web Development 2
Python 75 Apache Kafka 68 AWS 50 Docker 50 Apache Flink 34 Apache Beam 17 Apache Spark 16 Kafka Connect 15 Amazon MSK 14 AWS Lambda 14 dbt 13 Amazon EMR 11 Kubernetes 8 PyFlink 8 dynamic-des 7 Kotlin 7 PostgreSQL 7 Apache Airflow 6 Change Data Capture (CDC) 6 Debezium 6 Amazon DynamoDB 5 Apache Iceberg 5 Discrete Event Simulation 5 Factor House Local 5 PySpark 5 Amazon API Gateway 4 Amazon Athena 4 AWS Glue 4 AWS Glue Schema Registry 4 BigQuery 4 Digital Twin 4 FastAPI 4 Minikube 4 SimPy 4 Amazon EKS 3 Amazon QuickSight 3 Amazon S3 3 Apache Hudi 3 EMR on EKS 3 GCP 3 ALL 141
Kafka Development with Docker 11 Apache Beam Python Examples 10 Real Time Streaming with Kafka and Flink 7 dbt Pizza Shop Demo 6 Apache Beam Local Development with Python 5 Building Real-Time Digital Twins with dynamic-des 5 dbt for Effective Data Transformation on AWS 5 Getting Started with Real-Time Streaming in Kotlin 5 Kafka Connect for AWS Services Integration 5 Serverless Data Product 4 Data Lake Demo Using Change Data Capture 3 Getting Started with PyFlink on AWS 3 Kafka Development on Kubernetes 3 Parallel processing on single machine 3 Realtime Dashboard with FastAPI, Streamlit and Next.js 3 API development with R 2 dbt Guide for Production 2 Deploy Python Stream Processing App on Kubernetes 2 Download Stock Data 2 From Prototype to Production: Real-Time Product Recommendation with Contextual Bandits 2 ALL 25
2026 12 2025 13 2024 29 2023 39 2022 15 2021 7 2020 1 2019 5 2018 2 2017 6 2016 6 2015 15 2014 5
Posts
  • Learning MLOps with a Feature Store: A New Series
    Learning MLOps with a Feature Store: A New Series
    September 28, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • Introducing odctl: One CLI for a Local Open Data Stack
    Introducing odctl: One CLI for a Local Open Data Stack
    July 16, 2026
  • One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    May 25, 2026
  • Building an Event-Driven Hybrid Digital Twin with dynamic-des
    Building an Event-Driven Hybrid Digital Twin with dynamic-des
    April 29, 2026
  • Why Digital Twins Are Rewiring Industry 4.0
    Why Digital Twins Are Rewiring Industry 4.0
    April 22, 2026
  • Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    April 21, 2026
  • Slides as Code: Integrating Reveal.js into my Hugo Blog
    Slides as Code: Integrating Reveal.js into my Hugo Blog
    March 9, 2026
  • Productionizing an Online Product Recommender using Event Driven Architecture
    Productionizing an Online Product Recommender using Event Driven Architecture
    February 23, 2026
  • Learning MLOps with a Feature Store: A New Series
    Learning MLOps with a Feature Store: A New Series
    September 28, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • Introducing odctl: One CLI for a Local Open Data Stack
    Introducing odctl: One CLI for a Local Open Data Stack
    July 16, 2026
  • One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    May 25, 2026
  • Current London 2026: Building End-to-End Data Lineage
    Current London 2026: Building End-to-End Data Lineage
    May 22, 2026
  • Building an Event-Driven Hybrid Digital Twin with dynamic-des
    Building an Event-Driven Hybrid Digital Twin with dynamic-des
    April 29, 2026
  • Why Digital Twins Are Rewiring Industry 4.0
    Why Digital Twins Are Rewiring Industry 4.0
    April 22, 2026
  • Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    April 21, 2026
  • Slides as Code: Integrating Reveal.js into my Hugo Blog
    Slides as Code: Integrating Reveal.js into my Hugo Blog
    March 9, 2026
Actions
Go back Reload Copy URL

Jaehyeon Kim

Data Engineer | Data Streaming | Powering ML & AI in Real Time

Copyright © 2023-2026 Jaehyeon Kim. All Rights Reserved.