Jaehyeon Kim
Jaehyeon Kim

  • Blog
    • Categories

      List of categories.

    • Tags

      List of tags.

    • Series

      List of series.

    • Archives

  • Projects
  • Slides

/

  • Github Linkedin RSS

  • Font Size
  • Palette
  • Mode
  1. Home
  2. Tags
  3. Amazon EMR

Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images

Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images
December 7, 202317 min read Data-Engineering Data-Streaming DevelopmentAmazon EMRApache FlinkApache KafkaApache SparkAWSDocker

A local Flink and Spark environment built from EMR container images, where Flink ingests data in real time and Spark queries it via the Glue Data Catalog.

Read More: Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images

EMR on EKS - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 4

EMR on EKS - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 4
November 1, 202220 min read Data-Engineering Dbt for Effective Data Transformation on AWSAmazon EKSAmazon EMRApache SparkAWSDbtEMR on EKS

Transform IMDb data with dbt on Amazon EMR on EKS, through a Spark Thrift server kept running on Kubernetes by a small wrapper class.

Read More: EMR on EKS - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 4

EMR on EC2 - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 3

EMR on EC2 - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 3
October 19, 202219 min read Data-Engineering Dbt for Effective Data Transformation on AWSAmazon EMRAmazon QuickSightApache SparkAWSDbt

Transform IMDb data with dbt on Amazon EMR on EC2, running layered models through a Spark Thrift server.

Read More: EMR on EC2 - Data Build Tool (dbt) for Effective Data Transformation on AWS Part 3

Develop and Test Apache Spark Apps for EMR Remotely Using Visual Studio Code

Develop and Test Apache Spark Apps for EMR Remotely Using Visual Studio Code
September 7, 202216 min read Data-EngineeringAmazon EMRApache SparkAWSPySpark

Develop Spark apps on an EMR cluster in a private subnet over VPN and the VS Code remote SSH extension, with the cluster shared by several users.

Read More: Develop and Test Apache Spark Apps for EMR Remotely Using Visual Studio Code

Manage EMR on EKS with Terraform

Manage EMR on EKS with Terraform
August 26, 202212 min read Data-EngineeringAmazon EKSAmazon EMRApache SparkAWSEMR on EKSKarpenterTerraform

Provision EMR on EKS with Terraform and EKS Blueprints, then compare two Spark jobs with and without Dynamic Resource Allocation under Karpenter.

Read More: Manage EMR on EKS with Terraform

Data Warehousing ETL Demo with Apache Iceberg on EMR Local Environment

Data Warehousing ETL Demo with Apache Iceberg on EMR Local Environment
June 26, 202212 min read Data-EngineeringAmazon EMRApache IcebergApache SparkAWSPySparkPython

Run a data warehousing ETL job with Apache Iceberg for storage and PySpark for processing in an EMR local environment, then verify results in Athena.

Read More: Data Warehousing ETL Demo with Apache Iceberg on EMR Local Environment

Develop and Test Apache Spark Apps for EMR Locally Using Docker

Develop and Test Apache Spark Apps for EMR Locally Using Docker
May 8, 202218 min read Data-EngineeringAmazon EMRApache SparkAWSDockerPySpark

Create a Spark local development environment for Amazon EMR with Docker and Visual Studio Code, with Spark examples and Glue Catalog integration.

Read More: Develop and Test Apache Spark Apps for EMR Locally Using Docker

EMR on EKS by Example

EMR on EKS by Example
January 17, 202215 min read Data-EngineeringAmazon EKSAmazon EMRApache SparkAWSEMR on EKSKubernetes

Run Spark jobs on EMR on EKS, the EMR deployment option that provisions and manages open source big data frameworks, with simple and extended examples.

Read More: EMR on EKS by Example

Implement Data Lake - Data Lake Demo using Change Data Capture (CDC) on AWS Part 3

Implement Data Lake - Data Lake Demo using Change Data Capture (CDC) on AWS Part 3
December 19, 202112 min read Data-Engineering Data Integration Data-Streaming Data Lake Demo Using Change Data CaptureAmazon EMRAmazon MSKApache HudiApache KafkaAWSChange Data Capture (CDC)DebeziumKafka Connect

Run a Hudi DeltaStreamer application on Amazon EMR, then query the resulting Hudi table with Athena and build a QuickSight dashboard on it.

Read More: Implement Data Lake - Data Lake Demo using Change Data Capture (CDC) on AWS Part 3

Data Lake Demo using Change Data Capture (CDC) on AWS - Part 2 Implement CDC

Data Lake Demo using Change Data Capture (CDC) on AWS - Part 2 Implement CDC
December 12, 202117 min read Data-Engineering Data Integration Data-Streaming Data Lake Demo Using Change Data CaptureAmazon EMRAmazon MSKApache HudiApache KafkaAWSChange Data Capture (CDC)DebeziumKafka Connect

Build change data capture on AWS with Amazon MSK and MSK Connect, streaming PostgreSQL row changes into the topics that feed the data lake.

Read More: Data Lake Demo using Change Data Capture (CDC) on AWS - Part 2 Implement CDC
  • ««
  • «
  • 1
  • 2
  • »
  • »»
Profile
Jaehyeon Kim
Jaehyeon Kim
Data Engineer | Data Streaming | Powering ML & AI in Real Time
Taxonomies
Data Streaming 70 Data Engineering 43 Development 28 Data Analysis 17 Open Source 13 Data Integration 12 Machine Learning 7 Kubernetes 6 Security 5 Data Architecture 3 Data Processing 3 Big Data 2 System Architecture 2 Web Development 2
Python 77 Apache Kafka 73 AWS 50 Docker 50 Apache Flink 39 Apache Beam 17 Apache Spark 17 Kafka Connect 16 Amazon MSK 14 AWS Lambda 14 Benchtop 13 dbt 13 odctl 12 Amazon EMR 11 dynamic-des 11 PostgreSQL 10 Kubernetes 8 PyFlink 8 Apache Iceberg 7 Change Data Capture (CDC) 7 Debezium 7 Kotlin 7 Apache Airflow 6 Discrete Event Simulation 6 Kafka UI 6 Amazon DynamoDB 5 PySpark 5 SimPy 5 Amazon API Gateway 4 Amazon Athena 4 AWS Glue 4 AWS Glue Schema Registry 4 BigQuery 4 Digital Twin 4 FastAPI 4 Minikube 4 Amazon EKS 3 Amazon QuickSight 3 Amazon S3 3 Apache Hudi 3 ALL 143
Kafka Development with Docker 11 Apache Beam Python Examples 10 Real Time Streaming with Kafka and Flink 7 Building Real-Time Digital Twins with dynamic-des 6 dbt Pizza Shop Demo 6 Apache Beam Local Development with Python 5 dbt for Effective Data Transformation on AWS 5 Getting Started with Real-Time Streaming in Kotlin 5 Kafka Connect for AWS Services Integration 5 Serverless Data Product 4 Data Lake Demo Using Change Data Capture 3 Getting Started with PyFlink on AWS 3 Kafka Development on Kubernetes 3 Parallel processing on single machine 3 Realtime Dashboard with FastAPI, Streamlit and Next.js 3 API development with R 2 dbt Guide for Production 2 Deploy Python Stream Processing App on Kubernetes 2 Download Stock Data 2 From Prototype to Production: Real-Time Product Recommendation with Contextual Bandits 2 ALL 25
2026 17 2025 13 2024 29 2023 39 2022 15 2021 7 2020 1 2019 5 2018 2 2017 6 2016 6 2015 15 2014 5
Posts
  • odctl 1.0: Kafka, Flink, Spark, Iceberg, Trino and MLflow on Your Laptop with One Command
    odctl 1.0: Kafka, Flink, Spark, Iceberg, Trino and MLflow on Your Laptop with One Command
    October 12, 2026
  • Defining Data-Streaming Simulations in YAML, Without Writing Python
    Defining Data-Streaming Simulations in YAML, Without Writing Python
    October 6, 2026
  • Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    October 2, 2026
  • Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    October 1, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Current London 2026: Building End-to-End Data Lineage
    Current London 2026: Building End-to-End Data Lineage
    May 22, 2026
  • Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    April 21, 2026
  • Productionizing an Online Product Recommender using Event Driven Architecture
    Productionizing an Online Product Recommender using Event Driven Architecture
    February 23, 2026
  • Stream Processing with Flink in Kotlin
    Stream Processing with Flink in Kotlin
    December 10, 2025
  • Guide to Building Integrated Web Applications with FastAPI and NiceGUI
    Guide to Building Integrated Web Applications with FastAPI and NiceGUI
    November 19, 2025
  • odctl 1.0: Kafka, Flink, Spark, Iceberg, Trino and MLflow on Your Laptop with One Command
    odctl 1.0: Kafka, Flink, Spark, Iceberg, Trino and MLflow on Your Laptop with One Command
    October 12, 2026
  • Defining Data-Streaming Simulations in YAML, Without Writing Python
    Defining Data-Streaming Simulations in YAML, Without Writing Python
    October 6, 2026
  • Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    October 2, 2026
  • Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    October 1, 2026
  • Data Streaming and Machine Learning Projects That Run on Your Laptop
    Data Streaming and Machine Learning Projects That Run on Your Laptop
    September 30, 2026
  • Learning MLOps with a Feature Store: A New Series
    Learning MLOps with a Feature Store: A New Series
    September 28, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Dynamic DES: A Declarative API with Postgres and Redis Connectors
    Dynamic DES: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • Running Kafka, Flink, Spark, Trino and Iceberg Locally with One CLI
    Running Kafka, Flink, Spark, Trino and Iceberg Locally with One CLI
    July 16, 2026
  • One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES
    May 25, 2026
Actions
Go back Reload Copy URL

Jaehyeon Kim

Data Engineer | Data Streaming | Powering ML & AI in Real Time

Copyright © 2023-2026 Jaehyeon Kim. All Rights Reserved.