Jaehyeon Kim
Jaehyeon Kim

  • Blog
    • Categories

      List of categories.

    • Tags

      List of tags.

    • Series

      List of series.

    • Archives

  • Projects
  • Slides

/

  • Github Linkedin RSS

  • Font Size
  • Palette
  • Mode

Call RPC Service in Batch using Stateless DoFn - Apache Beam Python Examples Part 5

featured.png
September 18, 202412 min read Data-Streaming Apache Beam Python ExamplesApache BeamApache FlinkApache KafkaGRPCPython

Batching gRPC calls in a stateless DoFn so one request covers a whole bundle, cutting the time a Beam Python pipeline spends on enrichment.

Read More

Guide to Running DBT in Production

featured.png
September 13, 202423 min read Data-Engineering Dbt Guide for ProductionBigQueryContinuous DeliveryContinuous IntegrationDbtGitHub Actions

Deploying a dbt project to dev and prod on BigQuery, covering slim CI, unit tests and a write audit publish step that builds on a cloned dataset.

Read More

DBT CI/CD Demo with BigQuery and GitHub Actions

featured.png
September 5, 202418 min read Data-Engineering Dbt Guide for ProductionBigQueryContinuous DeliveryContinuous IntegrationDbtGitHub Actions

GitHub Actions gives a dbt project on BigQuery a slim CI run on pull requests and a deploy job that publishes the project as a container image.

Read More

Cache Data on Apache Beam Pipelines Using a Shared Object

featured.png
August 22, 20248 min read Data ProcessingApache BeamCachingData EnrichmentPython

The Shared class in the Beam Python SDK caches lookup data in memory for batch and streaming pipelines, with a periodic refresh for the latter.

Read More

Apache Beam Python Examples - Part 4 Call RPC Service for Data Augmentation

featured.png
August 15, 202414 min read Data-Streaming Apache Beam Python ExamplesApache BeamApache FlinkApache KafkaGRPCPython

Data augmentation in Beam Python by calling a gRPC service once per input element, running on a local Flink cluster with Kafka as the source.

Read More

Build Sport Activity Tracker with/without SQL - Apache Beam Python Examples Part 3

featured.png
August 1, 202420 min read Data-Streaming Apache Beam Python ExamplesApache BeamApache FlinkApache KafkaPython

A sport activity tracker in Beam Python, built first with native transforms and then with Beam SQL, showing the limits of Beam SQL in the Python SDK.

Read More

Calculate Average Word Length with/without Fixed Look back - Apache Beam Python Examples Part 2

featured.png
July 18, 202415 min read Data-Streaming Apache Beam Python ExamplesApache BeamApache FlinkApache KafkaPython

Two Beam Python pipelines compute average word length from a Kafka topic, one emitting a global average and one using a sliding time window.

Read More

Calculate K Most Frequent Words and Max Word Length - Apache Beam Python Examples Part 1

featured.png
July 4, 202423 min read Data-Streaming Apache Beam Python ExamplesApache BeamApache FlinkApache KafkaPython

Set up a local Apache Flink and Kafka environment, then build two Beam Python streaming pipelines for top K frequent words and longest word length.

Read More

Beam Pipeline on Flink Runner - Deploy Python Stream Processing App on Kubernetes Part 2

featured.png
June 6, 202416 min read Data-Streaming Kubernetes Deploy Python Stream Processing App on KubernetesApache BeamApache FlinkApache KafkaDockerKubernetesPython

Deploy an Apache Beam Python pipeline to a Flink session cluster on minikube, packaged as a Docker image and submitted as a Kubernetes job.

Read More

Deploy Python Stream Processing App on Kubernetes - Part 1 PyFlink Application

featured.png
May 30, 202413 min read Data-Streaming Kubernetes Deploy Python Stream Processing App on KubernetesApache FlinkApache KafkaDockerKubernetesPython

Deploy a PyFlink application to minikube with the Flink Kubernetes Operator, alongside a Kafka cluster that provides its source and sink topics.

Read More
  • ««
  • «
  • 2
  • 3
  • 4
  • 5
  • 6
  • »
  • »»
Profile
Jaehyeon Kim
Jaehyeon Kim
Data Engineer | Data Streaming Enthusiast | Powering ML & AI in Real Time
Taxonomies
Data Streaming 70 Data Engineering 38 Development 28 Data Analysis 17 Data Integration 12 Open Source 7 Kubernetes 6 Machine Learning 5 Security 5 Data Architecture 3 Data Processing 3 Big Data 2 System Architecture 2 Web Development 2
Python 75 Apache Kafka 68 AWS 50 Docker 50 R 37 Apache Flink 34 Kpow 22 Apache Beam 17 Apache Spark 16 Kafka Connect 15 Amazon MSK 14 AWS Lambda 14 dbt 13 Amazon EMR 11 Kubernetes 8 PyFlink 8 Kotlin 7 PostgreSQL 7 Change Data Capture (CDC) 6 Debezium 6 dynamic-des 6 Amazon DynamoDB 5 Apache Airflow 5 Discrete Event Simulation 5 Factor House Local 5 PySpark 5 Amazon API Gateway 4 Amazon Athena 4 Apache Iceberg 4 AWS Glue 4 AWS Glue Schema Registry 4 BigQuery 4 Digital Twin 4 FastAPI 4 Minikube 4 R Shiny 4 RServe 4 SimPy 4 Amazon EKS 3 Amazon QuickSight 3 ALL 137
Kafka Development with Docker 11 Apache Beam Python Examples 10 Real Time Streaming with Kafka and Flink 7 dbt Pizza Shop Demo 6 Tree Based Methods in R 6 Apache Beam Local Development with Python 5 Building Real-Time Digital Twins with dynamic-des 5 dbt for Effective Data Transformation on AWS 5 Getting Started with Real-Time Streaming in Kotlin 5 Kafka Connect for AWS Services Integration 5 Serverless Data Product 4 Data Lake Demo Using Change Data Capture 3 Getting Started with PyFlink on AWS 3 Kafka Development on Kubernetes 3 Parallel processing on single machine 3 Realtime Dashboard with FastAPI, Streamlit and Next.js 3 API development with R 2 dbt Guide for Production 2 Deploy Python Stream Processing App on Kubernetes 2 Download Stock Data 2 ALL 24
2026 11 2025 13 2024 29 2023 39 2022 15 2021 7 2020 1 2019 5 2018 2 2017 6 2016 6 2015 15 2014 5
Posts
  • featured.png
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • featured.png
    Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • featured.png
    Introducing odctl: One CLI for a Local Open Data Stack
    July 16, 2026
  • featured.png
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    May 25, 2026
  • featured.png
    Building an Event-Driven Hybrid Digital Twin with dynamic-des
    April 29, 2026
  • featured.png
    Why Digital Twins Are Rewiring Industry 4.0
    April 22, 2026
  • featured.png
    Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    April 21, 2026
  • featured.png
    Slides as Code: Integrating Reveal.js into my Hugo Blog
    March 9, 2026
  • featured.gif
    Productionizing an Online Product Recommender using Event Driven Architecture
    February 23, 2026
  • featured.png
    Stream Processing with Flink in Kotlin
    December 10, 2025
  • featured.png
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • featured.png
    Dynamic DES v0.11.1: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • featured.png
    Introducing odctl: One CLI for a Local Open Data Stack
    July 16, 2026
  • featured.png
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES v0.8.1
    May 25, 2026
  • featured.jfif
    Current London 2026: Building End-to-End Data Lineage
    May 22, 2026
  • featured.png
    Building an Event-Driven Hybrid Digital Twin with dynamic-des
    April 29, 2026
  • featured.png
    Why Digital Twins Are Rewiring Industry 4.0
    April 22, 2026
  • featured.png
    Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    April 21, 2026
  • featured.png
    Slides as Code: Integrating Reveal.js into my Hugo Blog
    March 9, 2026
  • featured.gif
    Productionizing an Online Product Recommender using Event Driven Architecture
    February 23, 2026
Actions
Go back Reload Copy URL

Jaehyeon Kim

Data Engineer | Data Streaming Enthusiast | Powering ML & AI in Real Time

Copyright © 2023-2026 Jaehyeon Kim. All Rights Reserved.