<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Benchtop on Jaehyeon Kim</title><link>https://jaehyeon.me/tags/benchtop/</link><description>Recent content in Benchtop on Jaehyeon Kim</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>Copyright © 2023-2026 Jaehyeon Kim. All Rights Reserved.</copyright><lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://jaehyeon.me/tags/benchtop/index.xml" rel="self" type="application/rss+xml"/><item><title>Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL</title><link>https://jaehyeon.me/blog/2026-10-02-game-leaderboard-flink-sql/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2026-10-02-game-leaderboard-flink-sql/</guid><description>&lt;p>A mobile game produces a score every time a player finishes a round, and players expect the leaderboards to move as they play. In this post, a simulation plays the game and sends each score to Kafka. Four Flink SQL jobs keep four top 10 leaderboards up to date in PostgreSQL, and a web dashboard shows them as they change. Everything runs on your own machine.&lt;/p></description><enclosure url="https://jaehyeon.me/blog/2026-10-02-game-leaderboard-flink-sql/featured.png" length="451195" type="image/png"/></item><item><title>Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect</title><link>https://jaehyeon.me/blog/2026-10-01-ecommerce-cdc-debezium-kafka-connect/</link><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2026-10-01-ecommerce-cdc-debezium-kafka-connect/</guid><description>&lt;p>Most databases change all day: users sign up, orders are placed, and each order moves from one status to the next. Change data capture (CDC) turns those changes into a stream of events that other systems can read as they happen. In this post, a simulated online shop writes to PostgreSQL in real time, Debezium streams every change to Kafka, and a sink connector saves the changes as files in object storage. Everything runs on your own machine.&lt;/p></description><enclosure url="https://jaehyeon.me/blog/2026-10-01-ecommerce-cdc-debezium-kafka-connect/featured.png" length="583526" type="image/png"/></item><item><title>Data Streaming and Machine Learning Projects That Run on Your Laptop</title><link>https://jaehyeon.me/blog/2026-09-30-introducing-benchtop/</link><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2026-09-30-introducing-benchtop/</guid><description><![CDATA[<p>Reading about a data tool only goes so far. The hard part is getting several tools to run together and seeing how data moves between them. <a href="https://github.com/jaehyeon-kim/benchtop" target="_blank" rel="noopener noreferrer">Benchtop<i class="fas fa-external-link-square-alt ms-1"></i></a> is a collection of small, hands-on demos for data engineering, stream processing, machine learning, AI engineering and MLOps. Each demo runs locally on the odctl stack, and works from a fresh clone of the repository, with nothing to install beforehand but Docker and uv. Each one builds a working system from open-source tools and explains the ideas behind it, so you learn a tool by running it rather than only reading about it.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2026-09-30-introducing-benchtop/featured.png" length="15991" type="image/png"/></item><item><title>Productionizing an Online Product Recommender using Event Driven Architecture</title><link>https://jaehyeon.me/blog/2026-02-23-productionize-recommender-with-eda/</link><pubDate>Mon, 23 Feb 2026 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2026-02-23-productionize-recommender-with-eda/</guid><description><![CDATA[<p>Real-world recommendation systems require low-latency inference for users and high-throughput training for model updates. A monolithic script cannot handle production scale, while it is effective for testing algorithms locally. In <a href="/blog/2026-01-29-prototype-recommender-with-python/"><strong>Part 1</strong></a>, we built a contextual bandit prototype using Python and <a href="https://github.com/fidelity/mab2rec" target="_blank" rel="noopener noreferrer"><code>Mab2Rec</code><i class="fas fa-external-link-square-alt ms-1"></i></a>.</p>
<p>This post demonstrates how to decouple these concerns using an event-driven architecture with Apache Flink, Kafka, and Valkey.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2026-02-23-productionize-recommender-with-eda/featured.png" length="451606" type="image/png"/></item><item><title>Prototyping an Online Product Recommender in Python</title><link>https://jaehyeon.me/blog/2026-01-29-prototype-recommender-with-python/</link><pubDate>Tue, 27 Jan 2026 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2026-01-29-prototype-recommender-with-python/</guid><description>Overview Traditional recommendation approaches such as Collaborative Filtering remain widely adopted, yet they come with notable constraints. They are particularly vulnerable to the cold-start problem, where new users lack sufficient interaction history, and they depend heavily on long-term behavioral data. As a result, they frequently overlook real-time contextual signals, including time of day, device type, location, or session intent. This can prevent them from capturing situational preferences, such as someone preferring coffee in the morning but pizza in the evening.</description><enclosure url="https://jaehyeon.me/blog/2026-01-29-prototype-recommender-with-python/featured.png" length="340249" type="image/png"/></item><item><title>Flink Table API - Declarative Analytics for Supplier Stats in Real Time</title><link>https://jaehyeon.me/blog/2025-06-17-kotlin-getting-started-flink-table/</link><pubDate>Tue, 17 Jun 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-06-17-kotlin-getting-started-flink-table/</guid><description><![CDATA[<p>In the last post, we explored the fine-grained control of Flink&rsquo;s DataStream API. Now, we&rsquo;ll approach the same problem from a higher level of abstraction using the <strong>Flink Table API</strong>. This post demonstrates how to build a declarative analytics pipeline that processes our continuous stream of Avro-formatted order events. We will define a <code>Table</code> on top of a <code>DataStream</code> and use SQL-like expressions to perform windowed aggregations. This example highlights the power and simplicity of the Table API for analytical tasks and showcases Flink&rsquo;s seamless integration between its different API layers to handle complex requirements like late data.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2025-06-17-kotlin-getting-started-flink-table/featured.png" length="144113" type="image/png"/></item><item><title>Flink DataStream API - Scalable Event Processing for Supplier Stats</title><link>https://jaehyeon.me/blog/2025-06-10-kotlin-getting-started-flink-datastream/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-06-10-kotlin-getting-started-flink-datastream/</guid><description><![CDATA[<p>Building on our exploration of stream processing, we now transition from Kafka&rsquo;s native library to <strong>Apache Flink</strong>, a powerful, general-purpose distributed processing engine. In this post, we&rsquo;ll dive into Flink&rsquo;s foundational <strong>DataStream API</strong>. We will tackle the same supplier statistics problem - analyzing a stream of Avro-formatted order events - but this time using Flink&rsquo;s robust features for stateful computation. This example will highlight Flink&rsquo;s sophisticated event-time processing with watermarks and its elegant, built-in mechanisms for handling late-arriving data through side outputs.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2025-06-10-kotlin-getting-started-flink-datastream/featured.png" length="142918" type="image/png"/></item><item><title>Kafka Streams - Lightweight Real-Time Processing for Supplier Stats</title><link>https://jaehyeon.me/blog/2025-06-03-kotlin-getting-started-kafka-streams/</link><pubDate>Tue, 03 Jun 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-06-03-kotlin-getting-started-kafka-streams/</guid><description>&lt;p>A Kotlin application analyses a continuous stream of Avro-formatted order events, calculates supplier statistics in tumbling windows, and handles late-arriving data. This post shifts our focus from basic Kafka clients to real-time stream processing with &lt;strong>Kafka Streams&lt;/strong>. The example demonstrates the power of Kafka Streams for building lightweight, yet robust, stream processing applications directly within your Kafka ecosystem, using event-time processing and custom logic.&lt;/p></description><enclosure url="https://jaehyeon.me/blog/2025-06-03-kotlin-getting-started-kafka-streams/featured.png" length="131804" type="image/png"/></item><item><title>Kafka Clients with Avro - Schema Registry and Order Events</title><link>https://jaehyeon.me/blog/2025-05-27-kotlin-getting-started-kafka-avro-clients/</link><pubDate>Tue, 27 May 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-05-27-kotlin-getting-started-kafka-avro-clients/</guid><description>&lt;p>A Kafka producer that generates mock order data and a consumer that processes these orders are built with Kotlin, Apache Avro for data serialization, and Gradle for build management. This example highlights best practices such as schema management with Avro, robust error handling, and graceful shutdown, providing a solid foundation for your own Kafka-based projects. We&amp;rsquo;ll dive into the build configuration, the Avro schema definition, utility functions for Kafka administration, and the core logic of both the producer and consumer applications.&lt;/p></description><enclosure url="https://jaehyeon.me/blog/2025-05-27-kotlin-getting-started-kafka-avro-clients/featured.png" length="73988" type="image/png"/></item><item><title>Kafka Clients with JSON - Producing and Consuming Order Events</title><link>https://jaehyeon.me/blog/2025-05-20-kotlin-getting-started-kafka-json-clients/</link><pubDate>Tue, 20 May 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-05-20-kotlin-getting-started-kafka-json-clients/</guid><description>&lt;p>A Kafka producer application generates and sends order data, and a Kafka consumer application receives and processes those orders. This post explores that Kotlin-based Kafka project, detailing the construction and operation of both applications. We&amp;rsquo;ll go through each component, from build configuration to message handling, to understand how they work together in an event-driven system.&lt;/p></description><enclosure url="https://jaehyeon.me/blog/2025-05-20-kotlin-getting-started-kafka-json-clients/featured.png" length="97922" type="image/png"/></item><item><title>Next.js Dashboard - Realtime Dashboard with FastAPI, Streamlit and Next.js Part 3</title><link>https://jaehyeon.me/blog/2025-03-04-realtime-dashboard-3/</link><pubDate>Tue, 04 Mar 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-03-04-realtime-dashboard-3/</guid><description><![CDATA[<p>A real-time monitoring dashboard connects to the WebSocket server from <a href="/blog/2025-02-18-realtime-dashboard-1">Part 1</a> to continuously fetch and visualize key metrics such as <strong>order counts</strong>, <strong>sales data</strong>, and <strong>revenue by traffic source and country</strong>. With interactive bar charts and dynamic metrics, users can monitor sales trends and other critical business KPIs in real-time. In this post, we build it using <a href="https://nextjs.org/" target="_blank" rel="noopener noreferrer">Next.js<i class="fas fa-external-link-square-alt ms-1"></i></a>, a React framework that supports server-side rendering, static site generation, and full-stack capabilities with built-in performance optimizations. It is similar to the <em>Streamlit</em> app we developed in <a href="/blog/2025-02-25-realtime-dashboard-2">Part 2</a>.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2025-03-04-realtime-dashboard-3/featured.png" length="65389" type="image/png"/></item><item><title>Streamlit Dashboard - Realtime Dashboard with FastAPI, Streamlit and Next.js Part 2</title><link>https://jaehyeon.me/blog/2025-02-25-realtime-dashboard-2/</link><pubDate>Tue, 25 Feb 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-02-25-realtime-dashboard-2/</guid><description><![CDATA[<p>A real-time monitoring dashboard is developed using <a href="https://streamlit.io/" target="_blank" rel="noopener noreferrer">Streamlit<i class="fas fa-external-link-square-alt ms-1"></i></a>, an open-source Python framework that allows data scientists and AI/ML engineers to create interactive data apps. The app connects to the WebSocket server we developed in <a href="/blog/2025-02-18-realtime-dashboard-1">Part 1</a> and continuously fetches data to visualize key metrics such as <strong>order counts</strong>, <strong>sales data</strong>, and <strong>revenue by traffic source and country</strong>. With interactive bar charts and dynamic metrics, users can monitor sales trends and other important business KPIs in real-time.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2025-02-25-realtime-dashboard-2/featured.png" length="60147" type="image/png"/></item><item><title>Realtime Dashboard with FastAPI, Streamlit and Next.js - Part 1 Data Producer</title><link>https://jaehyeon.me/blog/2025-02-18-realtime-dashboard-1/</link><pubDate>Tue, 18 Feb 2025 00:00:00 +0000</pubDate><guid>https://jaehyeon.me/blog/2025-02-18-realtime-dashboard-1/</guid><description><![CDATA[<p>A data generating app is created with Python, and it ingests the <a href="https://console.cloud.google.com/marketplace/product/bigquery-public-data/thelook-ecommerce" target="_blank" rel="noopener noreferrer">theLook eCommerce<i class="fas fa-external-link-square-alt ms-1"></i></a> data continuously into a PostgreSQL database. A WebSocket server, built by <a href="https://fastapi.tiangolo.com/" target="_blank" rel="noopener noreferrer">FastAPI<i class="fas fa-external-link-square-alt ms-1"></i></a>, periodically queries the data to serve its clients. In this series, we develop real-time monitoring dashboard applications, and this post walks through the data generation app and backend API. The monitoring dashboards will be developed using <a href="https://streamlit.io/" target="_blank" rel="noopener noreferrer">Streamlit<i class="fas fa-external-link-square-alt ms-1"></i></a> and <a href="https://nextjs.org/" target="_blank" rel="noopener noreferrer">Next.js<i class="fas fa-external-link-square-alt ms-1"></i></a>, with <a href="https://echarts.apache.org/en/index.html" target="_blank" rel="noopener noreferrer">Apache ECharts<i class="fas fa-external-link-square-alt ms-1"></i></a> for visualization. They will be discussed in later posts.</p>]]></description><enclosure url="https://jaehyeon.me/blog/2025-02-18-realtime-dashboard-1/featured.png" length="402206" type="image/png"/></item></channel></rss>