Jaehyeon Kim
Jaehyeon Kim

  • Blog
    • Categories

      List of categories.

    • Tags

      List of tags.

    • Series

      List of series.

    • Archives

  • Projects
  • Slides

/

  • Github Linkedin RSS

  • Font Size
  • Palette
  • Mode
  1. Home
  2. Tags
  3. R

Serverless Data Product POC Backend Part 1 - Packaging R ML Model for Lambda

Serverless Data Product POC Backend Part 1 - Packaging R ML Model for Lambda
April 8, 201712 min read Development Serverless Data ProductAmazon API GatewayAWSAWS LambdaPythonR

Package a logistic regression model built in R so it can run on AWS Lambda, after testing it locally.

Read More: Serverless Data Product POC Backend Part 1 - Packaging R ML Model for Lambda

Some Thoughts on Shiny Open Source - Render Multiple Pages

June 27, 20166 min read DevelopmentRR Shiny

Render multiple pages in an open source R Shiny application with htmlOutput and renderUI, including login and registration backed by a SQLite database.

Read More: Some Thoughts on Shiny Open Source - Render Multiple Pages

Some Thoughts on Shiny Open Source - Internal Load Balancing

May 23, 20166 min read DevelopmentRR Shiny

Spread users of an open source Shiny Server app across copies of the same app, by sending each new session to the copy with the fewest users.

Read More: Some Thoughts on Shiny Open Source - Internal Load Balancing

Asynchronous Processing Using Job Queue

May 12, 20163 min read DevelopmentR

Keep an R session responsive while long computations run in the background, by queuing them with the jobqueue package.

Read More: Asynchronous Processing Using Job Queue

Boost SparkR with Hive

April 30, 20168 min read Data-EngineeringApache HiveApache SparkHiveQLRSparkR

Run SparkR in Hive Context to reach the Hive UDFs and window functions the SQL Context lacks, compared against dplyr, plus building Spark with Hive.

Read More: Boost SparkR with Hive

Quick Start SparkR in Local and Cluster Mode

March 2, 20168 min read Data-EngineeringApache SparkRSparkR

Run SparkR in local and cluster mode from an R project, setting environment variables and the library path, then reading JSON and CSV files.

Read More: Quick Start SparkR in Local and Cluster Mode

Spark Cluster Setup on VirtualBox

February 22, 20164 min read Data-EngineeringApache SparkRSparkR

Set up a two node Spark standalone cluster on VirtualBox Ubuntu guests, covering machine preparation, copying the VDI image and password-less SSH.

Read More: Spark Cluster Setup on VirtualBox

Quick Test to Wrap Python in R

November 21, 20154 min read DevelopmentPythonR

Wrap Python boto calls for Amazon S3 in an R package, rs3helper, where the Python scripts return JSON so R parses it as vectors and data frames.

Read More: Quick Test to Wrap Python in R

Some Thoughts on Python for R Users

August 9, 20155 min read DevelopmentPythonR

Call a SOAP web service from Python with the suds library, a job R has no comprehensive client package for, demonstrated on the Sizmek MDX API.

Read More: Some Thoughts on Python for R Users

Some Thoughts on Python

August 8, 20154 min read DevelopmentPythonR

Write Python in an object-oriented style, shown with SOAP API client classes, and a reading order for the Introducing Python book by chapter group.

Read More: Some Thoughts on Python
  • ««
  • «
  • 1
  • 2
  • 3
  • 4
  • »
  • »»
Profile
Jaehyeon Kim
Jaehyeon Kim
Data Engineer | Data Streaming | Powering ML & AI in Real Time
Taxonomies
Data Streaming 70 Data Engineering 42 Development 28 Data Analysis 17 Data Integration 12 Open Source 12 Machine Learning 7 Kubernetes 6 Security 5 Data Architecture 3 Data Processing 3 Big Data 2 System Architecture 2 Web Development 2
Python 76 Apache Kafka 72 AWS 50 Docker 50 Apache Flink 38 Apache Beam 17 Apache Spark 16 Kafka Connect 16 Amazon MSK 14 AWS Lambda 14 Benchtop 13 dbt 13 Amazon EMR 11 dynamic-des 11 odctl 11 PostgreSQL 10 Kubernetes 8 PyFlink 8 Change Data Capture (CDC) 7 Debezium 7 Kotlin 7 Apache Airflow 6 Apache Iceberg 6 Discrete Event Simulation 6 Kafka UI 6 Amazon DynamoDB 5 PySpark 5 SimPy 5 Amazon API Gateway 4 Amazon Athena 4 AWS Glue 4 AWS Glue Schema Registry 4 BigQuery 4 Digital Twin 4 FastAPI 4 Minikube 4 Amazon EKS 3 Amazon QuickSight 3 Amazon S3 3 Apache Hudi 3 ALL 143
Kafka Development with Docker 11 Apache Beam Python Examples 10 Real Time Streaming with Kafka and Flink 7 Building Real-Time Digital Twins with dynamic-des 6 dbt Pizza Shop Demo 6 Apache Beam Local Development with Python 5 dbt for Effective Data Transformation on AWS 5 Getting Started with Real-Time Streaming in Kotlin 5 Kafka Connect for AWS Services Integration 5 Serverless Data Product 4 Data Lake Demo Using Change Data Capture 3 Getting Started with PyFlink on AWS 3 Kafka Development on Kubernetes 3 Parallel processing on single machine 3 Realtime Dashboard with FastAPI, Streamlit and Next.js 3 API development with R 2 dbt Guide for Production 2 Deploy Python Stream Processing App on Kubernetes 2 Download Stock Data 2 From Prototype to Production: Real-Time Product Recommendation with Contextual Bandits 2 ALL 25
2026 16 2025 13 2024 29 2023 39 2022 15 2021 7 2020 1 2019 5 2018 2 2017 6 2016 6 2015 15 2014 5
Posts
  • Defining Data-Streaming Simulations in YAML, Without Writing Python
    Defining Data-Streaming Simulations in YAML, Without Writing Python
    October 6, 2026
  • Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    October 2, 2026
  • Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    October 1, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Current London 2026: Building End-to-End Data Lineage
    Current London 2026: Building End-to-End Data Lineage
    May 22, 2026
  • Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    Building a Real-Time Industrial Digital Twin with Apache Flink and Online Machine Learning
    April 21, 2026
  • Productionizing an Online Product Recommender using Event Driven Architecture
    Productionizing an Online Product Recommender using Event Driven Architecture
    February 23, 2026
  • Stream Processing with Flink in Kotlin
    Stream Processing with Flink in Kotlin
    December 10, 2025
  • Guide to Building Integrated Web Applications with FastAPI and NiceGUI
    Guide to Building Integrated Web Applications with FastAPI and NiceGUI
    November 19, 2025
  • Self-service Data Platform via a Multi-tenant SQL Gateway
    Self-service Data Platform via a Multi-tenant SQL Gateway
    July 17, 2025
  • Defining Data-Streaming Simulations in YAML, Without Writing Python
    Defining Data-Streaming Simulations in YAML, Without Writing Python
    October 6, 2026
  • Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    Keeping Game Leaderboards Up to Date in Real Time with Kafka and Flink SQL
    October 2, 2026
  • Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    Change Data Capture on a Simulated Online Shop with Debezium and Kafka Connect
    October 1, 2026
  • Data Streaming and Machine Learning Projects That Run on Your Laptop
    Data Streaming and Machine Learning Projects That Run on Your Laptop
    September 30, 2026
  • Learning MLOps with a Feature Store: A New Series
    Learning MLOps with a Feature Store: A New Series
    September 28, 2026
  • Building an Agentic Analytics System over an Iceberg Lakehouse
    Building an Agentic Analytics System over an Iceberg Lakehouse
    July 18, 2026
  • Dynamic DES: A Declarative API with Postgres and Redis Connectors
    Dynamic DES: A Declarative API with Postgres and Redis Connectors
    July 17, 2026
  • Running Kafka, Flink, Spark, Trino and Iceberg Locally with One CLI
    Running Kafka, Flink, Spark, Trino and Iceberg Locally with One CLI
    July 16, 2026
  • One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES
    One Simulation, Two Pipelines: Batch Training and Live Inference with Dynamic DES
    May 25, 2026
  • Current London 2026: Building End-to-End Data Lineage
    Current London 2026: Building End-to-End Data Lineage
    May 22, 2026
Actions
Go back Reload Copy URL

Jaehyeon Kim

Data Engineer | Data Streaming | Powering ML & AI in Real Time

Copyright © 2023-2026 Jaehyeon Kim. All Rights Reserved.