Data Streaming

Apache Beam Python Examples - Part 2 Calculate Average Word Length With/Without Fixed Look Back

July 18, 202415 min read Apache Beam Data Streaming Apache Beam Python Examples Apache Beam Apache Flink Docker Docker Compose Python

In this post, we develop two Apache Beam pipelines that calculate average word lengths from input texts that are ingested by a Kafka topic. They obtain the statistics in different angles. The first pipeline emits the global average lengths whenever a new input text arrives while the latter triggers those values in a sliding time window.

July 4, 202422 min read Apache Beam Data Streaming Apache Beam Python Examples Apache Beam Apache Flink Docker Docker Compose Python

In this series, we develop Apache Beam Python pipelines. The majority of them are from Building Big Data Pipelines with Apache Beam by Jan Lukavský. Mainly relying on the Java SDK, the book teaches fundamentals of Apache Beam using hands-on tasks, and we convert those tasks using the Python SDK. We focus on streaming pipelines, and they are deployed on a local (or embedded) Apache Flink cluster using the Apache Flink Runner. Beginning with setting up the development environment, we build two pipelines that obtain top K most frequent words and the word that has the longest word length in this post.

June 6, 202416 min read Apache Beam Data Streaming Deploy Python Stream Processing App on Kubernetes Apache Beam Apache Flink Apache Kafka Docker Kubernetes Minikube Python

In this post, we develop an Apache Beam pipeline using the Python SDK and deploy it on an Apache Flink cluster via the Apache Flink Runner. Same as Part I, we deploy a Kafka cluster using the Strimzi Operator on a minikube cluster as the pipeline uses Apache Kafka topics for its data source and sink. Then, we develop the pipeline as a Python package and add the package to a custom Docker image so that Python user code can be executed externally. For deployment, we create a Flink session cluster via the Flink Kubernetes Operator, and deploy the pipeline using a Kubernetes job. Finally, we check the output of the application by sending messages to the input Kafka topic using a Python producer application.

May 30, 202413 min read Apache Flink Data Streaming Deploy Python Stream Processing App on Kubernetes Apache Flink Apache Kafka Docker Kubernetes Minikube Python

Flink Kubernetes Operator acts as a control plane to manage the complete deployment lifecycle of Apache Flink applications. With the operator, we can simplify deployment and management of Python stream processing applications. In this series, we discuss how to deploy a PyFlink application and Python Apache Beam pipeline on the Flink Runner on Kubernetes. In Part 1, we first deploy a Kafka cluster on a minikube cluster as the source and sink of the PyFlink application are Kafka topics. Then, the application source is packaged in a custom Docker image and deployed on the minikube cluster using the Flink Kubernetes Operator. Finally, the output of the application is checked by sending messages to the input Kafka topic using a Python producer application.

January 11, 20247 min read Apache Kafka Data Streaming Kafka Development on Kubernetes Apache Kafka Docker Kafka Connect Kubernetes Minikube Strimzi

Kafka Connect is a tool for scalably and reliably streaming data between Apache Kafka and other systems. In this post, we discuss how to set up a data ingestion pipeline using Kafka connectors. Fake customer and order data is ingested into Kafka topics using the MSK Data Generator. Also, we use the Confluent S3 sink connector to save the messages of the topics into a S3 bucket. The Kafka Connect servers and individual connectors are deployed using the custom resources of Strimzi on Kubernetes.

January 4, 20248 min read Apache Kafka Data Streaming Kafka Development on Kubernetes Apache Kafka Docker Kubernetes Minikube Python Strimzi

Apache Kafka has five core APIs, and we can develop applications to send/read streams of data to/from topics in a Kafka cluster using the producer and consumer APIs. While the main Kafka project maintains only the Java APIs, there are several open source projects that provide the Kafka client APIs in Python. In this post, we discuss how to develop Kafka client applications using the kafka-python package on Kubernetes.

December 21, 20237 min read Apache Kafka Data Streaming Kafka Development on Kubernetes Apache Kafka Docker Kubernetes Minikube Strimzi

Apache Kafka is one of the key technologies for implementing data streaming architectures. Strimzi provides a way to run an Apache Kafka cluster and related resources on Kubernetes in various deployment configurations. In this series of posts, we will discuss how to create a Kafka cluster, to develop Kafka client applications in Python and to build a data pipeline using Kafka connectors on Kubernetes.

December 14, 20236 min read Apache Kafka Data Streaming Real Time Streaming With Kafka and Flink Amazon MSK Apache Kafka AWS AWS Lambda Docker Docker Compose Python

Amazon MSK can be configured as an event source of a Lambda function. Lambda internally polls for new messages from the event source and then synchronously invokes the target Lambda function. With this feature, we can develop a Kafka consumer application in serverless environment where developers can focus on application logic. In this lab, we will discuss how to create a Kafka consumer using a Lambda function.

November 30, 20239 min read Apache Kafka Data Streaming Real Time Streaming With Kafka and Flink Amazon DynamoDB Amazon MSK Amazon MSK Connect Apache Kafka AWS Docker Docker Compose Kafka Connect

Kafka Connect is a tool for scalably and reliably streaming data between Apache Kafka and other systems. It makes it simple to quickly define connectors that move large collections of data into and out of Kafka. In this lab, we will discuss how to create a data pipeline that ingests data from a Kafka topic into a DynamoDB table using the Camel DynamoDB sink connector.

November 23, 202315 min read Apache Flink Data Streaming Real Time Streaming With Kafka and Flink Amazon MSK Amazon OpenSearch Service Apache Flink Apache Kafka AWS Docker Docker Compose OpenSearch Pyflink Python

The value of data can be maximised when it is used without delay. With Apache Flink, we can build streaming analytics applications that incorporate the latest events with low latency. In this lab, we will create a Pyflink application that writes accumulated taxi rides data into an OpenSearch cluster. It aggregates the number of trips/passengers and trip durations by vendor ID for a window of 5 seconds. The data is then used to create a chart that monitors the status of taxi rides in the OpenSearch Dashboard.

Apache Beam Python Examples - Part 2 Calculate Average Word Length With/Without Fixed Look Back

Apache Beam Python Examples - Part 1 Calculate K Most Frequent Words and Max Word Length

Deploy Python Stream Processing App on Kubernetes - Part 2 Beam Pipeline on Flink Runner

Deploy Python Stream Processing App on Kubernetes - Part 1 PyFlink Application

Kafka Development on Kubernetes - Part 3 Kafka Connect

Kafka Development on Kubernetes - Part 2 Producer and Consumer

Kafka Development on Kubernetes - Part 1 Cluster Setup

Real Time Streaming With Kafka and Flink - Lab 6 Consume Data From Kafka Using Lambda

Real Time Streaming With Kafka and Flink - Lab 5 Write Data to DynamoDB Using Kafka Connect

Real Time Streaming With Kafka and Flink - Lab 4 Clean, Aggregate, and Enrich Events With Flink