Docker

Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images

December 7, 202316 min read Data Engineering Data Streaming Development Amazon EMR Apache Flink Apache Kafka Apache Spark AWS Docker

Apache Flink became generally available for Amazon EMR on EKS from the EMR 6.15.0 releases. As it is integrated with the Glue Data Catalog, it can be particularly useful if we develop real time data ingestion/processing via Flink and build analytical queries using Spark (or any other tools or services that can access to the Glue Data Catalog). In this post, we will discuss how to set up a local development environment for Apache Flink and Spark using the EMR container images. After illustrating the environment setup, we will discuss a solution where data ingestion/processing is performed in real time using Apache Flink and the processed data is consumed by Apache Spark for analysis.

November 9, 202315 min read Data Streaming Real Time Streaming With Kafka and Flink Amazon S3 Apache Flink Apache Kafka AWS Docker Kpow Pyflink

In this lab, we will create a Pyflink application that reads records from S3 and sends them into a Kafka topic. A custom pipeline Jar file will be created as the Kafka cluster is authenticated by IAM, and it will be demonstrated how to execute the app in a Flink cluster deployed on Docker as well as locally as a typical Python app. We can assume the S3 data is static metadata that needs to be joined into another stream, and this exercise can be useful for data enrichment.

October 23, 202312 min read Data Integration Data Streaming Kafka Connect for AWS Services Integration Apache Kafka AWS Docker Kafka Connect Kpow OpenSearch

Kafka Connect can be an effective tool to ingest data from Apache Kafka into OpenSearch. In this post, we will discuss how to develop a data pipeline from Apache Kafka into OpenSearch locally using Docker while the pipeline will be deployed on AWS in the next post. Fake impressions and clicks data will be pushed into Kafka topics using a Kafka source connector and those records will be ingested into OpenSearch indexes using a sink connector for near-real time analytics.

October 19, 20236 min read Data Streaming Apache Flink Apache Kafka Docker Pyflink Python

Building Apache Flink Applications in Java by Confluent is a course to introduce Apache Flink through a series of hands-on exercises. Utilising the Flink DataStream API, the course develops three Flink applications from ingesting source data into calculating usage statistics. As part of learning the Flink DataStream API in Pyflink, I converted the Java apps into Python equivalent while performing the course exercises in Pyflink. This post summarises the progress of the conversion and shows the final output.

September 4, 202313 min read Data Streaming Getting Started With Pyflink on AWS Amazon MSK Apache Kafka AWS Docker Kpow Pyflink Python

In this series of posts, we discuss a Flink (Pyflink) application that reads/writes from/to Kafka topics. In the previous posts, I demonstrated a Pyflink app that targets a local Kafka cluster as well as a Kafka cluster on Amazon MSK. The app was executed in a virtual environment as well as in a local Flink cluster for improved monitoring. In this post, the app will be deployed via Amazon Managed Service for Apache Flink.

August 17, 202316 min read Data Streaming Getting Started With Pyflink on AWS Apache Flink Apache Kafka Docker Kpow Pyflink Python

Apache Flink is widely used for building real-time stream processing applications. On AWS, Amazon Managed Service for Apache Flink is the easiest option to develop a Flink app as it provides the underlying infrastructure. Updating a guide from AWS, this series of posts discuss how to develop and deploy a Flink (Pyflink) application via KDA where the data source and sink are Kafka topics. In part 1, the app will be developed locally targeting a Kafka cluster created by Docker. Furthermore, it will be executed in a virtual environment as well as in a local Flink cluster for improved monitoring.

August 10, 202316 min read Data Streaming Kafka, Flink and DynamoDB for Real Time Fraud Detection Amazon DynamoDB Apache Flink Apache Kafka AWS Docker Kpow Python

Apache Flink is widely used for building real-time stream processing applications. On AWS, Amazon Managed Service for Apache Flink is the easiest option to develop a Flink app as it provides the underlying infrastructure. Re-implementing a solution from an AWS workshop, this series of posts discuss how to develop and deploy a fraud detection app using Kafka, Flink and DynamoDB. Part 1 covers local development using Docker while deployment via KDA will be discussed in part 2.

July 20, 202314 min read Data Streaming Security Kafka Development With Docker Apache Kafka Docker Python

In the previous posts, we discussed how to implement client authentication by TLS (SSL or TLS/SSL) and SASL authentication. One of the key benefits of client authentication is achieving user access control. In this post, we will discuss how to configure Kafka authorization with Java and Python client examples while SASL is kept for client authentication.

July 13, 202311 min read Data Streaming Security Kafka Development With Docker Apache Kafka Docker Python SASL

In the previous post, we discussed TLS (SSL or TLS/SSL) authentication to improve security. It enforces two-way verification where a client certificate is verified by Kafka brokers. Client authentication can also be enabled by Simple Authentication and Security Layer (SASL), and we will discuss how to implement SASL authentication with Java and Python client examples in this post.

July 6, 202314 min read Data Streaming Security Kafka Development With Docker Apache Kafka Docker Python TLS

To improve security, we can extend TLS (SSL or TLS/SSL) encryption either by enforcing two-way verification where a client certificate is verified by Kafka brokers (SSL authentication). Or we can choose a separate authentication mechanism, which is typically Simple Authentication and Security Layer (SASL). In this post, we will discuss how to implement SSL authentication with Java and Python client examples while SASL authentication is covered in the next post.

Setup Local Development Environment for Apache Flink and Spark Using EMR Container Images

Real Time Streaming With Kafka and Flink - Lab 2 Write Data to Kafka From S3 Using Flink

Kafka Connect for AWS Services Integration - Part 4 Develop Aiven OpenSearch Sink Connector

Building Apache Flink Applications in Python

Getting Started With Pyflink on AWS - Part 3 AWS Managed Flink and MSK

Getting Started With Pyflink on AWS - Part 1 Local Flink and Local Kafka

Kafka, Flink and DynamoDB for Real Time Fraud Detection - Part 1 Local Development

Kafka Development With Docker - Part 11 Kafka Authorization

Kafka Development With Docker - Part 10 SASL Authentication

Kafka Development With Docker - Part 9 SSL Authentication