
A local Flink and Spark environment built from EMR container images, where Flink ingests data in real time and Spark queries it via the Glue Data Catalog.

A local Flink and Spark environment built from EMR container images, where Flink ingests data in real time and Spark queries it via the Glue Data Catalog.

Transform IMDb data with dbt on Amazon EMR on EKS, through a Spark Thrift server kept running on Kubernetes by a small wrapper class.

Transform IMDb data with dbt on Amazon EMR on EC2, running layered models through a Spark Thrift server.

Develop Spark apps on an EMR cluster in a private subnet over VPN and the VS Code remote SSH extension, with the cluster shared by several users.

Provision EMR on EKS with Terraform and EKS Blueprints, then compare two Spark jobs with and without Dynamic Resource Allocation under Karpenter.

Run a data warehousing ETL job with Apache Iceberg for storage and PySpark for processing in an EMR local environment, then verify results in Athena.

Create a Spark local development environment for Amazon EMR with Docker and Visual Studio Code, with Spark examples and Glue Catalog integration.

Run Spark jobs on EMR on EKS, the EMR deployment option that provisions and manages open source big data frameworks, with simple and extended examples.

Run a Hudi DeltaStreamer application on Amazon EMR, then query the resulting Hudi table with Athena and build a QuickSight dashboard on it.

Build change data capture on AWS with Amazon MSK and MSK Connect, streaming PostgreSQL row changes into the topics that feed the data lake.