
Amazon EMR on EC2 data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

Amazon EMR on EC2 data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

AWS Glue data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

Redshift Serverless data transformation pipelines with dbt. Subsets of IMDb data feed models developed in multiple layers following dbt best practices.

Develop Spark apps on an EMR cluster in a private subnet over VPN and the VS Code remote SSH extension, with the cluster shared by several users.

Provision EMR on EKS with Terraform and EKS Blueprints, then compare two Spark jobs with and without Dynamic Resource Allocation under Karpenter.

Extend the Apache Airflow Lambda invoke function operator with a correlation ID, so it reports the invocation result and the exact error message.

Run a data warehousing ETL job with Apache Iceberg for storage and PySpark for processing in an EMR local environment, then verify results in Athena.

Create a Spark local development environment for Amazon EMR with Docker and Visual Studio Code, with Spark examples and Glue Catalog integration.

Run Spark jobs on EMR on EKS, the EMR deployment option that provisions and manages open source big data frameworks, with simple and extended examples.

Run a Hudi DeltaStreamer application on Amazon EMR, then query the resulting Hudi table with Athena and build a QuickSight dashboard on it.