This repository contains source code for the AWS Database Blog Post Reduce data archiving costs for compliance by automating RDS snapshot exports to Amazon S3
-
Updated
Apr 26, 2023
This repository contains source code for the AWS Database Blog Post Reduce data archiving costs for compliance by automating RDS snapshot exports to Amazon S3
This repository has a collection of utilities for Glue Crawlers. These utilities come in the form of AWS CloudFormation templates or AWS CDK applications.
Automation framework to catalog AWS data sources using Glue
ETL Data pipeline using aws services
Smart City Realtime Data Engineering Project
Terraform configuration that creates several AWS services, uploads data in S3 and starts the Glue Crawler and Glue Job.
An end-to-end data pipeline built with AWS S3, Glue, Crawler, Athena, Tableau visulization
It is a project build using ETL(Extract, Transform, Load) pipeline using Spotify API on AWS.
Production ETL pipeline: Spotify API → S3 Parquet, orchestrated with Airflow 3.x (CeleryExecutor), fault-tolerant with data quality validation
AWS Athena, Glue Database, Glue Crawler and S3 buckets deployment through AWS GUI console.
Quality_Movie_Data_Analysis
This project automates the extraction, transformation, and loading (ETL) of Reddit data into a Redshift data warehouse using Airflow. Key technologies include Celery, PostgreSQL, S3, Glue, Athena, and Redshift, providing a complete data pipeline solution.
End-to-end AWS data analytics pipeline for product risk detection and customer dissatisfaction analysis.
Analyzed a multicategory e-commerce store using big data techniques on a Kaggle dataset with the help of AWS EC2, AWS S3, PySpark, AWS Glue ETL, AWS Athena, AWS CloudFormation, AWS Lambda and Power BI!
AWS Athena, Glue Database, Glue Crawler and S3 buckets deployment through CloudFormation stack on AWS console.
An end-to-end solution for managing and analyzing YouTube video data from Kaggle, leveraging AWS services and visualized through Quicksight and Tableau
Creating an audit table for a DynamoDB table using CloudTrail, Kinesis Data Stream, Lambda, S3, Glue and Athena and CloudFormation
End-to-end YouTube Data Engineering Pipeline built on AWS using Amazon S3, Lambda, Glue, Athena, and Python. Automated ETL transforms raw CSV/JSON data into Parquet for analytics, with interactive insights delivered through Power BI.
This project automates an ETL pipeline using AWS Glue, S3, Athena, and Step Functions to transform raw Airbnb data. It cleanses, enriches, and organizes the data into separate raw and transformed databases, enabling efficient querying and analysis via Athena, with automated notifications through SNS.
To associate your repository with the aws-glue-crawler topic, visit your repo's landing page and select "manage topics."