Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
-
Updated
Dec 21, 2024 - Rust
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
The simplest way to scale Python.
Data pipelines from re-usable components
The open-source Useful SDK. One python decorator in the Useful library allows for full observability of Python functions within an ETL.
Data Cleaning for Pyspark
A project structure for doing and sharing data engineer work.
Lien de l'application
Master the AWS Data Stack! 🚀 This repository features 15+ Industrial Data Engineering Projects covering Serverless ETL, Real-Time Streaming, & Data Warehousing. Hands-on labs for S3, Lambda, Spark, Airflow, Snowflake, Redshift, Kinesis, & Glue. Includes production-grade CICD pipelines. A complete roadmap to becoming a top Data Professional.
End To End MLOPS Project With ETL Pipelines- Building Network Security System
A smart parking decision-support platform for drivers unfamiliar with Melbourne CBD. Built for FIT5120 Industry Experience Studio (S1 2026) by Team FlaminGO.
A Data Pipeline and Data House For Euro Flight Data.
An end-to-end cloud data pipeline tracking 400+ AI models from OpenRouter. Features automated Python ETL transformations, robust relational validations, a Supabase PostgreSQL cloud warehouse, and a premium glassmorphic live analytics dashboard detailing model context and price benchmarking.
This repository contains my first end-to-end Data Engineering project, built using Microsoft Azure Cloud and Azure Databricks with PySpark.
Python ETL pipeline normalizing Ahmedabad transit APIs (BRTS, AMTS, Metro) into unified JSON/GTFS datasets for GraphHopper and OpenTripPlanner
DataSift auto applies a data pre-processing pipeline to Data Science Projects.
End-to-end Azure Data Factory project transforming raw sales data into customer-level insights using pivot transformation and storing results in Blob Storage.
Complete portfolio of data engineering projects from Udacity's Data Engineering with AWS Nanodegree.
A robust ETL pipeline using MySQL for data transformation and Python for programmatic visualization of the 2018 NYC Squirrel Census (3,000+ records).
🗄️ IBM Relational Database Administrator with GenAI Certificate Portfolio – A comprehensive collection of projects, labs, and assignments showcasing expertise in relational database administration, 🏘️data warehousing, 🔁ETL pipelines, and 🤖Generative AI integration for modern database management.
Modern Data Warehouse and Analytics Project implementing Medallion Architecture (Bronze, Silver, Gold) with ETL pipelines, SQL data modeling, and analytical reporting.
To associate your repository with the etl-pipelines topic, visit your repo's landing page and select "manage topics."