Skip to content
#

data-pipeline-automation

Here are 41 public repositories matching this topic...

End-to-end data pipeline transforming Olist e-commerce data through Azure cloud services. Implements medallion architecture (Bronze-Silver-Gold) with multi-source ingestion, Spark-based processing, and OLTP-to-OLAP optimization for analytics-ready datasets.

  • Updated Aug 27, 2026
  • Jupyter Notebook

Pulls Meta ad data automatically, syncs to Google Sheets, delivers a live analytics dashboard with rule-based insights, decision engine, and Telegram alerts.

  • Updated Jun 4, 2026
  • Python

A complete reference implementation of a local-first ecosystem for AI-powered analytics. This repository contains the source code for the SBDK.dev website, the central hub for the SBDK suite of open-source tools.

  • Updated Nov 25, 2025
  • TypeScript

Data Engineering project implementing a fully automated Medallion Architecture (Bronze, Silver, Gold) using Snowpipe, Streams, Tasks, and SQL to build scalable cloud-native data pipelines for analytics.

  • Updated Jul 6, 2026

This repository contains scripts to build and process a Real-Time GDP (RTD) dataset for Peru, focusing on extracting, cleaning, and analyzing GDP revisions from the BCRP's Weekly Reports. Future updates will include econometric models and visualizations.

  • Updated Apr 28, 2026
  • Python

VizFlow is a VS Code extension that brings a visual workflow builder, RBQL querying, database connectors (MongoDB, MySQL, PostgreSQL), HTTP/REST integration, file transformations, and built-in scheduling — all running locally on your machine. Build repeatable data pipelines with 35 activity types, no cloud account or vendor lock-in required.

  • Updated Aug 22, 2026
  • JavaScript

End-to-end data pipeline integrating manufacturing and financial data. Demonstrates ETL processes, a PostgreSQL data warehouse with a star schema, and business intelligence capabilities using PowerBI. Built with Python, Apache Airflow, and Docker.

  • Updated Feb 16, 2026
  • Python

Data automation involves automating the extraction, transformation, and loading (ETL) processes to streamline data workflows. GitHub Actions enables automated execution of tasks, such as building, testing, and deploying code, in response to events. This integration simplifies continuous deployment and ensures repeatable data pipeline operations

  • Updated Feb 25, 2025
  • HTML

Add this topic to your repo

To associate your repository with the data-pipeline-automation topic, visit your repo's landing page and select "manage topics."

Learn more