Tìm thấy 653 ứng dụng & công cụ phù hợp
DataTalks.Club's Data Engineering Zoomcamp Project
Verified Open Source software & tools ready for production.
DBConvert Streams: Database IDE, Federated SQL, Real-time CDC & AI assistants via MCP — explore, query, and replicate data across databases and files
dpq is an open-source python library that makes prompt-based data transformations and feature engineering easy
The Open Jobs Observatory public mirror repo
The definitive end-to-end machine learning (ML lifecycle) guide and tutorial for data engineers.
Kedro plugin to support running pipelines on Dagster
Data pipelines and notebooks for RAG tuning using Fondant
Documentation repository for the Egeria project.
Your own personal data engineer for OpenClaw.
SensApp, time-series with ease.
Download all content from a Telegram channel
This project demonstrates how to build and automate an ETL pipeline written in Python and schedule it using open source Apache Airflow orchestration tool on AWS EC2 instance.
Starter project for building an ETL pipeline using SSIS in Visual Studio 2019
Scheduling Big Data Workloads and Data Pipelines in the Cloud with pyDag
Skooldio: Data Pipelines with Airflow
This repository contains resources and materials for a data science bootcamp. The bootcamp is designed to teach individuals the fundamentals of data science.
A Complete Big Data Analytics course materials for CST-043 with theory, practicals, Hadoop, Spark, and student-friendly explanations. University curriculum aligned + industry ready
Open Data Stack Platform: a collection of projects and pipelines built with open data stack tools for scalable, observable data platform.
K2I - Kafka to Iceberg streaming ingestion engine. A Rust CLI tool inspired by Moonlink architecture that consumes from Kafka, buffers with Apache Arrow for sub-second query freshness, and writes to A
This project aims to leverage Amazon Web Services to create trending Youtube videos analytics service. Project contains different data engineering, data analysis and data science parts.
Marketing Attribution Data Model. SQL, Clickhouse, BigQuery
Blog post on ETL pipelines with Airflow
🐋 Docker image for AWS Glue Spark/Python