Data-Engineering
Tìm thấy 653 ứng dụng & công cụ phù hợp
data-engineering-zoomcamp
Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. Join the course here 👇🏼
prefect
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
taipy
Turns Data and AI algorithms into production-ready web applications in no time.
dagster
An orchestration platform for the development, production, and observation of data assets.
risingwave
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.
evidence
Business intelligence as code: build fast, interactive data visualizations in SQL and markdown
cloudquery
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
dlt
data load tool (dlt) is an open source Python library that makes data loading easy 🛠️
rudder-server
Privacy and Security focused Segment-alternative, in Golang and React
sql-translator
SQL Translator is a tool for converting natural language queries into SQL code using artificial intelligence. This project is 100% free and open source.
aws-sdk-pandas
pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Par
Data-Engineering-HowTo
A list of useful resources to learn Data Engineering from scratch
awesome-opensource-data-engineering
An Awesome List of Open-Source Data Engineering Projects
meltano
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
soda-core
Data Contracts engine for the modern data stack. https://www.soda.io
pyspark-example-project
Implementing best practices for PySpark ETL jobs and applications.