Tìm thấy 653 ứng dụng & công cụ phù hợp
DE or DIE meetup made by data engineers for data engineers. Currently in Russian only.
A data and analytics engineering platform designed for real-time sports betting analytics.
Community supported integrations for the Dagster platform.
A client for connecting and running DDLs on hive metastore.
100+ data engineering projects from scratch — streaming, CDC, table formats, query engines, consensus, governance. 2,500+ tests, mypy strict.
Workbench: An easy to use Python API for creating and deploying AWS SageMaker Models
Inspection of tabular (csv, xls-like, parquet) files to guess the columns' content
notes for DP203-Data Engineering on Azure
A self-hosted, lightweight, visual ETL (Extract, Transform, Load) platform inspired by Alteryx.
OHLC,news,economic event restfull api and WebSocket API for specifically designed for AI quantitative trading/training.Multiple Timeframes,Multiple-Symbols-Multiple-Timeframes
A portable Datamart and Business Intelligence suite built with Docker, Mage, dbt, DuckDB and Superset
A curated list of dagster code snippets for data engineers
DataOps Observability is part of DataKitchen's Open Source Data Observability. DataOps Observability monitors every data journey from data source to customer value, from any team development environm
Sample ELT project using Dagster, data load tool and Snowflake
Udacity Data Engineering Nanodegree Program
Containerized distributed programming framework for Python
Notes on Data Engineering with Pandas, PySpark, Dask, Ray, Arrow DataFusion, Polars etc.
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define m
Data Engineer with Python lecture notes from #datacamp.
Move whole tables between databases fast — Postgres, MySQL, ClickHouse, BigQuery. Rust engine, one-line Python API, bounded memory.
170+ curated resources every Databricks Data Engineer should bookmark - tools, courses, creators, labs, and communities
A repository defining a simple data pipleine for ETL jobs relating to media metadata.
The practical use-cases of how to make your Machine Learning Pipelines robust and reliable using Apache Airflow.
Toolbox for building Generative AI applications on top of Apache Spark.