Data-Engineering
Tìm thấy 653 ứng dụng & công cụ phù hợp
data-forge
Data Forge — a modern data stack playground to practice flows and best practices, not just tools. Spark, Trino, Kafka, Iceberg, ClickHouse, Airflow, MinIO, Superset — all wired together locally with D
jayvee
Jayvee is a domain-specific language and runtime for automated processing of data pipelines
Data-Science-For-EveryOne
Data Science boot camp aims to make the field of data science accessible and understandable to a wide range of individuals, regardless of their background or expertise.
Movalytics-Data-Warehouse
Data pipeline performing ETL to AWS Redshift using Spark, orchestrated with Apache Airflow
big-data-modeling
Big Data Modeling, MapReduce, Spark, PySpark @ Santa Clara University
raptor
Transform your pythonic research to an artifact that engineers can deploy easily.
FMD_FRAMEWORK
The Fabric Metadata-Driven (FMD) Framework is a community-driven accelerator for Microsoft Fabric that provides the foundation for automated, scalable, and enterprise-grade data platforms. By treating
job-board-aggregator
Job board aggregator indexing 1,000,000+ active positions from 20,000+ companies across Greenhouse, Lever, Ashby, Workday, and other major ATS platforms. Multithreaded Python ETL pipeline, daily autom
khaos
Kafka data generator and load testing tool - generate fake messages, simulate producers and consumers, test broker failures, and run chaos engineering scenarios
Databricks-Certified-Data-Engineer-Professional-Questions
This repo contains "Databricks Certified Data Engineer Professional" Questions and related docs.
pyspark-tutorial
PySpark Tutorial for Beginners - Practical Examples in Jupyter Notebook with Spark version 3.4.1. The tutorial covers various topics like Spark Introduction, Spark Installation, Spark RDD Transformati
dbt-sugar
dbt-sugar is a CLI tool that allows users of dbt to have fun and ease performing actions around dbt models
data-engineer-portfolio
This is a repository to demonstrate my details, skills, projects and to keep track of my progression in Data Analytics and Data Science topics.
opendatadiscovery-specification
ODD Specification is a universal open standard for collecting metadata.
dbmask
Discover, mask, and verify sensitive data in SQL databases — an auditable scan → mask → validate workflow for safe database copies.
accelerator
The Accelerator is a tool for fast and reproducible processing of large amounts of data.
docglow
Modern documentation site generator for dbt Core — lineage explorer, health scoring, full-text search. Live demo: https://demo.docglow.com