Tìm thấy 653 ứng dụng & công cụ phù hợp
Cool DE Projects
Production-ready PySpark ETL template for Databricks — medallion architecture, DABs, tests, DQX, CI/CD, and agentic development with Claude Code.
Data engineering interview prep - PySpark notebooks, theory docs, quizzes, and company-specific patterns. Built around Zephyr Coffee Co., a fictional 200-store chain with messy data.
A Python package that creates fine-grained dbt tasks on Apache Airflow
System Design, Solution Architecture, Data Systems Practice
Beneath is a serverless real-time data platform ⚡️
Magniv Core - A Python-decorator based job orchestration platform. Avoid responsibility handoffs by abstracting infra and DevOps.
Streaming Ethereum and Bitcoin blockchain data to Google Pub/Sub or Postgres in Kubernetes
Materials for the Deploy and Monitor ML Pipelines with Python, Docker and GitHub Actions workshop at the PyData NYC 2024 conference
PDF DataSource for Apache Spark, allow to read PDF files directly to the DataFrame and ocr it
A change-impact decision engine for dbt. Column-level lineage plus a policy-gated verdict on every PR — what breaks, what to rebuild — computed offline from your dbt artifacts. No warehouse, no dbt ru
Memory for your AI work across projects and sessions. Save decisions, rules and lessons with source and applicability. Connect over MCP.
Complete Roadmap For Data Science
Convert monolithic Jupyter notebooks 📙 into maintainable Ploomber pipelines. 📊
📡 Real-time data pipeline with Kafka, Flink, Iceberg, Trino, MinIO, and Superset. Ideal for learning data systems.
Interactive computing for complex data processing, modeling and analysis in Python 3
DataOps Data Quality TestGen is part of DataKitchen's Open Source Data Observability. DataOps TestGen delivers simple, fast data quality test generation and execution by data profiling, new dataset
A local Python/SQL notebook in the browser for exploring, transforming and visualizing data using WebAssembly.
A lightweight, declarative PySpark framework for data quality validation — check columns, rows, and entire datasets directly in your Spark pipelines
A Streamlit powered GPT-3 Application that allows you to chat with tabular data. In addition to AI Chart creation, insights are given too.
A robust (🐢) and fast (🐇) MLOps tool for managing data and pipelines in Rust (🦀)
breadroll 🥟 is a simple lightweight library for data processing operations written in Typescript and powered by Bun.
This documentation is like a quick snapshot of my project in the data field, showing off my skills and know-how in this area.
OpenSnowcat Collector, an open source fork of Snowplow (Apache 2.0 License)