Tìm thấy 653 ứng dụng & công cụ phù hợp
Open-source community study guide for all six Databricks certifications (Data Engineer Associate / Professional, Data Analyst Associate, ML Associate / Professional, GenAI Engineer Associate). Aligned
Awesome list of dataops products, open source and resources
Tidyverse-friendly R interface to DuckLake, the DuckDB lakehouse format
🎯 Building Scalable Cloud-Native DataHub Serverless Application ⛅
Metadata-driven framework for Databricks Spark Declarative Pipelines. Config-driven, pattern based approach to batch & streaming across the medallion architecture. Deploys via Declarative Automation B
A Python package of end-to-end weather data clients & raw data clients with VPN/Proxy Server support, data processors that decode variable keys from GRIB format into a plain-language format & various
Verified Open Source software & tools ready for production.
A modern Spark SQL DataSource V2 connector for Redis, focused on DataFrame and SQL support for Redis native data types.
Vendor-neutral Go CLI for the Open Knowledge Format. Create, validate, lint, index, search, and graph OKF bundles. Built to be driven by AI agents.
Autonomous multi-agent system for intelligent HDB resale search — combining geospatial analytics, MRT proximity, and XGBoost price valuation using DeepAgents, LangGraph, FastAPI, and Gradio UI. No dat
Challenge Data Engineer
Example code to create high-quality knowledge graphs using entity resolution with Kuzu and Senzing
This repository contains the necessary configuration files and DAGs (Directed Acyclic Graphs) for setting up a robust data engineering environment using Kubernetes and Apache Airflow
Self-hosted multi-user Apache Spark notebooks — PySpark + Scala, per-user kernel pods, MinIO IAM storage, OAuth. One Helm install. Apache-2.0.
DIT 638 project - Cyber Physical Systems and Systems of Systems, using C++ , Python , Docker File
Pipeline de Dados do Gov-Hub
Welcome to my data engineering projects repository! Here you will find a collection of data engineering projects that I have worked on.
A Data Engineering Project that implements an ETL data pipeline using Dagster, Apache Spark, Streamlit, MinIO, Metabase, Dbt, Polars, Docker. Data from kaggle and youtube-api
Schema-aware JSON compression with millisecond lookups — cut transfer/storage while enabling exists /pos queries. (Demo + wheels; core is binary-only)
Inspect Parquet metadata, encodings, compression, indexes, and Bloom filters locally
Methods for better worker data engineering in the human resources (HR) corporate domain. Designed for HR analytics practitioner to get value from common workforce-oriented data sets.
Decision provenance and data-flow tracking for coding agents: computes what changed in your code and data, records why it changed, and surfaces both before a silent failure ships.
A batch processing data pipeline, using AWS resources (S3, EMR, Redshift, EC2, IAM), provisioned via Terraform, and orchestrated from locally hosted Airflow containers. The end product is a Superset d
Dex is the agent-native analytics engineering toolkit. Point it at your warehouse and your dbt project. It learns the landscape, authors your transformations, and tells you exactly what to fix when th