Tìm thấy 653 ứng dụng & công cụ phù hợp
A project portfolio to accompany my resume
DataOps Observability Integration Agents are part of DataKitchen's Open Source Data Observability. They connect to various ETL, ELT, BI, data science, data visualization, data governance, and data ana
A data lineage tool detects table dependencies from rendered SQL statements.
An end-to-end data engineering pipeline that fetches data from Wikipedia, cleans and transforms it with Apache Airflow and saves it on Azure Data Lake. Other processing takes place on Azure Data Facto
Integrates LLMs as PTransform in Apache Beam pipelines using LangChain
A curated list of big data engineering tools, resources and communities.
By Smart Shaped s.r.l. (https://www.smartshaped.com/)
电商全域数据治理与 Multi-Agent 分析中台:基于 n8n + LLM 架构,实现多源电商数据 ETL、规范化数仓分层建模 (DWD/ADS) 及飞书端智能对话查询闭环。
Resources and projects from Udacity Data Engineering with AWS nano degree programme
Convert a SQLite database into a DuckDB database — tables, indexes and views, in one command
Full-stack air quality analytics platform built with FastAPI, React, and MySQL. Aggregates multi-source PM2.5/PM10 data, performs multi-city comparison and time-series forecasting (SARIMAX), and integ
I am using confluent Kafka cluster to produce and consume scraped data. In this project, I've created a real-time data pipeline that utilizes Kafka to scrape, process, and load data onto S3 in JSON f
Data Engineering Digest
Built a stream processing data pipeline to get data from disparate systems into a dashboard using Kafka as an intermediary.
A end-to-end real-time stock market data pipeline with Python, AWS EC2, Apache Kafka, and Cassandra Data is processed on AWS EC2 with Apache Kafka and stored in a local Cassandra database.
Use Gemini Pro LLM via VertexAI to create an engaging quiz game incorporating TMDB API data
60 Days of Data Science and ML
Data2Neo is a library that simplifies the conversion of data in relational format to a graph knowledge database.
Debussy is an opinionated Data Architecture and Engineering framework, enabling data analysts and engineers to build better platforms and pipelines.
Advanded Software Tools for Reliable Industrial Datasets
Semantic cost-linting and performance warnings extension for Databricks in VS Code
A fluent, scalable, and easy-to-use LLM data processing framework.
A Python package to load complex XML files into a relational database
Desafio Ingestão no Limite: Edição Hardware Leve: Provando que engenharia de dados de verdade se faz com código otimizado, não com máquina cara.