Awesome Data Engineering 
A curated list of awesome things related to Data Engineering.
A curated list of awesome things related to Data Engineering.
A Python library that facilitates the comparison of two DataFrames in Pandas, Polars, Spark and more. The library goes beyond basic equality checks by providing detailed insights into discrepancies at both row and column levels.
Data Validation Tool compares data from source and target tables to ensure that they match. It provides column validation, row validation, schema validation, custom query validation, and ad hoc SQL exploration.
A high-performance Python library for comparing large datasets (CSV, Parquet) locally using Rust and Polars. It features zero-copy streaming to prevent OOM errors and generates interactive HTML data quality reports.
Python SDK that dispatches parallel web-research agents across
table rows, synthesizing multi-agent findings into structured columns.
Data ingestion engine that connects 400+ Singer taps to Parquet files in cloud buckets (S3, GCS, Azure). Streaming, incremental, with auto-catalog.
Managed event ingestion service that converts JSON sent to a REST API into Hive-partitioned Parquet on Cloudflare R2, queryable from DuckDB, ClickHouse, BigQuery, Snowflake, and Python.
CLI tool to enrich CSV files with company data (financials, contacts, metadata) from 250M+ company records. Available on [npm](https://www.npmjs.com/package/enrich-companies).
CLI tool to copy data between databases with a single command. Supports 50+ sources including PostgreSQL, MySQL, MongoDB, Salesforce, Shopify to any data warehouse.
Publish-subscribe messaging rethought as a distributed commit log.
A fully managed, cloud-based service for real-time data processing over large, distributed data streams.
Robust messaging for applications.
A fast&simple pipeline building library for Python data devs, runs in notebooks, cloud functions, airflow, etc.
OSS Reverse ETL CLI. Sync data from warehouses to business tools via YAML.
An open source data collector for unified logging layer.
An open source bulk data loader that helps data transfer between various databases, storages, file formats, and cloud services.
A tool designed for efficiently transferring bulk data between Apache Hadoop and structured datastores such as relational databases.
Data Acquisition and Processing Made Easy. Deprecated.
Universal data ingestion framework for Hadoop from LinkedIn.
An open source event messaging platform that provides a REST API on top of Kafka-like queues.
Provides a new storage abstraction - a stream - for continuous and unbounded data.
An open-source distributed pub-sub messaging system.
Utility belt to handle data on AWS.
Open-source data integration for modern data teams.
self-hosted database migration and change data capture (CDC) tool with built-in SQL IDE.
Real-time data ingestion tool leveraging change data capture.
CLI data integration tool specialized in moving data between databases, as well as storage systems.
CLI & code-first ELT.
Live import all your Google Sheets to your data warehouse.
A delimited data preboarding framework that fills the gap between MFT and the data lake.
No/low-code data pipeline platform that handles both batch and real-time data ingestion.
Lightweight Node.js ETL framework for databases → data lakes/warehouses.
High-performance, streaming-first ETL engine for Node.js and TypeScript with constant memory footprint.
Polyglot document intelligence library with a Rust core and bindings for Python, TypeScript, Go, and more. Extracts text, tables, and metadata from 62+ document formats for data pipeline ingestion.
Python PDF-to-Markdown orchestrator. Classifies each page and routes to the optimal backend (PyMuPDF, Docling, RapidOCR, Gemini Flash), emitting Markdown plus a per-page confidence score so ingestion pipelines can quarantine low-trust pages before feeding LLMs or retrieval.
Managed cloud object storage transfers for ingestion workflows.
Real-time X (Twitter) data extraction platform with REST API (76 endpoints), 20 bulk extraction tools, account monitoring, HMAC-signed webhooks, and MCP server for AI agent integration.
High-speed CLI tools for database export, import, replication and migration with parallel streaming to CSV, Parquet, JSON and cloud storage, supporting PostgreSQL, MySQL, Oracle, SQL Server and 80+ sources.
A real-time B2B data API for company and people intelligence, providing firmographics, headcount signals, job listings, web traffic, and funding events via REST API and webhooks for data enrichment pipelines.
Conflict-free merge for DataFrames, JSON, ML models & distributed agents — powered by CRDTs.
Crawlee-based actor extracting structured LinkedIn job listings at scale without API keys.
Context-Aware RAG Processing Queue for high availability and adaptive rate-limiting.
Local-first, open-source desktop ETL/ELT studio: drag a pipeline onto a canvas (or describe it to a built-in on-device AI assistant) and run it at native speed through DuckDB. 290+ connectors, a scheduler, and an MCP server for driving pipelines from an LLM. No cloud, no servers.
Open-source self-hosted analytics pipeline that lands raw events as Parquet in your own object storage. Uses NATS JetStream for durable buffering and BigQuery external tables for querying. Designed for teams that want to own their raw event data.
Config-driven data-movement platform for Rust with pluggable source and sink connectors, running ETL, CDC, and streaming pipelines declaratively from YAML or embedded as a library.
A distributed file system designed to run on commodity hardware.
Object storage built to retrieve any amount of data from anywhere.
A memory-centric distributed storage system enabling reliable data sharing at memory-speed across cluster frameworks, such as Spark and MapReduce.
A unified, distributed storage system designed for excellent performance, reliability, and scalability.
A high-performance Cloud-Native file system driven by object storage for large-scale data storage.
Orange File System is a branch of the Parallel Virtual File System.
A bite-sized, lightweight HDFS compatible file system built over Cassandra.
Gluster Filesystem.
Fault-tolerant distributed file system for all storage needs.
Seaweed-FS is a simple and highly scalable distributed file system. There are two objectives: to store billions of files! to serve the files fast! Instead of supporting full POSIX file system semantics, Seaweed-FS choose to implement only a key~file mapping. Similar to the word "NoSQL", you can call it as "NoFS".
A file system that stores all its data online using storage services like Google Storage, Amazon S3, or OpenStack.
Software Defined Storage is a distributed, parallel, scalable, fault-tolerant, Geo-Redundant and highly available file system.
The AI native file format. Trust scores, source provenance, and compliance metadata that embed into 20+ formats (DOCX, PDF, images, code). EXIF for AI.
Apache Avro™ is a data serialization system.
A columnar storage format available to any project in the Hadoop ecosystem, regardless of the choice of data processing framework, data model or programming language.
The smallest, fastest columnar storage for Hadoop workloads.
The Apache Thrift software framework, for scalable cross-language services development.
Protocol Buffers - Google's data interchange format.
A flat file consisting of binary key/value pairs. It is extensively used in MapReduce as input/output formats.
A fast and efficient object graph serialization framework for Java.
Specialized JSONL log compressor with block-level timestamp indexing and DuckDB integration. Achieves ~9% compression ratio (better than gzip) with time-range random access queries.
Browser-based viewer, SQL workbench and converter for Parquet files powered by DuckDB-WASM. Fully client-side, no upload.
A unified programming model that implements both batch and streaming data processing jobs that run on many execution engines.
Makes it easy to build scalable fault-tolerant streaming applications.
A streaming dataflow engine that provides data distribution, communication, and fault tolerance for distributed computations over data streams.
A free and open source distributed realtime computation system.
A distributed stream processing framework.
An easy to use, powerful, and reliable system to process and distribute data.
An open source framework for managing storage for real time processing, one of the most interesting feature is the Upsert.
An open source ETL framework to build fresh index for AI.
An ACID-compliant RDBMS which uses a [shared nothing architecture](https://en.wikipedia.org/wiki/Shared-nothing_architecture).
The Streaming SQL Database.
Streaming and tasks execution between Spring Boot apps.
A data-processing toolkit for python 3.5+.
Forever scalable event processing & in-memory durable K/V store as a library with asyncio & static typing.
The streaming database built for IoT data storage and real-time processing.
An edge lightweight IoT data analytics/streaming software implemented by Golang, and it can be run at all kinds of resource-constrained edge devices.
- An API gateway built for event-driven architectures and streaming that supports standard protocols such as HTTP, SSE, gRPC, MQTT, and the native Kafka protocol.
A framework for building real-time streaming data processing applications that supports a wide range of ingestion sources.
Performant open-source Python ETL framework with Rust runtime, supporting 300+ data sources.
A software framework for easily writing applications which process vast amounts of data (multi-terabyte data-sets) - in-parallel on large clusters (thousands of nodes) - of commodity hardware in a reliable, fault-tolerant manner.
A multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.
A web service that makes it easy to quickly and cost-effectively process vast amounts of data.
A cloud-based platform deployed on Kubernetes making Apache Spark more developer-friendly and cost-effective.
An application framework which allows for a complex directed-acyclic-graph of tasks for processing data.
A light-weight engine for general-purpose data processing including both batch and stream analytics. It is based on a novel unique data model, which represents data via _functions_ and processes data via _columns operations_ as opposed to having only set operations in conventional approaches like MapReduce or SQL.
A cloud native data pipeline and transformation toolkit written in Go.
Personal genome analysis toolkit with Python scripts analyzing raw DNA data across 17 categories (health risks, ancestry, pharmacogenomics, nutrition, psychology, etc.) and generating a terminal-style single-page HTML visualization.
A charting library written in pure JavaScript, offering an easy way of adding interactive charts to your web site or web application.
Fast JavaScript charts for any data set.
D3-based reusable chart library.
A JavaScript library for manipulating documents based on data.
A JavaScript Charting Library for Streaming Data.
Python helpers for building dashboards using Flask and React.
Flask, JS, and CSS boilerplate for interactive, web-based visualization apps in Python.
A modern, enterprise-ready business intelligence web application.
Make Your Company Data Driven. Connect to any data source, easily visualize and share your data.
The easy, open source way for everyone in your company to ask questions and learn from data.
Open-source, self-hosted, warehouse-native product analytics. Runs funnels, retention, and paths directly on DuckDB, Postgres, Snowflake, or ClickHouse.
A pure-python graphics and GUI library built on PyQt4 / PySide and numpy. It is intended for use in mathematics / scientific / engineering applications.
A Python visualization library based on matplotlib. It provides a high-level interface for drawing attractive statistical graphics.
Natural language database query interface with automatic chart generation, supporting Chinese and English queries.
Agentic AI platform to connect any database (PostgreSQL, MySQL, MongoDB, etc.) and query in plain English; includes self-refreshing intelligent dashboards and action workflows triggered by data changes.
Open-source SQL to map platform for BigQuery, Snowflake, and PostGIS.
Open-source analytics notebook for reusable SQL workflows, interactive reports, and AI-assisted data exploration.
Governed, multi-tenant MCP access to your customers' data. Turn your warehouse, dbt, or semantic layer into a secure, per-customer MCP for AI agents.
Intent-as-code workflow engine for AI data pipelines: reviewable YAML DAGs statically checked (schema, permits, cost floor) before execution, with tamper-evident run traces.
Open-source semantic sidecar that compiles YAML-defined dimensions, measures, and metrics into optimized SQL across 8 engines (BigQuery, ClickHouse, Databricks, Dremio, DuckDB, MySQL, PostgreSQL, Snowflake). Unified REST, MCP, and Postgres wire protocol; one model powers AI agents, analytics, DQ rules, and KPIs.
End-to-end data pipeline tool that combines ingestion, transformation (SQL + Python), and data quality in a single CLI. Connects to BigQuery, Snowflake, PostgreSQL, Redshift, and more. Includes VS Code extension with live previews.
Open-source platform for data preparation, synthetic data generation, and AI/data pipelines. Includes reusable skills for automating workflow steps across data and AI tasks.
A Python module that helps you build complex pipelines of batch jobs.
An application cron-like system. [Used](https://chairnerd.seatgeek.com/building-out-the-seatgeek-data-pipeline/) w/Luigi. Deprecated.
Java based application development platform.
A system to programmatically author, schedule, and monitor data pipelines.
A batch workflow job scheduler created at LinkedIn to run Hadoop jobs. Azkaban resolves the ordering through job dependencies and provides an easy-to-use web user interface to maintain and track your workflows.
A workflow scheduler system to manage Apache Hadoop jobs.
DAG based workflow manager. Job flows are defined programmatically in Python. Support output passing between jobs.
An open-source Python library for building data applications.
A lightweight library to define data transformations as a directed-acyclic graph (DAG). If you like dbt for SQL transforms, you will like Hamilton for Python processing.
A framework that makes it easy to build robust and scalable data pipelines by providing uniform project templates, data abstraction, configuration and pipeline assembly.
An open-source framework and web based IDE to manage datasets and their dependencies. SQLX extends your existing SQL warehouse dialect to add features that support dependency management, testing, documentation and more.
A lightweight Python library for building execution pipelines with retry, parallel execution, cron scheduling, and async support.
A reverse-ETL tool that let you sync data from your cloud data warehouse to SaaS applications like Salesforce, Marketo, HubSpot, Zendesk, etc. No engineering favors required—just SQL.
A command line tool that enables data analysts and engineers to transform data in their warehouses more effectively.
Scalable, event-driven, language-agnostic orchestration and scheduling platform to manage millions of workflows declaratively in code.
A warehouse-first Customer Data Platform that enables you to collect data from every application, website and SaaS platform, and then activate it in your warehouse and business tools.
An open source framework that allows you to enforce agreements on how data should be accessed, used, and transformed, regardless of the data platform (Snowflake, BigQuery, DataBricks, etc.)
Self-hosted gateway for safe, auditable queries for agents across approved data sources.
An orchestration and observability platform. With it, developers can rapidly build and scale resilient code, and triage disruptions effortlessly.
The open-source reverse ETL, data activation platform for modern data teams.
Create automated workflows and logic using API's for your notification service. Add templates, batching, preferences, inapp inbox with workflows to trigger notifications directly from your data warehouse.
Open-source data pipeline tool for transforming and integrating data.
An open-source data transformation framework for managing, testing, and deploying SQL and Python-based data pipelines with version control, environment isolation, and automatic dependency resolution.
An open source platform that delivers resilience and manageability to object-storage based data lakes.
A Transactional Catalog for Data Lakes with Git-like semantics. Works with Apache Iceberg tables.
A modular Data Lakehouse platform that simplifies the management and monitoring of Apache Spark clusters across Kubernetes and Hadoop environments.
An open-source, unified metadata management for data lakes, data warehouses, and external catalogs.
FlightPath is a gateway to a data lake's bronze layer, protecting it from invalid external data file feeds as a trusted publisher.
Managed lakehouse platform on Apache Iceberg with DuckDB query compute, S3 storage, Postgres wire protocol, and SQL transforms.
A highly configurable Logstash (1.4.4) - Docker image running Elasticsearch (1.7.0) - and Kibana (3.1.2).
JDBC importer for Elasticsearch.
PostgreSQL Extension that allows creating an index backed by Elasticsearch.
Package golang service into minimal Docker containers.
Easily manage Docker containers & their data.
RancherOS is a 20mb Linux distro that runs the entire OS as Docker containers.
Application Containers for Masses.
Weaving Docker containers into applications.
A lightweight tool for easy deployment and rollback of dockerized applications.
Analyzes resource usage and performance characteristics of running containers.
Docker microservice for saving/restoring volume data to S3.
Docker composition tool with idempotency features for deploying apps composed of multiple containers. Deprecated.
A cluster manager, designed for both long-lived services and short-lived batch processing workloads.
Visualize Docker images and the layers that compose them.
Free real-time DEX data via SSE streaming across 34 blockchains. 30M+ pools, 27M+ tokens, ~1 second price updates. No API key, no rate limits. [Docs](https://docs.dexpaprika.com)
Remote MCP server for real-time financial data, 3.2M+ news articles, ML options pricing, and news bias analysis. Free, no API key. [MCP](https://heliumtrades.com/mcp)
The Streaming APIs give developers low latency access to Twitter's global stream of Tweet data.
Real-time X (Twitter) data API providing tweets, profiles, search, communities and engagement metrics. Up to 50x cheaper than the official X API with 20 req/sec rate limit, JSON output.
Event data simulator. Generates a stream of pseudo-random events from a set of users, designed to simulate web traffic.
Data generation platform for producing synthetic event streams with complex correlations.
Real-time data is available including comments, submissions and links posted to reddit.
GitHub's public timeline since 2011, updated every hour.
Open source repository of web crawl data.
Wikipedia's complete copy of all wikis, in the form of Wikitext source and metadata embedded in XML. A number of raw database tables in SQL form are also available.
A 30-metro composite of US household cost burdens (housing, taxes, childcare, healthcare, transport) aggregated from Census ACS, BLS Consumer Expenditure Survey, and HUD Fair Market Rents. Open methodology, free, no email gate.
The world's most comprehensive authoritative data source knowledge base. 160+ curated sources from governments, international organizations, and research institutions with MCP integration.
42-table synthetic SME dataset with double-entry accounting, tax compliance (AU/US/UK), and temporal realism. CSV, SQL, Parquet, SQLite. Ideal for ETL pipeline testing.
Synthetic financial savings behavior generator for Latin America: users, savings goals, and transactions calibrated on 506K real records (2015–2024). Reproducible by seed, 100% synthetic.
An open-source service monitoring system and time series database.
Simple server that scrapes HAProxy stats and exports them via HTTP for Prometheus consumption.
Intent signal monitoring CLI. Track LinkedIn engagers, keyword posters, job changers, funding events. JSON output for data pipelines.
The DataProfiler is a Python library designed to make data analysis, monitoring, and sensitive data detection easy.
A general-purpose open-source data profiler for high-level analysis of a dataset.
An open-source data profiler specifically focused on discovery and validation of complex patterns in data.
Open-source and free relational database schema discovery and comprehension tool. Documents and diagrams relational database schemas from your Java programs, build tools and the command line. Find database design issues with lint, and write scripts against the database. Includes an MCP Server for use by AI agents.
Open-source agentic data quality framework with LLM-powered diagnosis, root-cause analysis, SQL auto-fix proposals, and 31 rule types — DuckDB, Postgres, BigQuery, Databricks, Athena, Snowflake.
A data catalog tool that integrates into your CI system exposing downstream impact testing of data changes. These tests prevent data changes which might break data pipelines or BI dashboards from making it to production.
An open-source data quality platform for the whole data platform lifecycle from profiling new data sources to applying full automation of data quality monitoring.
Open Source Data Observability for end-to-end Data Journey Observability, data profiling, anomaly detection, and auto-created data quality validation tests.
Open Source data validation framework to manage data quality. Users can define and document “expectations” rules about how data should look and behave.
A vendor-neutral, declarative data quality engine. Define checks in YAML, run anywhere. Includes 16 built-in check types, SQL batch optimizer, anomaly detection, and data contracts.
Zero-config data quality CLI. Profiles every table on first run, then auto-detects anomalies (volume drops, schema drift, freshness misses, distribution shifts) on subsequent runs. No YAML, no rules to write. Works with Postgres, BigQuery, Snowflake, and dbt.
Free online SQL playground for MySQL, PostgreSQL, and SQL Server. Create database structures, run queries, and share results instantly.
Write, run, and test PySpark code on Spark Playground's online compiler. Access real-world sample datasets & solve interview questions to enhance your PySpark skills for data engineering roles.
Decorator-first DataFrame contracts/validation (columns/dtypes/constraints) at function boundaries. Supports Pandas/Polars/PyArrow/Modin.
A Snowflake-compatible emulator for local development and testing.
Real-time data quality firewall for pipelines and APIs. Screens rows in milliseconds for schema drift, null spikes, type mismatches, and data anomalies with PASS / WARN / BLOCK decisions.
Interview practice with SQL query execution, Python, and data modeling exercises.
JSON/XML validation and API contract monitoring tool for debugging and testing structured data.
News, tips, and background on Data Engineering.
Subreddit focused on ETL.
Job board focused on AI, ML, and data engineering roles with 7,400+ listings, salary data, and a free REST API.
The first technical conference that bridges the gap between data scientists, data engineers and data analysts.
Interviews with AI and data infrastructure leaders on building production systems.
The show about modern data infrastructure.
Technical deep dives on AI engineering, from model training to deployment.
Making AI practical, productive, and accessible to everyone.
Daily interviews about technical software topics, including data infrastructure.
How analytics engineers build and maintain data pipelines at scale.
A show where they talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.
A practical introduction to data engineering on the Snowflake cloud data platform.
This blog offers a curated list of top data science books, categorized by topics and learning stages, to aid readers in building foundational knowledge and staying updated with industry trends.
A guide to designing an Apache Iceberg lakehouse from scratch.
A fast, friendly guide to integrating large language models into your data workflows.