Upserts, Deletes And Incremental Processing on Big Data.
-
Updated
Sep 29, 2026 - Java
Upserts, Deletes And Incremental Processing on Big Data.
汇总Apache Hudi相关资料
Incremental processing and maintaining data freshness with CocoIndex and LanceDB
A modern banking data pipeline built with Dagster and DBT!
Autonomous email agent on one Postgres. Inbox rows and pgvector embeddings in the same table, rules decide, the model only drafts, dry-run by default.
Monitor AI agent memory, skills, and behavior in a live terminal HUD for Hermes.
Incrementally parse and process structured InputStream content one item at a time without materializing the complete input in memory.
Scalable data engineering pipeline processing 38M+ NYC Yellow Taxi trips with Databricks, PySpark, Delta Lake and incremental processing.
Reusable data matching system with incremental processing for efficient reuse of historical results.
Real-time CDC pipeline using Snowpipe Streaming and Dynamic Tables for continuous ingestion, transformation, and analytics in Snowflake.
Automated incremental retail data pipeline built with Databricks, SQL, Delta Lake, and Medallion Architecture.
A production-grade cryptocurrency data pipeline built on GCP that ingests real-time market data from the CoinGecko API, implements Medallion architecture (Raw → Staging → Curated), and supports idempotent backfill and metadata-driven incremental processing for reliable, scalable analytics.
Production-style Enterprise Sales Lakehouse using PySpark, Delta Lake and Medallion Architecture with incremental processing, data quality, monitoring and business analytics.
A declarative SQL data pipeline built with Snowflake Dynamic Tables, using a layered RAW → enrichment → fact → business metrics architecture, incremental refresh testing, monitoring, and a Semantic View layer.
Delta check state machine design for MiFID II regulatory reporting — NEWT / REPL / CANC lifecycle | Phase A → D pipeline
❄️ 🔨End-to-end data engineering project built in Snowflake using a Medallion Architecture (🟫 Bronze → 🟦 Silver → 🟨 Gold). The project demonstrates ELT pipeline design, data ingestion from AWS S3, data cleaning and transformation, incremental processing, and dimensional modelling using a star schema.
End-to-end data engineering pipeline on Databricks with Delta Lake, incremental watermarking, and Dockerized Airflow orchestration (Postgres-backed, SLA-enabled).
Reproducible lakehouse benchmark lab exploring Parquet file layouts, incremental processing, and Apache Iceberg snapshots.
End- to-End Performance-optimized sales data pipeline using Medallion Architecture with broadcast joins, fact/dimension modeling, Autoloader & incremental processing
End-to-end data pipeline that ingests job market data from the Adzuna API and processes it using a Medallion Architecture (Bronze→Silver→Gold). Includes data quality validation, incremental loading into DuckDB, and Airflow orchestration. Generates analytics-ready datasets for role demand, salary trends, and skill demand across multiple countries.
To associate your repository with the incremental-processing topic, visit your repo's landing page and select "manage topics."