Data Engineering
Data pipelines, processing engines, lakehouses, warehouses, streaming platforms, and analytical infrastructure.

Apache Druid
A high-performance, real-time analytical database designed for rapid queries on massive transactional and event-driven datasets.

Delta Lake
An open-source storage framework that enables building a Lakehouse architecture on top of existing cloud object stores, bringing ACID transactions, scalable metadata handling, and unified stream and batch data processing.

ClickHouse
A high-performance, open-source column-oriented database management system designed for real-time online analytical processing (OLAP).

Apache Spark
A unified, multi-language analytics engine designed for large-scale distributed data processing, machine learning, and stream processing.

Trino
A highly parallel, distributed SQL query engine designed for fast analytical queries against diverse data sources ranging from gigabytes to petabytes.

DuckDB
An embedded, high-performance analytical SQL database engine designed for in-process OLAP workloads and seamless data ecosystem integration.

Apache Iceberg
An open-source, high-performance table format for massive analytic datasets, bringing ACID transactions and SQL-like table behavior to object storage.