druid by apache

Apache Druid: a high performance real-time analytics database.

updated at May 5, 2024, 9:16 a.m.

Java

592 -1

13,204 +9

3,637 +3

GitHub
nessie by projectnessie

Nessie: Transactional Catalog for Data Lakes with Git-like semantics

updated at May 4, 2024, 9:33 a.m.

Java

27 -1

841 +7

116 +1

GitHub
kairosdb by kairosdb

Fast scalable time series database

updated at May 3, 2024, 7:05 p.m.

Java

118 +0

1,726 +2

344 -1

GitHub
dqo by dqops

Data Quality and Observability platform for the whole data lifecycle, from profiling new data sources to full automation with Data Observability. Configure data quality checks from the UI or in YAML files, let DQOps run the data quality checks daily to detect data quality issues.

updated at May 3, 2024, 4:46 p.m.

Java

5 +0

54 +1

11 +0

GitHub
gobblin by apache

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

updated at May 2, 2024, 9:14 p.m.

Java

167 +0

2,190 +0

742 +0

GitHub
zilla by aklivity

🦎 A multi-protocol, event-native proxy. Securely interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

updated at May 2, 2024, 3:48 p.m.

Java

9 +0

486 +0

47 +0

GitHub
Gaffer by gchq

A large-scale entity and relation database supporting aggregation of properties

updated at May 1, 2024, 11:33 a.m.

Java

142 +0

1,734 +1

354 +0

GitHub
opentsdb by OpenTSDB

A scalable, distributed Time Series Database.

updated at April 29, 2024, 1:49 p.m.

Java

337 +0

4,951 +2

1,253 +0

GitHub
elasticsearch-jdbc by jprante

JDBC importer for Elasticsearch

updated at April 23, 2024, 2:40 a.m.

Java

231 +0

2,838 +0

711 -1

GitHub
secor by pinterest

Secor is a service implementing Kafka log persistence

updated at April 22, 2024, 8:31 a.m.

Java

70 +0

1,835 +0

541 +0

GitHub
incubator-hivemall by apache

Mirror of Apache Hivemall (incubating)

updated at April 6, 2024, 6:43 a.m.

Java

32 +0

310 +0

119 +0

GitHub
blueflood by rax-maas

A distributed system designed to ingest and process time series data

updated at April 3, 2024, 8:32 p.m.

Java

95 +0

592 +0

102 +0

GitHub
heroic by spotify

The Heroic Time Series Database

updated at April 2, 2024, 5:42 p.m.

Java

58 +0

843 +0

109 +0

GitHub
bistro by asavinov

A general-purpose data analysis engine radically changing the way batch and stream data is processed

updated at Feb. 7, 2024, 7:30 p.m.

Java

2 +0

7 +0

0 +0

GitHub
deep-spark by Stratio

Connecting Apache Spark with different data stores [DEPRECATED]

updated at Jan. 1, 2024, 6:17 p.m.

Java

115 +0

197 +0

42 +0

GitHub