awesome-spark/awesome-spark

spark-jobserver by spark-jobserver

REST job server for Apache Spark

updated at May 9, 2024, 3:16 a.m.

Scala

221 +0

2,841 +1

1,004 +0

GitHub

spark-cassandra-connector by datastax

DataStax Connector for Apache Spark to Apache Cassandra

updated at May 9, 2024, 3:23 a.m.

Scala

162 +0

1,931 +1

913 -1

GitHub

flint by twosigma

A Time Series Library for Apache Spark

updated at May 9, 2024, 3:30 a.m.

Scala

77 +0

992 +0

184 +0

GitHub

spark-daria by MrPowers

Essential Spark extensions and helper methods ✨😲

updated at May 9, 2024, 4:48 p.m.

Scala

33 +0

743 +1

148 +0

GitHub

sparklyr by sparklyr

R interface for Apache Spark

updated at May 9, 2024, 5:06 p.m.

R

73 +0

926 +2

302 +0

GitHub

incubator-livy by apache

Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.

updated at May 10, 2024, 3:34 a.m.

Scala

57 +0

857 +1

594 +0

GitHub

spark-xml by databricks

XML data source for Spark SQL and DataFrames

updated at May 10, 2024, 3:38 a.m.

Scala

40 +0

489 +1

223 +0

GitHub

dist-keras by cerndb

Distributed Deep Learning, with a focus on distributed training, using Keras and Apache Spark.

updated at May 10, 2024, 5:12 a.m.

Python

49 +0

623 -1

170 +0

GitHub

SynapseML by Microsoft

Simple and Distributed Machine Learning

updated at May 10, 2024, 10:34 a.m.

Scala

146 +0

4,975 +3

815 +0

GitHub

graphframes by graphframes

None

updated at May 10, 2024, 11:48 a.m.

Scala

58 +0

972 +1

232 +0

GitHub

neo4j-spark-connector by neo4j-contrib

Neo4j Connector for Apache Spark, which provides bi-directional read/write access to Neo4j from Spark, using the Spark DataSource APIs

updated at May 10, 2024, 1:50 p.m.

Scala

35 +0

304 +1

114 +0

GitHub

dplyr by tidyverse

dplyr: A grammar of data manipulation

updated at May 10, 2024, 1:59 p.m.

R

247 +1

4,665 +6

2,119 +1

GitHub

adam by bigdatagenomics

ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.

updated at May 10, 2024, 3:25 p.m.

Scala

100 +0

968 +1

304 -1

GitHub

Mobius by Microsoft

C# and F# language binding and extensions to Apache Spark

updated at May 10, 2024, 9:08 p.m.

C#

145 +0

939 +2

212 +0

GitHub

cromwell by broadinstitute

Scientific workflow engine designed for simplicity & scalability. Trivially transition between one off use cases to massive scale production environments

updated at May 10, 2024, 10:06 p.m.

Scala

112 +0

959 -1

350 +0

GitHub

delta by delta-io

An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs

updated at May 10, 2024, 11 p.m.

Scala

215 +0

6,935 +13

1,583 +3

GitHub

koalas by databricks

Koalas: pandas API on Apache Spark

updated at May 11, 2024, 3:34 a.m.

Python

316 +0

3,321 +0

355 +0

GitHub

spark-testing-base by holdenk

Base classes to use when writing tests with Spark

updated at May 11, 2024, 6 a.m.

Scala

78 +0

1,497 +4

358 +0

GitHub

sedona by apache

A cluster computing framework for processing large-scale geospatial data

updated at May 11, 2024, 6:20 a.m.

Java

96 +0

1,784 +4

646 +1

GitHub

kyuubi by apache

Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.

updated at May 11, 2024, 7:40 a.m.

Scala

62 -1

1,947 +6

860 +1

GitHub