How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
| Status | Stable | Latest | Source code|Spark compatibility| |:----------:|:-------------:|:------|:------:|:------| | GeoSpark | | | |Spark 2.X, 1.X| | GeoSparkSQL | | | | Spark SQL 2.1, 2.2| | GeoSparkViz | | | |Spark 2.X, 1.X|
The Archives Unleashed Toolkit is an open-source platform for analyzing web archives using Apache Spark, and makes use of Sparkling for parsing W/ARC records. The toolkit provides powerful tools for analytics and data processing. It is part of the Archives Unleashed Project.
ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.
Cloud-native genomic dataframes and batch computing
The main features of graphframes/graphframes are: Domain Specific Processing.
Projects with overlapping indexed features include: apache/incubator-sedona — | Status | Stable | Latest | Source code|Spark compatibility| |:----------:|:-------------:|:------|:------:|:------|… archivesunleashed/aut — The Archives Unleashed Toolkit is an open-source platform for analyzing web archives using Apache Spark, and makes use… bigdatagenomics/adam — ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache… hail-is/hail — Cloud-native genomic dataframes and batch computing. neo4j-contrib/neo4j-spark-connector — This repository contains the Neo4j Connector for Apache Spark.