18 个仓库
Tools for data manipulation, statistical analysis, and large-scale data processing.
Explore 18 awesome GitHub repositories matching part of an awesome list · Data Processing and Analytics. Refine with filters or upvote what's useful.
D3 is a modular library providing low-level primitives for creating data-driven visualizations. It functions as a flexible framework that allows for direct control over visual presentation by mapping abstract data dimensions to graphical properties, such as position, color, and size, without imposing predefined chart abstractions. The library distinguishes itself by offering specialized tools for complex data representation, including algorithmic layouts for hierarchical structures and geographic projection utilities for mapping spherical coordinates. It also includes a comprehensive suite fo
Library for bringing data to life with web-based visualizations.
Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e
Unified analytics engine for large-scale data processing.
Prisma1 is a TypeScript object-relational mapper and type-safe database client designed for interacting with relational databases. It functions as a system for declarative schema modeling, where database structures are defined in a single schema file that automatically synchronizes with the underlying database. The project provides a type-safe query builder that generates a custom client to ensure database queries match defined schema types at compile time. It also includes a database GUI administrator, providing a visual web interface for browsing, editing, and managing relational database r
Database tools including ORM and migration support.
CMAK is a Kafka cluster management tool and web interface designed for the administration of brokers, topics, and partitions. It provides a centralized system for Kafka cluster governance, encompassing resource administration, access control, and data distribution optimization. The project features a management UI that allows for the creation, deletion, and update of topic configurations and partition counts. It includes a partition rebalancer for executing data reassignment and preferred replica elections to balance load across cluster nodes. The system provides observability through broker
Management tool for Apache Kafka clusters.
OpenRefine is a data cleaning tool and wrangling platform used to transform raw, messy datasets into consistent and structured formats. It operates as a Java-based data processor that runs a local server and provides a web browser interface for managing and manipulating data. The platform includes a data reconciliation engine for matching local entries against external knowledge bases to standardize entities. It also functions as a web data augmentation tool, allowing users to fetch and integrate information from external web sources to enrich their datasets. The system provides a transforma
Power tool for cleaning and transforming messy data.
Azure Docs is the official technical documentation repository for Microsoft Azure, the cloud computing platform. It provides comprehensive guidance on the full spectrum of Azure services, covering everything from core infrastructure components like virtual machines, Kubernetes clusters, and serverless computing to platform services for AI, machine learning, data analytics, and storage. The documentation details how to provision, manage, and govern cloud resources at scale, including policy enforcement, identity management, and cost optimization. The documentation distinguishes Azure through i
Covers ingesting, storing, processing, and analyzing structured and unstructured data at scale.
Storm is a distributed stream processing framework and fault-tolerant compute engine designed for executing real-time continuous computations across a cluster of machines. It functions as a stateful stream processor and cluster topology manager, enabling the deployment and monitoring of distributed data flow configurations. The system ensures exactly-once semantics by utilizing transactional state management to guarantee that every message in a data stream is processed exactly one time. It further operates as a distributed RPC system, allowing for the integration of non-native languages throu
Distributed system for real-time stream processing.
This project is an Android RPA framework designed for automating user interfaces and system tasks on rooted Android devices using Python and ADB. It provides a suite of tools for rooted device management, allowing for programmatic control of system settings, application lifecycles, and shell command execution via a remote API. The framework distinguishes itself through a combination of dynamic instrumentation and AI integration. It can inject scripts into running processes to hook Java interfaces and modifies application behavior in real time. Additionally, it supports large language model in
Provides pre-installed modules for image processing, XML parsing, JSON serialization, and scientific computing.
ggplot2 is a data visualization library for R based on a formal grammar of graphics. It provides a declarative plotting framework that allows users to create complex graphics by combining geometric objects, statistical summaries, and coordinate systems. The system is distinguished by a layered approach to composition, where visualizations are built incrementally by stacking independent geometric, statistical, and coordinate layers. It utilizes a hierarchical styling engine to manage non-data elements such as backgrounds, fonts, and margins, and includes a multi-panel faceting tool for splitti
Implementation of the grammar of graphics for data visualization.
这是一个用于射电干涉成像的软件套件,专门用于处理、分析和重建甚长基线干涉测量(VLBI)观测数据。它提供了使用正则化最大似然法从干涉数据中重建图像的工具,并管理从原始可见度数据到最终图像的端到端数据处理流水线。 该软件以其专用的星际散射模拟器脱颖而出,该模拟器对薄屏散射效应进行建模,并将散射核应用于射电图像。它还具有射电图像合成流水线,能够生成合成 VLBI 数据并执行参数化调查以优化成像配置。 该系统涵盖了广泛的功能,包括偏振和多频图像重建、射电天文校准以及望远镜阵列模拟。它为循环分布、自助法(bootstrap)置信区间估计以及生成观测摘要图以评估图像可靠性和质量提供了全面的数据分析工具。 该工具集支持导入 FITS、UVFITS 和 HDF5 等格式的射电天文数据。
Provides a system for calibrating, manipulating, and analyzing Very Long Baseline Interferometry observations and visibility data.
dplyr 是一个 R 语言数据处理库,为转换表格数据框提供了语法。它充当内存数据框处理器和关系数据代数工具,使用一组一致的动词来过滤、选择和汇总数据。 该项目包含一个 SQL 翻译引擎,可将高级数据处理表达式转换为优化后的查询。这允许用户直接在远程关系数据库和云存储上执行转换,而无需将数据拉取到本地。 该库涵盖了广泛的表格操作,包括列变异、行子集化和关系数据连接。它还提供了分组数据分析功能,允许对数据集进行分区以进行独立的聚合和汇总。
Grammar-based toolkit for efficient data manipulation.
A High Dynamic Range (HDR) Histogram
Tool for recording and analyzing high dynamic range histograms.
Stream summarizer and cardinality estimator.
Library for stream summarization and cardinality estimation.
DeepDive
System for extracting structured data from unstructured text.
Machine Learning Platform and Recommendation Engine built on Kubernetes
Platform for machine learning predictions and recommendations.
Netflix's distributed Data Pipeline
Data pipeline service for collecting and dispatching events.
Web-based notebook for interactive data analytics.
Core components for real-time data pipeline analytics.