awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
activeloopai avatar

activeloopai/Hub

0
View on GitHub↗
9,177 نجوم·715 تفرعات·C++·Apache-2.0·8 مشاهداتdeeplake.ai↗

Hub

Hub is a multimodal AI data lake and vector database designed for storing and querying embeddings, text, audio, and images. It functions as a dataset version control system and a machine learning data streaming engine to support large-scale model training.

The system utilizes a serverless PostgreSQL vector store to index high-dimensional embeddings for semantic search. It provides a visual interface for inspecting multimodal datasets and viewing annotations such as bounding boxes and masks.

The platform handles cloud-agnostic storage synchronization and implements lazy, compressed data streaming to move datasets from remote sources into deep learning frameworks. It maintains dataset lineage and versioning to track iterations across the development lifecycle.

Features

  • Dataset Versioning Systems - Provides a comprehensive system for tracking and managing versions of massive datasets used for machine learning.
  • Data Lakes - Provides a scalable multimodal data lake for organizing and retrieving large datasets for AI training.
  • Data Lineage - Tracks data lineage and transformation history of multimodal datasets throughout the machine learning development lifecycle.
  • Large Scale Training - Streams massive datasets from cloud storage into deep learning frameworks to avoid local memory exhaustion.
  • Multimodal - Functions as a multimodal AI data lake that organizes diverse data types into a unified storage format.
  • GPU-Accelerated Data Streams - Streams compressed data arrays directly from cloud storage into deep learning frameworks to optimize training.
  • PostgreSQL Vector Stores - Utilizes serverless PostgreSQL vector stores to index and store high-dimensional embeddings for semantic retrieval.
  • Large Dataset Streaming - Supports lazy streaming of massive datasets from remote storage to prevent system memory exhaustion during training.
  • Streaming Compression Engines - Implements streaming compression engines to transmit data arrays to ML frameworks with reduced memory overhead.
  • Training Sample Streaming - Streams individual training samples from large multimodal datasets to increase fine-tuning speed.
  • Multimodal Data Storage - Saves embeddings, audio, text, and images in a unified format optimized for deep learning applications.
  • Multimodal Databases - Organizes and stores text, images, audio, and embeddings in a unified format optimized for deep learning.
  • Vector Indexing - Implements vector indexing to enable fast semantic search and retrieval of relevant information for LLMs.
  • Vector Search - Performs vector search to retrieve relevant information based on mathematical similarity in high-dimensional spaces.
  • Vector Embedding Indexes - Indexes vector embeddings to enable high-performance similarity search across large multimodal datasets.
  • Bounding Box Visualizers - Renders bounding boxes and masks over multimodal data for immediate visual inspection of annotations.
  • AI Dataset Visualizers - Provides a visual interface for inspecting multimodal datasets and viewing spatial annotations like bounding boxes.
  • Cloud Synchronization Services - Synchronizes multimodal datasets across different cloud providers and local storage using a unified interface.
  • Dataset Annotations - Offers a visual interface for inspecting multimodal datasets and viewing annotations like bounding boxes and masks.
  • Remote Data Fetching - Implements remote data fetching to stream multimodal training data from cloud sources to local frameworks.
  • Vector Databases - Indexes and searches high-dimensional vector embeddings to enable semantic retrieval for LLM applications.
  • Cloud-Agnostic Synchronization - Provides a unified interface to synchronize and stream data across diverse cloud storage providers.
  • Cross-Cloud Synchronization - Enables the movement and streaming of datasets across different cloud providers using a single interface.
  • Multimodal Visualizers - Ships a visual tool for inspecting multimodal datasets by rendering diverse data types in a synchronized view.
  • Deep Learning and Computer Vision - Version-controlled dataset management for deep learning.
  • General Machine Learning - Dataset management for TensorFlow and PyTorch.
  • إدارة البيانات - Version-controlled dataset management for machine learning workflows.
  • Data Management and Catalogues - Version-controlled dataset management for deep learning workflows.
  • أدوات المطور - Dataset management and versioning for deep learning.
  • Image Processing and Manipulation - Manages and versions large datasets for machine learning pipelines.

سجل النجوم

مخطط تاريخ النجوم لـ activeloopai/hubمخطط تاريخ النجوم لـ activeloopai/hub

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Hub

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Hub.
  • activeloopai/deeplakeالصورة الرمزية لـ activeloopai

    activeloopai/deeplake

    9,175عرض على GitHub↗

    DeepLake is AI data infrastructure consisting of a multimodal data lake, a hybrid search engine, and a serverless vector database. It provides a PostgreSQL-based AI data runtime that combines multimodal storage with streaming pipelines to load and shuffle datasets from cloud storage directly into deep learning training pipelines. The system utilizes lazy indexing to store and slice images, audio, and video without loading entire files into memory. It enables retrieval-augmented generation by persisting high-dimensional embeddings in a serverless vector store and implementing hybrid search tha

    C++agentagentic-ragai
    عرض على GitHub↗9,175
  • lancedb/lancedbالصورة الرمزية لـ lancedb

    lancedb/lancedb

    9,031عرض على GitHub↗

    LanceDB is a vector database and columnar data store designed to function as a versioned dataset manager and vector search engine. It serves as a high-performance backend for indexing and retrieving high-dimensional embeddings, providing the foundation for machine learning data pipelines. The system distinguishes itself through a combination of cloud-native object storage and immutable version tracking, allowing for data time-travel and reproducible AI experiments. It integrates hybrid search capabilities, merging dense vector similarity with BM25 full-text search and SQL-like scalar filters

    HTMLapproximate-nearest-neighbor-searchimage-searchnearest-neighbor-search
    عرض على GitHub↗9,031
  • chroma-core/chromaالصورة الرمزية لـ chroma-core

    chroma-core/chroma

    26,198عرض على GitHub↗

    Chroma is a specialized vector database designed to index and retrieve high-dimensional data representations for semantic similarity search. It functions as a comprehensive platform for information retrieval, enabling the storage and management of unstructured documents alongside structured metadata. By mapping data into numerical representations, the system facilitates rapid similarity lookups across large datasets. The platform distinguishes itself through a hybrid search infrastructure that combines dense vector embeddings with sparse keyword and regular expression matching to balance sema

    Rustaidatabasedocument-retrieval
    عرض على GitHub↗26,198
  • infiniflow/infinityالصورة الرمزية لـ infiniflow

    infiniflow/infinity

    4,570عرض على GitHub↗

    Infinity is a distributed vector database and multimodal vector store designed to manage large-scale datasets for retrieval and similarity search. It serves as a backend for large language model applications and retrieval augmented generation pipelines by storing and retrieving dense vectors, sparse vectors, and full-text data. The system functions as a hybrid search engine, combining vector embeddings and full-text search with reranking algorithms to identify the most relevant documents. It supports multimodal data storage, allowing the maintenance of diverse data types including tensors, st

    C++ai-nativeapproximate-nearest-neighbor-searchbm25
    عرض على GitHub↗4,570
عرض جميع البدائل الـ 30 لـ Hub→

الأسئلة الشائعة

ما هي وظيفة activeloopai/hub؟

Hub is a multimodal AI data lake and vector database designed for storing and querying embeddings, text, audio, and images. It functions as a dataset version control system and a machine learning data streaming engine to support large-scale model training.

ما هي الميزات الرئيسية لـ activeloopai/hub؟

الميزات الرئيسية لـ activeloopai/hub هي: Dataset Versioning Systems, Data Lakes, Data Lineage, Large Scale Training, Multimodal, GPU-Accelerated Data Streams, PostgreSQL Vector Stores, Large Dataset Streaming.

ما هي البدائل مفتوحة المصدر لـ activeloopai/hub؟

تشمل البدائل مفتوحة المصدر لـ activeloopai/hub: activeloopai/deeplake — DeepLake is AI data infrastructure consisting of a multimodal data lake, a hybrid search engine, and a serverless… lancedb/lancedb — LanceDB is a vector database and columnar data store designed to function as a versioned dataset manager and vector… chroma-core/chroma — Chroma is a specialized vector database designed to index and retrieve high-dimensional data representations for… ryancodrai/turbovec — TurboVec is a high-performance Rust vector database and quantized search index designed for storing and retrieving… infiniflow/infinity — Infinity is a distributed vector database and multimodal vector store designed to manage large-scale datasets for… databendlabs/databend — Databend is a cloud-native data warehouse and OLAP database designed for large-scale analytics. It functions as a…