awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

15 रिपॉजिटरी

Awesome GitHub RepositoriesData Catalogs

Collaborative repositories for indexing and sharing structured data sources.

Distinct from Open Source Software: Distinct from general open-source software: focuses on the collaborative maintenance of data-centric registries rather than code libraries.

Explore 15 awesome GitHub repositories matching development tools & productivity · Data Catalogs. Refine with filters or upvote what's useful.

Awesome Data Catalogs GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • awesome-selfhosted/awesome-selfhostedawesome-selfhosted का अवतार

    awesome-selfhosted/awesome-selfhosted

    299,516GitHub पर देखें↗

    यह प्रोजेक्ट निजी सर्वर वातावरण और होम लैब में डिप्लॉयमेंट के लिए डिज़ाइन किए गए ओपन-सोर्स सॉफ्टवेयर की एक समुदाय-क्यूरेटेड निर्देशिका है। यह मुख्यधारा की क्लाउड सेवाओं के स्वतंत्र, स्व-होस्ट किए गए विकल्पों को खोजने के लिए एक व्यापक संसाधन के रूप में कार्य करता है, जिससे उपयोगकर्ता अपने डिजिटल इंफ्रास्ट्रक्चर पर पूर्ण डेटा स्वामित्व और नियंत्रण बनाए रख सकते हैं। निर्देशिका को एक पदानुक्रमित वर्गीकरण के माध्यम से संरचित किया गया है जो अनुप्रयोगों के एक विशाल संग्रह को तार्किक श्रेणियों में व्यवस्थित करता है, जो मीडिया प्रबंधन और डेटा एनालिटिक्स से लेकर निजी संचार और टीम उत्पादकता टूल तक फैला हुआ है। यह एक सहयोगात्मक पीयर-रिव्यू प्रक्रिया के माध्यम से खुद को अलग करती है, जहाँ समुदाय के सदस्य निर्देशिका को सटीक और विश्वसनीय सुनिश्चित करने के लिए प्रत्येक सबमिशन की गुणवत्ता और प्रासंगिकता को मान्य करते हैं। प्रोजेक्ट इंफ्रास्ट्रक्चर ऑटोमेशन, कंटेनर-आधारित सर्विस डिप्लॉयमेंट और घोषणात्मक कॉन्फ़िगरेशन प्रबंधन सहित क्षमताओं के एक व्यापक क्षेत्र को कवर करता है। ये टूल उपयोगकर्ताओं को पुनरुत्पादनीय सर्वर वातावरण बनाए रखने और निजी हार्डवेयर पर जटिल सर्विस निर्भरताओं को प्रबंधित करने में सहायता करते हैं। निर्देशिका को एक वर्ज़न-कंट्रोल रिपॉजिटरी के रूप में बनाए रखा जाता है, यह सुनिश्चित करते हुए कि सभी अपडेट और समुदाय-संचालित परिवर्तन ट्रैक किए जाते हैं और पारदर्शी हैं।

    Stores and organizes recorded trail paths and associated metadata into a searchable database for outdoor enthusiasts.

    awesomeawesome-listcloud
    GitHub पर देखें↗299,516
  • dagster-io/dagsterdagster-io का अवतार

    dagster-io/dagster

    14,974GitHub पर देखें↗

    Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality. The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows.

    Maintains a centralized view of data assets, workflows, and lineage to help teams discover and reuse components.

    Pythonanalyticsdagsterdata-engineering
    GitHub पर देखें↗14,974
  • aws/aws-cdkaws का अवतार

    aws/aws-cdk

    12,817GitHub पर देखें↗

    The AWS Cloud Development Kit is an infrastructure-as-code framework that enables developers to define and provision cloud resources using familiar programming languages. By utilizing construct-based synthesis, it translates high-level, object-oriented code into declarative templates, allowing for the automated management of complex cloud environments through a centralized, code-driven control plane. The framework distinguishes itself through its ability to model infrastructure as a dependency-aware resource graph, ensuring that components are provisioned and updated in the correct order. It

    Registers data moved into storage within a central metadata repository to simplify discovery and access for analytics and machine learning tools.

    TypeScriptawscloud-infrastructurehacktoberfest
    GitHub पर देखें↗12,817
  • datahub-project/datahubdatahub-project का अवतार

    datahub-project/datahub

    12,141GitHub पर देखें↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Maintains a centralized repository of metadata to provide transparency into data sources and usage patterns.

    Pythondata-catalogdata-discoverydata-governance
    GitHub पर देखें↗12,141
  • insforge/insforgeInsForge का अवतार

    InsForge/InsForge

    11,794GitHub पर देखें↗

    InsForge is a backend-as-a-service platform that provides an integrated suite of tools for managing relational databases, identity provision, object storage, and serverless compute. It functions as an open-source identity provider and a PostgreSQL database manager featuring integrated vector storage and row-level security. The platform serves as an LLM orchestration gateway, offering a unified endpoint to route requests across various AI providers through an OpenAI-compatible interface. It enables AI-driven application generation and connects AI agents to backend resources using a standardize

    Maintains a standardized catalog of available AI models to identify their capabilities and modalities.

    TypeScriptaiai-agentscoding
    GitHub पर देखें↗11,794
  • quay/clairquay का अवतार

    quay/clair

    11,012GitHub पर देखें↗

    Clair is a container image vulnerability scanner and security analyzer. It performs static analysis of container images by matching package contents against vulnerability databases to identify security risks across different package formats and architectures. The project functions as both an image indexer and a vulnerability database manager. It processes container layers into intermediate representations to enable fast security lookups and synchronizes security metadata from multiple external sources to maintain a local registry. Capability areas include continuous security monitoring, whic

    Periodically synchronizes security metadata from remote vulnerability databases to maintain a local searchable registry.

    Goclaircontainersdocker
    GitHub पर देखें↗11,012
  • delta-io/deltadelta-io का अवतार

    delta-io/delta

    8,596GitHub पर देखें↗

    Delta is a lakehouse table format that brings ACID transactions and data warehouse consistency to large scale data lakes on cloud object storage. It serves as an ACID transaction manager, coordinating atomic commits and serializable isolation for concurrent reads and writes across distributed compute engines. The project provides a multi-engine interoperability layer that uses format translation to allow diverse SQL engines and processing frameworks to read and write the same tables. It functions as a data versioning system, utilizing a transaction log to enable time travel, historical snapsh

    Generates compatible metadata for external table formats at commit time to enable cross-engine interoperability.

    Scalaacidanalyticsbig-data
    GitHub पर देखें↗8,596
  • alibaba/higressalibaba का अवतार

    alibaba/higress

    7,558GitHub पर देखें↗

    Higress is an AI API gateway and cloud-native traffic manager that functions as a Kubernetes ingress controller. It provides a centralized system for routing, securing, and optimizing traffic directed toward large language models, AI agents, and microservice architectures. The project distinguishes itself through deep AI orchestration, including the ability to host and manage Model Context Protocol servers that transform REST APIs into tools for AI agents. It features specialized AI infrastructure for model request proxying, protocol translation across multiple providers, and semantic-based c

    Distributes models and agents through a centralized artifact catalog featuring versioning and grayscale releases.

    Goai-gatewayai-nativeapi-gateway
    GitHub पर देखें↗7,558
  • cloudquery/cloudquerycloudquery का अवतार

    cloudquery/cloudquery

    6,438GitHub पर देखें↗

    CloudQuery is a cloud infrastructure ETL tool and multi-cloud data pipeline designed to collect, synchronize, and normalize resource metadata from various cloud providers and SaaS platforms. It functions as a centralized asset inventory manager and security posture manager, extracting configuration and state data into relational databases, data lakes, or data warehouses. The system distinguishes itself by transforming complex, nested cloud API responses into flat relational tables, enabling the use of standard SQL for asset querying and analysis. It employs a modular plugin system for data ex

    Automates the synchronization of resource configurations from multiple cloud providers into a centralized destination.

    Goairbyteattack-surface-managementaws
    GitHub पर देखें↗6,438
  • ethanniser/nextfasterethanniser का अवतार

    ethanniser/NextFaster

    4,860GitHub पर देखें↗

    NextFaster is a Next.js e-commerce storefront that uses artificial intelligence to generate a complete product catalog, including categories, descriptions, and images. Product images are stored in cloud object storage and served directly via CDN, offloading delivery from the origin server. Pages are delivered with pre-rendered shells from the edge, with dynamic content streamed in afterward for fast initial interactivity. All write operations—such as orders and catalog updates—are handled through server-side functions for data consistency and security. The project differentiates itself by com

    Generates a complete product catalog with categories, descriptions, and images using AI.

    TypeScript
    GitHub पर देखें↗4,860
  • hanydd/bilibilisponsorblockhanydd का अवतार

    hanydd/BilibiliSponsorBlock

    4,862GitHub पर देखें↗

    BilibiliSponsorBlock is a content filtering system and API server designed to identify and remove sponsored segments and filler content from Bilibili video playback. It utilizes a crowdsourced segment database where users contribute and vote on timestamps to create a shared repository of skippable video sections. The project features a video metadata synchronizer that links equivalent videos across different platforms, allowing skip markers and timing data to be shared between mirrored content. It implements a reputation-based permission system to manage submissions and edits, alongside a pri

    Links videos across different platforms to synchronize skip markers and timing data.

    TypeScriptadblockbilibilibrowser-extension
    GitHub पर देखें↗4,862
  • awslabs/aws-data-wranglerawslabs का अवतार

    awslabs/aws-data-wrangler

    4,107GitHub पर देखें↗

    यह प्रोजेक्ट एक AWS pandas एकीकरण लाइब्रेरी और डेटा पाइपलाइन फ्रेमवर्क है जिसे स्थानीय मेमोरी और AWS स्टोरेज और एनालिटिक्स सेवाओं के बीच डेटा की आवाजाही और रूपांतरण को सरल बनाने के लिए डिज़ाइन किया गया है। यह एक क्लाउड डेटा लेक टूलकिट और स्टोरेज फाइल मैनेजर के रूप में कार्य करता है, जो उपयोगकर्ताओं को विभिन्न क्लाउड वातावरणों में संरचित डेटा को पढ़ने, लिखने और बदलने की अनुमति देता है। लाइब्रेरी एक डिस्ट्रीब्यूटेड कंप्यूट ऑर्केस्ट्रेटर के रूप में खुद को अलग करती है जो EMR जैसे वातावरण में क्लस्टर्स को प्रबंधित करने में सक्षम है ताकि उन डेटासेट्स को प्रोसेस किया जा सके जो एक मशीन की मेमोरी सीमा से अधिक हैं। यह वेक्टर इंडेक्स को प्रबंधित करने और क्लाउड स्टोरेज बकेट्स के भीतर समानता खोज (similarity searches) करने के लिए विशेष क्षमताएं भी प्रदान करती है। इसकी व्यापक क्षमता सतह DynamoDB, RDS और Timestream जैसी सेवाओं के लिए क्लाउड डेटाबेस ETL, और AWS Glue के माध्यम से क्लाउड डेटा कैटलॉग प्रबंधन को कवर करती है। यह Athena और Redshift के माध्यम से सर्वरलेस डेटा एनालिटिक्स का समर्थन करती है, और S3 ऑब्जेक्ट्स को प्रबंधित करने, OpenSearch में दस्तावेजों को इंडेक्स करने और CloudWatch लॉग्स का विश्लेषण करने के लिए यूटिलिटीज प्रदान करती है।

    Automatically synchronizes table metadata with central catalogs when writing data frames to cloud storage.

    Python
    GitHub पर देखें↗4,107
  • open-wanderer/wandereropen-wanderer का अवतार

    open-wanderer/wanderer

    3,688GitHub पर देखें↗

    Wanderer is a self-hosted trail database and federated social network for storing GPS tracks and hiking routes. It functions as a GPX route manager and geographic route planner that allows users to catalog outdoor explorations in a private database. The system utilizes the ActivityPub protocol to enable decentralized federation, allowing users to follow others and share activities across independent server instances. This social layer includes collaborative route exchange, engagement management for likes and comments, and granular content visibility controls. The platform provides tools for

    Provides a self-hosted database for cataloging and searching outdoor travel tracks and hiking route metadata.

    Gogeographygpsmeilisearch
    GitHub पर देखें↗3,688
  • 5rahim/seanime5rahim का अवतार

    5rahim/seanime

    2,864GitHub पर देखें↗

    This project is a self-hosted media server for organizing, streaming, and tracking anime and manga collections. It functions as a BitTorrent streaming client that allows video content to be played directly from torrents and cloud storage, a manga reader and tracker, and a media processing system using hardware-accelerated transcoding to ensure browser compatibility. The system distinguishes itself through synchronized media viewing, enabling users to host watch parties by coordinating playback in real time across multiple devices. It also features an extensible framework with a JavaScript-bas

    Automatically fetches and updates series and episode metadata from external data platforms.

    Goanilistanimeanime-downloader
    GitHub पर देखें↗2,864
  • apache/gravitinoapache का अवतार

    apache/gravitino

    2,866GitHub पर देखें↗

    Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across diverse data sources and cloud storage. It serves as a centralized interface for governing schemas, access controls, and tagging across relational databases, messaging queues, and object stores. The project distinguishes itself by unifying the management of AI assets, such as machine learning models and their version lineages, alongside traditional tabular data. It also implements the Iceberg REST specification to provide a standardized metadata server and proxy for lakehouse

    Maintains structured catalogs for AI-specific artifacts, including model metadata and governance.

    Javaai-catalogdata-catalogdatalake
    GitHub पर देखें↗2,866
  1. Home
  2. Development Tools & Productivity
  3. Open Source Software
  4. Data Catalogs

सब-टैग एक्सप्लोर करें

  • AI Artifact Catalogs1 सब-टैगMaintains records of models, prompts, and agents with data-to-model lineage and performance metrics. **Distinct from Data Catalogs:** Distinct from Data Catalogs: focuses on AI-specific artifacts like prompts and models rather than structured data sources.
  • Managed Catalog ProvisioningAutomated deployment of hosted data catalogs to reduce operational overhead. **Distinct from Data Catalogs:** Focuses on the provisioning of managed catalog services rather than collaborative repository maintenance.
  • Metadata Synchronizers1 सब-टैगAutomated services for bi-directional metadata synchronization between external data platforms and centralized catalogs. **Distinct from Data Catalogs:** Distinct from Data Catalogs: focuses on the synchronization mechanism specifically rather than the catalog repository itself.
  • Trail DatabasesRegistries for cataloging and searching outdoor travel tracks and metadata. **Distinct from Data Catalogs:** Distinct from Data Catalogs: focuses on geographic trail and route data rather than general-purpose structured data.