awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 रिपॉजिटरी

Awesome GitHub RepositoriesEnterprise Data Platforms

Comprehensive data management systems designed for high-scale, regulated production environments with built-in security and auditing.

Distinct from Enterprise Data Services: Distinct from Enterprise Data Services: focuses on the platform as a complete database system rather than just connectors or integration services.

Explore 14 awesome GitHub repositories matching data & databases · Enterprise Data Platforms. Refine with filters or upvote what's useful.

Awesome Enterprise Data Platforms GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • pankod/refinepankod का अवतार

    pankod/refine

    34,909GitHub पर देखें↗

    Refine is a React-based framework for building data-intensive internal tools, admin panels, and B2B applications. It functions as a data-driven UI library and a headless admin panel generator that connects frontends to external backend services using standardized logic for state management and network request handling. The project decouples business logic from the presentation layer, allowing any custom design system or interface library to be applied to the application. It includes a CRUD application generator that automatically creates user interfaces for managing records based on the struc

    Enables the construction of centralized tools for monitoring record changes and managing large-scale enterprise datasets.

    TypeScript
    GitHub पर देखें↗34,909
  • mongodb/mongomongodb का अवतार

    mongodb/mongo

    28,158GitHub पर देखें↗

    This project is a distributed, document-oriented database system designed to store information in flexible, hierarchical structures. It supports horizontal scaling through automated sharding and maintains high availability across global clusters using a multi-node replication protocol. By executing multi-document operations as atomic units, the system ensures data integrity and consistency across distributed environments. The platform distinguishes itself by integrating advanced vector-based indexing, which enables semantic similarity searches alongside traditional geospatial and lexical quer

    Provides a secure, enterprise-grade data platform with granular access controls, encryption, and auditing for regulated production environments.

    C++c-plus-plusdatabasemongodb
    GitHub पर देखें↗28,158
  • nocobase/nocobasenocobase का अवतार

    nocobase/nocobase

    21,542GitHub पर देखें↗

    This platform is a modular, metadata-driven framework designed for building custom business applications and data management systems without traditional coding. It functions as a low-code environment where data models, user interfaces, and business logic are defined through visual configurations rather than hardcoded views. The architecture supports multi-tenant isolation, allowing multiple independent applications to run within a single shared memory space while maintaining strict logical separation of data and configurations. What distinguishes this system is its deep integration of artific

    Provides a centralized platform for defining relational data models, connecting external databases, and visualizing information through interactive dashboards.

    TypeScriptadmin-dashboardairtableapp-builder
    GitHub पर देखें↗21,542
  • airbytehq/airbyteairbytehq का अवतार

    airbytehq/airbyte

    21,472GitHub पर देखें↗

    Airbyte is a data integration platform designed to synchronize information between diverse applications, databases, and data warehouses. It functions as an extract, transform, and load orchestrator that manages automated data movement workflows across cloud, on-premise, and hybrid environments. The platform provides a standardized interface for connectors, enabling the movement of structured and unstructured data while maintaining stateful checkpoints for reliable incremental syncing. The platform distinguishes itself through a containerized architecture that isolates connectors to prevent de

    Synchronizes data between diverse applications, databases, and warehouses using a library of pre-built and custom connectors.

    Pythonbigquerychange-data-capturedata
    GitHub पर देखें↗21,472
  • semi-technologies/weaviatesemi-technologies का अवतार

    semi-technologies/weaviate

    16,337GitHub पर देखें↗

    Weaviate is a cloud-native vector database and distributed vector store designed to save high-dimensional vectors alongside structured data. It functions as a hybrid search engine that combines vector similarity, keyword matching, and structured metadata filtering within a single query. The system is optimized for retrieval-augmented generation, integrating vector search with generative AI and reranking to power question-and-answer workflows. It distinguishes itself through the ability to merge semantic search with traditional keyword queries and structured metadata filters to improve result

    Provides an enterprise-grade platform with role-based access control and multi-tenancy for secure organizational search.

    Go
    GitHub पर देखें↗16,337
  • datahub-project/datahubdatahub-project का अवतार

    datahub-project/datahub

    12,141GitHub पर देखें↗

    DataHub is a metadata management platform designed to unify technical, operational, and business context across diverse data ecosystems. By utilizing a graph-based metadata model and an event-driven ingestion architecture, it creates a centralized source of truth that maps complex data relationships, lineage, and ownership. This foundational framework enables organizations to maintain a synchronized view of their data landscape, supporting both human-led discovery and automated data operations. The platform distinguishes itself through its focus on grounding artificial intelligence and autono

    Ingests and structures metadata from diverse data stores and query histories to create a centralized, searchable knowledge base for organizational data.

    Pythondata-catalogdata-discoverydata-governance
    GitHub पर देखें↗12,141
  • pachyderm/pachydermpachyderm का अवतार

    pachyderm/pachyderm

    6,292GitHub पर देखें↗

    Pachyderm is a containerized, versioned, and lineage-tracked data pipeline platform that runs natively on Kubernetes. It combines a distributed file system backend with immutable data versioning, so every commit to a data repository creates an auditable snapshot, and every pipeline step executes as an isolated container. The platform is defined by a data-centric pipeline model where pipelines are specified by their input and output data repositories rather than explicit task sequences, and provenance is recorded as a directed acyclic graph of commits linking output data to its input sources an

    Provisioning a complete, scalable data pipeline infrastructure on cloud Kubernetes for enterprise-grade data processing workloads.

    Go
    GitHub पर देखें↗6,292
  • fleetdm/fleetfleetdm का अवतार

    fleetdm/fleet

    6,058GitHub पर देखें↗

    Fleet is an open-source device management platform that provides centralized control over computing devices running macOS, Linux, Windows, Chromebooks, iOS, and Android. It enables organizations to enroll devices, collect real-time telemetry, enforce security compliance policies, and manage software remotely from a single system. The platform can be deployed as a single binary, run locally for testing, or scaled horizontally across cloud infrastructure on AWS, Kubernetes, GCP, or Render, with support for high availability through database replication and load balancing. The platform distingui

    Exports data to enterprise platforms like Snowflake, Splunk, GitHub Actions, and Jira for workflow automation.

    Gobinary-authorizationconfiguration-managementdevice-management
    GitHub पर देखें↗6,058
  • infinyon/fluvioinfinyon का अवतार

    infinyon/fluvio

    5,231GitHub पर देखें↗

    Fluvio एक वितरित इवेंट स्ट्रीमिंग प्लेटफ़ॉर्म और क्लाउड-नेटिव स्ट्रीमिंग इंजन है जिसे वितरित क्लस्टर में रीयल-टाइम डेटा स्ट्रीम को एकत्र करने, बनाए रखने और दोहराने के लिए डिज़ाइन किया गया है। यह बाहरी स्रोतों और सिंक के बीच डेटा को इनजेस्ट, समृद्ध और निर्यात करने वाले स्टेटफ़ुल वर्कफ़्लो बनाने के लिए रीयल-टाइम डेटा पाइपलाइन के रूप में कार्य करता है। यह प्लेटफ़ॉर्म इन-लाइन डेटा परिवर्तनों और फ़िल्टरिंग के लिए संकलित मॉड्यूल को निष्पादित करने के लिए WebAssembly के उपयोग से प्रतिष्ठित है। यह क्लस्टर को पुनरारंभ करने की आवश्यकता के बिना जानकारी को फिर से आकार देने के लिए कस्टम व्यावसायिक तर्क के निष्पादन की अनुमति देता है। सिस्टम बाहरी प्रोटोकॉल से कनेक्टर-आधारित डेटा इंजेक्शन, ज़ीरो-कॉपी IO के साथ लॉग-स्ट्रक्चर्ड अपरिवर्तनीय भंडारण और क्षैतिज क्लस्टर स्केलिंग सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह जटिल इवेंट-संचालित पाइपलाइनों के निर्माण का समर्थन करता है जो स्टेटफ़ुल प्रोसेसिंग, विंडो-आधारित एकत्रीकरण और विभाजन-आधारित डेटा वितरण का उपयोग करते हैं। इंजन को एज डेटा प्रोसेसिंग के लिए ARM64 IoT उपकरणों सहित विविध सिस्टम आर्किटेक्चर पर एक हल्के बाइनरी के रूप में तैनात किया जा सकता है।

    Pushes processed information to external databases, object storage, and search engines via outbound connectors.

    Rust
    GitHub पर देखें↗5,231
  • jaypyles/scraperrjaypyles का अवतार

    jaypyles/Scraperr

    4,897GitHub पर देखें↗

    Scraperr एक सेल्फ-होस्टेड वेब स्क्रैपिंग और क्रॉलिंग प्लेटफॉर्म है जिसे XPath सिलेक्टर्स का उपयोग करके वेबसाइटों से स्ट्रक्चर्ड डेटा निकालने के लिए डिज़ाइन किया गया है। यह एक कतार (queue) के माध्यम से स्क्रैपिंग जॉब्स को प्रबंधित करने और आर्टिफिशियल इंटेलिजेंस का उपयोग करके परिणामी कंटेंट का विश्लेषण करने के लिए एक कंटेनराइज्ड सिस्टम के रूप में कार्य करता है। यह प्रोजेक्ट अपने Kubernetes-नेटिव आर्किटेक्चर के माध्यम से खुद को अलग करता है, जो पैकेज मैनेजर के माध्यम से स्केलेबल डिप्लॉयमेंट और प्रबंधन की अनुमति देता है। इसमें लिंक्ड पेजों को खोजने के लिए डोमेन-लेवल स्पाइडरिंग में सक्षम एक क्रॉलिंग इंजन और निकाले गए वेब कंटेंट को क्वेरी करने के लिए आर्टिफिशियल इंटेलिजेंस का उपयोग करने वाला एक डेटा एनालाइजर शामिल है। यह प्लेटफॉर्म ऑटोमेटेड डेटा एक्सट्रैक्शन, बल्क वेब क्रॉलिंग और मीडिया फाइल डाउनलोडिंग सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह स्क्रैप किए गए डेटा को टेबल में विज़ुअलाइज़ करने, ब्राउज़र पहचान की नकल करने के लिए कस्टम रिक्वेस्ट हेडर कॉन्फ़िगर करने और परिणामों को CSV या Markdown प्रारूपों में निर्यात करने के लिए टूल्स प्रदान करता है। एप्लिकेशन Kubernetes डिप्लॉयमेंट कॉन्फ़िगरेशन के माध्यम से अनुकूलन योग्य इंस्टॉलेशन पैरामीटर और वर्जन अपडेट का समर्थन करता है।

    Implements a containerized scraping platform designed specifically for scalable orchestration on Kubernetes.

    TypeScriptdockerhelmkubernetes
    GitHub पर देखें↗4,897
  • briefercloud/brieferbriefercloud का अवतार

    briefercloud/briefer

    4,308GitHub पर देखें↗

    Briefer is an interactive data notebook platform and business intelligence dashboard tool used for collaborative data analysis and reporting. It provides a containerized environment for building reports that combine SQL, Python, and Markdown with native visualizations. The platform features an integrated code assistant that uses large language models to generate SQL and Python snippets from natural language prompts. It is designed as a Kubernetes data application, deploying via Helm charts to manage isolated compute environments and ensure separate resources per page through pod-based isolati

    Deployable via Helm as a Kubernetes-native data platform that manages isolated compute environments.

    TypeScriptanalyticsbibigquery
    GitHub पर देखें↗4,308
  • vdaas/valdvdaas का अवतार

    vdaas/vald

    1,706GitHub पर देखें↗

    Vald is a distributed, cloud-native search engine designed for high-dimensional vector data. It functions as an approximate nearest neighbor search platform, enabling the identification of similar data points across massive datasets through horizontal scaling and distributed indexing. The system is built for container orchestration environments, utilizing custom resource controllers to automate cluster lifecycle management and infrastructure state. It employs graph-based indexing to perform rapid similarity lookups and supports zero-downtime operations by decoupling index construction from qu

    Provides a resilient, cloud-native infrastructure for managing and querying vector indices with zero-downtime updates.

    Goanngapproximate-nearest-neighbor-searchcloud
    GitHub पर देखें↗1,706
  • tduckcloud/tduck-survey-formTDuckCloud का अवतार

    TDuckCloud/tduck-survey-form

    1,224GitHub पर देखें↗

    यह प्रोजेक्ट एंटरप्राइज-ग्रेड डेटा संग्रह, सर्वेक्षण निर्माण और ऑटोमेटेड परिचालन वर्कफ़्लो के लिए डिज़ाइन किया गया एक ओपन-सोर्स, सेल्फ-होस्टेड प्लेटफॉर्म है। यह संगठनों को अपने स्वयं के निजी बुनियादी ढांचे के भीतर इंटरैक्टिव प्रश्नावली, ऑनलाइन परीक्षा और जटिल सूचना-एकत्रीकरण परियोजनाओं का प्रबंधन करते हुए पूर्ण डेटा संप्रभुता बनाए रखने के लिए एक व्यापक वातावरण प्रदान करता है। यह प्लेटफॉर्म उच्च-समवर्ती, स्केलेबल डिप्लॉयमेंट और दानेदार संगठनात्मक नियंत्रण पर ध्यान केंद्रित करके खुद को अलग करता है। इसमें एक डायनामिक ड्रैग-एंड-ड्रॉप बिल्डर है जो तर्क-आधारित ब्रांचिंग और AI-सहायता प्राप्त कंटेंट निर्माण का समर्थन करता है, जिससे ऐसे परिष्कृत फॉर्म बनाना संभव होता है जो उपयोगकर्ता इनपुट के अनुकूल होते हैं। साधारण डेटा प्रविष्टि से परे, सिस्टम बहु-चरणीय अनुमोदन वर्कफ़्लो, सुरक्षित भुगतान प्रसंस्करण और ऑटोमेटेड दस्तावेज़ निर्माण जैसी उन्नत व्यावसायिक क्षमताओं को इंटीग्रेट करता है, जो सभी एक केंद्रीकृत, भूमिका-आधारित एक्सेस कंट्रोल सिस्टम के माध्यम से प्रबंधित होते हैं। सिस्टम रीयल-टाइम डेटा प्रबंधन, बहु-प्लेटफ़ॉर्म वितरण और प्रतिभा मूल्यांकन और संसाधन शेड्यूलिंग के लिए विशेष टूल सहित एक व्यापक क्षमता सतह को कवर करता है। यह विविध स्टोरेज बैकएंड का समर्थन करता है और मजबूत सबमिशन ट्रैकिंग प्रदान करता है, यह सुनिश्चित करता है कि एकत्र की गई जानकारी सुरक्षित और विश्लेषण के लिए सुलभ रहे। एप्लिकेशन को कंटेनरीकृत डिप्लॉयमेंट के लिए डिज़ाइन किया गया है, जो निजी क्लाउड या ऑन-प्रिमाइसेस वातावरण के लिए सेटअप प्रक्रिया को सरल बनाता है।

    Provides an enterprise-grade system for managing form submissions, approvals, and reporting within private infrastructure.

    Javaquestionnairesurveysurvey-form
    GitHub पर देखें↗1,224
  • rheosoph/flow-likeRheosoph का अवतार

    Rheosoph/flow-like

    899GitHub पर देखें↗

    Flow-like is a workflow orchestration engine designed for building and executing strictly typed automated processes. It provides a secure, sandboxed runtime environment that supports the integration of local artificial intelligence models, allowing for the processing of data entirely on host hardware without reliance on external cloud services. The platform distinguishes itself through its event-sourced execution tracing, which records every state change and data movement to enable full auditability and the replay of past processes. It combines this with a hybrid storage system that integrate

    Manages large-scale information processing with built-in audit logging and execution tracing for complex logic flows.

    Rustagentsaiapis
    GitHub पर देखें↗899
  1. Home
  2. Data & Databases
  3. Enterprise Data Services
  4. Enterprise Data Platforms

सब-टैग एक्सप्लोर करें

  • Data Collection SystemsBusiness-grade platforms for managing high-volume form submissions and automated operational workflows. **Distinct from Enterprise Data Platforms:** Distinct from general data platforms: focuses on the application-level management of form-based data collection and approval workflows.
  • Data Export ConnectorsMechanisms for exporting data to enterprise platforms like Snowflake, Splunk, and Jira for workflow automation. **Distinct from Enterprise Data Platforms:** Distinct from Enterprise Data Platforms: focuses on outbound data export connectors rather than the platforms themselves.
  • Kubernetes-Native Data PlatformsEnterprise-grade data pipeline platforms deployed on Kubernetes for scalable, production data processing. **Distinct from Enterprise Data Platforms:** Distinct from Enterprise Data Platforms: focuses on Kubernetes-native orchestration and pipeline automation rather than general database management.
  • Metadata Knowledge BasesIngests and structures metadata from diverse data stores and query histories into a centralized knowledge base. **Distinct from Enterprise Data Platforms:** Focuses on the creation of a searchable knowledge base from metadata rather than the platform infrastructure itself.