awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

4 रिपॉजिटरी

Awesome GitHub RepositoriesKubernetes-Native Data Platforms

Enterprise-grade data pipeline platforms deployed on Kubernetes for scalable, production data processing.

Distinct from Enterprise Data Platforms: Distinct from Enterprise Data Platforms: focuses on Kubernetes-native orchestration and pipeline automation rather than general database management.

Explore 4 awesome GitHub repositories matching data & databases · Kubernetes-Native Data Platforms. Refine with filters or upvote what's useful.

Awesome Kubernetes-Native Data Platforms GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • pachyderm/pachydermpachyderm का अवतार

    pachyderm/pachyderm

    6,292GitHub पर देखें↗

    Pachyderm is a containerized, versioned, and lineage-tracked data pipeline platform that runs natively on Kubernetes. It combines a distributed file system backend with immutable data versioning, so every commit to a data repository creates an auditable snapshot, and every pipeline step executes as an isolated container. The platform is defined by a data-centric pipeline model where pipelines are specified by their input and output data repositories rather than explicit task sequences, and provenance is recorded as a directed acyclic graph of commits linking output data to its input sources an

    Provisioning a complete, scalable data pipeline infrastructure on cloud Kubernetes for enterprise-grade data processing workloads.

    Go
    GitHub पर देखें↗6,292
  • jaypyles/scraperrjaypyles का अवतार

    jaypyles/Scraperr

    4,897GitHub पर देखें↗

    Scraperr एक सेल्फ-होस्टेड वेब स्क्रैपिंग और क्रॉलिंग प्लेटफॉर्म है जिसे XPath सिलेक्टर्स का उपयोग करके वेबसाइटों से स्ट्रक्चर्ड डेटा निकालने के लिए डिज़ाइन किया गया है। यह एक कतार (queue) के माध्यम से स्क्रैपिंग जॉब्स को प्रबंधित करने और आर्टिफिशियल इंटेलिजेंस का उपयोग करके परिणामी कंटेंट का विश्लेषण करने के लिए एक कंटेनराइज्ड सिस्टम के रूप में कार्य करता है। यह प्रोजेक्ट अपने Kubernetes-नेटिव आर्किटेक्चर के माध्यम से खुद को अलग करता है, जो पैकेज मैनेजर के माध्यम से स्केलेबल डिप्लॉयमेंट और प्रबंधन की अनुमति देता है। इसमें लिंक्ड पेजों को खोजने के लिए डोमेन-लेवल स्पाइडरिंग में सक्षम एक क्रॉलिंग इंजन और निकाले गए वेब कंटेंट को क्वेरी करने के लिए आर्टिफिशियल इंटेलिजेंस का उपयोग करने वाला एक डेटा एनालाइजर शामिल है। यह प्लेटफॉर्म ऑटोमेटेड डेटा एक्सट्रैक्शन, बल्क वेब क्रॉलिंग और मीडिया फाइल डाउनलोडिंग सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह स्क्रैप किए गए डेटा को टेबल में विज़ुअलाइज़ करने, ब्राउज़र पहचान की नकल करने के लिए कस्टम रिक्वेस्ट हेडर कॉन्फ़िगर करने और परिणामों को CSV या Markdown प्रारूपों में निर्यात करने के लिए टूल्स प्रदान करता है। एप्लिकेशन Kubernetes डिप्लॉयमेंट कॉन्फ़िगरेशन के माध्यम से अनुकूलन योग्य इंस्टॉलेशन पैरामीटर और वर्जन अपडेट का समर्थन करता है।

    Implements a containerized scraping platform designed specifically for scalable orchestration on Kubernetes.

    TypeScriptdockerhelmkubernetes
    GitHub पर देखें↗4,897
  • briefercloud/brieferbriefercloud का अवतार

    briefercloud/briefer

    4,308GitHub पर देखें↗

    Briefer is an interactive data notebook platform and business intelligence dashboard tool used for collaborative data analysis and reporting. It provides a containerized environment for building reports that combine SQL, Python, and Markdown with native visualizations. The platform features an integrated code assistant that uses large language models to generate SQL and Python snippets from natural language prompts. It is designed as a Kubernetes data application, deploying via Helm charts to manage isolated compute environments and ensure separate resources per page through pod-based isolati

    Deployable via Helm as a Kubernetes-native data platform that manages isolated compute environments.

    TypeScriptanalyticsbibigquery
    GitHub पर देखें↗4,308
  • vdaas/valdvdaas का अवतार

    vdaas/vald

    1,706GitHub पर देखें↗

    Vald is a distributed, cloud-native search engine designed for high-dimensional vector data. It functions as an approximate nearest neighbor search platform, enabling the identification of similar data points across massive datasets through horizontal scaling and distributed indexing. The system is built for container orchestration environments, utilizing custom resource controllers to automate cluster lifecycle management and infrastructure state. It employs graph-based indexing to perform rapid similarity lookups and supports zero-downtime operations by decoupling index construction from qu

    Provides a resilient, cloud-native infrastructure for managing and querying vector indices with zero-downtime updates.

    Goanngapproximate-nearest-neighbor-searchcloud
    GitHub पर देखें↗1,706
  1. Home
  2. Data & Databases
  3. Enterprise Data Services
  4. Enterprise Data Platforms
  5. Kubernetes-Native Data Platforms