awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

3 Repos

Awesome GitHub RepositoriesDatabricks Connectors

Integrations for retrieving unstructured files from data volumes.

Distinct from Data Ingestion: Focuses on Databricks-specific file ingestion, distinct from general data ingestion.

Explore 3 awesome GitHub repositories matching data & databases · Databricks Connectors. Refine with filters or upvote what's useful.

Awesome Databricks Connectors GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • vectordotdev/vectorAvatar von vectordotdev

    vectordotdev/vector

    22,071Auf GitHub ansehen↗

    Vector is a high-performance observability data pipeline designed to collect, transform, and route logs, metrics, and traces across distributed infrastructure. It functions as a modular engine that decouples data ingestion from processing and transmission, utilizing a component-based architecture to connect diverse sources to multiple destinations. The project distinguishes itself through a focus on reliability and flow control. It implements backpressure-aware data movement to prevent data loss during traffic spikes and utilizes disk-backed event buffering to ensure durability during network

    Streams observability data into catalog tables with automatic schema discovery.

    Rusteventsforwarderhacktoberfest
    Auf GitHub ansehen↗22,071
  • prefecthq/prefectAvatar von PrefectHQ

    PrefectHQ/prefect

    21,640Auf GitHub ansehen↗

    Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as Python code. It functions as a container-native engine that wraps individual tasks in isolated environments, ensuring consistent dependencies and resource allocation across diverse infrastructure. By utilizing a state-machine-based orchestration model, the system tracks execution progress through discrete transitions and persistent event logs to maintain reliable and observable task processing. The platform distinguishes itself through a decoupled worker-API architecture, which sep

    Manages secure connections to data environments using personal access tokens or service principal credentials.

    Pythonautomationdatadata-engineering
    Auf GitHub ansehen↗21,640
  • unstructured-io/unstructuredAvatar von Unstructured-IO

    Unstructured-IO/unstructured

    14,019Auf GitHub ansehen↗

    Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t

    Connects to data volumes to retrieve and process unstructured files for AI applications.

    HTMLdata-pipelinesdeep-learningdocument-image-analysis
    Auf GitHub ansehen↗14,019
  1. Home
  2. Data & Databases
  3. Data Ingestion
  4. Databricks Connectors

Unter-Tags erkunden

  • Databricks Authentication ManagersTools for managing secure connections to Databricks environments using tokens or service principals. **Distinct from Databricks Connectors:** Distinct from Databricks Connectors: focuses on the authentication layer rather than data ingestion.
  • Databricks ExportersIntegrations for writing processed document data into Databricks volumes. **Distinct from Databricks Connectors:** Distinct from Databricks Connectors: focuses on data export/egress rather than ingestion.