awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

15 Repos

Awesome GitHub RepositoriesCross-Modal Representations

Techniques for mapping different data types into a shared latent space.

Distinguishing note: Focuses on bridging linguistic and sonic domains via shared latent spaces.

Explore 15 awesome GitHub repositories matching artificial intelligence & ml · Cross-Modal Representations. Refine with filters or upvote what's useful.

Awesome Cross-Modal Representations GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • suno-ai/barkAvatar von suno-ai

    suno-ai/bark

    39,159Auf GitHub ansehen↗

    Bark is a generative audio engine and machine learning inference library designed to convert written text into high-fidelity speech and sound effects. It functions as a text-to-audio transformer, utilizing multi-stage neural network architectures to map semantic input tokens into detailed audio codebooks for synthesis. The system distinguishes itself through a hierarchical transformer stacking approach that separates semantic understanding from acoustic realization. By employing autoregressive token prediction and vector quantized codebook mapping, the engine bridges linguistic and sonic doma

    Projects text and audio data into a shared mathematical space to bridge linguistic and sonic domains.

    Jupyter Notebook
    Auf GitHub ansehen↗39,159
  • microsoft/visual-chatgptAvatar von microsoft

    microsoft/visual-chatgpt

    34,079Auf GitHub ansehen↗

    Visual-ChatGPT is a visual orchestration framework and multimodal AI pipeline designed to coordinate large language models with visual foundation models. It functions as an integration layer that enables the exchange of text and images between different AI models to automate image analysis and editing tasks without requiring additional model training. The system differentiates itself through model-chain orchestration and prompt-based task dispatching, allowing natural language instructions to trigger specific vision models or tools. It utilizes coordinate-based region mapping and iterative ma

    Exchanges image embeddings and text tokens between models to refine visual analysis and generation.

    Python
    Auf GitHub ansehen↗34,079
  • openbmb/minicpm-oAvatar von OpenBMB

    OpenBMB/MiniCPM-o

    23,850Auf GitHub ansehen↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Maps disparate text, image, and audio inputs into a shared vector representation for cross-modal interaction.

    Pythonminicpmminicpm-vmulti-modal
    Auf GitHub ansehen↗23,850
  • deepseek-ai/deepseek-coderAvatar von deepseek-ai

    deepseek-ai/DeepSeek-Coder

    22,804Auf GitHub ansehen↗

    DeepSeek-Coder is a large language model and foundational neural network architecture designed specifically for software development tasks. It functions as an artificial intelligence assistant capable of interpreting complex programming instructions to generate, transpile, and structure source code. The system distinguishes itself through its ability to perform project-level code generation, analyzing broader context and patterns across entire software projects rather than isolated files. It supports multimodal input processing, allowing for the integration of text and visual data to inform i

    Projects visual and textual inputs into a shared vector space to enable unified reasoning.

    Python
    Auf GitHub ansehen↗22,804
  • microsoft/unilmAvatar von microsoft

    microsoft/unilm

    22,030Auf GitHub ansehen↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Applies iterative word alignment and contrastive loss functions during training to synchronize semantic representations across different languages.

    Pythonbeitbeit-3bitnet
    Auf GitHub ansehen↗22,030
  • huggingface/sentence-transformersAvatar von huggingface

    huggingface/sentence-transformers

    18,817Auf GitHub ansehen↗

    This project is a transformer-based framework for generating dense and sparse vector embeddings of text and multimodal data. It serves as a library for fine-tuning models to perform semantic similarity tasks, retrieval, and reranking. The system is distinguished by its support for diverse architectural patterns, including bi-encoders for fast similarity search and cross-encoders for high-precision reranking. It provides dedicated pipelines for multimodal embeddings, mapping text and images into a shared vector space, and implements knowledge distillation to compress large models into smaller,

    Maps text and images into a shared embedding space to enable cross-modal retrieval.

    Python
    Auf GitHub ansehen↗18,817
  • qwenlm/qwen3-vlAvatar von QwenLM

    QwenLM/Qwen3-VL

    18,329Auf GitHub ansehen↗

    Qwen3-VL is a multimodal vision-language model designed to process and reason across images, videos, and text. It functions as a computer vision framework capable of identifying objects, extracting structured data from documents, and interpreting spatial elements within visual media. The system operates as an automated user interface interaction agent, interpreting screen data to navigate software and mobile applications. By utilizing a unified transformer architecture, it performs complex visual reasoning to execute user-defined tasks without manual input. Beyond interface navigation, the m

    Maps visual and textual data into shared vector spaces to enable unified cross-modal reasoning.

    Jupyter Notebook
    Auf GitHub ansehen↗18,329
  • wan-video/wan2.1Avatar von Wan-Video

    Wan-Video/Wan2.1

    15,350Auf GitHub ansehen↗

    Wan2.1 is a generative video synthesis framework that provides foundation models for creating high-fidelity video sequences and static images from descriptive text prompts. The system utilizes a unified architecture trained on both static and dynamic datasets, allowing it to function as a comprehensive tool for visual media creation. The framework distinguishes itself through a transformer-based temporal modeling approach that ensures structural coherence and consistent motion across video frames. It supports multi-resolution latent scaling, enabling the generation of content in various aspec

    Maps text-based semantic inputs into shared latent spaces to guide the generative synthesis process.

    Pythonaigcvideogeneration
    Auf GitHub ansehen↗15,350
  • hanxiao/bert-as-serviceAvatar von hanxiao

    hanxiao/bert-as-service

    12,831Auf GitHub ansehen↗

    Dieses Projekt ist ein BERT-Einbettungsdienst mit hoher Leistung und ein Inferenzserver, der darauf ausgelegt ist, Textsequenzen in numerische Vektoren fester Länge abzubilden. Es fungiert als Microservice für maschinelles Lernen und verteilter Modellserver, der die Anforderungsbehandlung von rechenintensiven Aufgaben entkoppelt. Das System nutzt eine ZeroMQ-Messaging-Infrastruktur, um eine Kommunikation mit geringer Latenz zwischen verteilten Clients und dem Inferenzserver bereitzustellen. Es integriert serverseitige Batch-Verarbeitung und GPU-Workload-Skalierung, um die Hardwareauslastung zu maximieren und hohe Anforderungsvolumina zu verwalten. Die Plattform unterstützt die Infrastruktur für semantische Suche durch die Generierung modalübergreifender Einbettungen für Text und Bilder innerhalb eines gemeinsamen Vektorraums. Dies ermöglicht modalübergreifende Suche, Relevanz-Ranking von Inhalten und das Re-Ranking von Ergebnissen basierend auf der semantischen Ausrichtung zwischen visuellem Inhalt und Textbeschreibungen. Der Dienst kann als elastischer Microservice bereitgestellt werden, der über gRPC-, HTTP- oder WebSocket-Protokolle zugänglich ist, und bietet nicht-blockierendes Duplex-Streaming für die Handhabung großer Datensätze.

    Transforms images and text into a shared latent space to enable cross-modal vectorization.

    Python
    Auf GitHub ansehen↗12,831
  • salesforce/lavisAvatar von salesforce

    salesforce/LAVIS

    11,236Auf GitHub ansehen↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    Converts visual and textual data into unified representations for use in classification and similarity tasks.

    Jupyter Notebook
    Auf GitHub ansehen↗11,236
  • facebookresearch/dinov3Avatar von facebookresearch

    facebookresearch/dinov3

    9,613Auf GitHub ansehen↗

    This project is a self-supervised vision foundation model based on a vision transformer architecture. It is designed to learn dense visual representations from unlabeled images, serving as a general-purpose backbone for a wide variety of downstream vision tasks. The system is distinguished by its use of self-distillation and masked image modeling to extract semantic and geometric features. It also incorporates an image-text alignment model that maps visual embeddings to textual descriptions, enabling zero-shot image recognition, zero-shot segmentation, and cross-modal retrieval. The project

    Maps visual embeddings to textual descriptions to support cross-modal retrieval and zero-shot vision tasks.

    Jupyter Notebook
    Auf GitHub ansehen↗9,613
  • apple/ml-ferretAvatar von apple

    apple/ml-ferret

    8,680Auf GitHub ansehen↗

    ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts

    Aligns visual region embeddings with linguistic tokens to enable natural language pointing to image objects.

    Python
    Auf GitHub ansehen↗8,680
  • ucas-haoranwei/got-ocr2.0Avatar von Ucas-HaoranWei

    Ucas-HaoranWei/GOT-OCR2.0

    8,141Auf GitHub ansehen↗

    GOT-OCR2.0 is an end-to-end optical character recognition system and document text extractor. It utilizes a unified transformer architecture to recognize and extract plain and formatted text from diverse images and documents. The system features a multi-crop processing method that divides high-resolution or dense documents into smaller sections to maintain recognition detail. It also includes a renderer that transforms recognized text into HTML to preserve the original structure and layout of the document. The project provides a framework for fine-tuning pre-trained models on custom datasets

    Maps image features and text tokens into a shared latent space to correlate visual structure with linguistic meaning.

    Python
    Auf GitHub ansehen↗8,141
  • microsoft/muzicAvatar von microsoft

    microsoft/muzic

    4,928Auf GitHub ansehen↗

    Muzic ist eine Deep-Learning-Plattform und ein Framework für KI-gestützte Musikanalyse, Komposition und Synthese. Es fungiert als Musikgenerierungs-Framework und Analysetool, das große Sprachmodelle und autonome Agenten nutzt, um die Erstellung und Interpretation symbolischer und auditiver Musik zu orchestrieren. Das Projekt zeichnet sich durch seine cross-modale Fähigkeiten aus, bei denen natürliche Sprache und symbolische Musik in einen gemeinsamen Embedding-Raum für Zero-Shot-Klassifizierung und Informationsabruf abgebildet werden. Es verwendet eine Vielzahl spezialisierter Architekturen, einschließlich Diffusions-Frameworks für die Audiosynthese, Dual-Grain-Aufmerksamkeitsmechanismen für strukturelle Konsistenz bei langen Sequenzen und ein hybrides System, das musiktheoretische Regeln mit neuronalen Netzwerken kombiniert. Die Plattform deckt ein breites Spektrum an Funktionen ab, einschließlich der Generierung von MIDI-Sequenzen aus Text und Liedtexten, neuronaler Gesangssynthese und automatisierter Liedtext-Transkription. Sie bietet zudem Tools für die Modellierung von Musikstrukturen, attributbasierte symbolische Generierung und die Orchestrierung externer Musiktools über autonome Agenten. Unterstützende Dienstprogramme umfassen Data-Engineering-Pipelines für die MIDI-Binarisierung im großen Maßstab, Datensatz-Kodierung und Audiosignalverarbeitung für die Extraktion von Melodienoten und die Ausrichtung von Sprache zu Phonemen.

    Maps symbolic music and natural language into a shared joint embedding space via contrastive learning.

    Pythonai-musicdeep-learningmusic
    Auf GitHub ansehen↗4,928
  • evolvinglmms-lab/otterAvatar von EvolvingLMMs-Lab

    EvolvingLMMs-Lab/Otter

    3,331Auf GitHub ansehen↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Maps visual encoder embeddings into the textual token space using a learned projection layer for unified multimodal processing.

    Pythonartificial-inteligencechatgptdeep-learning
    Auf GitHub ansehen↗3,331
  1. Home
  2. Artificial Intelligence & ML
  3. Cross-Modal Representations

Unter-Tags erkunden

  • Visual-Textual AlignmentsMechanisms for projecting visual and textual data into shared vector spaces for unified reasoning. **Distinct from Cross-Modal Representations:** Distinct from Cross-Modal Representations: focuses specifically on the vision-text alignment for code-aware models.