10 रिपॉजिटरी
Architectures that use specialized encoders for different modalities before merging them through a central combiner.
Distinct from Encoder-Decoder Architectures: Distinct from standard encoder-decoders: focuses on multi-input fusion via a central combiner rather than sequence-to-sequence mapping.
Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Encoder-Combiner Architectures. Refine with filters or upvote what's useful.
Ludwig is a multimodal machine learning platform and low-code framework designed for building, training, and deploying neural networks. It enables the construction of models that process text, images, audio, and tabular data through a unified interface using declarative configuration files rather than custom code. The system features a specialized low-code framework for large language models, supporting supervised fine-tuning, preference alignment, and a constrained decoding tool to force structured data output via logit extraction. It also includes an automated model architecture search to i
Processes diverse data types using specialized feature extractors and a central fusion layer for multimodal predictions.
Ludwig is a declarative machine learning framework designed for training neural networks and large language models using configuration files instead of manual coding. It functions as a multimodal model builder and a low-code tool for supervised fine-tuning, allowing users to build models that process mixed inputs of text, images, audio, and tabular data. The project distinguishes itself through an automated hyperparameter optimizer and a system for large language model fine-tuning using parameter-efficient adapters. It features a multimodal data pipeline and the ability to automatically gener
Utilizes an architecture that separates diverse input processing into specialized encoders and merges them via a central combiner.
Cosmos is an open platform of world models, datasets, and tools for building physical AI systems such as robots and autonomous vehicles. It provides video generation and video understanding models that can generate synthetic videos and world simulations from text, image, video, or action inputs, and analyze videos to produce captions, event timestamps, spatial bounding boxes, and next-action predictions. The platform includes a world simulation generator that produces images, videos, synchronized audio, and action-conditioned rollouts for synthetic data, alongside a visual content analyzer th
Combines text, image, video, and action inputs into a unified latent space using cross-attention layers for flexible conditioning.
AutoGluon is an automated machine learning framework designed to optimize model selection and hyperparameter tuning across tabular, text, image, and time series data. It functions as an ensemble learning library and a tabular data prediction engine, aiming to build high-accuracy predictive models without manual algorithm selection. The framework integrates multimodal machine learning pipelines that combine disparate data types into a single representation using specialized encoders. It also includes a probabilistic time series forecaster that fits multiple statistical and deep learning models
Employs encoder-combiner architectures to map text, images, and tables into a single fused representation.
ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr
Uses dedicated modality-specific encoders to process raw inputs before merging them into a common space.
DUSt3R is a geometric vision transformer model that predicts dense 3D pointmaps directly from one or more uncalibrated images, without requiring prior camera intrinsics, extrinsics, or known camera positions. Its core identity is an end-to-end approach to 3D reconstruction that bypasses traditional depth estimation and camera calibration pipelines, instead outputting metric-scale 3D coordinates from RGB inputs. The model processes image pairs through a shared dual-image encoder architecture, using cross-attention feature fusion in the decoder to merge features from two images into a unified p
Uses cross-attention layers in the decoder to merge features from two images into a unified pointmap.
System-bus-radio is a software-defined radio transmitter that generates AM radio signals by modulating the electromagnetic emissions from a computer's processor and memory bus, without requiring any dedicated radio hardware or physical antennas. It functions as a CPU electromagnetic emissions tool and processor-based signal generator, enabling radio transmission through precise control of CPU instructions and memory bus operations. The project encodes musical notes as sequences of frequency and duration pairs, then synthesizes the AM radio waveform in real-time by executing a tight loop of CP
Encodes musical notes as frequency-duration pairs for playback via the modulated carrier.
This project is a high-performance C++ and CUDA neural network library designed for fast training and inference of small networks on NVIDIA GPUs. It serves as a specialized backend for neural radiance fields and coordinate-based networks, providing a fused GPU kernel library and a hash grid encoder for transforming raw input dimensions into high-dimensional representations. The library distinguishes itself through the use of C++ template metaprogramming and fused-kernel execution, which merge neural network layers into single GPU device functions to eliminate memory bottlenecks. It leverages
Encodes each input dimension into a set of bins using a quartic kernel for accurate fitting with limited dynamic range.
VisualGLM-6B is a multimodal large language model and vision-language system designed to process and generate text based on combined textual and visual inputs. It functions as a bilingual conversational AI capable of maintaining natural language interactions in both English and Chinese. The project utilizes quantized model weights to reduce memory requirements, enabling the deployment of the neural network on consumer-grade hardware. These compressed parameters allow for lower VRAM usage while maintaining the model's ability to analyze visual content and generate corresponding natural languag
Combines a visual encoder and language model using a projection layer to align image features with text embeddings.
Open Flamingo एक मल्टीमॉडल लार्ज लैंग्वेज मॉडल ट्रेनिंग फ्रेमवर्क है जिसे प्रीट्रेन्ड विजन एनकोडर को लैंग्वेज मॉडल्स के साथ एकीकृत करने के लिए डिज़ाइन किया गया है। यह एक विजन-लैंग्वेज आर्किटेक्चर को लागू करता है जो छवियों और टेक्स्ट के इंटरलीव्ड अनुक्रमों को प्रोसेस करने के लिए क्रॉस-अटेंशन लेयर्स का उपयोग करता है। सिस्टम अपनी फ्यू-शॉट मल्टीमॉडल लर्निंग क्षमताओं द्वारा विशेषता है, जो मॉडल को प्रॉम्प्ट में प्रदान किए गए इमेज-टेक्स्ट उदाहरणों के एक छोटे सेट का उपयोग करके नए विज़ुअल कार्यों के अनुकूल होने की अनुमति देता है। यह विज़ुअल क्वेश्चन आंसरिंग और कैप्शनिंग जैसे कार्यों के लिए इन-कॉन्टेक्स्ट लर्निंग और मल्टीमॉडल टेक्स्ट जनरेशन का समर्थन करता है। फ्रेमवर्क में एक डिस्ट्रीब्यूटेड मॉडल ट्रेनर शामिल है जो कई GPU में मेमोरी ऑप्टिमाइज़ेशन के लिए डेटा पैरेललिज़्म और ग्रेडिएंट चेकपॉइंटिंग का उपयोग करता है। यह शार्ड मल्टीमॉडल डेटासेट लोडिंग, पैरेललाइज़्ड मॉडल मूल्यांकन, और इन्फरेंस के लिए बड़े पैमाने पर मॉडल्स को होस्ट करने के लिए इंफ्रास्ट्रक्चर भी प्रदान करता है।
Uses specialized encoders for different modalities and merges them through a central combiner to create a unified architecture.