awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

8 个仓库

Awesome GitHub RepositoriesMultimodal Perception Models

Models designed to interpret and analyze visual data, charts, or cross-modal inputs alongside text.

Explore 8 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Perception Models. Refine with filters or upvote what's useful.

Awesome Multimodal Perception Models GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • abi/screenshot-to-codeabi 的头像

    abi/screenshot-to-code

    72,926在 GitHub 上查看↗

    This project is an artificial intelligence-powered frontend generator that translates visual design inputs into functional source code. It functions as a workflow engine that interprets graphical user interfaces, mapping layout structures and styling rules to structured markup and programming language syntax. The tool distinguishes itself by supporting both static design mockups and dynamic video recordings. It processes temporal and spatial information from screen captures to reconstruct interaction flows and state transitions, enabling the creation of functional software prototypes from vis

    Processes visual design inputs through neural networks to interpret layout structures and translate them into functional source code.

    Python
    在 GitHub 上查看↗72,926
  • getomni-ai/zeroxgetomni-ai 的头像

    getomni-ai/zerox

    12,241在 GitHub 上查看↗

    Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting. The system differentiates itself through the use of coordinate-based element mapping and multimodal layout analysis to identify structural elements like tables, charts, and headers. It utilizes rasterization to convert vector PDF pages into high-resolution bitmaps, ensuring consistent input for t

    Identifies structural elements like tables and headers by processing document images through a large multimodal model.

    TypeScriptocrpdf
    在 GitHub 上查看↗12,241
  • yichuan-w/leannyichuan-w 的头像

    yichuan-w/LEANN

    11,985在 GitHub 上查看↗

    LEANN is a framework for local retrieval augmented generation and vector indexing. It functions as a system for building local knowledge bases and source code search engines that combine large language models with retrieved private data to generate context-aware responses. The project distinguishes itself through a vision-model based document layout extractor for parsing complex PDF figures and diagrams, and a source code search engine that employs structure-aware chunking to preserve function and class boundaries. It also implements the Model Context Protocol to integrate real-time data sour

    Employs multimodal vision models to interpret complex PDF layouts and diagrams for structural text extraction.

    Pythonaifaissgpt-oss
    在 GitHub 上查看↗11,985
  • nielsrogge/transformers-tutorialsNielsRogge 的头像

    NielsRogge/Transformers-Tutorials

    11,641在 GitHub 上查看↗

    This is a collection of tutorials and practical demonstrations for implementing machine learning tasks using the HuggingFace Transformers library. It serves as a guide for applying transformer architectures across computer vision, natural language processing, and audio analysis. The repository provides implementation examples for multimodal model deployment, including the combination of text, image, and audio inputs. It includes resources for optimizing pre-trained models through fine-tuning on custom datasets and provides examples for preparing PyTorch datasets by converting raw files into t

    Provides implementation examples for deploying models that combine text, image, and audio inputs.

    Jupyter Notebookbertgpt-2layoutlm
    在 GitHub 上查看↗11,641
  • lostruins/koboldcppLostRuins 的头像

    LostRuins/koboldcpp

    9,511在 GitHub 上查看↗

    KoboldCPP is a local large language model inference engine and GGUF model runner designed to execute quantized models on personal hardware. It functions as a multimodal AI server and API gateway, providing OpenAI-compatible endpoints that allow third-party clients to interact with locally hosted models. The project distinguishes itself as an AI storytelling backend, featuring dedicated tools for long-form narrative management through persistent memory, world lore tracking, and character state management. It further extends its capabilities as a multimodal server capable of processing text, im

    Loads multimodal projectors to process and react to image or audio inputs.

    C++gemmaggmlgguf
    在 GitHub 上查看↗9,511
  • paddlepaddle/erniePaddlePaddle 的头像

    PaddlePaddle/ERNIE

    7,717在 GitHub 上查看↗

    ERNIE is a development toolkit for training, fine-tuning, and deploying large language models built on the PaddlePaddle deep learning platform. It provides a comprehensive suite of core components, including an inference server for vision and language models, a training and fine-tuning toolkit, and a framework for building retrieval-augmented generation systems using private knowledge bases. The project features multimodal AI models capable of reasoning across text, images, and video to perform complex visual understanding and information extraction. It distinguishes itself through specialize

    Ships models designed to interpret and analyze visual data, charts, and cross-modal inputs alongside text.

    Pythonernieernie-45ernie-45-vl
    在 GitHub 上查看↗7,717
  • huggingface/smol-coursehuggingface 的头像

    huggingface/smol-course

    6,661在 GitHub 上查看↗

    This project is an educational program focused on the alignment of small language models. It provides a technical curriculum and a series of courses designed to teach how to align models with human preferences and behaviors. The material covers the implementation of preference optimization algorithms and the adaptation of vision-language models to process both text and image data simultaneously. It also includes instructional guides on synthetic data generation to improve model performance in specialized domains. The curriculum encompasses supervised fine-tuning workflows, the use of chat te

    Configures multimodal vision models to interpret visual inputs alongside text for complex tasks.

    Jupyter Notebook
    在 GitHub 上查看↗6,661
  • google-research/big_visiongoogle-research 的头像

    google-research/big_vision

    3,363在 GitHub 上查看↗

    This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces. The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide

    Performs complex visual perception tasks by integrating normalizing flow encoders within transformer architectures.

    Jupyter Notebook
    在 GitHub 上查看↗3,363
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Architectures
  5. Multimodal Perception Models

探索子标签

  • Multimodal Vision Models1 个子标签Neural networks capable of processing and interpreting visual inputs alongside other data modalities.