awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to microsoft/mm-react

Open-source alternatives to MM REACT

30 open-source projects similar to microsoft/mm-react, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best MM REACT alternative.

  • aigc-audio/audiogptAIGC-Audio 的头像

    AIGC-Audio/AudioGPT

    10,174在 GitHub 上查看↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Pythonaudiogptmusic
    在 GitHub 上查看↗10,174
  • microsoft/jarvismicrosoft 的头像

    microsoft/JARVIS

    24,854在 GitHub 上查看↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Python
    在 GitHub 上查看↗24,854
  • baaivision/emubaaivision 的头像

    baaivision/Emu

    1,775在 GitHub 上查看↗

    Emu Series: Generative Multimodal Models from BAAI

    Python
    在 GitHub 上查看↗1,775
  • cvlab-columbia/vipercvlab-columbia 的头像

    cvlab-columbia/viper

    1,717在 GitHub 上查看↗

    Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"

    Jupyter Notebook
    在 GitHub 上查看↗1,717

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • simular-ai/agent-ssimular-ai 的头像

    simular-ai/Agent-S

    11,855在 GitHub 上查看↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    在 GitHub 上查看↗11,855
  • othersideai/self-operating-computerOthersideAI 的头像

    OthersideAI/self-operating-computer

    10,153在 GitHub 上查看↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    在 GitHub 上查看↗10,153
  • 11cafe/jaaz11cafe 的头像

    11cafe/jaaz

    6,384在 GitHub 上查看↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    在 GitHub 上查看↗6,384
  • baptistearno/typebot.iobaptisteArno 的头像

    baptisteArno/typebot.io

    10,042在 GitHub 上查看↗

    Typebot is a visual chatbot builder and conversational platform designed for lead generation and data collection. It provides a drag-and-drop workflow designer that converts visual nodes into structured conversation logic, allowing users to build interactive forms and chatbots with conditional routing. The platform is designed as a self-hosted conversational infrastructure, enabling the deployment of the entire application stack on private servers using Docker and PostgreSQL. This allows for complete control over data storage and server maintenance. The system integrates with external servic

    TypeScriptchat-applicationchatbotconversational-bots
    在 GitHub 上查看↗10,042
  • qwenlm/qwen3-omniQwenLM 的头像

    QwenLM/Qwen3-Omni

    3,843在 GitHub 上查看↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    在 GitHub 上查看↗3,843
  • sfyc23/everydaywechatsfyc23 的头像

    sfyc23/EverydayWechat

    10,296在 GitHub 上查看↗

    EverydayWechat is a WeChat automation bot designed to automate messages, scheduled reminders, and automatic replies within the WeChat messaging ecosystem. It functions as a multi-purpose system that combines the roles of a scheduled message sender, an auto-reply bot, and a chatbot assistant. The project enables the delivery of customized recurring messages to specific users and group chats on a fixed timetable. It also provides automated individual replies based on preconfigured rules and group chat assistance that fetches real-time data for weather, logistics, and calendars. The system inco

    Pythonaiautoreplybot
    在 GitHub 上查看↗10,296
  • alirezadir/machine-learning-interviewsalirezadir 的头像

    alirezadir/Machine-Learning-Interviews

    8,455在 GitHub 上查看↗

    This project is a comprehensive machine learning interview guide and technical study resource designed for individuals preparing for machine learning and AI engineering roles. It provides a collection of materials and practice problems covering core algorithms, theoretical fundamentals, and the implementation of neural network architectures. The resource serves as a technical reference for generative AI development, focusing on the design and optimization of large language models and diffusion systems. It includes frameworks for system design, covering the architecture of production machine l

    Jupyter Notebookagenticaiai-agents
    在 GitHub 上查看↗8,455
  • facebookresearch/fairseqfacebookresearch 的头像

    facebookresearch/fairseq

    32,228在 GitHub 上查看↗

    Fairseq is a PyTorch toolkit for sequence-to-sequence modeling, specializing in neural machine translation, automatic speech recognition, and large-scale language model training. It provides a framework for processing and aligning diverse data sources, including text, audio, and video, to support tasks such as speech-to-text conversion and multimodal sequence learning. The project is distinguished by its distributed training capabilities, which utilize parameter sharding, mixed-precision training, and CPU offloading to handle models that exceed single-device memory. It also includes specializ

    Python
    在 GitHub 上查看↗32,228
  • blob42/instruktblob42 的头像

    blob42/Instrukt

    327在 GitHub 上查看↗

    Integrated AI environment in the terminal. Build, test and instruct agents.

    Pythonagent-executoragentsai
    在 GitHub 上查看↗327
  • google-research/albertgoogle-research 的头像

    google-research/ALBERT

    3,279在 GitHub 上查看↗

    ALBERT

    Python
    在 GitHub 上查看↗3,279
  • eth-ait/multiplyeth-ait 的头像

    eth-ait/MultiPly

    254在 GitHub 上查看↗

    Official Repository for CVPR 2024 paper MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild.

    Python
    在 GitHub 上查看↗254
  • ej0cl6/texteeej0cl6 的头像

    ej0cl6/TextEE

    60在 GitHub 上查看↗

    Updates | Datasets | Models | Environment | Running | Results | Website | Paper

    Python
    在 GitHub 上查看↗60
  • artpli/codeieartpli 的头像

    artpli/CodeIE

    41在 GitHub 上查看↗

    This is the official repository for "CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors" (ACL 2023).

    Python
    在 GitHub 上查看↗41
  • akshata29/chatpdfakshata29 的头像

    akshata29/chatpdf

    866在 GitHub 上查看↗

    Chat and Ask on your own data. Accelerator to quickly upload your own enterprise data and use OpenAI services to chat to that uploaded data and ask questions

    TypeScript
    在 GitHub 上查看↗866
  • alphasecio/langchain-text-summarizerA

    alphasecio/langchain-text-summarizer

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • csunny/db-gptcsunny 的头像

    csunny/DB-GPT

    19,006在 GitHub 上查看↗

    DB-GPT is an AI-driven database management system that uses agentic reasoning to execute data tasks. It converts natural language prompts into executable database queries and combines structured database records with unstructured knowledge bases to provide grounded analysis. The system orchestrates multi-step reasoning chains that integrate database queries, custom scripts, and external tool calls. It allows for the packaging of domain knowledge into reusable analysis skills and executes generated code within sandboxed environments for system safety. The platform covers data orchestration ac

    Python
    在 GitHub 上查看↗19,006
  • cliport/cliportcliport 的头像

    cliport/cliport

    545在 GitHub 上查看↗

    CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

    Jupyter Notebook
    在 GitHub 上查看↗545
  • dlyuangod/tinygpt-vDLYuanGod 的头像

    DLYuanGod/TinyGPT-V

    1,315在 GitHub 上查看↗

    TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

    Python
    在 GitHub 上查看↗1,315
  • allenai/visprogallenai 的头像

    allenai/visprog

    773在 GitHub 上查看↗

    Official code for VisProg (CVPR 2023 Best Paper!)

    Python
    在 GitHub 上查看↗773
  • emma1066/self-improve-zero-shot-nerEmma1066 的头像

    Emma1066/Self-Improve-Zero-Shot-NER

    53在 GitHub 上查看↗

    This is the github repository for the paper to be appeared at NAACL 2024 main conference: Self-Improving for Zero-Shot Named Entity Recognition with Large Language Models.

    Python
    在 GitHub 上查看↗53
  • cheshire-cat-ai/corecheshire-cat-ai 的头像

    cheshire-cat-ai/core

    3,045在 GitHub 上查看↗

    AI agent microservice

    Python
    在 GitHub 上查看↗3,045
  • facebookresearch/detrfacebookresearch 的头像

    facebookresearch/detr

    15,305在 GitHub 上查看↗

    This project provides a transformer-based object detection model that treats the task as a direct set prediction problem. It implements a vision system capable of predicting bounding boxes and class labels for objects within an image, as well as frameworks for instance and panoptic segmentation. The architecture utilizes a transformer encoder and decoder to perform end-to-end set prediction, employing a Hungarian matcher to assign predicted boxes to ground truth objects. It incorporates a convolutional backbone for feature extraction and a system of learnable object queries to probe image loc

    Python
    在 GitHub 上查看↗15,305
  • chenfei-wu/taskmatrixchenfei-wu 的头像

    chenfei-wu/TaskMatrix

    34,082在 GitHub 上查看↗

    TaskMatrix is a multimodal AI chat interface and visual task orchestrator. It combines language models with visual recognition to enable the exchange, analysis, and modification of images within a conversational environment. The system coordinates multiple foundation models through orchestration pipelines that chain language, detection, and segmentation models. This allows for complex visual operations, such as using text instructions to guide image masking and executing modular inpainting workflows to edit specific image regions. The project includes a computer vision toolset for object det

    Python
    在 GitHub 上查看↗34,082
  • fraserxu/book-gptfraserxu 的头像

    fraserxu/book-gpt

    436在 GitHub 上查看↗

    Drop a book, start asking question.

    TypeScript
    在 GitHub 上查看↗436
  • allenai/unified-io-2allenai 的头像

    allenai/unified-io-2

    649在 GitHub 上查看↗

    This repo contains code for Unified-IO 2, including code to run a demo, do training, and do inference. This codebase is modified from T5X.

    Python
    在 GitHub 上查看↗649
  • chen700564/metaner-iclchen700564 的头像

    chen700564/metaner-icl

    40在 GitHub 上查看↗

    An implementation for ACL 2023 paper Learning In-context Learning for Named Entity Recognition

    Python
    在 GitHub 上查看↗40