awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to ofa-sys/one-peace

Open-source alternatives to ONE PEACE

30 open-source projects similar to ofa-sys/one-peace, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best ONE PEACE alternative.

  • qwenlm/qwen3-omniالصورة الرمزية لـ QwenLM

    QwenLM/Qwen3-Omni

    3,843عرض على GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    عرض على GitHub↗3,843
  • simular-ai/agent-sالصورة الرمزية لـ simular-ai

    simular-ai/Agent-S

    11,855عرض على GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    عرض على GitHub↗11,855
  • othersideai/self-operating-computerالصورة الرمزية لـ OthersideAI

    OthersideAI/self-operating-computer

    10,153عرض على GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    عرض على GitHub↗10,153

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • 11cafe/jaazالصورة الرمزية لـ 11cafe

    11cafe/jaaz

    6,384عرض على GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    عرض على GitHub↗6,384
  • cliport/cliportالصورة الرمزية لـ cliport

    cliport/cliport

    545عرض على GitHub↗

    CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

    Jupyter Notebook
    عرض على GitHub↗545
  • cvlab-columbia/viperالصورة الرمزية لـ cvlab-columbia

    cvlab-columbia/viper

    1,717عرض على GitHub↗

    Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"

    Jupyter Notebook
    عرض على GitHub↗1,717
  • dlyuangod/tinygpt-vالصورة الرمزية لـ DLYuanGod

    DLYuanGod/TinyGPT-V

    1,315عرض على GitHub↗

    TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

    Python
    عرض على GitHub↗1,315
  • eth-ait/multiplyالصورة الرمزية لـ eth-ait

    eth-ait/MultiPly

    254عرض على GitHub↗

    Official Repository for CVPR 2024 paper MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild.

    Python
    عرض على GitHub↗254
  • gpt-omni/mini-omniG

    gpt-omni/mini-omni

    0عرض على GitHub↗

    Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

    عرض على GitHub↗0
  • hackiey/visual-chatgptالصورة الرمزية لـ hackiey

    hackiey/visual-chatgpt

    1عرض على GitHub↗

    Visual ChatGPT connects ChatGPT and a series of Visual Foundation Models to enable sending and receiving images during chatting.

    عرض على GitHub↗1
  • llava-vl/llava-nextالصورة الرمزية لـ LLaVA-VL

    LLaVA-VL/LLaVA-NeXT

    4,695عرض على GitHub↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    عرض على GitHub↗4,695
  • lyuchenyang/macaw-llmالصورة الرمزية لـ lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590عرض على GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Pythondeep-learninglanguage-modelmachine-learning
    عرض على GitHub↗1,590
  • meituan-automl/mobilevlmالصورة الرمزية لـ Meituan-AutoML

    Meituan-AutoML/MobileVLM

    1,359عرض على GitHub↗

    MobileVLM: Vision Language Model for Mobile Devices

    Python
    عرض على GitHub↗1,359
  • microsoft/i-codeالصورة الرمزية لـ microsoft

    microsoft/i-Code

    1,706عرض على GitHub↗

    The ambition of the i-Code project is to build integrative and composable multimodal Artificial Intelligence. The "i" stands for integrative multimodal learning.

    Jupyter Notebook
    عرض على GitHub↗1,706
  • microsoft/mm-reactالصورة الرمزية لـ microsoft

    microsoft/MM-REACT

    966عرض على GitHub↗

    Official repo for MM-REACT

    Python
    عرض على GitHub↗966
  • microsoft/omniparserالصورة الرمزية لـ microsoft

    microsoft/OmniParser

    24,377عرض على GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Jupyter Notebook
    عرض على GitHub↗24,377
  • mshukor/univalالصورة الرمزية لـ mshukor

    mshukor/UnIVAL

    236عرض على GitHub↗

    TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

    Jupyter Notebook
    عرض على GitHub↗236
  • next-gpt/next-gptالصورة الرمزية لـ NExT-GPT

    NExT-GPT/NExT-GPT

    3,636عرض على GitHub↗

    Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence )

    Python
    عرض على GitHub↗3,636
  • openbmb/minicpm-vالصورة الرمزية لـ OpenBMB

    OpenBMB/MiniCPM-V

    25,653عرض على GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Python
    عرض على GitHub↗25,653
  • openrobotlab/pointllmالصورة الرمزية لـ OpenRobotLab

    OpenRobotLab/PointLLM

    1,026عرض على GitHub↗

    ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds

    Python
    عرض على GitHub↗1,026
  • peract/peractP

    peract/peract

    0عرض على GitHub↗

    Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2022

    عرض على GitHub↗0
  • phellonchen/x-llmالصورة الرمزية لـ phellonchen

    phellonchen/X-LLM

    318عرض على GitHub↗

    X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

    Python
    عرض على GitHub↗318
  • pku-yuangroup/languagebindالصورة الرمزية لـ PKU-YuanGroup

    PKU-YuanGroup/LanguageBind

    885عرض على GitHub↗

    【ICLR 2024 🔥】LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment If you like our project, please give us a star ⭐ on GitHub for latest update.

    Python
    عرض على GitHub↗885
  • qwenlm/qwen2.5-vlالصورة الرمزية لـ QwenLM

    QwenLM/Qwen2.5-VL

    19,480عرض على GitHub↗

    Qwen2.5-VL is an autoregressive multimodal transformer designed to process interleaved sequences of text and visual tokens. It integrates visual feature embeddings into a shared language model space to perform cross-modal reasoning and generate coherent responses or structured layout code. The project distinguishes itself through vision-language-action mapping, allowing it to perceive visual interfaces and translate that perception into actionable commands for operating digital screens and robotic hardware. It employs dynamic-resolution image encoding and temporal-frame video indexing to hand

    Jupyter Notebook
    عرض على GitHub↗19,480
  • qwenlm/qwen2-audioQ

    QwenLM/Qwen2-Audio

    0عرض على GitHub↗

    中文 &nbsp| &nbsp English&nbsp&nbsp

    عرض على GitHub↗0
  • tangyuan96/minigpt-3dالصورة الرمزية لـ TangYuan96

    TangYuan96/MiniGPT-3D

    130عرض على GitHub↗

    MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors Yuan Tang  Xu Han  Xianzhi Li*  Qiao Yu  Yixue Hao  Long Hu  Min Chen Huazhong University of Science and Technology South China University of Technology

    Python
    عرض على GitHub↗130
  • thudm/cogvlm2الصورة الرمزية لـ THUDM

    THUDM/CogVLM2

    2,437عرض على GitHub↗

    中文版README

    Python
    عرض على GitHub↗2,437
  • vision-cair/minigpt-4الصورة الرمزية لـ Vision-CAIR

    Vision-CAIR/MiniGPT-4

    25,679عرض على GitHub↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    Python
    عرض على GitHub↗25,679
  • xinke-wang/modaverseالصورة الرمزية لـ xinke-wang

    xinke-wang/ModaVerse

    28عرض على GitHub↗

    🎆🎆🎆 ~~Visit our online demo here.~~

    Python
    عرض على GitHub↗28
  • yangdongchao/llm-codecالصورة الرمزية لـ yangdongchao

    yangdongchao/LLM-Codec

    147عرض على GitHub↗

    This Repository provides an LLM-driven audio codec model, which can be used to build multi-modal LLMs (text and audio modalities). More details will be introduced as soon as. You can find the paper from https://arxiv.org/pdf/2406.10056

    Python
    عرض على GitHub↗147