awesome-repositories.comCategoríasBlog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to peract/peract

Open-source alternatives to Peract

30 open-source projects similar to peract/peract, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Peract alternative.

  • cliport/cliportAvatar de cliport

    cliport/cliport

    545Ver en GitHub↗

    CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

    Jupyter Notebook
    Ver en GitHub↗545
  • othersideai/self-operating-computerAvatar de OthersideAI

    OthersideAI/self-operating-computer

    10,153Ver en GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    Ver en GitHub↗10,153
  • qwenlm/qwen3-omniAvatar de QwenLM

    QwenLM/Qwen3-Omni

    3,843Ver en GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    Ver en GitHub↗3,843
  • 11cafe/jaazAvatar de 11cafe

    11cafe/jaaz

    6,384Ver en GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    Ver en GitHub↗6,384

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
  • simular-ai/agent-sAvatar de simular-ai

    simular-ai/Agent-S

    11,855Ver en GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    Ver en GitHub↗11,855
  • baaivision/emuAvatar de baaivision

    baaivision/Emu

    1,775Ver en GitHub↗

    Emu Series: Generative Multimodal Models from BAAI

    Python
    Ver en GitHub↗1,775
  • changhaonan/a3vlmC

    changhaonan/A3VLM

    0Ver en GitHub↗
    Ver en GitHub↗0
  • columbia-ai-robotics/scalingupC

    columbia-ai-robotics/scalingup

    0Ver en GitHub↗
    Ver en GitHub↗0
  • craftjarvis/mc-plannerC

    CraftJarvis/MC-Planner

    0Ver en GitHub↗
    Ver en GitHub↗0
  • cvlab-columbia/viperAvatar de cvlab-columbia

    cvlab-columbia/viper

    1,717Ver en GitHub↗

    Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"

    Jupyter Notebook
    Ver en GitHub↗1,717
  • dlyuangod/tinygpt-vAvatar de DLYuanGod

    DLYuanGod/TinyGPT-V

    1,315Ver en GitHub↗

    TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

    Python
    Ver en GitHub↗1,315
  • dongyh20/octopusD

    dongyh20/Octopus

    0Ver en GitHub↗
    Ver en GitHub↗0
  • eth-ait/multiplyAvatar de eth-ait

    eth-ait/MultiPly

    254Ver en GitHub↗

    Official Repository for CVPR 2024 paper MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild.

    Python
    Ver en GitHub↗254
  • facebookresearch/r3mF

    facebookresearch/r3m

    0Ver en GitHub↗
    Ver en GitHub↗0
  • gen-robot/rl4vlaAvatar de gen-robot

    gen-robot/RL4VLA

    274Ver en GitHub↗

    This repository contains the code for the paper What Can RL Bring to VLA Generalization? An Empirical Study. The pretrained checkpoints are available at HuggingFace.

    Python
    Ver en GitHub↗274
  • gpt-omni/mini-omniG

    gpt-omni/mini-omni

    0Ver en GitHub↗

    Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

    Ver en GitHub↗0
  • hackiey/visual-chatgptAvatar de hackiey

    hackiey/visual-chatgpt

    1Ver en GitHub↗

    Visual ChatGPT connects ChatGPT and a series of Visual Foundation Models to enable sending and receiving images during chatting.

    Ver en GitHub↗1
  • llava-vl/llava-nextAvatar de LLaVA-VL

    LLaVA-VL/LLaVA-NeXT

    4,695Ver en GitHub↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    Ver en GitHub↗4,695
  • lyuchenyang/macaw-llmAvatar de lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590Ver en GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Pythondeep-learninglanguage-modelmachine-learning
    Ver en GitHub↗1,590
  • meituan-automl/mobilevlmAvatar de Meituan-AutoML

    Meituan-AutoML/MobileVLM

    1,359Ver en GitHub↗

    MobileVLM: Vision Language Model for Mobile Devices

    Python
    Ver en GitHub↗1,359
  • microsoft/i-codeAvatar de microsoft

    microsoft/i-Code

    1,706Ver en GitHub↗

    The ambition of the i-Code project is to build integrative and composable multimodal Artificial Intelligence. The "i" stands for integrative multimodal learning.

    Jupyter Notebook
    Ver en GitHub↗1,706
  • microsoft/mm-reactAvatar de microsoft

    microsoft/MM-REACT

    966Ver en GitHub↗

    Official repo for MM-REACT

    Python
    Ver en GitHub↗966
  • microsoft/omniparserAvatar de microsoft

    microsoft/OmniParser

    24,377Ver en GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Jupyter Notebook
    Ver en GitHub↗24,377
  • minedojo/voyagerAvatar de MineDojo

    MineDojo/Voyager

    6,987Ver en GitHub↗

    Voyager is an autonomous embodied agent and lifelong learning framework that uses a large language model to explore virtual environments. It functions as a code-based action controller, translating natural language instructions into executable scripts to interact with its surroundings. The system features an automatic curriculum generator that creates sequences of exploration goals to discover new items and behaviors without human intervention. It maintains a skill library manager that stores learned behaviors as reusable code fragments, which can be composed to execute complex tasks. The fr

    JavaScriptembodied-learninglarge-language-modelsminecraft
    Ver en GitHub↗6,987
  • mshukor/univalAvatar de mshukor

    mshukor/UnIVAL

    236Ver en GitHub↗

    TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

    Jupyter Notebook
    Ver en GitHub↗236
  • next-gpt/next-gptAvatar de NExT-GPT

    NExT-GPT/NExT-GPT

    3,636Ver en GitHub↗

    Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence )

    Python
    Ver en GitHub↗3,636
  • notmahi/clip-fieldsAvatar de notmahi

    notmahi/clip-fields

    189Ver en GitHub↗

    [Paper](https://arxiv.org/abs/2210.05663) [Website](https://mahis.life/clip-fields/) [Code](https://github.com/notmahi/clip-fields) [Data](https://osf.io/famgv) [Video](https://youtu.be/bKu7GvRiSQU)

    Python
    Ver en GitHub↗189
  • ofa-sys/one-peaceAvatar de OFA-Sys

    OFA-Sys/ONE-PEACE

    1,063Ver en GitHub↗

    A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    Pythonaudio-languagecontrastive-lossfoundation-models
    Ver en GitHub↗1,063
  • openbmb/minicpm-vAvatar de OpenBMB

    OpenBMB/MiniCPM-V

    25,653Ver en GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Python
    Ver en GitHub↗25,653
  • opendrivelab/agibot-worldAvatar de OpenDriveLab

    OpenDriveLab/AgiBot-World

    2,786Ver en GitHub↗

    AgiBot-World is a suite of software pipelines and tools designed for robotic policy training, dataset standardization, embodiment transfer, and performance benchmarking. It provides infrastructure for developing bimanual manipulation policies using foundation models and human-reference trajectory data. The project features a robot embodiment transfer suite that adapts pre-trained models to different robot bodies without requiring new multi-embodiment training data. It also includes a specialized evaluation framework for validating vision-language-action models through open-loop testing and ph

    Pythonpretraining-for-roboticsrobotic-foundation-modelrobotic-manipulation
    Ver en GitHub↗2,786