awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to phellonchen/x-llm

Open-source alternatives to X LLM

30 open-source projects similar to phellonchen/x-llm, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best X LLM alternative.

  • lyuchenyang/macaw-llmlyuchenyang avatar

    lyuchenyang/Macaw-LLM

    1,590View on GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Pythondeep-learninglanguage-modelmachine-learning
    View on GitHub↗1,590
  • simular-ai/agent-ssimular-ai avatar

    simular-ai/Agent-S

    11,855View on GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    View on GitHub↗11,855
  • othersideai/self-operating-computerOthersideAI avatar

    OthersideAI/self-operating-computer

    10,153View on GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    View on GitHub↗10,153

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • 11cafe/jaaz11cafe avatar

    11cafe/jaaz

    6,384View on GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    View on GitHub↗6,384
  • qwenlm/qwen3-omniQwenLM avatar

    QwenLM/Qwen3-Omni

    3,843View on GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    View on GitHub↗3,843
  • plexpt/chatgpt-corpusPlexPt avatar

    PlexPt/chatgpt-corpus

    964View on GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    View on GitHub↗964
  • cluebenchmark/cluecorpus2020CLUEbenchmark avatar

    CLUEbenchmark/CLUECorpus2020

    1,012View on GitHub↗

    Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

    albertbertchinese
    View on GitHub↗1,012
  • cliport/cliportcliport avatar

    cliport/cliport

    545View on GitHub↗

    CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

    Jupyter Notebook
    View on GitHub↗545
  • ai-chef/hugginggptAI-Chef avatar

    AI-Chef/HuggingGPT

    24View on GitHub↗

    The mission of JARVIS is to explore artificial general intelligence (AGI) and deliver cutting-edge research to the whole community.

    Python
    View on GitHub↗24
  • eth-ait/multiplyeth-ait avatar

    eth-ait/MultiPly

    254View on GitHub↗

    Official Repository for CVPR 2024 paper MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild.

    Python
    View on GitHub↗254
  • fuxiaoliu/lrv-instructionFuxiaoLiu avatar

    FuxiaoLiu/LRV-Instruction

    297View on GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    View on GitHub↗297
  • gpt-omni/mini-omniG

    gpt-omni/mini-omni

    0View on GitHub↗

    Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

    View on GitHub↗0
  • guoyang9/unk-vqaguoyang9 avatar

    guoyang9/UNK-VQA

    7View on GitHub↗

    A VQA dataset that includes unanswerable questions TPAMI 2024.

    View on GitHub↗7
  • hackiey/visual-chatgpthackiey avatar

    hackiey/visual-chatgpt

    1View on GitHub↗

    Visual ChatGPT connects ChatGPT and a series of Visual Foundation Models to enable sending and receiving images during chatting.

    View on GitHub↗1
  • esbatmop/mnbvcesbatmop avatar

    esbatmop/MNBVC

    4,123View on GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    View on GitHub↗4,123
  • baaivision/emubaaivision avatar

    baaivision/Emu

    1,775View on GitHub↗

    Emu Series: Generative Multimodal Models from BAAI

    Python
    View on GitHub↗1,775
  • kakaobrain/coyo-datasetkakaobrain avatar

    kakaobrain/coyo-dataset

    1,256View on GitHub↗

    COYO-700M: Large-scale Image-Text Pair Dataset

    Python
    View on GitHub↗1,256
  • inst-it/inst-itinst-it avatar

    inst-it/inst-it

    40View on GitHub↗

    NeurIPS 2025 The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning"

    Pythoninstruction-tuninglarge-multimodal-modelsmultimodal
    View on GitHub↗40
  • llava-vl/llava-nextLLaVA-VL avatar

    LLaVA-VL/LLaVA-NeXT

    4,695View on GitHub↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    View on GitHub↗4,695
  • luodian/otterLuodian avatar

    Luodian/Otter

    3,410View on GitHub↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Python
    View on GitHub↗3,410
  • dlyuangod/tinygpt-vDLYuanGod avatar

    DLYuanGod/TinyGPT-V

    1,315View on GitHub↗

    TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

    Python
    View on GitHub↗1,315
  • mbzuai-oryx/video-chatgptmbzuai-oryx avatar

    mbzuai-oryx/Video-ChatGPT

    1,504View on GitHub↗

    ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    Pythonchatbotclipgpt-4
    View on GitHub↗1,504
  • meituan-automl/mobilevlmMeituan-AutoML avatar

    Meituan-AutoML/MobileVLM

    1,359View on GitHub↗

    MobileVLM: Vision Language Model for Mobile Devices

    Python
    View on GitHub↗1,359
  • microsoft/i-codemicrosoft avatar

    microsoft/i-Code

    1,706View on GitHub↗

    The ambition of the i-Code project is to build integrative and composable multimodal Artificial Intelligence. The "i" stands for integrative multimodal learning.

    Jupyter Notebook
    View on GitHub↗1,706
  • microsoft/llava-medmicrosoft avatar

    microsoft/LLaVA-Med

    2,214View on GitHub↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    Python
    View on GitHub↗2,214
  • microsoft/mm-reactmicrosoft avatar

    microsoft/MM-REACT

    966View on GitHub↗

    Official repo for MM-REACT

    Python
    View on GitHub↗966
  • microsoft/omniparsermicrosoft avatar

    microsoft/OmniParser

    24,377View on GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Jupyter Notebook
    View on GitHub↗24,377
  • mobvoi/seq-monkey-datamobvoi avatar

    mobvoi/seq-monkey-data

    178View on GitHub↗
    View on GitHub↗178
  • mshukor/univalmshukor avatar

    mshukor/UnIVAL

    236View on GitHub↗

    TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

    Jupyter Notebook
    View on GitHub↗236
  • icoz69/stablellavaicoz69 avatar

    icoz69/StableLLAVA

    95View on GitHub↗

    Official repo for StableLLAVA

    Python
    View on GitHub↗95