awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
phellonchen avatar

phellonchen/X-LLM

0
View on GitHub↗
318 stars·18 forks·Python·Apache-2.0·7 vuesx-llm.github.io↗

X LLM

X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Features

  • Multimodal Agents - Treating multi-modalities as foreign languages for LLMs.
  • Pre-training Datasets - Chinese multimodal instruction dataset for foreign language treatment.

Historique des stars

Graphique de l'historique des stars pour phellonchen/x-llmGraphique de l'historique des stars pour phellonchen/x-llm

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à X LLM

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec X LLM.
  • lyuchenyang/macaw-llmAvatar de lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590Voir sur GitHub↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Pythondeep-learninglanguage-modelmachine-learning
    Voir sur GitHub↗1,590
  • othersideai/self-operating-computerAvatar de OthersideAI

    OthersideAI/self-operating-computer

    10,153Voir sur GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    Voir sur GitHub↗10,153
  • 11cafe/jaazAvatar de 11cafe

    11cafe/jaaz

    6,384Voir sur GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    Voir sur GitHub↗6,384
  • qwenlm/qwen3-omniAvatar de QwenLM

    QwenLM/Qwen3-Omni

    3,843Voir sur GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    Voir sur GitHub↗3,843
Voir les 30 alternatives à X LLM→

Questions fréquentes

Que fait phellonchen/x-llm ?

X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Quelles sont les fonctionnalités principales de phellonchen/x-llm ?

Les fonctionnalités principales de phellonchen/x-llm sont : Multimodal Agents, Pre-training Datasets.

Quelles sont les alternatives open-source à phellonchen/x-llm ?

Les alternatives open-source à phellonchen/x-llm incluent : lyuchenyang/macaw-llm — Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration. simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… qwenlm/qwen3-omni — Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video… plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of…