awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to mshukor/unival

Open-source alternatives to UnIVAL

30 open-source projects similar to mshukor/unival, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best UnIVAL alternative.

  • baaivision/emuالصورة الرمزية لـ baaivision

    baaivision/Emu

    1,775عرض على GitHub↗

    Emu Series: Generative Multimodal Models from BAAI

    Python
    عرض على GitHub↗1,775
  • openrobotlab/pointllmالصورة الرمزية لـ OpenRobotLab

    OpenRobotLab/PointLLM

    1,026عرض على GitHub↗

    ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds

    Python
    عرض على GitHub↗1,026
  • 11cafe/jaazالصورة الرمزية لـ 11cafe

    11cafe/jaaz

    6,384عرض على GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    عرض على GitHub↗6,384
  • qwenlm/qwen3-omniالصورة الرمزية لـ QwenLM

    QwenLM/Qwen3-Omni

    3,843عرض على GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    عرض على GitHub↗3,843

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • othersideai/self-operating-computerالصورة الرمزية لـ OthersideAI

    OthersideAI/self-operating-computer

    10,153عرض على GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    عرض على GitHub↗10,153
  • simular-ai/agent-sالصورة الرمزية لـ simular-ai

    simular-ai/Agent-S

    11,855عرض على GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    عرض على GitHub↗11,855
  • allenai/unified-io-2الصورة الرمزية لـ allenai

    allenai/unified-io-2

    649عرض على GitHub↗

    This repo contains code for Unified-IO 2, including code to run a demo, do training, and do inference. This codebase is modified from T5X.

    Python
    عرض على GitHub↗649
  • baichuan-inc/baichuan2الصورة الرمزية لـ baichuan-inc

    baichuan-inc/Baichuan2

    4,098عرض على GitHub↗

    Baichuan2 is a collection of pre-trained large language models, including base and chat variants, designed for natural language generation and multi-turn conversational AI. It provides an inference engine and a fine-tuning framework to adapt these models to custom datasets and specialized domains. The project features a quantization toolkit and an inference engine that enable model execution across diverse hardware, including graphics processors, central processors, and specialized accelerators. These tools support low-bit weight quantization to reduce memory usage and increase inference spee

    Pythonartificial-intelligencebenchmarkceval
    عرض على GitHub↗4,098
  • alembics/disco-diffusionالصورة الرمزية لـ alembics

    alembics/disco-diffusion

    7,407عرض على GitHub↗

    This project is a diffusion-based AI art generator and animation framework used to create digital images and motion graphics from text prompts. It functions as a system for producing stylized videos and AI art through iterative diffusion sampling and neural network models. The framework distinguishes itself through specialized tools for 3D depth animation, using depth-map transformations to create spatial movement. It also includes neural style transfer capabilities to apply specific artistic looks, such as watercolor or pixel art, and utilizes optical flow frame blending to reduce flickering

    Jupyter Notebook
    عرض على GitHub↗7,407
  • baichuan-inc/baichuan-7bالصورة الرمزية لـ baichuan-inc

    baichuan-inc/Baichuan-7B

    5,654عرض على GitHub↗

    Baichuan-7B is an open-source 7 billion parameter bilingual Transformer model designed for text generation and few-shot learning across Chinese and English. It is built on a large Transformer architecture trained on a bilingual corpus, enabling it to produce coherent text in both languages from a single model. The model incorporates several optimization techniques that distinguish it from standard large language models. It uses rotary position embeddings that can extrapolate to longer sequences than seen during training, allowing context extension beyond the original 4096-token training lengt

    Pythonartificial-intelligencecevalchatgpt
    عرض على GitHub↗5,654
  • eleutherai/gpt-neoxالصورة الرمزية لـ EleutherAI

    EleutherAI/gpt-neox

    7,392عرض على GitHub↗

    gpt-neox is a distributed training system and framework for building large-scale autoregressive language models. It implements the transformer architecture and provides a toolkit for training models with billions of parameters by distributing weights across compute clusters. The framework distinguishes itself through extensive support for distributed model parallelism, including pipeline and sequence parallelism, to overcome single-device memory limits. It further supports sparse model architectures using a mixture of experts system with Sinkhorn-based routing. The project covers a broad ran

    Pythondeepspeed-librarygpt-3language-model
    عرض على GitHub↗7,392
  • biomap-research/scfoundationB

    biomap-research/scFoundation

    0عرض على GitHub↗
    عرض على GitHub↗0
  • bowang-lab/scgptالصورة الرمزية لـ bowang-lab

    bowang-lab/scGPT

    1,585عرض على GitHub↗

    This is the official codebase for scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI.

    Jupyter Notebookfoundation-modelgptsingle-cell
    عرض على GitHub↗1,585
  • camenduru/text-to-video-synthesis-colabالصورة الرمزية لـ camenduru

    camenduru/text-to-video-synthesis-colab

    1,515عرض على GitHub↗

    Text To Video Synthesis Colab

    Jupyter Notebookcolabcolab-notebookcolaboratory
    عرض على GitHub↗1,515
  • baichuan-inc/baichuan-13bالصورة الرمزية لـ baichuan-inc

    baichuan-inc/Baichuan-13B

    2,931عرض على GitHub↗

    A 13B large language model developed by Baichuan Intelligent Technology

    Pythonartificial-intelligencebenchmarkceval
    عرض على GitHub↗2,931
  • ai-chef/hugginggptالصورة الرمزية لـ AI-Chef

    AI-Chef/HuggingGPT

    24عرض على GitHub↗

    The mission of JARVIS is to explore artificial general intelligence (AGI) and deliver cutting-edge research to the whole community.

    Python
    عرض على GitHub↗24
  • damo-nlp-sg/videollama3الصورة الرمزية لـ DAMO-NLP-SG

    DAMO-NLP-SG/VideoLLaMA3

    1,105عرض على GitHub↗
    Jupyter Notebook
    عرض على GitHub↗1,105
  • cvlab-columbia/viperالصورة الرمزية لـ cvlab-columbia

    cvlab-columbia/viper

    1,717عرض على GitHub↗

    Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"

    Jupyter Notebook
    عرض على GitHub↗1,717
  • databrickslabs/dollyالصورة الرمزية لـ databrickslabs

    databrickslabs/dolly

    10,795عرض على GitHub↗

    Dolly is an instruction-tuned large language model designed to follow complex natural language directions. It operates as a causal language model that predicts the next token in a sequence to generate coherent conversational responses and perform tasks such as brainstorming, classification, and question answering. The project focuses on the development of models using open datasets suitable for commercial application. It enables the creation of instruction-following models by utilizing curated collections of human-generated instruction-response pairs. The repository provides capabilities for

    Python
    عرض على GitHub↗10,795
  • datadog/totoالصورة الرمزية لـ DataDog

    DataDog/toto

    482عرض على GitHub↗

    Toto 2.0: Model Weights | Blog Toto 1.0: Paper | Blog | Model Card

    Jupyter Notebook
    عرض على GitHub↗482
  • deepseek-ai/deepseek-llmالصورة الرمزية لـ deepseek-ai

    deepseek-ai/deepseek-LLM

    7,100عرض على GitHub↗

    DeepSeek-LLM is a large language model and causal language model designed for natural language generation. It functions as a multi-lingual system capable of predicting the next token in a sequence to perform text completion and conversational generation. The model is specialized for logical reasoning, specifically as a code and math LLM. This enables it to perform complex problem solving, which includes generating executable code and solving mathematical equations through step-by-step analysis. The system's broader capabilities cover conversational AI, including the generation of chat comple

    Makefile
    عرض على GitHub↗7,100
  • dlyuangod/tinygpt-vالصورة الرمزية لـ DLYuanGod

    DLYuanGod/TinyGPT-V

    1,315عرض على GitHub↗

    TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

    Python
    عرض على GitHub↗1,315
  • baaivision/emu3الصورة الرمزية لـ baaivision

    baaivision/Emu3

    2,417عرض على GitHub↗

    Next-Token Prediction is All You Need

    Python
    عرض على GitHub↗2,417
  • eth-ait/multiplyالصورة الرمزية لـ eth-ait

    eth-ait/MultiPly

    254عرض على GitHub↗

    Official Repository for CVPR 2024 paper MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild.

    Python
    عرض على GitHub↗254
  • facebookresearch/codellamaالصورة الرمزية لـ facebookresearch

    facebookresearch/codellama

    16,307عرض على GitHub↗

    Code Llama is a large language model based on Llama 2 trained specifically for programming tasks and software development. It provides specialized model types optimized for general code generation, instruction following, and context-aware infilling. The project includes an instruction-tuned programming model for executing technical tasks via natural language prompts and a code infilling model that predicts missing sections based on surrounding source context. A large context code model is also provided to analyze extensive blocks of source code for improved coherence. The system covers capab

    Python
    عرض على GitHub↗16,307
  • facebookresearch/llamaالصورة الرمزية لـ facebookresearch

    facebookresearch/llama

    59,466عرض على GitHub↗

    Llama is a large language model runtime and inference engine designed to load and execute autoregressive transformer models. It enables the generation of natural language text completions from prompts using pretrained weights. The system features multi-GPU model parallelism, which distributes model weights and workloads across multiple graphics processors to support larger parameter counts. It also incorporates a content safety filter that uses classifiers to intercept and block unsafe inputs or outputs during the inference process. The project covers broad capabilities in distributed model

    Python
    عرض على GitHub↗59,466
  • facebookresearch/moviegenbenchF

    facebookresearch/MovieGenBench

    0عرض على GitHub↗
    عرض على GitHub↗0
  • facebookresearch/segment-anythingالصورة الرمزية لـ facebookresearch

    facebookresearch/segment-anything

    54,353عرض على GitHub↗

    This project provides a deep learning architecture designed to identify and isolate distinct objects within images by generating precise pixel-level masks. It functions as a browser-based inference engine, enabling the execution of complex machine learning models directly within web environments without requiring server-side processing. The system distinguishes itself by utilizing hardware-accelerated execution and parallel processing to achieve real-time segmentation speeds. It supports prompt-based mask decoding, allowing users to generate spatial masks by providing specific points or boxes

    Jupyter Notebook
    عرض على GitHub↗54,353
  • flagai-open/flagaiالصورة الرمزية لـ FlagAI-Open

    FlagAI-Open/FlagAI

    3,870عرض على GitHub↗

    FlagAI is a distributed deep learning framework and platform designed for the end-to-end lifecycle of large-scale foundation models. It provides a toolkit for training, fine-tuning, and deploying large language models and multi-modal systems across multi-node computing clusters. The project features hardware-agnostic compute abstractions to ensure consistent execution across different accelerators. It includes a dedicated library for parameter-efficient fine-tuning, allowing large neural networks to be adapted to specific tasks with minimal parameter updates and reduced computational overhead

    Python
    عرض على GitHub↗3,870
  • compvis/stable-diffusionالصورة الرمزية لـ CompVis

    CompVis/stable-diffusion

    73,125عرض على GitHub↗

    Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I

    Jupyter Notebook
    عرض على GitHub↗73,125