awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
ยฉ 2026 Bringes Technology SRLยทVAT RO45896025ยทhello@awesome-repositories.com
Back to openrobotlab/pointllm

Open-source alternatives to PointLLM

30 open-source projects similar to openrobotlab/pointllm, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best PointLLM alternative.

  • mshukor/univalmshukor avatar

    mshukor/UnIVAL

    236View on GitHubโ†—

    TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

    Jupyter Notebook
    View on GitHubโ†—236
  • baaivision/emubaaivision avatar

    baaivision/Emu

    1,775View on GitHubโ†—

    Emu Series: Generative Multimodal Models from BAAI

    Python
    View on GitHubโ†—1,775
  • next-gpt/next-gptNExT-GPT avatar

    NExT-GPT/NExT-GPT

    3,636View on GitHubโ†—

    Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence )

    Python
    View on GitHubโ†—3,636
  • yuliang-liu/monkeyYuliang-Liu avatar

    Yuliang-Liu/Monkey

    1,948View on GitHubโ†—

    Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)

    Python
    View on GitHubโ†—1,948
  • simular-ai/agent-ssimular-ai avatar

    simular-ai/Agent-S

    11,855View on GitHubโ†—

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    View on GitHubโ†—11,855

AI search

Explore more awesome repositories

Describe what you need in plain English โ€” the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • othersideai/self-operating-computerOthersideAI avatar

    OthersideAI/self-operating-computer

    10,153View on GitHubโ†—

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    View on GitHubโ†—10,153
  • 11cafe/jaaz11cafe avatar

    11cafe/jaaz

    6,384View on GitHubโ†—

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    View on GitHubโ†—6,384
  • qwenlm/qwen3-omniQwenLM avatar

    QwenLM/Qwen3-Omni

    3,843View on GitHubโ†—

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    View on GitHubโ†—3,843
  • alembics/disco-diffusionalembics avatar

    alembics/disco-diffusion

    7,407View on GitHubโ†—

    This project is a diffusion-based AI art generator and animation framework used to create digital images and motion graphics from text prompts. It functions as a system for producing stylized videos and AI art through iterative diffusion sampling and neural network models. The framework distinguishes itself through specialized tools for 3D depth animation, using depth-map transformations to create spatial movement. It also includes neural style transfer capabilities to apply specific artistic looks, such as watercolor or pixel art, and utilizes optical flow frame blending to reduce flickering

    Jupyter Notebook
    View on GitHubโ†—7,407
  • damo-nlp-sg/videollama3DAMO-NLP-SG avatar

    DAMO-NLP-SG/VideoLLaMA3

    1,105View on GitHubโ†—
    Jupyter Notebook
    View on GitHubโ†—1,105
  • baaivision/emu3baaivision avatar

    baaivision/Emu3

    2,417View on GitHubโ†—

    Next-Token Prediction is All You Need

    Python
    View on GitHubโ†—2,417
  • baaivision/uni3dbaaivision avatar

    baaivision/Uni3D

    674View on GitHubโ†—

    Uni3D: Exploring Unified 3D Representation at Scale

    Python
    View on GitHubโ†—674
  • baichuan-inc/baichuan-13bbaichuan-inc avatar

    baichuan-inc/Baichuan-13B

    2,931View on GitHubโ†—

    A 13B large language model developed by Baichuan Intelligent Technology

    Pythonartificial-intelligencebenchmarkceval
    View on GitHubโ†—2,931
  • baichuan-inc/baichuan-7bbaichuan-inc avatar

    baichuan-inc/Baichuan-7B

    5,654View on GitHubโ†—

    Baichuan-7B is an open-source 7 billion parameter bilingual Transformer model designed for text generation and few-shot learning across Chinese and English. It is built on a large Transformer architecture trained on a bilingual corpus, enabling it to produce coherent text in both languages from a single model. The model incorporates several optimization techniques that distinguish it from standard large language models. It uses rotary position embeddings that can extrapolate to longer sequences than seen during training, allowing context extension beyond the original 4096-token training lengt

    Pythonartificial-intelligencecevalchatgpt
    View on GitHubโ†—5,654
  • ai-chef/hugginggptAI-Chef avatar

    AI-Chef/HuggingGPT

    24View on GitHubโ†—

    The mission of JARVIS is to explore artificial general intelligence (AGI) and deliver cutting-edge research to the whole community.

    Python
    View on GitHubโ†—24
  • compvis/stable-diffusionCompVis avatar

    CompVis/stable-diffusion

    73,125View on GitHubโ†—

    Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I

    Jupyter Notebook
    View on GitHubโ†—73,125
  • chat-3d/chat-3dChat-3D avatar

    Chat-3D/Chat-3D

    57View on GitHubโ†—

    This is a repo for paper "Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes". paper, project page

    Python
    View on GitHubโ†—57
  • baai-dcai/visual-instruction-tuningBAAI-DCAI avatar

    BAAI-DCAI/Visual-Instruction-Tuning

    168View on GitHubโ†—

    Scale up visual instruction tuning to millions by GPT-4.

    Python
    View on GitHubโ†—168
  • abenechehab/adaptsabenechehab avatar

    abenechehab/AdaPTS

    51View on GitHubโ†—

    Our paper has been accepted at ICML 2025! ๐ŸŽ‰๐Ÿ“„๐ŸŽ“ ๐Ÿ‘‰ AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series Forecasting

    Jupyter Notebook
    View on GitHubโ†—51
  • cliport/cliportcliport avatar

    cliport/cliport

    545View on GitHubโ†—

    CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

    Jupyter Notebook
    View on GitHubโ†—545
  • camenduru/text-to-video-synthesis-colabcamenduru avatar

    camenduru/text-to-video-synthesis-colab

    1,515View on GitHubโ†—

    Text To Video Synthesis Colab

    Jupyter Notebookcolabcolab-notebookcolaboratory
    View on GitHubโ†—1,515
  • cvlab-columbia/vipercvlab-columbia avatar

    cvlab-columbia/viper

    1,717View on GitHubโ†—

    Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"

    Jupyter Notebook
    View on GitHubโ†—1,717
  • bowang-lab/scgptbowang-lab avatar

    bowang-lab/scGPT

    1,585View on GitHubโ†—

    This is the official codebase for scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI.

    Jupyter Notebookfoundation-modelgptsingle-cell
    View on GitHubโ†—1,585
  • databrickslabs/dollydatabrickslabs avatar

    databrickslabs/dolly

    10,795View on GitHubโ†—

    Dolly is an instruction-tuned large language model designed to follow complex natural language directions. It operates as a causal language model that predicts the next token in a sequence to generate coherent conversational responses and perform tasks such as brainstorming, classification, and question answering. The project focuses on the development of models using open datasets suitable for commercial application. It enables the creation of instruction-following models by utilizing curated collections of human-generated instruction-response pairs. The repository provides capabilities for

    Python
    View on GitHubโ†—10,795
  • datadog/totoDataDog avatar

    DataDog/toto

    482View on GitHubโ†—

    Toto 2.0: Model Weights | Blog Toto 1.0: Paper | Blog | Model Card

    Jupyter Notebook
    View on GitHubโ†—482
  • deepmind/deepmind-researchdeepmind avatar

    deepmind/deepmind-research

    15,024View on GitHubโ†—

    This project is an AI research implementation library and machine learning research repository. It provides a collection of reference code, illustrative implementations, and open-source research datasets used to verify hypotheses and build upon existing models in artificial intelligence. The repository focuses on scientific research reproduction by translating theoretical findings from published papers into executable code. It includes specialized scientific simulation environments designed to test the behavior of autonomous agents and models within controlled settings. The project covers AI

    Jupyter Notebook
    View on GitHubโ†—15,024
  • deepseek-ai/deepseek-llmdeepseek-ai avatar

    deepseek-ai/deepseek-LLM

    7,100View on GitHubโ†—

    DeepSeek-LLM is a large language model and causal language model designed for natural language generation. It functions as a multi-lingual system capable of predicting the next token in a sequence to perform text completion and conversational generation. The model is specialized for logical reasoning, specifically as a code and math LLM. This enables it to perform complex problem solving, which includes generating executable code and solving mathematical equations through step-by-step analysis. The system's broader capabilities cover conversational AI, including the generation of chat comple

    Makefile
    View on GitHubโ†—7,100
  • deepseek-ai/janusdeepseek-ai avatar

    deepseek-ai/Janus

    17,746View on GitHubโ†—

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    View on GitHubโ†—17,746
  • djiajunustc/3d-llavadjiajunustc avatar

    djiajunustc/3D-LLaVA

    98View on GitHubโ†—

    3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

    Python
    View on GitHubโ†—98
  • anjiecheng/spatialrgptAnjieCheng avatar

    AnjieCheng/SpatialRGPT

    330View on GitHubโ†—

    arxiv / Huggingface

    Python
    View on GitHubโ†—330