awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to openrobotlab/pointllm

Projects sharing features with PointLLM

30 open-source projects similar to openrobotlab/pointllm, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • mshukor/univalmshukor avatar

    mshukor/UnIVAL

    236View on GitHub↗

    TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

    Jupyter Notebook
    View on GitHub↗236
  • baaivision/emubaaivision avatar

    baaivision/Emu

    1,775View on GitHub↗

    Emu Series: Generative Multimodal Models from BAAI

    Python
    View on GitHub↗1,775
  • next-gpt/next-gptNExT-GPT avatar

    NExT-GPT/NExT-GPT

    3,636View on GitHub↗

    Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence )

    Python
    View on GitHub↗3,636
  • yuliang-liu/monkeyYuliang-Liu avatar

    Yuliang-Liu/Monkey

    1,948View on GitHub↗

    Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)

    Python
    View on GitHub↗1,948
  • simular-ai/agent-ssimular-ai avatar

    simular-ai/Agent-S

    11,855View on GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    View on GitHub↗11,855

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • othersideai/self-operating-computerOthersideAI avatar

    OthersideAI/self-operating-computer

    10,153View on GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    View on GitHub↗10,153
  • 11cafe/jaaz11cafe avatar

    11cafe/jaaz

    6,384View on GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    View on GitHub↗6,384
  • qwenlm/qwen3-omniQwenLM avatar

    QwenLM/Qwen3-Omni

    3,843View on GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    View on GitHub↗3,843
  • alembics/disco-diffusionalembics avatar

    alembics/disco-diffusion

    7,407View on GitHub↗

    This project is a diffusion-based AI art generator and animation framework used to create digital images and motion graphics from text prompts. It functions as a system for producing stylized videos and AI art through iterative diffusion sampling and neural network models. The framework distinguishes itself through specialized tools for 3D depth animation, using depth-map transformations to create spatial movement. It also includes neural style transfer capabilities to apply specific artistic looks, such as watercolor or pixel art, and utilizes optical flow frame blending to reduce flickering

    Jupyter Notebook
    View on GitHub↗7,407
  • damo-nlp-sg/videollama3DAMO-NLP-SG avatar

    DAMO-NLP-SG/VideoLLaMA3

    1,105View on GitHub↗
    Jupyter Notebook
    View on GitHub↗1,105
  • baaivision/emu3baaivision avatar

    baaivision/Emu3

    2,417View on GitHub↗

    Next-Token Prediction is All You Need

    Python
    View on GitHub↗2,417
  • baaivision/uni3dbaaivision avatar

    baaivision/Uni3D

    674View on GitHub↗

    Uni3D: Exploring Unified 3D Representation at Scale

    Python
    View on GitHub↗674
  • baichuan-inc/baichuan-13bbaichuan-inc avatar

    baichuan-inc/Baichuan-13B

    2,931View on GitHub↗

    A 13B large language model developed by Baichuan Intelligent Technology

    Pythonartificial-intelligencebenchmarkceval
    View on GitHub↗2,931
  • baichuan-inc/baichuan-7bbaichuan-inc avatar

    baichuan-inc/Baichuan-7B

    5,654View on GitHub↗

    Baichuan-7B is an open-source 7 billion parameter bilingual Transformer model designed for text generation and few-shot learning across Chinese and English. It is built on a large Transformer architecture trained on a bilingual corpus, enabling it to produce coherent text in both languages from a single model. The model incorporates several optimization techniques that distinguish it from standard large language models. It uses rotary position embeddings that can extrapolate to longer sequences than seen during training, allowing context extension beyond the original 4096-token training lengt

    Pythonartificial-intelligencecevalchatgpt
    View on GitHub↗5,654
  • ai-chef/hugginggptAI-Chef avatar

    AI-Chef/HuggingGPT

    24View on GitHub↗

    The mission of JARVIS is to explore artificial general intelligence (AGI) and deliver cutting-edge research to the whole community.

    Python
    View on GitHub↗24
  • compvis/stable-diffusionCompVis avatar

    CompVis/stable-diffusion

    73,125View on GitHub↗

    Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I

    Jupyter Notebook
    View on GitHub↗73,125
  • chat-3d/chat-3dChat-3D avatar

    Chat-3D/Chat-3D

    57View on GitHub↗

    This is a repo for paper "Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes". paper, project page

    Python
    View on GitHub↗57
  • baai-dcai/visual-instruction-tuningBAAI-DCAI avatar

    BAAI-DCAI/Visual-Instruction-Tuning

    168View on GitHub↗

    Scale up visual instruction tuning to millions by GPT-4.

    Python
    View on GitHub↗168
  • abenechehab/adaptsabenechehab avatar

    abenechehab/AdaPTS

    51View on GitHub↗

    Our paper has been accepted at ICML 2025! 🎉📄🎓 👉 AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series Forecasting

    Jupyter Notebook
    View on GitHub↗51
  • cliport/cliportcliport avatar

    cliport/cliport

    545View on GitHub↗

    CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

    Jupyter Notebook
    View on GitHub↗545
  • camenduru/text-to-video-synthesis-colabcamenduru avatar

    camenduru/text-to-video-synthesis-colab

    1,515View on GitHub↗

    Text To Video Synthesis Colab

    Jupyter Notebookcolabcolab-notebookcolaboratory
    View on GitHub↗1,515
  • cvlab-columbia/vipercvlab-columbia avatar

    cvlab-columbia/viper

    1,717View on GitHub↗

    Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"

    Jupyter Notebook
    View on GitHub↗1,717
  • bowang-lab/scgptbowang-lab avatar

    bowang-lab/scGPT

    1,585View on GitHub↗

    This is the official codebase for scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI.

    Jupyter Notebookfoundation-modelgptsingle-cell
    View on GitHub↗1,585
  • databrickslabs/dollydatabrickslabs avatar

    databrickslabs/dolly

    10,795View on GitHub↗

    Dolly is an instruction-tuned large language model designed to follow complex natural language directions. It operates as a causal language model that predicts the next token in a sequence to generate coherent conversational responses and perform tasks such as brainstorming, classification, and question answering. The project focuses on the development of models using open datasets suitable for commercial application. It enables the creation of instruction-following models by utilizing curated collections of human-generated instruction-response pairs. The repository provides capabilities for

    Python
    View on GitHub↗10,795
  • datadog/totoDataDog avatar

    DataDog/toto

    482View on GitHub↗

    Toto 2.0: Model Weights | Blog Toto 1.0: Paper | Blog | Model Card

    Jupyter Notebook
    View on GitHub↗482
  • deepmind/deepmind-researchdeepmind avatar

    deepmind/deepmind-research

    15,024View on GitHub↗

    This project is an AI research implementation library and machine learning research repository. It provides a collection of reference code, illustrative implementations, and open-source research datasets used to verify hypotheses and build upon existing models in artificial intelligence. The repository focuses on scientific research reproduction by translating theoretical findings from published papers into executable code. It includes specialized scientific simulation environments designed to test the behavior of autonomous agents and models within controlled settings. The project covers AI

    Jupyter Notebook
    View on GitHub↗15,024
  • deepseek-ai/deepseek-llmdeepseek-ai avatar

    deepseek-ai/deepseek-LLM

    7,100View on GitHub↗

    DeepSeek-LLM is a large language model and causal language model designed for natural language generation. It functions as a multi-lingual system capable of predicting the next token in a sequence to perform text completion and conversational generation. The model is specialized for logical reasoning, specifically as a code and math LLM. This enables it to perform complex problem solving, which includes generating executable code and solving mathematical equations through step-by-step analysis. The system's broader capabilities cover conversational AI, including the generation of chat comple

    Makefile
    View on GitHub↗7,100
  • deepseek-ai/janusdeepseek-ai avatar

    deepseek-ai/Janus

    17,746View on GitHub↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    View on GitHub↗17,746
  • djiajunustc/3d-llavadjiajunustc avatar

    djiajunustc/3D-LLaVA

    98View on GitHub↗

    3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

    Python
    View on GitHub↗98
  • anjiecheng/spatialrgptAnjieCheng avatar

    AnjieCheng/SpatialRGPT

    330View on GitHub↗

    arxiv / Huggingface

    Python
    View on GitHub↗330