awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to lxtgh/omg-seg

Projects sharing features with OMG Seg

13 open-source projects similar to lxtgh/omg-seg, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • atcold/pytorch-deep-learning-minicourseAtcold avatar

    Atcold/pytorch-Deep-Learning-Minicourse

    6,810View on GitHub↗

    This is an educational curriculum for building and training neural networks using PyTorch. It serves as a deep learning training guide and resource, providing a structured series of lessons on tensor computation and architecture development. The course uses an interactive learning model that synchronizes academic theory with practice. It pairs theoretical lecture slides with exercise-driven notebooks, requiring students to implement model logic within predefined templates to validate their conceptual understanding. The curriculum covers a broad range of deep learning capabilities, including

    Jupyter Notebook
    View on GitHub↗6,810
  • huggingface/coursehuggingface avatar

    huggingface/course

    3,715View on GitHub↗

    This project is an educational course and learning curriculum for implementing and fine-tuning transformer models using the Hugging Face ecosystem. It serves as a structured guide and technical walkthrough for processing multimodal data, adapting pre-trained neural networks, and deploying models. The material includes a guide for managing, versioning, and distributing model weights and datasets through a centralized asset hub. It also provides a practical tutorial on adapting models to specific datasets using parameter-efficient methods and an implementation guide for solving natural language

    MDXdeep-learninghacktoberfestnlp
    View on GitHub↗3,715
  • microsoft/llava-medmicrosoft avatar

    microsoft/LLaVA-Med

    2,214View on GitHub↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    Python
    View on GitHub↗2,214
  • openbmb/minicpm-oOpenBMB avatar

    OpenBMB/MiniCPM-o

    23,850View on GitHub↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Pythonminicpmminicpm-vmulti-modal
    View on GitHub↗23,850

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • openbmb/minicpm-vOpenBMB avatar

    OpenBMB/MiniCPM-V

    25,653View on GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Python
    View on GitHub↗25,653
  • openrobotlab/pointllmOpenRobotLab avatar

    OpenRobotLab/PointLLM

    1,026View on GitHub↗

    ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds

    Python
    View on GitHub↗1,026
  • tencent/vitaTencent avatar

    Tencent/VITA

    159View on GitHub↗

    The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.

    Python
    View on GitHub↗159
  • vita-mllm/vitaVITA-MLLM avatar

    VITA-MLLM/VITA

    2,518View on GitHub↗

    ✨✨NeurIPS 2025 VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

    Python
    View on GitHub↗2,518
  • weihuanglin/inf-llavaWeihuangLin avatar

    WeihuangLin/INF-LLaVA

    42View on GitHub↗

    INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

    Python
    View on GitHub↗42
  • wentaoyuan/robopointwentaoyuan avatar

    wentaoyuan/RoboPoint

    224View on GitHub↗

    A Vision-Language Model for Spatial Affordance Prediction in Robotics

    Python
    View on GitHub↗224
  • yuliang-liu/monkeyYuliang-Liu avatar

    Yuliang-Liu/Monkey

    1,948View on GitHub↗

    Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)

    Python
    View on GitHub↗1,948
  • zzzhang-jx/dockylinZZZHANG-jx avatar

    ZZZHANG-jx/DocKylin

    36View on GitHub↗

    AAAI 2025 DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

    Python
    View on GitHub↗36
  • dvlab-research/lisadvlab-research avatar

    dvlab-research/LISA

    2,649View on GitHub↗

    Project Page for "LISA: Reasoning Segmentation via Large Language Model"

    Python
    View on GitHub↗2,649