awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 个仓库

Awesome GitHub RepositoriesPhrase-Specific Isolation

Isolating specific objects by matching visual regions to high-similarity scores for individual words within a phrase.

Distinct from Object Detection: Moves beyond general object detection to target specific sub-phrases for precise isolation.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Phrase-Specific Isolation. Refine with filters or upvote what's useful.

Awesome Phrase-Specific Isolation GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • idea-research/groundingdinoIDEA-Research 的头像

    IDEA-Research/GroundingDINO

    9,738在 GitHub 上查看↗

    GroundingDINO is a deep learning vision model and open-vocabulary object detector designed to map natural language prompts to spatial coordinates. It functions as a text-to-bounding-box framework that enables zero-shot image localization, allowing the system to identify and locate arbitrary objects without requiring predefined classes or specific training for those categories. The project distinguishes itself by matching visual features to natural language descriptions to achieve open-set visual recognition. It supports text-guided image localization and the isolation of specific objects base

    Extracts precise object locations by targeting the highest text similarity scores for specific words within a sentence.

    Pythonobject-detectionopen-worldopen-world-detection
    在 GitHub 上查看↗9,738
  • facebookresearch/multimodalfacebookresearch 的头像

    facebookresearch/multimodal

    1,723在 GitHub 上查看↗

    Multimodal is a machine learning library built on PyTorch for training large-scale models that combine text, image, audio, and video data streams. It functions as a deep learning framework dedicated to generative diffusion models, multi-task training, and vision-language tasks. The library supplies modular building blocks, discrete latent codebook quantization, shared-space embeddings, and stackable adapter layers to handle diverse conditional inputs during training and inference. The framework supports specific architectures for diffusion models, text-to-video generation, image-text retrieva

    Locates and boxes specific regions in an image corresponding to noun phrases found in text queries.

    Python
    在 GitHub 上查看↗1,723
  1. Home
  2. Artificial Intelligence & ML
  3. Computer Vision Systems
  4. Computer Vision
  5. Object Detection and Tracking
  6. Object Detection
  7. Phrase-Specific Isolation