awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
openimages avatar

openimages/dataset

0
View on GitHub↗
4,366 Stars·606 Forks·Python·Apache-2.0·8 Aufrufestorage.googleapis.com/openimages/web/index.html↗

Dataset

Dieses Projekt ist ein Computer-Vision-Datensatz und ein Repository für Bildannotationen, das für das Training und die Evaluierung von Machine-Learning-Modellen entwickelt wurde. Es bietet eine große Sammlung beschrifteter Bilder und dient als Benchmark für Objekterkennung sowie als Quelle für Pixel-Level-Segmentierungsdaten.

Das Repository zeichnet sich als multimodaler visueller Datensatz aus, indem es Bilder mit synchronisierten Sprach-, Text- und Mausspuren kombiniert, um das narrative Verständnis zu unterstützen. Es ermöglicht zudem die Analyse der Modellgerechtigkeit durch die Einbeziehung demografischer Attribute und erschöpfender Annotationen.

Der Datensatz deckt ein breites Spektrum an Computer-Vision-Funktionen ab, einschließlich Objekterkennung über Bounding-Boxen, Bildinstanzsegmentierung mittels Pixelmasken und visuelle Beziehungszuordnung durch Objekt-Attribut-Tripletts. Er unterstützt zudem punktbasierte Klassifizierung, hierarchische Texterkennung und den Abruf kuratierter Datensatz-Teilmengen basierend auf Klassen- oder Attributfilterung.

Features

  • Diverse Visual Datasets - Provides a broad collection of global images featuring diverse objects and people to improve model generalization.
  • Image Labeling - Provides millions of labeled images with bounding boxes and point locations to generate ground truth for computer vision.
  • Visual Object Grounding - Maps visual entities to geometric coordinates using bounding boxes and pixel-level masks for object localization.
  • Computer Vision Benchmarks - Serves as a standardized benchmark for computing precision and recall in object detection and classification models.
  • Instance Segmentation Engines - Provides pixel-level masks that delineate exact object boundaries to separate individual instances from the background.
  • Object Detection - Locates object instances using precise spatial bounding boxes across hundreds of classes.
  • Segmentation Object Management - Delineates exact pixel-level boundaries of individual objects using high-resolution masks.
  • Computer Vision Training - Provides a benchmark dataset and evaluation scripts for training supervised machine learning models for visual recognition.
  • Dense Visual Annotations - Provides detailed labels including bounding boxes, segmentation masks, and point-level annotations for image subsets.
  • Image Annotation Datasets - Provides an extensive library of bounding boxes, segmentation masks, and labels to teach models scene perception.
  • Multimedia Content Analyzers - Processes multimodal data, including synchronized voice and text, to provide detailed descriptions of visual content.
  • Image Segmentation Datasets - Provides detailed pixel masks and point-level annotations for high-precision instance and semantic segmentation.
  • Multimodal Narrative Datasets - Provides synchronized voice, text, and mouse traces paired with image regions to train narrative understanding models.
  • Multimodal Narrative Synchronizations - Pairs visual regions with synchronized voice, text, and mouse traces to link natural language to specific image areas.
  • Multimodal Visual Understanding - Integrates visual and language data, linking voice traces and narratives to image regions for complex reasoning.
  • Coordinate-Based Spatial Mappings - Represents visual entities as geometric coordinates using bounding boxes and pixel masks for spatial mapping.
  • Visual Relationship Modeling - Maps the relationships between different objects and their attributes within images to support complex scene understanding.
  • Visual Relationship Triplets - Represents scene interactions by linking two objects and their specific interaction or an object and its attribute.
  • Image Annotation - Provides a comprehensive repository of bounding boxes, masks, and relationship labels across thousands of classes.
  • Media Metadata Indexes - Indexes images against a structured hierarchy of categories and attributes to enable efficient dataset subset filtering.
  • Visual Relationship Triplets - Identifies triplets consisting of two objects and their interaction or an object and its physical property.
  • Dataset Category Filters - Provides a system to filter and retrieve image subsets based on a taxonomy of thousands of human-verified categories.
  • AI Ethics and Fairness - Includes demographic attributes and exhaustive annotations to evaluate bias and fairness in machine learning models.
  • Dataset Splitting Utilities - Includes utilities for dividing large-scale visual data into training, validation, and test sets based on label distribution.
  • Dataset Subset Extractions - Provides tools and scripts to extract specific subsets of the dataset based on classes, attributes, or metadata.
  • Pre-trained Model Zoos - Provides model checkpoints for image classification and object detection to facilitate immediate inference or fine-tuning.
  • Model Performance Evaluators - Enables quantification of model accuracy and reliability via mean Average Precision and precision-recall curves.
  • Vision Model Training - Supplies large-scale visual data including bounding boxes and instance segmentations for training supervised models.
  • Scene Text Recognition - Extracts hierarchical text annotations from images of natural scenes and documents at the word, line, and paragraph levels.
  • Point-Level Classifications - Identifies and categorizes specific pixels or regions using dense point-level annotations across thousands of classes.
  • Visual Relationship Datasets - Includes object-attribute triplets that map interactions and relationships between different visual entities in a scene.
  • Zero-Shot Segmentations - Assigns semantic labels to pixel coordinates to facilitate zero-shot or few-shot semantic segmentation.
  • Bias and Fairness - Includes demographic attributes and exhaustive annotations to help detect and mitigate biases in multimodal representations.
  • Multimodal Training Datasets - Pairs images with synchronized voice, text, and mouse traces to support multimodal narrative understanding.
  • Point-Level Semantic Annotations - Assigns discrete category labels to individual pixel coordinates to enable dense classification and zero-shot segmentation.
  • Hierarchical Text Annotations - Organizes optical character recognition data into nested levels of words, lines, and paragraphs for structured analysis.
  • Image Classifiers - Provides human-verified positive and negative labels across thousands of classes for image categorization.
  • Dataset Filtering - Allows isolating specific data slices based on object classes, annotation types, or dataset splits.
  • Batch Image Downloads - Enables the batch retrieval of raw images, thumbnails, and associated metadata like rotation and licensing.
  • Computer Vision Datasets - Large-scale image collection annotated with thousands of object categories.

Star-Verlauf

Star-Verlauf für openimages/datasetStar-Verlauf für openimages/dataset

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Dataset

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Dataset.
  • paddlepaddle/paddledetectionAvatar von PaddlePaddle

    PaddlePaddle/PaddleDetection

    14,243Auf GitHub ansehen↗

    PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of computer vision models. It provides a comprehensive library of modular neural network architectures and pipelines that support object detection, instance segmentation, and multi-object tracking tasks. The project distinguishes itself through a configuration-driven approach that decouples model components like backbones and heads, allowing for the flexible assembly of custom vision workflows. It incorporates advanced techniques such as anchor-free detection logic, joint detecti

    Pythonblazefacedeepsortdetr
    Auf GitHub ansehen↗14,243
  • wongkinyiu/yolov9Avatar von WongKinYiu

    WongKinYiu/yolov9

    9,534Auf GitHub ansehen↗

    YOLOv9 is a real-time computer vision framework and deep learning model designed for image classification, object detection, and instance segmentation. It functions as both a vision model and a trainer, allowing for the optimization of neural network weights on custom datasets using single or multiple GPUs. The framework utilizes programmable gradient information to perform high-speed identification and location of multiple objects within images and video streams. It extends beyond bounding box detection to provide instance segmentation and panoptic segmentation, which labels every pixel in a

    Pythonyolov9
    Auf GitHub ansehen↗9,534
  • open-mmlab/mmpretrainAvatar von open-mmlab

    open-mmlab/mmpretrain

    3,842Auf GitHub ansehen↗

    mmpretrain is a modular PyTorch computer vision framework designed for developing, training, and benchmarking deep learning architectures. It serves as a comprehensive toolkit for vision tasks, providing a specialized platform for multimodal machine learning and self-supervised learning. The project features a computer vision model zoo containing architectural definitions and pre-trained weights for backbones such as ViT, ConvNeXt, and Swin Transformer. It distinguishes itself through a dedicated self-supervised learning toolkit that implements algorithms like MAE and DINO to train models wit

    Pythonbeitclipconstrastive-learning
    Auf GitHub ansehen↗3,842
  • facebookresearch/detectron2Avatar von facebookresearch

    facebookresearch/detectron2

    34,548Auf GitHub ansehen↗

    Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying models for object detection, image segmentation, and visual recognition. It provides a research-oriented environment for training complex vision models with multi-GPU acceleration. The project includes a specialized object detection library for identifying and locating multiple objects via bounding boxes, as well as an image segmentation toolkit for creating pixel-level masks through instance, semantic, and panoptic segmentation. Additionally, it features a human pose estimati

    Python
    Auf GitHub ansehen↗34,548
Alle 30 Alternativen zu Dataset anzeigen→

Häufig gestellte Fragen

Was macht openimages/dataset?

Dieses Projekt ist ein Computer-Vision-Datensatz und ein Repository für Bildannotationen, das für das Training und die Evaluierung von Machine-Learning-Modellen entwickelt wurde. Es bietet eine große Sammlung beschrifteter Bilder und dient als Benchmark für Objekterkennung sowie als Quelle für Pixel-Level-Segmentierungsdaten.

Was sind die Hauptfunktionen von openimages/dataset?

Die Hauptfunktionen von openimages/dataset sind: Diverse Visual Datasets, Image Labeling, Visual Object Grounding, Computer Vision Benchmarks, Instance Segmentation Engines, Object Detection, Segmentation Object Management, Computer Vision Training.

Welche Open-Source-Alternativen gibt es zu openimages/dataset?

Open-Source-Alternativen zu openimages/dataset sind unter anderem: paddlepaddle/paddledetection — PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of… wongkinyiu/yolov9 — YOLOv9 is a real-time computer vision framework and deep learning model designed for image classification, object… open-mmlab/mmpretrain — mmpretrain is a modular PyTorch computer vision framework designed for developing, training, and benchmarking deep… facebookresearch/detectron2 — Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying… dmlc/gluon-cv — Gluon-CV is an MXNet computer vision library that provides a comprehensive collection of pre-implemented vision… apple/corenet — Corenet is a deep learning training framework and computer vision model library designed for developing neural…