awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
openimages avatar

openimages/dataset

0
View on GitHub↗
4,366 stars·606 forks·Python·Apache-2.0·13 viewsstorage.googleapis.com/openimages/web/index.html↗

Dataset

This project is a computer vision dataset and image annotation repository designed for training and evaluating machine learning models. It provides a large collection of labeled images, serving as an object detection benchmark and a source of pixel-level segmentation data.

The repository distinguishes itself as a multimodal visual dataset by pairing images with synchronized voice, text, and mouse traces to support narrative understanding. It further enables the analysis of model fairness through the inclusion of demographic attributes and exhaustive annotations.

The dataset covers a broad range of computer vision capabilities, including object detection via bounding boxes, image instance segmentation using pixel masks, and visual relationship mapping through object-attribute triplets. It also supports point-level classification, hierarchical text recognition, and the retrieval of curated dataset subsets based on class or attribute filtering.

Features

  • Diverse Visual Datasets - Provides a broad collection of global images featuring diverse objects and people to improve model generalization.
  • Image Labeling - Provides millions of labeled images with bounding boxes and point locations to generate ground truth for computer vision.
  • Visual Object Grounding - Maps visual entities to geometric coordinates using bounding boxes and pixel-level masks for object localization.
  • Computer Vision Benchmarks - Serves as a standardized benchmark for computing precision and recall in object detection and classification models.
  • Instance Segmentation Engines - Provides pixel-level masks that delineate exact object boundaries to separate individual instances from the background.
  • Object Detection - Locates object instances using precise spatial bounding boxes across hundreds of classes.
  • Segmentation Object Management - Delineates exact pixel-level boundaries of individual objects using high-resolution masks.
  • Computer Vision Training - Provides a benchmark dataset and evaluation scripts for training supervised machine learning models for visual recognition.
  • Dense Visual Annotations - Provides detailed labels including bounding boxes, segmentation masks, and point-level annotations for image subsets.
  • Image Annotation Datasets - Provides an extensive library of bounding boxes, segmentation masks, and labels to teach models scene perception.
  • Multimedia Content Analyzers - Processes multimodal data, including synchronized voice and text, to provide detailed descriptions of visual content.
  • Image Segmentation Datasets - Provides detailed pixel masks and point-level annotations for high-precision instance and semantic segmentation.
  • Multimodal Narrative Datasets - Provides synchronized voice, text, and mouse traces paired with image regions to train narrative understanding models.
  • Multimodal Narrative Synchronizations - Pairs visual regions with synchronized voice, text, and mouse traces to link natural language to specific image areas.
  • Multimodal Visual Understanding - Integrates visual and language data, linking voice traces and narratives to image regions for complex reasoning.
  • Coordinate-Based Spatial Mappings - Represents visual entities as geometric coordinates using bounding boxes and pixel masks for spatial mapping.
  • Visual Relationship Modeling - Maps the relationships between different objects and their attributes within images to support complex scene understanding.
  • Visual Relationship Triplets - Represents scene interactions by linking two objects and their specific interaction or an object and its attribute.
  • Image Annotation - Provides a comprehensive repository of bounding boxes, masks, and relationship labels across thousands of classes.
  • Media Metadata Indexes - Indexes images against a structured hierarchy of categories and attributes to enable efficient dataset subset filtering.
  • Visual Relationship Triplets - Identifies triplets consisting of two objects and their interaction or an object and its physical property.
  • Dataset Category Filters - Provides a system to filter and retrieve image subsets based on a taxonomy of thousands of human-verified categories.
  • AI Ethics and Fairness - Includes demographic attributes and exhaustive annotations to evaluate bias and fairness in machine learning models.
  • Dataset Splitting Utilities - Includes utilities for dividing large-scale visual data into training, validation, and test sets based on label distribution.
  • Dataset Subset Extractions - Provides tools and scripts to extract specific subsets of the dataset based on classes, attributes, or metadata.
  • Pre-trained Model Zoos - Provides model checkpoints for image classification and object detection to facilitate immediate inference or fine-tuning.
  • Model Performance Evaluators - Enables quantification of model accuracy and reliability via mean Average Precision and precision-recall curves.
  • Vision Model Training - Supplies large-scale visual data including bounding boxes and instance segmentations for training supervised models.
  • Scene Text Recognition - Extracts hierarchical text annotations from images of natural scenes and documents at the word, line, and paragraph levels.
  • Point-Level Classifications - Identifies and categorizes specific pixels or regions using dense point-level annotations across thousands of classes.
  • Visual Relationship Datasets - Includes object-attribute triplets that map interactions and relationships between different visual entities in a scene.
  • Zero-Shot Segmentations - Assigns semantic labels to pixel coordinates to facilitate zero-shot or few-shot semantic segmentation.
  • Bias and Fairness - Includes demographic attributes and exhaustive annotations to help detect and mitigate biases in multimodal representations.
  • Multimodal Training Datasets - Pairs images with synchronized voice, text, and mouse traces to support multimodal narrative understanding.
  • Point-Level Semantic Annotations - Assigns discrete category labels to individual pixel coordinates to enable dense classification and zero-shot segmentation.
  • Hierarchical Text Annotations - Organizes optical character recognition data into nested levels of words, lines, and paragraphs for structured analysis.
  • Image Classifiers - Provides human-verified positive and negative labels across thousands of classes for image categorization.
  • Dataset Filtering - Allows isolating specific data slices based on object classes, annotation types, or dataset splits.
  • Batch Image Downloads - Enables the batch retrieval of raw images, thumbnails, and associated metadata like rotation and licensing.
  • Computer Vision Datasets - Large-scale image collection annotated with thousands of object categories.

Star history

Star history chart for openimages/datasetStar history chart for openimages/dataset

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does openimages/dataset do?

This project is a computer vision dataset and image annotation repository designed for training and evaluating machine learning models. It provides a large collection of labeled images, serving as an object detection benchmark and a source of pixel-level segmentation data.

What are the main features of openimages/dataset?

The main features of openimages/dataset are: Diverse Visual Datasets, Image Labeling, Visual Object Grounding, Computer Vision Benchmarks, Instance Segmentation Engines, Object Detection, Segmentation Object Management, Computer Vision Training.

What are some open-source alternatives to openimages/dataset?

Open-source alternatives to openimages/dataset include: paddlepaddle/paddledetection — PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of… wongkinyiu/yolov9 — YOLOv9 is a real-time computer vision framework and deep learning model designed for image classification, object… open-mmlab/mmpretrain — mmpretrain is a modular PyTorch computer vision framework designed for developing, training, and benchmarking deep… facebookresearch/detectron2 — Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying… dmlc/gluon-cv — Gluon-CV is an MXNet computer vision library that provides a comprehensive collection of pre-implemented vision… apple/corenet — Corenet is a deep learning training framework and computer vision model library designed for developing neural…

Open-source alternatives to Dataset

Similar open-source projects, ranked by how many features they share with Dataset.
  • paddlepaddle/paddledetectionPaddlePaddle avatar

    PaddlePaddle/PaddleDetection

    14,243View on GitHub↗

    PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of computer vision models. It provides a comprehensive library of modular neural network architectures and pipelines that support object detection, instance segmentation, and multi-object tracking tasks. The project distinguishes itself through a configuration-driven approach that decouples model components like backbones and heads, allowing for the flexible assembly of custom vision workflows. It incorporates advanced techniques such as anchor-free detection logic, joint detecti

    Pythonblazefacedeepsortdetr
    View on GitHub↗14,243
  • wongkinyiu/yolov9WongKinYiu avatar

    WongKinYiu/yolov9

    9,534View on GitHub↗

    YOLOv9 is a real-time computer vision framework and deep learning model designed for image classification, object detection, and instance segmentation. It functions as both a vision model and a trainer, allowing for the optimization of neural network weights on custom datasets using single or multiple GPUs. The framework utilizes programmable gradient information to perform high-speed identification and location of multiple objects within images and video streams. It extends beyond bounding box detection to provide instance segmentation and panoptic segmentation, which labels every pixel in a

    Pythonyolov9
    View on GitHub↗9,534
  • open-mmlab/mmpretrainopen-mmlab avatar

    open-mmlab/mmpretrain

    3,842View on GitHub↗

    mmpretrain is a modular PyTorch computer vision framework designed for developing, training, and benchmarking deep learning architectures. It serves as a comprehensive toolkit for vision tasks, providing a specialized platform for multimodal machine learning and self-supervised learning. The project features a computer vision model zoo containing architectural definitions and pre-trained weights for backbones such as ViT, ConvNeXt, and Swin Transformer. It distinguishes itself through a dedicated self-supervised learning toolkit that implements algorithms like MAE and DINO to train models wit

    Pythonbeitclipconstrastive-learning
    View on GitHub↗3,842
  • facebookresearch/detectron2facebookresearch avatar

    facebookresearch/detectron2

    34,548View on GitHub↗

    Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying models for object detection, image segmentation, and visual recognition. It provides a research-oriented environment for training complex vision models with multi-GPU acceleration. The project includes a specialized object detection library for identifying and locating multiple objects via bounding boxes, as well as an image segmentation toolkit for creating pixel-level masks through instance, semantic, and panoptic segmentation. Additionally, it features a human pose estimati

    Python
    View on GitHub↗34,548
  • See all 30 alternatives to Dataset→