For data labeling interfaces, the strongest matches are inception-project/inception (INCEpTION is a semantic annotation platform tailored for text), jiesutd/yedda (YEDDA is a lightweight collaborative text annotation tool designed) and rtiinternational/smart (This repository provides a self-hostable tool for text and). heartexlabs/label-studio and chakki-works/doccano round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Hand-picked data labeling tools to annotate images, text, and audio. Compare top open-source GitHub repositories to find the best fit.
A semantic annotation platform offering intelligent assistance and knowledge management. Homepage · Usage · Demo · FAQ
INCEpTION is a semantic annotation platform tailored for text and NLP tasks, offering collaborative workflows and intelligent assistance, making it a strong fit for text labeling despite lacking image or video capabilities.
YEDDA: A Lightweight Collaborative Text Span Annotation Tool. Code for ACL 2018 Best Demo Paper Nomination.
YEDDA is a lightweight collaborative text annotation tool designed for manual and entity span labeling, fitting the text-annotation aspect of your search though it lacks image, audio, or video capabilities.
Smarter Manual Annotation for Resource-constrained collection of Training data
This repository provides a self-hostable tool for text and image annotation to assist in training machine learning models, fitting the core category despite lacking audio, video, and advanced active learning features.
Label Studio is a multi-type data labeling tool and data annotation workspace designed to prepare datasets for machine learning training. It functions as a cloud-integrated data pipeline that imports raw data from storage, manages the annotation process, and exports labels into standardized formats. The platform features a machine learning model integration framework that connects to external model servers. This enables model-assisted annotation and active learning, allowing the system to perform pre-labeling and refine predictions based on human feedback. The software provides project manag
Label Studio is a self-hostable, collaborative data labeling platform supporting multimodal annotations like images, text, audio, and video alongside active learning and customizable interfaces.
Doccano is a collaborative labeling platform and text annotation tool designed to create training data for machine learning. It provides a specialized interface for performing sequence labeling and text classification on natural language datasets. The system functions as a supervised learning dataset manager, allowing multiple users to coordinate within a shared workspace to label datasets for natural language processing tasks. It supports the preparation of raw text data for model training by converting unstructured documents into structured labeled examples. The platform includes capabilit
Doccano is a collaborative labeling platform built for natural language processing, matching the intent for data annotation tools despite focusing solely on text rather than images or video.
CVAT is an open-source, web-based platform designed for annotating images, videos, and 3D point clouds to create high-quality training datasets for machine learning. It functions as a containerized server that orchestrates the entire lifecycle of computer vision data, from initial task creation and manual labeling to quality assurance and final dataset export. The platform distinguishes itself through deep integration with machine learning models, allowing users to deploy custom AI models as serverless functions for automated object detection, tracking, and skeleton annotation. It supports co
CVAT is a comprehensive, self-hostable data annotation platform supporting image, video, and 3D data with collaborative workflows and automated AI-assisted labeling.
Doccano is a collaborative data labeling platform and machine learning dataset management system. It provides a web-based interface for teams to import raw text, mark datasets, and export structured annotations for model training. The project specifically supports text annotation for classification and named entity recognition tasks. It enables teams to coordinate multiple users on a single project to maintain consistent labeling guidelines and increase the speed of dataset creation. The system includes tools for data management and team coordination, providing the ability to import raw data
Doccano is a collaborative web-based data labeling platform tailored specifically for text annotation tasks, making it a great fit for team-based text projects despite lacking native image and audio annotation features.
Label Studio is a multi-modal data annotation platform designed to create and manage high-quality training datasets for machine learning. It functions as a self-hosted, containerized environment that supports secure, private deployments, including air-gapped configurations. The platform provides a centralized workspace for labeling diverse media types, such as images, text, audio, and time-series data, to support supervised and reinforcement learning workflows. The platform distinguishes itself through deep integration with machine learning backends, enabling active learning loops, automated
Label Studio is a self-hostable multi-modal data annotation platform that provides collaborative workspaces and active learning workflows for images, text, audio, and video training data.
  Quick Start   |   Documentation   |   Join Slack  
Label Sleuth is a self-hostable text annotation platform featuring active learning workflows tailored for building machine learning models, though it is limited to text data and lacks support for images, audio, or video.
Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop dataset platform designed to coordinate workforce annotators and domain experts in labeling, rating, and refining data samples for machine learning projects. The platform focuses on large language model dataset curation and reinforcement learning from human feedback workflows. It provides a shared workspace for integrating human expertise into AI development to validate model outputs and correct data errors. The system manages the end-to-end machine learning data pipeline, includ
Argilla is a collaborative AI feedback tool and data curation platform designed for human-in-the-loop data labeling and reinforcement learning workflows, directly fulfilling the requirements for an annotating and labeling tool.
CVAT is an open-source computer vision annotation tool and visual dataset management platform. It provides a self-hosted interface for labeling images, videos, and 3D data to create datasets for vision AI models. The platform features AI-assisted data labeling to automate the creation of masks and bounding boxes, utilizing a plug-in system to connect external machine learning models. It includes a consensus-based quality assurance system that verifies label accuracy by comparing independent annotations. The system covers collaborative team management, project organization through task decomp
CVAT is an open-source, self-hosted data labeling platform that provides collaborative workflows, AI-assisted image and video annotation, and quality assurance tools for computer vision training.
Labelme is a Python-based image annotation tool used to create computer vision datasets. It serves as a visual editor for semantic segmentation, allowing users to define object boundaries using polygons, rectangles, points, and circles. The application also functions as a multispectral image annotator, supporting high-bit depth TIFF files used in satellite and scientific imagery. The tool incorporates AI-assisted labeling capabilities to automate the creation of masks and polygons. These features allow for shape generation driven by text prompts or interactive point selections, which propose
Labelme is a Python-based image annotation tool specialized in image segmentation and computer vision datasets, though it lacks text, audio, and video annotation features.
VoTT is a computer vision annotation software and machine learning dataset preparation tool. It is a desktop application designed for drawing bounding boxes and assigning tags to objects in images and videos to create training datasets for object detection models. The application utilizes a cross-platform desktop interface to manage image and video assets. It features a local-first storage integration to handle large media assets directly from the host machine's file system and includes frame-rate controlled video sampling to extract specific images from video streams for labeling. The softw
VoTT is a desktop application for annotating images and videos with bounding boxes to prepare computer vision training datasets, which fits the category well despite missing text and audio modalities.
Cloud Annotations is a web-based platform designed for collaborative image annotation and the preparation of computer vision datasets. It provides an interface for teams to draw bounding boxes and polygons over digital media, transforming raw images into structured training data for machine learning models. The platform distinguishes itself through a real-time synchronization engine that allows multiple users to edit the same image simultaneously. By utilizing browser-based local storage and standardized data serialization, it supports offline workflows and ensures that exported annotations r
Cloud Annotations is a web-based collaborative image and vision dataset annotation tool that supports team workflows, though it lacks text, audio, and video labeling features.
labelImg is a computer vision labeling tool and image bounding box annotator used to create training datasets for machine learning models. It functions as a desktop utility for drawing rectangular labels on images and saving object coordinates and class names in common machine learning formats. The tool is specifically designed to generate and edit PascalVOC formatted XML files and create image labels in the text-based format required by YOLO object detection pipelines. The software covers object detection annotation and training data preparation, including the ability to manage label catego
LabelImg is a dedicated desktop utility for image bounding box annotation and dataset preparation, though its narrow focus on rectangular bounding boxes means it lacks advanced features like segmentation, video support, or collaborative workflows.
labelImg is a desktop image annotation tool and dataset preparation utility used to create labeled datasets for computer vision training. It provides a graphical interface for drawing bounding boxes around objects in images and assigning them class labels to build ground truth data for machine learning models. The software specifically supports the Pascal VOC XML annotation format, exporting image coordinates and class names into standard XML or text structures. It allows users to load predefined class lists from text files to standardize naming across an entire project. Beyond initial label
This repository is a desktop image annotation tool used to create bounding-box datasets for computer vision training, making it a valid option within the category even though it is limited to images and lacks collaborative or video features.
FiftyOne is a visual tool for curating, analyzing, and managing image and video datasets for machine learning model training. It serves as a platform for identifying annotation errors, refining ground truth labels, and evaluating vision model performance by comparing predictions against ground truth to identify failure modes. The system functions as a containerized data platform that supports team collaboration on large-scale visual datasets in a cloud environment. It includes specialized capabilities for exploring high-dimensional embeddings to discover data clusters and retrieve correspondi
FiftyOne is a visual dataset curation and evaluation platform tailored for machine learning, featuring image and video analysis capabilities that support annotation project management and collaboration, though it focuses more on dataset exploration and error identification than a dedicated end-to-end labeling tool.
Source code for the LabelMe annotation tool.
This repository provides the source code for the classic LabelMe annotation tool, which fits the data labeling category for images even though it lacks advanced features like active learning or collaborative workflows.
Label App is a free, simple application designed to assist in manually editing, visualizing and labelling your moderate-sized datasets.
Label App is a self-contained data labeling tool for manual dataset annotation, fitting the category well despite having a narrower feature set focused on moderate-sized datasets.
| 仓库 | Star 数 | 语言 | 许可证 | 最后推送 |
|---|---|---|---|---|
| inception-project/inception | 700 | Java | Apache-2.0 | |
| jiesutd/yedda | 1.1K | Python | Apache-2.0 | |
| rtiinternational/smart | 230 | Python | MIT | |
| heartexlabs/label-studio | 27.6K | TypeScript | Apache-2.0 | |
| chakki-works/doccano | 10.7K | Python | MIT | |
| cvat-ai/cvat | 15.3K | Python | mit | |
| doccano/doccano | 10.7K | Python | MIT | |
| humansignal/label-studio | 27.6K | TypeScript | Apache-2.0 | |
| label-sleuth/label-sleuth | 273 | Python | Apache-2.0 | |
| argilla-io/argilla | 5K | Python | Apache-2.0 |