40 रिपॉजिटरी
Collections of video and audio data for training and evaluating facial animation models.
Explore 40 awesome GitHub repositories matching part of an awesome list · Datasets. Refine with filters or upvote what's useful.
Alphabetical list of free/public domain datasets with text data for use in Natural Language Processing (NLP). Most stuff here is just raw unstructured text data, if you are looking for annotated corpora or Treebanks refer to the sources at the bottom.
Curated collection of various natural language processing datasets.
Github of the FaceForensics dataset
Large-scale dataset for facial manipulation and deepfake detection.
DeepFashion2 is a comprehensive fashion dataset. It contains 491K diverse images of 13 popular clothing categories from both commercial shopping stores and consumers. It totally has 801K clothing clothing items, where each item in an image is labeled with scale, occlusion, zoom-in, viewpoint,…
Comprehensive dataset for fashion-related image retrieval.
Offers an interactive 3D environment for visual AI research.
pip install medmnist 18x Standardized Datasets for 2D and 3D Biomedical Image Classification
Biomedical image classification datasets.
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces. Each reconstruction has clean dense geometry, high resolution and high dynamic range textures, glass and mirror surface information, planar segmentation as well as semantic class and instance…
Contains high-fidelity digital replicas of indoor spaces.
The Matterport3D V1.0 dataset contains data captured throughout 90 properties with a Matterport Pro Camera.
Provides large-scale RGB-D data for indoor environment learning.
Research datasets regularly disappear, change over time, become obsolete or come without a sane implementation to handle the data format reading and processing.
Repository for pretrained models and linguistic corpora.
Gibson Environments: एम्बॉडीड एजेंट्स के लिए रियल-वर्ल्ड परसेप्शन।
Delivers real-world perception data for training embodied agents.
Baca README ini dalam Bahasa Indonesia.
Benchmark suite for Indonesian natural language understanding.
This is a movie review dataset in the Korean language. Reviews were scraped from Naver Movies.
Korean sentiment analysis dataset based on movie reviews.
Stopwords for various languages in JSON format. Per Wikipedia:
JSON-formatted collection of multilingual stopword lists.
أكبر قائمة لمستبعدات الفهرسة العربية على جيت هاب
Comprehensive list of Arabic language stopwords.
This Tutorial contains installation instructions for the packages released along with Ford Multi AV Dataset. For more details please visit the website.
Time-stamped sensor data including LIDAR and calibration for autonomous driving.
CropHarvest is an open source remote sensing dataset for agriculture with benchmarks. It collects data from a variety of agricultural land use datasets and remote sensing products.
Remote sensing dataset for global crop type mapping.
State-of-the-Art Language Modeling and Text Classification in Hindi Language
Hindi language dataset for classification tasks.
This repository contains a pytorch implementation for the Interspeech 2024 paper, MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset. MultiTalk generates 3D talking head with enhanced multilingual performance.
Dataset for multi-modal talking head synthesis and interaction.
TalkingHead-1KH is a talking-head dataset consisting of YouTube videos, originally created as a benchmark for face-vid2vid:
High-quality dataset for talking head generation research.
Navigation and localisation dataset for self driving cars and autonomous robots.
Navigation and localization dataset for autonomous robots.
This repository contains State of the Art Language models and Classifier for Hindi language (spoken in Indian sub-continent).
Sentiment analysis dataset for Hindi language tasks.