awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

40 repositorios

Awesome GitHub RepositoriesDatasets

Collections of video and audio data for training and evaluating facial animation models.

Explore 40 awesome GitHub repositories matching part of an awesome list · Datasets. Refine with filters or upvote what's useful.

Awesome Datasets GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • niderhoff/nlp-datasetsAvatar de niderhoff

    niderhoff/nlp-datasets

    5,987Ver en GitHub↗

    Alphabetical list of free/public domain datasets with text data for use in Natural Language Processing (NLP). Most stuff here is just raw unstructured text data, if you are looking for annotated corpora or Treebanks refer to the sources at the bottom.

    Curated collection of various natural language processing datasets.

    Ver en GitHub↗5,987
  • ondyari/faceforensicsAvatar de ondyari

    ondyari/FaceForensics

    2,738Ver en GitHub↗

    Github of the FaceForensics dataset

    Large-scale dataset for facial manipulation and deepfake detection.

    Python
    Ver en GitHub↗2,738
  • switchablenorms/deepfashion2Avatar de switchablenorms

    switchablenorms/DeepFashion2

    2,615Ver en GitHub↗

    DeepFashion2 is a comprehensive fashion dataset. It contains 491K diverse images of 13 popular clothing categories from both commercial shopping stores and consumers. It totally has 801K clothing clothing items, where each item in an image is labeled with scale, occlusion, zoom-in, viewpoint,…

    Comprehensive dataset for fashion-related image retrieval.

    Jupyter Notebook
    Ver en GitHub↗2,615
  • allenai/ai2thorAvatar de allenai

    allenai/ai2thor

    1,668Ver en GitHub↗

    Offers an interactive 3D environment for visual AI research.

    C#artificial-intelligencecomputer-visioninteraction
    Ver en GitHub↗1,668
  • medmnist/medmnistAvatar de MedMNIST

    MedMNIST/MedMNIST

    1,374Ver en GitHub↗

    pip install medmnist 18x Standardized Datasets for 2D and 3D Biomedical Image Classification

    Biomedical image classification datasets.

    Python2d3dautoml
    Ver en GitHub↗1,374
  • facebookresearch/replica-datasetAvatar de facebookresearch

    facebookresearch/Replica-Dataset

    1,290Ver en GitHub↗

    The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces. Each reconstruction has clean dense geometry, high resolution and high dynamic range textures, glass and mirror surface information, planar segmentation as well as semantic class and instance…

    Contains high-fidelity digital replicas of indoor spaces.

    C++
    Ver en GitHub↗1,290
  • niessner/matterportAvatar de niessner

    niessner/Matterport

    1,224Ver en GitHub↗

    The Matterport3D V1.0 dataset contains data captured throughout 90 properties with a Matterport Pro Camera.

    Provides large-scale RGB-D data for indoor environment learning.

    C++
    Ver en GitHub↗1,224
  • rare-technologies/gensim-dataAvatar de RaRe-Technologies

    RaRe-Technologies/gensim-data

    1,054Ver en GitHub↗

    Research datasets regularly disappear, change over time, become obsolete or come without a sane implementation to handle the data format reading and processing.

    Repository for pretrained models and linguistic corpora.

    Python
    Ver en GitHub↗1,054
  • stanfordvl/gibsonenvAvatar de StanfordVL

    StanfordVL/GibsonEnv

    940Ver en GitHub↗

    Entornos Gibson: Percepción del mundo real para agentes encarnados

    Delivers real-world perception data for training embodied agents.

    C
    Ver en GitHub↗940
  • indobenchmark/indonluAvatar de indobenchmark

    indobenchmark/indonlu

    648Ver en GitHub↗

    Baca README ini dalam Bahasa Indonesia.

    Benchmark suite for Indonesian natural language understanding.

    Jupyter Notebook
    Ver en GitHub↗648
  • e9t/nsmcAvatar de e9t

    e9t/nsmc

    601Ver en GitHub↗

    This is a movie review dataset in the Korean language. Reviews were scraped from Naver Movies.

    Korean sentiment analysis dataset based on movie reviews.

    Python
    Ver en GitHub↗601
  • 6/stopwords-jsonAvatar de 6

    6/stopwords-json

    430Ver en GitHub↗

    Stopwords for various languages in JSON format. Per Wikipedia:

    JSON-formatted collection of multilingual stopword lists.

    JavaScript
    Ver en GitHub↗430
  • mohataher/arabic-stop-wordsAvatar de mohataher

    mohataher/arabic-stop-words

    333Ver en GitHub↗

    أكبر قائمة لمستبعدات الفهرسة العربية على جيت هاب

    Comprehensive list of Arabic language stopwords.

    Ver en GitHub↗333
  • ford/avdataAvatar de Ford

    Ford/AVData

    307Ver en GitHub↗

    This Tutorial contains installation instructions for the packages released along with Ford Multi AV Dataset. For more details please visit the website.

    Time-stamped sensor data including LIDAR and calibration for autonomous driving.

    C++
    Ver en GitHub↗307
  • nasaharvest/cropharvestAvatar de nasaharvest

    nasaharvest/cropharvest

    234Ver en GitHub↗

    CropHarvest is an open source remote sensing dataset for agriculture with benchmarks. It collects data from a variety of agricultural land use datasets and remote sensing products.

    Remote sensing dataset for global crop type mapping.

    Jupyter Notebook
    Ver en GitHub↗234
  • nirantk/hindi2vecAvatar de NirantK

    NirantK/hindi2vec

    219Ver en GitHub↗

    State-of-the-Art Language Modeling and Text Classification in Hindi Language

    Hindi language dataset for classification tasks.

    Jupyter Notebook
    Ver en GitHub↗219
  • postech-ami/multitalkAvatar de postech-ami

    postech-ami/MultiTalk

    196Ver en GitHub↗

    This repository contains a pytorch implementation for the Interspeech 2024 paper, MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset. MultiTalk generates 3D talking head with enhanced multilingual performance.

    Dataset for multi-modal talking head synthesis and interaction.

    Python
    Ver en GitHub↗196
  • tcwang0509/talkinghead-1khAvatar de tcwang0509

    tcwang0509/TalkingHead-1KH

    183Ver en GitHub↗

    TalkingHead-1KH is a talking-head dataset consisting of YouTube videos, originally created as a benchmark for face-vid2vid:

    High-quality dataset for talking head generation research.

    Python
    Ver en GitHub↗183
  • robotics-but/brno-urban-datasetAvatar de Robotics-BUT

    Robotics-BUT/Brno-Urban-Dataset

    164Ver en GitHub↗

    Navigation and localisation dataset for self driving cars and autonomous robots.

    Navigation and localization dataset for autonomous robots.

    Ver en GitHub↗164
  • goru001/nlp-for-hindiAvatar de goru001

    goru001/nlp-for-hindi

    123Ver en GitHub↗

    This repository contains State of the Art Language models and Classifier for Hindi language (spoken in Indian sub-continent).

    Sentiment analysis dataset for Hindi language tasks.

    Jupyter Notebook
    Ver en GitHub↗123
Ant.12Siguiente
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Datasets