awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mobvoi avatar

mobvoi/seq-monkey-data

0
View on GitHub↗
178 estrellas·8 forks·Apache-2.0·7 vistas

Seq Monkey Data

Features

  • Pre-training Datasets - Large-scale dataset used for training general-purpose language models.

Historial de estrellas

Gráfico del historial de estrellas de mobvoi/seq-monkey-dataGráfico del historial de estrellas de mobvoi/seq-monkey-data

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Seq Monkey Data

Proyectos open-source similares, clasificados según cuántas características comparten con Seq Monkey Data.
  • plexpt/chatgpt-corpusAvatar de PlexPt

    PlexPt/chatgpt-corpus

    964Ver en GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    Ver en GitHub↗964
  • fuxiaoliu/lrv-instructionAvatar de FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297Ver en GitHub↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    Ver en GitHub↗297
  • guoyang9/unk-vqaAvatar de guoyang9

    guoyang9/UNK-VQA

    7Ver en GitHub↗

    A VQA dataset that includes unanswerable questions TPAMI 2024.

    Ver en GitHub↗7
  • esbatmop/mnbvcAvatar de esbatmop

    esbatmop/MNBVC

    4,123Ver en GitHub↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    Ver en GitHub↗4,123
Ver las 27 alternativas a Seq Monkey Data→

Preguntas frecuentes

¿Cuáles son las características principales de mobvoi/seq-monkey-data?

Las características principales de mobvoi/seq-monkey-data son: Pre-training Datasets.

¿Qué alternativas de código abierto existen para mobvoi/seq-monkey-data?

Las alternativas de código abierto para mobvoi/seq-monkey-data incluyen: plexpt/chatgpt-corpus — This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of… fuxiaoliu/lrv-instruction — [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. guoyang9/unk-vqa — A VQA dataset that includes unanswerable questions [TPAMI 2024]. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. icoz69/stablellava — Official repo for StableLLAVA. esbatmop/mnbvc — MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and…