awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
esbatmop avatar

esbatmop/MNBVC

0
View on GitHub↗
4,123 estrellas·287 forks·mit·2 vistas

MNBVC

MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers.

The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JSON or Markdown. It also features specialized alignment algorithms to create paired multilingual training datasets and a text data cleaning toolkit for character encoding detection and noise removal.

The software covers a broad range of data engineering tasks, including large-scale dataset cleaning, deduplication, and the normalization of dialogue and question-answer formats. It also includes security utilities for private information sanitization and sensitive content filtering to ensure data privacy and compliance.

The project further supports multimodal dataset construction and provides access to vast collections of internet-sourced Chinese text and raw PDF data.

Features

  • LLM Dataset Collection - Gathers large scale text, code, and academic datasets from web archives and repositories to train language models.
  • Chinese Natural Language Processing - Supplies a large collection of diverse internet-sourced Chinese data to train natural language processing models.
  • Chinese Text Corpora - Provides access to large collections of internet-sourced Chinese text across multiple domains for NLP model training.
  • Chinese Language Model Repositories - Supplies a vast collection of diverse internet-sourced Chinese text data to support model development.
  • Language Corpora - Provides access to cleaned internet-sourced Chinese text and QA datasets categorized by domain.
  • Semantic Deduplication - Provides advanced semantic deduplication and filtering to remove redundancy from large-scale AI training sets.
  • Multilingual Text Processing - Gathers large-scale multilingual text from forums, eBooks, academic papers, and web archives.
  • Parallel Text Alignment - Implements specialized alignment algorithms to create paired multilingual training datasets.
  • Parallel Text Alignments - Implements specialized alignment algorithms to create paired multilingual training datasets.
  • Programming Dataset Collection - Downloads source code, commit history, and issue data from diverse code hosting platforms.
  • Text Dataset Preparation - Cleans and formats large-scale text corpora by removing sensitive info, copyright notices, and advertisements.
  • Text Formatting Standardization - Standardizes content by converting traditional characters to simplified, removing whitespace, and fixing punctuation.
  • Text Normalization Tools - Transforms raw text from varied sources into a consistent and standard format for NLP training.
  • Document Processing and Conversion - Transforms complex PDF layouts and formulas into structured JSON or Markdown via a multi-stage conversion pipeline.
  • Conversation Format Normalization - Converts specialized test data into a consistent multi-turn conversation format.
  • Intra-Dataset Deduplication - Removes noise, boilerplate, and duplicate content from massive text corpora to improve data quality for machine learning.
  • Source Code Extractions - Parses repository archives and raw data to extract clean source code corpora into structured formats.
  • Dataset Cleaning - Provides a pipeline for cleaning, deduplicating, and normalizing massive text and code corpora for model training.
  • Massive Corpora Acquisition - Retrieves terabytes of internet-sourced Chinese text data through direct links and cloud storage.
  • Text Cleaning Pipelines - Provides a toolkit for removing duplicates, detecting character encodings, and filtering noise from scraped data.
  • Textual Deduplication - Identifies and deletes identical or near-identical files and sentences to reduce redundancy in the corpus.
  • Repository Metadata Harvesting - Collects repository metadata and clone addresses from multiple platforms for large-scale data acquisition.
  • Character Set Detection - Analyzes byte patterns and character frequencies to automatically identify and normalize text encodings.
  • Corpus Noise Filtering - Uses predefined patterns and machine learning classifiers to remove sensitive information and boilerplate noise.
  • Repository Corpus Extractors - Ships a specialized tool for harvesting source code and commit histories from version control platforms.
  • Extraction Pipelines - Extracts version control commit histories and source code using platform-specific APIs and raw archive parsing.
  • Chinese QA Datasets - Supplies a vast collection of internet-sourced Chinese question-and-answer pairs across diverse academic and professional domains.
  • Code Snippet Extraction - Recognizes and extracts code snippets from non-code documents such as textbooks.
  • QA Format Normalization - Converts diverse question-and-answer datasets into a standardized corpus format suitable for training language models.
  • Multimodal Document Processing - Extracts metadata and converts complex, mixed-media documents into structured formats like JSON and Parquet.
  • Web Content Noise Reduction - Filters noise from web-scraped data by removing boilerplate, malformed characters, and irrelevant identifiers.
  • PDF Text Extraction - Converts complex PDF documents and academic papers into structured text or JSON for data analysis.
  • PDF to Markdown Converters - Transforms complex PDF layouts and formulas into structured Markdown or JSON formats.
  • Data Categorization - Organizes collected corpora into hierarchical taxonomies based on language, genre, and content type.
  • Data Extraction - Cleans and parses raw data from specialized sources including legal documents and academic papers.
  • Taxonomies - Organizes scraped content into hierarchical structures based on language, genre, and domain for dataset management.
  • Parallel Translation Corpora - Provides access to multilingual datasets where Chinese text is aligned with equivalent translations in multiple other languages.
  • Document Classification - Categorizes PDF files based on language, size, or metadata using classification algorithms.
  • Multimodal Dataset Loaders - Extracts text, audio, video, and academic paper data to create diverse multimodal datasets for model training.
  • Raw Document Retrieval - Supplies vast collections of internet-sourced Chinese text and raw PDF data for training models.
  • Document-to-Text Conversions - Transforms legacy document formats into clean text through intermediate conversion and processing steps.
  • Content Filtering Rules - Detects and removes prohibited material using rule-based filters and machine learning classifiers.
  • Data Anonymization - Identifies and removes personally identifiable information and prohibited content to ensure data privacy compliance.
  • Data Sanitization - Scans text to identify and remove personally identifiable information such as ID numbers and email addresses.
  • Academic Paper Crawlers - Collects research papers and source files from online archives to build comprehensive datasets of scholarly work.
  • Llama Model Ecosystem - Massive cleaned Chinese dataset for training and evaluating language models.
  • Pretraining and Large Corpora - Large-scale, continuously updated Chinese pretraining corpus.
  • Pretraining Corpora - Large-scale Chinese pretraining corpus with continuous updates.
  • Datasets and Corpora - Massive collection of cleaned Chinese text data for training.
  • Pre-training Datasets - Massive, diverse Chinese text corpus from internet sources.

Historial de estrellas

Gráfico del historial de estrellas de esbatmop/mnbvcGráfico del historial de estrellas de esbatmop/mnbvc

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a MNBVC

Proyectos open-source similares, clasificados según cuántas características comparten con MNBVC.
  • brightmart/nlp_chinese_corpusAvatar de brightmart

    brightmart/nlp_chinese_corpus

    9,903Ver en GitHub↗

    This is a large-scale collection of curated Chinese text corpora designed for training natural language processing models. The project provides a variety of datasets, including a deduplicated archive of millions of news articles with titles and keywords, high-quality categorized question-and-answer pairs, and parallel translation corpora. The collection includes millions of aligned Chinese and English sentence pairs used for cross-lingual model training and machine translation development. It also contains filtered question-and-answer data organized by label for the construction of knowledge-

    bertchinesechinese-corpus
    Ver en GitHub↗9,903
  • mikechongcan/scyllaAvatar de MikeChongCan

    MikeChongCan/scylla

    4,019Ver en GitHub↗

    Scylla is a system for managing HTTP proxy pools and automating web extraction. It provides a specialized data acquisition pipeline designed for gathering large-scale internet datasets for training and fine-tuning large language models. The project features a proxy rotation gateway that assigns fresh proxy addresses to incoming requests to mask origin traffic and avoid IP blocking. It includes a proxy pool manager that handles the collection, functional validation, and orchestration of proxy servers, complemented by a web dashboard for monitoring the health and geographic distribution of the

    Pythoncrawlerproxy-poolpython
    Ver en GitHub↗4,019
  • sloria/textblobAvatar de sloria

    sloria/TextBlob

    9,516Ver en GitHub↗

    TextBlob is a natural language processing library that provides a unified interface for common linguistic tasks. It operates as a wrapper-based API, simplifying the use of complex processing libraries by delegating core operations to specialized external frameworks. The project features a pluggable processing pipeline that allows for the integration of custom logic and alternative language engines. It supports the extension of processing models through plugins to add specific language support or custom data processing. The library covers a broad range of linguistic capabilities, including se

    Pythonnatural-language-processingnlpnltk
    Ver en GitHub↗9,516
  • facico/chinese-vicunaAvatar de Facico

    Facico/Chinese-Vicuna

    4,121Ver en GitHub↗

    Chinese-Vicuna is a Chinese large language model and instruction-following AI based on the LLaMA architecture. It is specifically designed for natural language understanding and generation in the Chinese language, utilizing an instruction-tuned model to follow complex user prompts across conversations. The project provides a LoRA fine-tuning framework and quantization systems to enable model adaptation and inference on consumer hardware. It implements quantized inference to reduce memory usage on both CPUs and GPUs, supported by a low-level C++ implementation to minimize system resource requi

    Calpacachinesellama
    Ver en GitHub↗4,121
Ver las 30 alternativas a MNBVC→

Preguntas frecuentes

¿Qué hace esbatmop/mnbvc?

MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers.

¿Cuáles son las características principales de esbatmop/mnbvc?

Las características principales de esbatmop/mnbvc son: LLM Dataset Collection, Chinese Natural Language Processing, Chinese Text Corpora, Chinese Language Model Repositories, Language Corpora, Semantic Deduplication, Multilingual Text Processing, Parallel Text Alignment.

¿Qué alternativas de código abierto existen para esbatmop/mnbvc?

Las alternativas de código abierto para esbatmop/mnbvc incluyen: brightmart/nlp_chinese_corpus — This is a large-scale collection of curated Chinese text corpora designed for training natural language processing… mikechongcan/scylla — Scylla is a system for managing HTTP proxy pools and automating web extraction. It provides a specialized data… sloria/textblob — TextBlob is a natural language processing library that provides a unified interface for common linguistic tasks. It… facico/chinese-vicuna — Chinese-Vicuna is a Chinese large language model and instruction-following AI based on the LLaMA architecture. It is… grangier/python-goose — python-goose is a Python library for web scraping and content extraction. It functions as an HTML boilerplate remover… opendcai/dataflow — DataFlow is an agent-based workflow orchestrator and data pipeline designed to synthesize, clean, and augment…