awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
huggingface avatar

huggingface/datatrove

0
View on GitHub↗
3,092 स्टार्स·273 फोर्क्स·Python·Apache-2.0·9 व्यूज़

Datatrove

Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.

Features

  • Data Curation and Filtering - Library for building scalable, platform-agnostic text processing pipelines.
  • Data Pipelines - Processes, filters, and deduplicates large-scale text data.
  • Data Processing - Library for large-scale text data processing and deduplication.
  • Data Processing Tools - Library for large-scale text data processing and deduplication.

स्टार हिस्ट्री

huggingface/datatrove के लिए स्टार हिस्ट्री चार्टhuggingface/datatrove के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Datatrove के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Datatrove के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • minishlab/semhashMinishLab का अवतार

    MinishLab/semhash

    936GitHub पर देखें↗

    Fast Multimodal Semantic Deduplication & Filtering

    Pythondatasetsdeduplicationimage-dataset-cleaning
    GitHub पर देखें↗936
  • jpmens/jojpmens का अवतार

    jpmens/jo

    4,868GitHub पर देखें↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    GitHub पर देखें↗4,868
  • argilla-io/distilabelargilla-io का अवतार

    argilla-io/distilabel

    3,277GitHub पर देखें↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    GitHub पर देखें↗3,277
  • 599yongyang/datasetloom5

    599yongyang/DatasetLoom

    0GitHub पर देखें↗
    GitHub पर देखें↗0
Datatrove के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

huggingface/datatrove क्या करता है?

Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.

huggingface/datatrove की मुख्य विशेषताएं क्या हैं?

huggingface/datatrove की मुख्य विशेषताएं हैं: Data Curation and Filtering, Data Pipelines, Data Processing, Data Processing Tools।

huggingface/datatrove के कुछ ओपन-सोर्स विकल्प क्या हैं?

huggingface/datatrove के ओपन-सोर्स विकल्पों में शामिल हैं: minishlab/semhash — Fast Multimodal Semantic Deduplication & Filtering. jpmens/jo — Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell… catchthetornado/pdf-extract-api. allenai/olmocr — Olmocr is a distributed document processing framework designed to convert PDF and image files into structured… bytedance/dolphin — Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital… 599yongyang/datasetloom.