awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Document deduplication

排名更新于 2026年7月23日

For document duplication, the strongest matches are arsenetar/dupeguru (DupeGuru is a content-based file deduplicator that finds and), qarmin/czkawka (Czkawka is a cross-platform duplicate file detector that uses) and google-research/deduplicate-text-datasets (This repository provides specialized tools to detect and remove). Each is ranked by relevance to your query, popularity and recent activity.

Compare the best open-source document deduplication tools on GitHub, ranked by stars and activity, and find the right fit.

Document deduplication

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • arsenetar/dupeguruarsenetar 的头像

    arsenetar/dupeguru

    7,631在 GitHub 上查看↗

    dupeguru is a content-based file deduplicator and cross-platform disk cleanup tool. It functions as a local file management utility designed to identify and remove redundant files to recover disk space. The application identifies identical files across a file system by using content hashing and metadata comparison. This allows it to detect duplicates regardless of their filenames or directory locations. The software covers a range of data management capabilities, including directory scanning, content-based data deduplication, and the consolidation of redundant copies into single versions. It

    DupeGuru is a content-based file deduplicator that finds and removes redundant files using hashing and filename comparison, though it lacks dedicated fuzzy text similarity matching for knowledge bases.

    PythonDeduplicationData Deduplication Tools
    在 GitHub 上查看↗7,631
  • qarmin/czkawkaqarmin 的头像

    qarmin/czkawka

    31,526在 GitHub 上查看↗

    Czkawka is a cross-platform utility designed for storage optimization and filesystem maintenance. It functions as a comprehensive file analysis engine that identifies redundant data, including duplicate files, empty directories, broken symbolic links, and temporary files. By utilizing hash-based content verification, the tool ensures accurate identification of duplicates regardless of file names or metadata. The project distinguishes itself by offering both a native graphical user interface and a command-line interface, allowing for both interactive management and automated, headless system m

    Czkawka is a cross-platform duplicate file detector that uses hash-based verification and similarity scanning to find redundant data across storage systems via both a command-line interface and a graphical interface.

    FluentCommand Line Interfaces
    在 GitHub 上查看↗31,526
  • google-research/deduplicate-text-datasetsgoogle-research 的头像

    google-research/deduplicate-text-datasets

    1,273在 GitHub 上查看↗

    Deduplicate-text-datasets is a data cleaning pipeline and text indexing engine built in Rust for large language model data preparation and corpus decontamination. It constructs massive uncompressed suffix arrays on disk using single-threaded and parallel chunk-based algorithms, enabling memory-efficient pattern searching and indexing across multi-gigabyte text corpora. The software identifies overlapping byte ranges from repeated substring matches to strip duplicate content and produce cleaned text datasets. Its search and indexing capabilities support querying corpus occurrences, locating d

    This repository provides specialized tools to detect and remove duplicate text documents for language model datasets using efficient exact and near-duplicate matching techniques.

    RustData Preparation PipelinesCorpus Decontamination UtilitiesDataset Deduplication
    在 GitHub 上查看↗1,273

Related searches

  • a library for processing digital documents
  • Document processing libraries
  • 重复照片查找工具
  • a javascript library for drag and drop
  • Docs as code
  • a static site generator for technical documentation
  • 数据库间的实时复制
  • a tool for aggregating technical documentation