awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

File format detection

Ranking updated Jul 27, 2026

For format detection, the first results are ahupp/python-magic (Python-magic is a Python wrapper for libmagic that identifies file formats and MIME types through magic number inspection, though it relies on external system libraries rather than a zero-dependency design), sindresorhus/file-type (This library detects file formats and MIME types by inspecting magic numbers in binary data and streams, making it a fitting choice though it lacks zero-dependency design and broad archive support) and file/file (This repository hosts the canonical libmagic implementation used to determine file types via magic number inspection and MIME type resolution, making it the definitive tool for this exact capability). apache/tika and google/magika round out the shortlist. Compare the match explanations and check the project documentation against your requirements.

Hand-picked open-source file format detection libraries ranked by GitHub stars and activity, with alternatives compared to help you find the best fit.

File format detection

Find the best repos with AI.We'll search the best matching repositories with AI.
  • ahupp/python-magicahupp avatar

    ahupp/python-magic

    2,886View on GitHub↗

    python-magic is a C-binding wrapper that provides a Python interface for the libmagic system library. It functions as a file signature analyzer and MIME type detector, identifying file formats by comparing header bytes against a database of known binary signatures. The library enables the identification of file types from both file paths and raw data buffers. It supports custom file signature matching through the injection of user-provided magic databases, allowing for the detection of specialized or proprietary formats. The project covers binary data analysis and MIME type mapping to transl

    Python-magic is a Python wrapper for libmagic that identifies file formats and MIME types through magic number inspection, though it relies on external system libraries rather than a zero-dependency design.

    PythonFile Signature AnalyzersMIME Type Detectors
    View on GitHub↗2,886
  • sindresorhus/file-typesindresorhus avatar

    sindresorhus/file-type

    4,297View on GitHub↗

    file-type is a binary file type detector that identifies file extensions and MIME types by analyzing magic numbers and signature bytes in binary data. It functions as a magic number parser and MIME type resolver, mapping binary signatures to standardized media type strings. The project is an extensible file format identifier that allows for the addition of custom detector plugins to recognize uncommon or non-binary file formats. The engine supports binary format identification across various data sources, including buffers and data streams. It utilizes a supported format registry and provide

    This library detects file formats and MIME types by inspecting magic numbers in binary data and streams, making it a fitting choice though it lacks zero-dependency design and broad archive support.

    JavaScriptMIME Type Detectors
    View on GitHub↗4,297
  • file/filefile avatar

    file/file

    1,632View on GitHub↗

    File is a command-line utility and C library designed to determine the precise file format and media type of unknown files by inspecting their internal byte sequences, binary signatures, and hierarchical pattern rules rather than relying on file extensions. The software evaluates magic number signature matching and complex multi-byte offsets to classify files accurately. The engine supports recursive archive inspection, automatically unpacking compressed archive structures on the fly to analyze and identify internal nested file payloads. Human-readable pattern definitions can be compiled into

    This repository hosts the canonical libmagic implementation used to determine file types via magic number inspection and MIME type resolution, making it the definitive tool for this exact capability.

    CFile Content AnalysisBinary Signature EnginesC Shared Libraries
    View on GitHub↗1,632
  • apache/tikaapache avatar

    apache/tika

    3,572View on GitHub↗

    Tika is a content analysis toolkit and Java library designed for detecting and extracting metadata and text from thousands of different file types. It functions as a universal document text extractor and metadata extraction engine, converting complex files into plain text or XHTML. The system employs a specialized MIME type detector that identifies document formats using magic bytes and metadata to determine the correct parser. It serves as an OCR integration gateway, connecting to external text recognition tools to extract content from image files. The project covers a broad range of extrac

    Apache Tika is a comprehensive Java library and content analysis toolkit that performs MIME type and format detection using magic bytes and metadata inspection, though it is far heavier than a zero-dependency design.

    JavaMIME Type Detectors
    View on GitHub↗3,572
  • google/magikagoogle avatar

    google/magika

    17,139View on GitHub↗

    Magika is an AI content type classifier and MIME type prediction engine that uses deep learning to identify file formats based on binary data. It analyzes byte sequences through a neural network to predict the content type of a file and provide associated confidence scores. The system features a foreign function interface that allows the core detection logic to be integrated across different programming languages. It includes a mechanism for configuring detection sensitivity and per-type thresholds to balance precision and recall. The project provides capabilities for bulk file analysis via

    Magika is a deep learning-based content type classifier that predicts MIME types from byte sequences, offering an AI-driven alternative to traditional magic number inspection for automated file analysis.

    PythonContent Type DetectionDeep Learning ClassifiersFile Type Validators
    View on GitHub↗17,139
  • h2non/filetypeh2non avatar

    h2non/filetype

    2,295View on GitHub↗

    Fast, dependency-free Go package to infer binary file types based on the magic numbers header signature

    This Go package provides fast, dependency-free file type inference using magic numbers, fitting the core detection capability well though it lacks streaming detection and some broader document formats.

    GoFile Storage SystemsGeneral UtilitiesGeneral Utility Libraries
    View on GitHub↗2,295
  • thephpleague/mime-type-detectionthephpleague avatar

    thephpleague/mime-type-detection

    1,338View on GitHub↗

    Mime Type Detection is a PHP library for determining and mapping media types from file contents, paths, or extensions using fallback lookup strategies. It provides utilities for inspecting file metadata and resolving file formats within PHP applications. The library supports bidirectional mapping between file extensions and media types through static community-curated mapping tables alongside content-based inspection. It includes extension-to-media-type mapping, reverse lookups from media types to valid file extensions, and fallback resolution chains that attempt content inspection before fal

    This PHP library provides mime type detection using various adapters, making it a solid choice for resolving file formats despite lacking some advanced streaming or zero-dependency features.

    PHPMIME Type DetectionCommunity-Curated Mapping TablesExtension Fallback Chains
    View on GitHub↗1,338
  • mholt/archivermholt avatar

    mholt/archiver

    4,467View on GitHub↗

    Archiver is a multi-format archive library and command-line tool for creating, extracting, and managing compressed archives. It provides a unified interface for working with formats including gzip, bzip2, zip, tar, rar, 7zip, and zstandard, and can automatically detect archive formats by inspecting binary byte headers rather than relying solely on file extensions. The library uses interface-based abstractions and a multi-format codec registry to support format-agnostic operations, while its stream-based compression pipeline processes archive data continuously without loading entire archives i

    Archiver detects archive formats by inspecting byte headers, but it is an archive management and compression tool rather than a general-purpose MIME type or file format detection library.

    GoArchive Format Detection
    View on GitHub↗4,467
  • bioruebe/uniextract2Bioruebe avatar

    Bioruebe/UniExtract2

    4,344View on GitHub↗

    UniExtract2 is a suite of tools designed for universal archive extraction, batch decompression, and file format analysis. It retrieves files from various compressed formats, software installers, disk images, and game archives into local directories. The project includes a file format analyzer that identifies file types by scanning internal contents and headers without requiring full extraction. It also features an archive password decrypter that attempts to recover access to protected archives using a predefined list of common passwords. The tool supports bulk decompression workflows through

    UniExtract2 is a universal archive extractor and file analysis suite rather than a dedicated programmatic file format detection library you can integrate into code.

    AutoItFile Signature Analyzers
    View on GitHub↗4,344
  • oliver-moran/jimpoliver-moran avatar

    oliver-moran/jimp

    14,621View on GitHub↗

    Jimp is a JavaScript image processing library and Node.js manipulation tool designed to perform image transformations and edits entirely within a JavaScript environment. It is a zero-dependency image library that operates without requiring native binaries or external system software dependencies. The project provides a programmatic interface for automated image transformations, including resizing, cropping, and filtering. It supports the creation of custom image pipelines and server-side image editing by processing data without relying on native system tools.

    Jimp is an image processing library rather than a general file format detector, so while it handles image decoding internally, it is not designed to identify arbitrary MIME types across various file formats.

    TypeScriptZero-Dependency Libraries
    View on GitHub↗14,621
  • oblac/joddoblac avatar

    oblac/jodd

    4,059View on GitHub↗

    Jodd is a suite of lightweight Java extensions and standard library utilities designed for application configuration, database mapping, dependency injection, and HTML parsing. It provides a consolidated set of core tools to facilitate Java development with a zero-dependency core to ensure compatibility and a small footprint across environments. The project features a pragmatic dependency injection container for managing object lifecycles and a database mapper that uses SQL templates to map result sets directly to Java objects. It includes a specialized configuration manager supporting profile

    Jodd is a broad utility library for Java rather than a dedicated file format or MIME type detection tool, making it the wrong category despite offering general-purpose utility functions.

    JavaZero-Dependency Libraries
    View on GitHub↗4,059
  • francisrstokes/super-expressivefrancisrstokes avatar

    francisrstokes/super-expressive

    4,615View on GitHub↗

    Super-expressive is a zero-dependency JavaScript library and domain-specific language used to construct complex regular expressions. It functions as a pattern generator that uses natural language syntax to produce native regular expression objects or strings, supporting international text standards through Unicode property matching. The library replaces manual string manipulation and escaping with a method-chaining fluent interface. It allows for modular expression composition, enabling the creation of reusable pattern hierarchies where existing expression instances can be nested as subexpres

    Super-expressive is a JavaScript regular expression builder library that shares a zero-dependency design but belongs to a completely different domain than file format and MIME type detection.

    JavaScriptZero-Dependency Libraries
    View on GitHub↗4,615
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
ahupp/python-magic2.9KPythonotherDec 2, 2025
sindresorhus/file-type4.3KJavaScriptMITApr 9, 2026
file/file
1.6K
C
NOASSERTION
Jun 9, 2026
apache/tika3.6KJavaapache-2.0Feb 20, 2026
google/magika17.1KPythonApache-2.0Jun 11, 2026
h2non/filetype2.3KGoMITJun 27, 2025
thephpleague/mime-type-detection1.3KPHPMITMar 13, 2025
mholt/archiver4.5KGomitNov 19, 2024
bioruebe/uniextract24.3KAutoItGPL-2.0Jul 6, 2024
oliver-moran/jimp14.6KTypeScriptMITApr 7, 2026

Related searches

  • a middleware for handling http content negotiation
  • Data interchange formats
  • Data format samples
  • an open source rich text editor library
  • User agent parser
  • a lightweight JavaScript library for date formatting
  • a library for parsing unstructured data
  • Terminal output formatter