awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to html5lib/html5lib-python

Projects sharing features with Html5lib Python

17 open-source projects similar to html5lib/html5lib-python, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • charmparticle/xpeC

    charmparticle/xpe

    0View on GitHub↗
    View on GitHub↗0
  • danburzo/hredD

    danburzo/hred

    0View on GitHub↗
    View on GitHub↗0
  • emilstenstrom/justhtmlEmilStenstrom avatar

    EmilStenstrom/justhtml

    1,143View on GitHub↗

    A pure Python HTML5 parser that just works. No C extensions to compile. No system dependencies to install. No complex API to learn.

    Python
    View on GitHub↗1,143
  • engali94/xmljsonE

    engali94/XMLJson

    0View on GitHub↗
    View on GitHub↗0
  • ericchiang/pupericchiang avatar

    ericchiang/pup

    8,427View on GitHub↗

    Pup is a command line tool for extracting and filtering data from HTML documents using CSS selectors. It functions as a parser and selector engine that isolates specific elements based on tags, IDs, classes, and attributes. The project provides utilities for converting selected HTML nodes into plain text, attribute values, or structured JSON objects. It includes a markup formatter that corrects missing tags and applies consistent indentation to improve the readability of HTML documents. The tool handles the retrieval of text content and attributes through a CSS selector engine, supporting co

    HTML
    View on GitHub↗8,427
  • gawel/pyquerygawel avatar

    gawel/pyquery

    2,380View on GitHub↗

    A jquery-like library for python

    Python
    View on GitHub↗2,380

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • kennethreitz/requests-htmlkennethreitz avatar

    kennethreitz/requests-html

    326View on GitHub↗

    Pythonic HTML Parsing for Humans™

    Python
    View on GitHub↗326
  • kozea/tinycss2Kozea avatar

    Kozea/tinycss2

    186View on GitHub↗

    A tiny CSS parser

    Python
    View on GitHub↗186
  • lxml/lxmllxml avatar

    lxml/lxml

    3,037View on GitHub↗

    The lxml XML toolkit for Python

    Python
    View on GitHub↗3,037
  • martinblech/xmltodictmartinblech avatar

    martinblech/xmltodict

    5,741View on GitHub↗

    xmltodict is a Python library that provides bidirectional serialization between XML documents and dictionaries. It functions as a parser that converts marked-up input into key-value pairs and a serialization utility that transforms dictionaries back into structured XML documents. The project includes an incremental stream processor that uses depth-based callbacks to handle large XML files while maintaining constant memory usage. It features a namespace manager for mapping prefixes and declarations, as well as a security sanitizer that blocks external entity expansion and validates element nam

    Python
    View on GitHub↗5,741
  • mgdm/htmlqmgdm avatar

    mgdm/htmlq

    7,552View on GitHub↗

    htmlq is a suite of command-line utilities for querying and extracting data from HTML documents using CSS selectors. It functions as a query language tool for HTML structures and attributes, providing a way to retrieve specific information from documents via the terminal. The tool provides capabilities for extracting text content, specific HTML attributes, and document fragments. It includes an HTML document formatter for cleaning and reformatting output with consistent indentation, as well as utilities for stripping tags to isolate plain text. The software handles structural HTML processing

    Rust
    View on GitHub↗7,552
  • pallets/markupsafepallets avatar

    pallets/markupsafe

    691View on GitHub↗

    Safely add untrusted strings to HTML/XML markup.

    Python
    View on GitHub↗691
  • plainas/tqP

    plainas/tq

    0View on GitHub↗
    View on GitHub↗0
  • shinima/temmeS

    shinima/temme

    0View on GitHub↗
    View on GitHub↗0
  • sinelaw/xml-to-json-fastS

    sinelaw/xml-to-json-fast

    0View on GitHub↗
    View on GitHub↗0
  • stchris/untanglestchris avatar

    stchris/untangle

    631View on GitHub↗

    Converts XML to Python objects

    Python
    View on GitHub↗631
  • xhtml2pdf/xhtml2pdfxhtml2pdf avatar

    xhtml2pdf/xhtml2pdf

    2,390View on GitHub↗

    A library for converting HTML into PDFs using ReportLab

    Python
    View on GitHub↗2,390