awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
goose3 avatar

goose3/goose3

0
View on GitHub↗
910 stars·106 forks·HTML·Apache-2.0·13 views

Goose3

A Python 3 compatible version of goose http://goose3.readthedocs.io/en/latest/index.html

Features

  • Content Extraction - HTML article content extractor for modern runtimes.

Star history

Star history chart for goose3/goose3Star history chart for goose3/goose3

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does goose3/goose3 do?

A Python 3 compatible version of goose http://goose3.readthedocs.io/en/latest/index.html

What are the main features of goose3/goose3?

The main features of goose3/goose3 are: Content Extraction.

Which projects share features with goose3/goose3?

Projects with overlapping indexed features include: executeautomation/mcp-playwright — This project is a Model Context Protocol server that enables Large Language Models to control Playwright browsers for… alir3z4/python-sanitize — Bringing sanity to world of messed-up data. buriy/python-readability — fast python port of arc90's readability tool, updated to match latest readability.js! codelucas/newspaper — Newspaper is a Python library designed for scraping, parsing, and analyzing web-based information. It functions as a… coleifer/micawber — a small library for extracting rich content from urls. deanmalmgren/textract — Textract is a multi-format text extraction tool and parser. It provides a unified interface to extract plain text from…

Projects sharing features with Goose3

These projects share indexed features with Goose3. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • executeautomation/mcp-playwrightexecuteautomation avatar

    executeautomation/mcp-playwright

    5,237View on GitHub↗

    This project is a Model Context Protocol server that enables Large Language Models to control Playwright browsers for web automation, scraping, and end-to-end testing. It functions as a programmable interface for executing JavaScript, capturing screenshots, and interacting with web elements across multiple browser engines. The server exposes browser automation capabilities as a set of standardized tools that models can discover and invoke. It supports session-based browser isolation to ensure unique contexts for each client connection and provides a transport layer using either standard input

    TypeScript
    View on GitHub↗5,237
  • buriy/python-readabilityburiy avatar

    buriy/python-readability

    2,895View on GitHub↗

    fast python port of arc90's readability tool, updated to match latest readability.js!

    Python
    View on GitHub↗2,895
  • codelucas/newspapercodelucas avatar

    codelucas/newspaper

    14,982View on GitHub↗

    Newspaper is a Python library designed for scraping, parsing, and analyzing web-based information. It functions as a framework for automated news aggregation and large-scale web content extraction, providing tools to download, clean, and structure text, metadata, and media from diverse online sources. The project distinguishes itself through a pipeline-oriented architecture that combines heuristic-based content extraction with natural language processing. It automatically identifies and isolates article bodies from web page boilerplate while simultaneously performing language detection, keywo

    HTMLcrawlercrawlingnews
    View on GitHub↗14,982
  • alir3z4/python-sanitizeAlir3z4 avatar

    Alir3z4/python-sanitize

    66View on GitHub↗

    Bringing sanity to world of messed-up data

    Python
    View on GitHub↗66
Compare all 12 related projects→