For a read-later service, the first results are wallabag/wallabag (Wallabag is a self-hosted read-it-later service that extracts article content for offline reading and organizes saved pages with tags and search, making it a direct Pocket alternative with the exact self-hosted setup you want), omnivore-app/omnivore (Omnivore is a self-hostable, open-source read-it-later application that captures web articles and stores them for offline reading, with article extraction, cross-device sync, and highlighting, directly matching the core features needed to replicate Pocket) and karakeep-app/karakeep. archivebox/archivebox and do-say-go/dn round out the shortlist. Compare the match explanations and check the project documentation against your requirements.
Open-source alternatives to Pocket for saving, organizing, and reading web articles on your own server.
Wallabag is a self-hosted, open-source bookmark manager designed to archive web content for later reading. It functions as a personal knowledge management tool, allowing users to collect, store, and organize web pages into a centralized, searchable library. The platform provides a distraction-free reading experience by extracting the primary text and images from web pages while removing advertisements and navigation menus. This process ensures that saved articles remain accessible for offline reading, preserving the content even if the original source is removed from the internet. The system
Wallabag is a self-hosted read-it-later service that extracts article content for offline reading and organizes saved pages with tags and search, making it a direct Pocket alternative with the exact self-hosted setup you want.
Omnivore is an open-source, self-hostable read-it-later application designed to centralize web articles, newsletters, and digital documents into a personal library. It functions as a comprehensive content archiver that captures web pages and stores them locally, ensuring permanent access and readability regardless of internet connectivity. The platform distinguishes itself through an event-sourced synchronization engine that maintains a consistent state across multiple devices by replaying user actions. It utilizes a headless web scraping service to extract clean text and metadata from raw we
Omnivore is a self-hostable, open-source read-it-later application that captures web articles and stores them for offline reading, with article extraction, cross-device sync, and highlighting, directly matching the core features needed to replicate Pocket.
Karakeep is a self-hosted, open-source platform designed for personal knowledge management and web content archiving. It functions as a centralized repository where users can capture, organize, and preserve bookmarks, notes, and media files, ensuring long-term access to digital information even if original sources are removed or modified. The system distinguishes itself through its automated content processing and security-focused architecture. It utilizes headless browser crawling and optical character recognition to ingest and index web content, while a modular artificial intelligence pipel
Karakeep is a self-hosted, open-source read-it-later service that captures, archives, and organizes web pages with headless browser content extraction, full-text search, and tagging — directly matching the core capability and most of the listed features, including self-hosting and content preservation for later reading.
ArchiveBox is a self-hosted archiving tool designed for personal digital preservation and research data management. It functions as an automated web preservation engine that monitors URL inputs from bookmarks, browser history, or manual entries to capture and store permanent, offline copies of web content. By utilizing headless browser automation, the system renders dynamic web pages to ensure that captured snapshots, PDFs, and media assets remain accurate and accessible even if the original source disappears. The project distinguishes itself through a modular extractor pipeline and a task-qu
ArchiveBox is a self-hosted web archiving tool that can serve as a read-it-later service by capturing full page snapshots, PDFs, and media, with support for tagging, full-text search, and offline storage – it squarely fits the category but is more focused on permanent preservation than a streamlined reading experience with highlights and notes.
dn is a self-hosted personal web archiving system that automatically intercepts and stores web pages on a local device. It uses a proxy-based request interception model to capture browser traffic and save content for offline access without an internet connection. The system features a local full-text search engine that indexes all saved page content for information retrieval across the collection. It includes a dedicated browser interface that simulates online connectivity to serve archived files, mimicking the original live web environment. Administrative control is provided through a web-b
dn is a self-hosted personal web archiving system that intercepts and stores pages for offline access and full-text search, but its proxy-based capture and lack of a browser clipper, tags, or highlights make it more of a web archive than a dedicated read-it-later bookmarking service like Pocket.
Bleve is a search indexing engine library written in Go, designed to provide full-text search and document retrieval capabilities for embedded application data. It functions as a framework for indexing structured or unstructured information, allowing developers to build searchable collections that support complex query logic and data analysis. The engine distinguishes itself through a pluggable analysis pipeline that normalizes text before indexing, alongside support for vector similarity search to identify semantically related content. It utilizes finite-state transducer automata for efficie
bleve is a full-text search library, not a ready-to-use read-it-later service — it lacks article extraction, offline reading, a web clipper, and the user-facing application layer needed to replicate Pocket’s functionality.
Orama is a search engine and vector database that provides full-text indexing, geospatial calculations, and semantic vector storage. It functions as an LLM retrieval engine designed to provide grounded context to language models for conversational interfaces. The project implements hybrid search by combining dense vector embeddings with inverted keyword indices to retrieve documents based on both semantic meaning and exact text matches. It utilizes a WebAssembly module to execute search logic across different JavaScript environments and platforms. The system covers a broad range of retrieval
Orama is a search engine and vector database, not a read-it-later bookmarking service; it could serve as a search backend but lacks the article-saving, offline reading, and browser extension features you need.
WCDB is a cross-platform storage layer and embedded database engine that serves as a framework for SQLite. It functions as an object relational mapper, linking application classes to database tables to enable data operations via objects rather than raw queries. The project is distinguished by an integrated encryption layer for securing data at rest and a full-text search engine that uses language-specific tokenizers for text lookups. It also features transparent field compression to reduce storage footprints and a connection-pooling model to coordinate simultaneous read and write operations a
WCDB is an embedded database engine and ORM framework for SQLite — it provides a storage and full-text search foundation, but it is not a self-contained read-it-later bookmarking service with article extraction, offline reading, or a browser extension, so it falls short of the intended application category.
RediSearch is a Redis module that adds secondary indexing, full-text search, aggregation, and vector similarity search directly into the in-memory data store. It operates as an in-process search engine, extending the core key-value store with capabilities for indexing hash and JSON documents, enabling fast field-level lookups beyond primary key access. The module provides a full-text search engine built on inverted indexes, supporting stemming, fuzzy matching, and relevance scoring via tf-idf. It also includes a vector similarity search engine using a Hierarchical Navigable Small World graph
RediSearch is a full-text search engine module for Redis, not a read-it-later service—it lacks core features like article extraction, offline reading, and a web clipper that you need for Pocket-style functionality.
lunr.js is a JavaScript full-text search library and client-side search engine. It creates in-memory search indexes for fast keyword retrieval and ranked document matching within browser or Node.js environments. The library utilizes a JSON serializable search index, allowing the search structure to be converted to and from JSON for storage and distribution of pre-built search data. This enables search functionality for static websites by indexing content into portable files. The system supports advanced querying capabilities, including fuzzy text matching to account for typos, field-scoped i
lunr.js is a client-side full-text search library, not a self-hosted read-it-later bookmarking service—it provides search indexing but lacks article saving, extraction, offline reading, or web clipper capabilities.
Winds is an open-source RSS and podcast reader that aggregates content from both feed types into a single, unified reading and listening interface. The application is designed to be self-hosted, allowing users to deploy the full stack on their own servers with configurable dependencies and environment variables. The platform integrates with Getstream.io to deliver real-time, personalized activity feeds that adapt to individual content preferences using machine learning. It includes a full-text search engine that indexes all subscribed articles and podcast episodes for fast, query-based retrie
Winds is a self-hosted RSS and podcast reader with full-text search, but it aggregates subscribed feeds rather than letting you save arbitrary web pages via a browser clipper—so it is a neighboring consumption tool rather than a read-it-later bookmarking service like Pocket.
NewsBlur is an RSS feed aggregator and social news reader that collects and organizes stories from feeds, newsletters, and websites into a single interface. It functions as a feed synchronization service that maintains read states and subscription data across multiple devices and third-party applications. The platform distinguishes itself with AI-powered content summarization to generate briefings and answer questions about articles, alongside a system for training content classifiers. These classifiers learn user preferences for authors and tags to automatically highlight preferred topics or
NewsBlur is an RSS feed aggregator and feed reader, not a dedicated read-later bookmarker for saving individual web pages—it excels at organizing subscribed feeds but lacks the core "save any page" workflow and explicit offline-reading and annotation features that define a Pocket replacement.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| wallabag/wallabag | 12.8K | PHP | MIT | |
| omnivore-app/omnivore | 15.9K | TypeScript | agpl-3.0 | |
| 26.2K |
| TypeScript |
| AGPL-3.0 |
| archivebox/archivebox | 26.9K | Python | mit |
| do-say-go/dn | 3.9K | JavaScript | — |
| blevesearch/bleve | 11K | Go | apache-2.0 |
| oramasearch/orama | 10.4K | TypeScript | NOASSERTION |
| tencent/wcdb | 11.5K | C | NOASSERTION |
| redisearch/redisearch | 6.2K | C | NOASSERTION |
| olivernn/lunr.js | 9.2K | JavaScript | MIT |