For a tool for organizing academic research papers, the strongest matches are zotero/zotero (Zotero is a full-featured reference manager that collects, organizes), retorquere/zotero-better-bibtex (This is a plugin that enhances BibTeX and citation) and xournalpp/xournalpp (Xournalpp is a PDF annotation and note-taking application, not). l0o0/jasminum and timescale/pg_textsearch round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
We curate open-source GitHub repositories matching “research paper collections”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
Zotero is reference management software designed for collecting, organizing, and citing bibliographic research sources and digital documents for academic work. It functions as a web bibliographic collector, a citation generator, and a collaborative research platform. The system integrates tools for capturing metadata and archiving web pages into a centralized research library. It provides a specialized environment for reading and marking up PDF and EPUB files with highlights and notes linked directly to research sources. The software covers a broad range of capabilities including bibliograph
Zotero is a full-featured reference manager that collects, organizes, and annotates research papers with PDF support, metadata extraction, tagging, BibTeX export, browser integration, and cross-platform clients — it covers nearly all the requested features, though official sync is cloud-based rather than fully self-hostable out of the box, which narrows it slightly but still makes it the right kind of tool for this search.
Make Zotero effective for us LaTeX holdouts
This is a plugin that enhances BibTeX and citation key management for Zotero, not a standalone reference manager—you would need Zotero itself for PDF storage, annotation, and other core features.
Xournalpp is a digital note-taking and annotation application designed for capturing natural handwriting and sketching. It functions as a vector graphics editor that treats individual strokes, shapes, and text as discrete, editable objects, allowing users to refine and manipulate their work after it has been placed on the canvas. The application provides a specialized environment for overlaying handwritten notes and drawings onto existing PDF documents. By utilizing pressure-sensitive stylus input, it simulates a natural writing experience, while its layered canvas composition enables users t
Xournalpp is a PDF annotation and note-taking application, not a reference manager—it lets you write on papers but has no metadata extraction, BibTeX export, or library organization features for managing a collection of research papers.
Jasminum is a Zotero plugin designed for the management of Chinese bibliographic data. It serves as a metadata integration tool that automates the extraction of publication details from the China National Knowledge Infrastructure database and provides utilities for editing PDF outlines and bookmarks directly within the reference manager. The project focuses on Chinese academic citation standards, providing specialized tools to format and parse personal names to meet specific regional requirements. It also manages the integration of language-specific translators and citation styles sourced fro
Jasminum is a plugin for Zotero that enriches Chinese bibliographic metadata and PDF outline editing, but it is not a standalone reference management tool — it extends one, so it only partially matches the search for a full application to collect and organize research papers.
pg_textsearch is a full-text search integration for PostgreSQL that provides large-scale text indexing and BM25 relevance ranking. It implements a scalable indexing architecture that uses a memtable system to spill data to disk segments, allowing for the processing of massive datasets. The project distinguishes itself through support for multilingual search via language-specific partial indexes and the ability to index complex expressions, such as JSONB fields or concatenated columns. It ensures high availability by utilizing PostgreSQL-native streaming replication and write-ahead logs to syn
pg_textsearch is a full-text search extension for PostgreSQL, not a reference management tool — it handles only the search component, missing paper organization, PDF storage, metadata extraction, and export features the visitor needs.
zotero-pdf-translate is a translation extension for Zotero that converts PDF text, annotations, and bibliographic metadata into target languages using external services. It functions as an academic PDF translator and a bibliographic metadata translator, enabling the conversion of research papers, EPubs, item titles, and abstracts. The tool distinguishes itself as a multi-provider translation client that allows users to connect to various language models and APIs using custom secret keys. It features a translation comparison view that renders outputs from multiple services side-by-side to eval
zotero-pdf-translate is a translation plugin for the Zotero reference manager, not a standalone tool to collect, organize, and manage research papers — it extends an existing application rather than being the full category you are searching for.
Embed PDF Viewer is a browser-based PDF rendering library that uses a WebAssembly port of the PDFium engine to display documents entirely on the client side, with no server-side processing required. It provides a framework-agnostic core engine layer that manages the PDF document lifecycle, memory allocation, and WebAssembly resource cleanup, with dedicated integration hooks for React and Vue 3 that handle initialization, document loading, and reactive state management. The library offers both a pre-built, embeddable viewer that can be inserted into any web page with a single initialization ca
Embed PDF Viewer is a client-side PDF rendering library, not a reference manager — it handles document display but offers none of the paper collection, metadata extraction, BibTeX export, or organizational features needed to manage research papers.
This project is a software development kit and cluster management tool for PHP. It serves as a full-text search SDK and vector search interface, enabling applications to perform lexical, fuzzy, and semantic searches against indexed data. The library implements a PSR 7 HTTP client to ensure cross-environment compatibility through standardized messaging interfaces. It provides a specialized interface for retrieving embeddings and performing semantic retrieval workflows using vector data. Its capability surface covers a wide range of administrative and operational tasks, including search index
Elasticsearch-php is a PHP client library for the Elasticsearch search engine, not a reference management tool—it provides powerful search infrastructure but lacks the paper-specific features like PDF annotation, metadata extraction, and BibTeX export that this search requires.
go-astilectron is a cross-platform GUI framework and binding that enables the creation of desktop software by combining a compiled Go backend with an Electron frontend. It functions as an inter-process communication bridge, utilizing an asynchronous messaging system to exchange JSON events and synchronize state between the Go process and the JavaScript user interface. The project provides a native desktop API wrapper to orchestrate system-level features from the backend. This includes the ability to manage browser windows, construct native application menus, and control system tray icons and
go-astilectron is a cross-platform GUI framework that pairs Go with Electron to build desktop apps, not a reference management tool — it provides no paper collection, PDF handling, or citation features, so it can only serve as a foundation for building one, not as the application itself.
DesktopEditors is an office suite application designed for creating and editing text documents, spreadsheets, and presentations across different operating systems. It serves as an OOXML compatible editor, ensuring that files are read and written according to Office Open XML standards for cross-platform document exchange. The suite functions as a collaborative document platform featuring real-time co-authoring, version tracking, and integrated communication tools. It also acts as an AI-powered document assistant and PDF editor, providing capabilities for content generation, automated spreadshe
OnlyOffice Desktop Editors is an office suite for editing documents and PDFs, but it lacks core reference management features like metadata extraction, BibTeX/CSL export, browser integration for capturing citations, and dedicated paper organization — so while it annotates PDFs, it is not a research-paper management tool itself.
zvec is an embedded vector database engine and indexing library designed for high-dimensional similarity search. It functions as a hybrid search engine and a retrieval-augmented generation knowledge base, allowing for the storage and retrieval of dense and sparse vectors. The system is distinguished by its hybrid retrieval pipeline, which fuses vector similarity, full-text keyword matching, and scalar metadata filtering into single query operations. It supports a plugin-based model integration system for registering custom embedding models and rerankers, as well as language bindings for nativ
zvec is an embedded vector database engine for similarity search, not a reference manager — it lacks paper-specific features like PDF storage, annotation, metadata extraction, and export, making it a potential building block for search rather than the self-hosted research paper tool you need.
TagSpaces is an offline-first file tagging and organization platform that lets you manage local files with portable metadata stored directly in filenames or sidecar JSON files, eliminating the need for a central database. It functions as a full-text file search engine, a Kanban board file organizer, a local AI file assistant, an S3-compatible cloud file manager, and a web clipper and bookmark manager, all within a single application. The project distinguishes itself through a local-first architecture where all file operations, indexing, and AI processing run entirely on the device, with cloud
TagSpaces is a general-purpose file tagging and organization platform with search and browser clipping, but it lacks the research-specific features like PDF annotation, metadata extraction from academic papers, and BibTeX/CSL export that define reference management software.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| zotero/zotero | 14.5K | JavaScript | NOASSERTION | |
| retorquere/zotero-better-bibtex | 6.8K | TypeScript | MIT | |
| xournalpp/xournalpp | 14.9K | C++ | GPL-2.0 | |
| l0o0/jasminum | 7K | TypeScript | AGPL-3.0 | |
| timescale/pg_textsearch | 3.1K | C | postgresql | |
| windingwind/zotero-pdf-translate | 11.1K | TypeScript | AGPL-3.0 | |
| embedpdf/embed-pdf-viewer | 3.3K | TypeScript | mit | |
| elastic/elasticsearch-php | 5.3K | PHP | MIT | |
| asticode/go-astilectron | 4.9K | Go | MIT | |
| onlyoffice/desktopeditors | 4.4K | — | other |