awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
HaujetZhao avatar

HaujetZhao/CapsWriter-Offline

0
View on GitHub↗
4,770 stars·436 forks·Python·25 views

CapsWriter Offline

CapsWriter-Offline is a suite of desktop tools that operates without an internet connection, combining local media browsing, voice dictation, audio and video transcription, and 360-degree media viewing into a single application. The project's core identity centers on providing offline functionality for both media handling and speech-to-text workflows.

What distinguishes it is the integration of voice dictation with a persistent local storage layer that saves every audio recording and daily transcript logs, along with a rule-based text normalization engine that converts spoken number phrases and user-defined substitutions using phonetic matching and regex. Recognized speech can be routed to a language model for polishing or role-specific processing based on predefined names. For media, the tool offers a transactional file operation manager for moving, renaming, and deleting files with undo support, and a panoramic media rendering engine that displays equirectangular 360-degree video and images with draggable viewport and device tilt interactions.

Additional capabilities include thumbnail generation with caching and manual refresh or purge, a customizable grid display for browsing images with adjustable sorting and column count, and EXIF metadata display. For audio and video files, speech can be extracted to produce subtitles, plain text, and timestamps for offline analysis.

Features

  • Audio and Video File Transcription - Extracts speech from audio and video files to produce subtitles, plain text, and timestamps for offline analysis.
  • Audio Transcription - Extracts speech from audio and video files to produce subtitles, plain text, and timestamps.
  • Hold-to-Dictate Mechanisms - Captures speech while a key is held and inserts the transcription into the active application, saving all audio locally.
  • Offline Media Transcribers - Extracts speech from media files and generates subtitles, plain text, and timestamp data offline.
  • Audio and Transcript Logs - Saves every voice recording as an audio file and maintains daily transcript logs for offline reference.
  • Key-Held Activations - Captures speech while a key is held and inserts the transcription into the active application.
  • 360 Media - Loads and displays 360-degree videos and images with draggable viewport and tilt interaction.
  • Voice and Transcript Log Persistence - Records every voice input as an audio file and maintains daily transcript logs for offline reference.
  • Offline Media Browsers with Dictation - Combines offline media browsing with voice dictation for captioning and transcripts.
  • Audio Persistence Speech Pipelines - Captures microphone input, passes it to a speech-to-text engine, saves raw audio files, and writes daily transcription logs to local storage.
  • Panoramic Media and Image Browsing - Browses and views 360-degree photos and videos on a local filesystem with EXIF metadata and customizable grid display.
  • LLM Delegation Mechanisms - Ships a mechanism that routes recognized speech to an LLM for polishing based on predefined role names.
  • Custom Text Normalizers - Applies user-defined rules with phonetic matching or regex to convert spoken text, including numeral normalization.
  • Phonetic and Regex Text Normalizers - Applies phonetic fuzzy-matching and regex to convert spoken number phrases and user-defined substitutions into normalized text.
  • Role-Based Text Polishing - Routes recognized speech to a language model for polishing or role-specific processing based on predefined names.
  • Role-Based Delegations - Routes recognized speech to a language model for role-specific polishing based on predefined names.
  • Numeral Normalizations - Converts spoken number phrases into numeric equivalents using pattern matching rules.
  • Comprehensive Media File Organizers - Moves, renames, and deletes media files with undo support and thumbnail caching to speed up repeated access.
  • Grid Display Customizations - Ships a configurable grid display with adjustable sorting, column count, and thumbnail quality.
  • Directory Browsing - Provides a file tree browser for navigating local directories and selecting media files.
  • Manual File Organizers with Undo - Moves, renames, and deletes files safely with undo support for accidental actions.
  • Thumbnail Caches - Generates and caches smaller image previews to accelerate repeated browsing access.
  • Cache Management Operations - Provides manual thumbnail cache management with reload, redraw, and purge operations.
  • Phonetic and Regex Replacements - Applies user-defined phonetic and regex replacement rules to normalize recognized speech text.
  • Transactional File Operation Managers - Executes file moves, renames, and deletions as atomic transactions with an undo log to revert accidental changes.
  • Caching Thumbnail Generation Layers - Generates smaller image previews on demand, stores them in a cache with expiry, and supports manual refresh or purge.
  • Panoramic Media Viewers - Provides an immersive viewer that loads and navigates 360-degree videos and images with drag and tilt controls.
  • Folder Browsers - Loads and displays images from a chosen local directory for immediate browsing.
  • Panoramic Media Rendering Engines - Renders equirectangular 360-degree video and images with draggable viewport and device tilt interaction using WebGL.
  • Customizable File Grids - Ships a configurable file grid that supports adjustable sorting, column count, thumbnail quality, and scroll sensitivity.

Star history

Star history chart for haujetzhao/capswriter-offlineStar history chart for haujetzhao/capswriter-offline

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with CapsWriter Offline

These projects share indexed features with CapsWriter Offline. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • chenyme/chenyme-aavtchenyme avatar

    chenyme/Chenyme-AAVT

    2,928View on GitHub↗

    Chenyme-AAVT is an AI-powered video transcription tool and translation platform. It converts speech from media files into editable text transcripts using speech recognition models and voice activity detection to ensure accurate phrasing and timing. The system functions as a content generator that transforms video transcripts into structured blog posts and marketing graphics using large language models. It also includes a subtitle formatting editor that allows for the modification of subtitle styles with a real-time video preview. The platform provides multilingual translation capabilities th

    Pythonfaster-whispergpt-4gpt-4o
    View on GitHub↗2,928
  • steipete/summarizesteipete avatar

    steipete/summarize

    3,771View on GitHub↗

    Summarize is a command line tool and multimodal content extractor designed to generate concise summaries from web pages, documents, and media files. It functions as an orchestrator that connects developer tools to various language model providers to process and condense information. The system provides specialized capabilities for audio and video processing, including transcription with speaker identification and the extraction of timestamped visual markers from video slides. It also includes a translation utility to convert generated summaries and extracted text into different target languag

    TypeScriptaiclisummarize
    View on GitHub↗3,771
  • cybertimon/rapidrawCyberTimon avatar

    CyberTimon/RapidRAW

    5,234View on GitHub↗

    RapidRAW is a non-destructive RAW photo editor and digital asset manager designed for decoding manufacturer RAW formats and applying tonal and color adjustments. It functions as a professional image processor that ensures original source data remains unmodified by saving all edits, masks, and crops to sidecar files. The software features a specialized color grading suite using 3D LUTs, color wheels, and HSL mixers, alongside AI-powered utilities for subject isolation, automatic masking, and generative inpainting for object removal. It distinguishes itself with AI-assisted photo retouching and

    TypeScriptcolor-gradingeditingimage-processing
    View on GitHub↗5,234
  • vaibhavs10/insanely-fast-whisperVaibhavs10 avatar

    Vaibhavs10/insanely-fast-whisper

    12,969View on GitHub↗

    This project is a high-throughput transcription engine and PyTorch inference wrapper designed to convert spoken audio files into text using the OpenAI Whisper model. It functions as a hardware-accelerated speech-to-text transcriber that runs locally on a user's machine. The system focuses on AI model performance tuning to maximize hardware throughput. It utilizes GPU acceleration, half-precision floating point tensors, and Flash-Attention to reduce processing time and memory overhead during transcription. The implementation covers large-scale transcription workflows and local speech-to-text

    Jupyter Notebook
    View on GitHub↗12,969
Compare all 30 related projects→

Frequently asked questions

What does haujetzhao/capswriter-offline do?

CapsWriter-Offline is a suite of desktop tools that operates without an internet connection, combining local media browsing, voice dictation, audio and video transcription, and 360-degree media viewing into a single application. The project's core identity centers on providing offline functionality for both media handling and speech-to-text workflows.

What are the main features of haujetzhao/capswriter-offline?

The main features of haujetzhao/capswriter-offline are: Audio and Video File Transcription, Audio Transcription, Hold-to-Dictate Mechanisms, Offline Media Transcribers, Audio and Transcript Logs, Key-Held Activations, 360 Media, Voice and Transcript Log Persistence.

Which projects share features with haujetzhao/capswriter-offline?

Projects with overlapping indexed features include: chenyme/chenyme-aavt — Chenyme-AAVT is an AI-powered video transcription tool and translation platform. It converts speech from media files… steipete/summarize — Summarize is a command line tool and multimodal content extractor designed to generate concise summaries from web… cybertimon/rapidraw — RapidRAW is a non-destructive RAW photo editor and digital asset manager designed for decoding manufacturer RAW… zackriya-solutions/meeting-minutes — This project is a self-hosted meeting transcription and summarization tool that converts audio recordings into text… vaibhavs10/insanely-fast-whisper — This project is a high-throughput transcription engine and PyTorch inference wrapper designed to convert spoken audio… m-bain/whisperx — WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining…