awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 个仓库

Awesome GitHub RepositoriesLLM-Based Text Processing

Using large language models to perform complex text transformations like translation and summarization.

Distinct from Text Summarization: Unlike specific tasks like summarization, this covers the architectural use of LLMs for multiple text processing types.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · LLM-Based Text Processing. Refine with filters or upvote what's useful.

Awesome LLM-Based Text Processing GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • yetone/openai-translatoryetone 的头像

    yetone/openai-translator

    24,926在 GitHub 上查看↗

    openai-translator is a cross-platform translation tool and language learning utility available as a browser extension and desktop application. It functions as a client for large language model APIs to translate, summarize, and polish text across different digital environments. The project differentiates itself by integrating optical character recognition to translate text extracted from images and screenshots. It also includes a language learning workflow that allows users to save new vocabulary to a digital book and use text-to-speech synthesis for pronunciation. The tool provides broad tex

    Implements the core engine for translation, summarization, and polishing using large language model APIs.

    TypeScript
    在 GitHub 上查看↗24,926
  • nextai-translator/nextai-translatornextai-translator 的头像

    nextai-translator/nextai-translator

    24,920在 GitHub 上查看↗

    Nextai-translator is an AI-powered text processor and cross-platform translation application. Available as a desktop app and browser extension, it uses large language model APIs to translate, summarize, and refine multilingual content in real time. The tool integrates with clipboard managers and text selection utilities to trigger automated translations immediately after content is copied or highlighted. It also functions as an OCR translation utility, extracting and translating text from screenshots and non-selectable image content. Additional capabilities include a vocabulary management sy

    Uses large language models to perform context-aware text transformations including translation and summarization.

    TypeScriptbrowser-extensionchatgptchrome-extension
    在 GitHub 上查看↗24,920
  • kaixindelele/chatpaperkaixindelele 的头像

    kaixindelele/ChatPaper

    19,594在 GitHub 上查看↗

    ChatPaper is a suite of AI agents and utilities designed for academic literature automation, manuscript editing, and research assistance. The system functions as a research assistant that summarizes, translates, and analyzes scholarly papers, while providing specialized tools for converting academic PDFs into structured markdown to preserve formulas for analysis. The project features a literature survey automator that crawls research repositories and synthesizes domain reports, alongside a research mind map generator that transforms linear document content into non-linear node-based maps. It

    Pipes extracted text from academic papers into LLMs for summarization, translation, and critical analysis.

    Python
    在 GitHub 上查看↗19,594
  • anxcye/anx-readerAnxcye 的头像

    Anxcye/anx-reader

    7,613在 GitHub 上查看↗

    anx-reader is a cross-platform e-book reader and cloud-synced library manager. It renders various electronic book formats into a standardized HTML view with customizable themes and fonts for a consistent experience across different operating systems. The project integrates a large language model as a reading assistant to summarize text and answer questions about book content. It also functions as a digital annotation tool for creating color-coded highlights and detailed notes for external research export. The system includes capabilities for organizing digital library collections, synchroniz

    Integrates large language models to perform text processing tasks such as summarization and content analysis.

    Dartdartebook-readerflutter
    在 GitHub 上查看↗7,613
  • mit-han-lab/streaming-llmmit-han-lab 的头像

    mit-han-lab/streaming-llm

    7,232在 GitHub 上查看↗

    This project is a long context inference engine and optimizer designed to process infinite text streams using large language models without memory growth or performance degradation. It serves as a system for maintaining constant memory usage during the generation of text from arbitrarily long input sequences. The implementation utilizes a rolling key-value cache manager and attention sink mechanisms to stabilize the attention process during continuous stream processing. By retaining initial tokens and employing a sliding window of key-value pairs, the system enables constant-time inference an

    Uses large language models to process and maintain context over massive volumes of text.

    Python
    在 GitHub 上查看↗7,232
  • richasy/bili.copilotRichasy 的头像

    Richasy/Bili.Copilot

    5,220在 GitHub 上查看↗

    Bili.Copilot is a native Windows desktop client for Bilibili that integrates large language models to provide AI-enhanced media browsing and video summarization. It functions as a media browser and video summarizer, enabling users to generate concise overviews of videos and articles by processing subtitles and text. The application allows users to interact with AI models to query specific information and evaluate content within videos. These capabilities are delivered through a native Windows interface built with the Windows App SDK and WinUI. The software covers media management and content

    Integrates large language models for complex text processing of subtitles and content.

    GLSLbilibiliwindows-app-sdkwinui3
    在 GitHub 上查看↗5,220
  • katanaml/sparrowkatanaml 的头像

    katanaml/sparrow

    5,162在 GitHub 上查看↗

    Sparrow 是一个 LLM 文档提取平台和基于视觉的推理引擎,旨在将图像和 PDF 转换为经过验证的结构化数据。它作为代理工作流编排器,将分类、提取和验证任务串联成多步流水线。 该系统的特色在于其后端无关的推理层,可管理本地 GPU、Apple Silicon 和云服务商上的模型。它利用基于坐标的视觉定位将提取的文本映射到精确的边界框坐标,并使用基于提示的模型引导来引导注意力并规范化数据格式。 该平台涵盖了文档智能工作流,包括用于保持结构完整性的专业图像表格处理,以及用于验证提取字段正确性的模式驱动验证。它还提供了一个用于监控 API 性能、使用分析和系统健康状况的文档分析仪表板。 架构包含一个基于插件的扩展系统,用于集成索引和编排中使用的第三方库。

    Executes text-based analysis, validation, and decision-making tasks using an inference API without document input.

    Pythonagentic-aicomputer-visiondocumentai
    在 GitHub 上查看↗5,162
  • blader/humanizerblader 的头像

    blader/humanizer

    5,012在 GitHub 上查看↗

    Humanizer is a text processing system designed to remove machine-generated patterns from writing to make it sound more natural and conversational. It functions as an auditor and rewriter that identifies robotic signatures, formulaic tropes, and mechanical formatting in machine output. The project features a style-matching system that analyzes provided writing samples to replicate a user's specific sentence rhythms, vocabulary, and punctuation habits. This allows the tool to mirror a personal voice and apply a calibrated tone to the rewritten text. The system covers a broad range of linguisti

    Removes machine patterns from LLM output to make AI-generated text sound more natural and conversational.

    在 GitHub 上查看↗5,012
  • hismax/redinkHisMax 的头像

    HisMax/RedInk

    4,860在 GitHub 上查看↗

    RedInk is an AI content automation tool designed to generate coordinated social media posts, including titles, body text, and matching visual assets, from a single user-provided topic. It functions as a stateless content pipeline that uses large language models to transform topics into structured marketing copy and image prompts. The system utilizes prompt-template orchestration to combine static instructions with dynamic inputs, guiding artificial intelligence toward specific output formats. Users can manage these AI behaviors and API preferences through a web-based settings interface that t

    Implements a pipeline using large language models to transform topics into structured marketing text and image prompts.

    Python
    在 GitHub 上查看↗4,860
  • haujetzhao/capswriter-offlineHaujetZhao 的头像

    HaujetZhao/CapsWriter-Offline

    4,770在 GitHub 上查看↗

    CapsWriter-Offline is a suite of desktop tools that operates without an internet connection, combining local media browsing, voice dictation, audio and video transcription, and 360-degree media viewing into a single application. The project's core identity centers on providing offline functionality for both media handling and speech-to-text workflows. What distinguishes it is the integration of voice dictation with a persistent local storage layer that saves every audio recording and daily transcript logs, along with a rule-based text normalization engine that converts spoken number phrases a

    Routes recognized speech to a language model for role-specific polishing based on predefined names.

    Python
    在 GitHub 上查看↗4,770
  • 201206030/novel-plus201206030 的头像

    201206030/novel-plus

    4,648在 GitHub 上查看↗

    Novel-Plus is a content management system and online reading platform designed for hosting, managing, and distributing web novels across PC and mobile web interfaces. It provides a comprehensive environment for web novel publishing, featuring a multi-platform reader with bookshelves, customizable themes, and community commenting systems. The platform integrates an automated content crawler to gather and update literary data from external remote sources and employs a scalable text storage system using database sharding and flat files to handle large volumes of content. It also includes an inte

    Uses large language models to perform complex text transformations including expansion and polishing within the editor.

    Javabookcrawlnovel
    在 GitHub 上查看↗4,648
  • dataabc/weibo-crawlerdataabc 的头像

    dataabc/weibo-crawler

    4,541在 GitHub 上查看↗

    这是一个新浪微博网页爬虫和社交媒体数据管道,旨在提取用户资料、帖子、评论和多媒体资源。它作为一个容器化的数据爬虫,自动化收集社交媒体内容和互动指标,并将其存储在本地。 该系统包含一个处理层,利用大语言模型分析抓取的文本,生成摘要和情感分析。它通过一个部署就绪的容器模型脱颖而出,该模型具有用于管理提取任务和监控作业进度的 HTTP 界面。 该爬虫涵盖了广泛的功能,包括通过定时增量更新进行社交媒体监控、将多媒体资源归档到本地磁盘,以及向平面文件或数据库进行多格式数据导出。它还能捕获详细的社交互动,如一级评论和转发。

    Utilizes large language models to perform text transformations such as summarization and sentiment analysis.

    Pythoncrawlerweiboweibo-spider
    在 GitHub 上查看↗4,541
  • comfyanonymous/comfyui_examplescomfyanonymous 的头像

    comfyanonymous/ComfyUI_examples

    3,918在 GitHub 上查看↗

    This repository is a collection of node-based pipeline configurations, examples, and templates for generating AI media. It provides a workflow library and a curated gallery of blueprints designed for creating images, videos, and 3D assets using diffusion models. The project specifically offers a set of pre-configured node graphs for implementing advanced image generation and refinement techniques, with a focus on Stable Diffusion workflows. These examples demonstrate how to interconnect processing nodes to define complex generative logic without writing code. The available templates cover a

    Utilizes external large language models to perform text-based chat and analysis.

    HTML
    在 GitHub 上查看↗3,918
  • op7418/humanizer-zhop7418 的头像

    op7418/Humanizer-zh

    3,020在 GitHub 上查看↗

    Humanizer-zh is a tool designed to transform AI-generated content into natural-sounding writing by removing repetitive patterns and mechanical formatting. It functions as an AI content de-optimizer and text humanizer specifically focused on making Chinese text sound more human-authored. The project specializes in Chinese language refinement, replacing corporate jargon and artificial linguistic markers with native expressions and varied sentence rhythms. It employs a system to strip formulaic signatures and repetitive structures common in large language model outputs to increase perceived auth

    Specializes in removing machine patterns and introducing conversational fluency to humanize AI text.

    在 GitHub 上查看↗3,020
  1. Home
  2. Artificial Intelligence & ML
  3. LLM-Based Text Processing

探索子标签

  • LLM Text Humanizers1 个子标签Specialized LLM processing for removing machine patterns and introducing conversational fluency. **Distinct from LLM-Based Text Processing:** Distinct from general LLM-Based Text Processing by focusing specifically on humanization rather than summarization/translation.
  • Role-Based DelegationsMatches recognized text against predefined role names and routes it to an LLM for polishing or assistance. **Distinct from LLM-Based Text Processing:** Distinct from LLM-Based Text Processing: adds role-based routing logic that delegates text to different LLM processing based on matched role names.