awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

701 repository-uri

Awesome GitHub RepositoriesContent Processing and Transformation

Middleware and utility layers that handle the conversion, formatting, and programmatic manipulation of content between different formats or schemas.

Explore 701 awesome GitHub repositories matching content management & publishing · Content Processing and Transformation. Refine with filters or upvote what's useful.

Awesome Content Processing and Transformation GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • public-apis/public-apisAvatar public-apis

    public-apis/public-apis

    441,986Vezi pe GitHub↗

    Acest proiect este un director curatoriat de comunitate cu endpoint-uri de servicii REST și GraphQL, conceput pentru a ajuta dezvoltatorii să descopere și să integreze surse de date terțe. Funcționează ca un registru centralizat unde serviciile externe sunt organizate pe domenii pentru a facilita prototiparea rapidă a software-ului și dezvoltarea aplicațiilor. Registrul se bazează pe un model de contribuție peer-reviewed, utilizând controlul distribuit al versiunilor pentru a gestiona actualizările și a asigura acuratețea endpoint-urilor listate. Pentru a menține o calitate ridicată a datelor, proiectul folosește validarea bazată pe schemă pentru toate trimiterile primite și compilează datele structurate într-un site web static, ușor de căutat, pentru o regăsire eficientă. Directorul acoperă un spectru larg de capabilități de integrare, inclusiv regăsirea datelor financiare, servicii de geolocalizare și diverse API-uri utilitare pentru sarcini precum detectarea limbajului, procesarea media și verificarea identității. Prin furnizarea unui index centralizat al acestor servicii, proiectul sprijină dezvoltatorii în identificarea furnizorilor de date fiabili pentru diverse cerințe funcționale.

    Transforms HTML content into high-quality PDF documents for reports and documentation.

    Pythonapiapisdataset
    Vezi pe GitHub↗441,986
  • awesome-selfhosted/awesome-selfhostedAvatar awesome-selfhosted

    awesome-selfhosted/awesome-selfhosted

    299,516Vezi pe GitHub↗

    Acest proiect este un director curatoriat de comunitate cu software open-source conceput pentru implementarea în medii de server private și laboratoare de acasă (home labs). Servește drept resursă cuprinzătoare pentru descoperirea alternativelor independente, auto-găzduite, la serviciile cloud mainstream, permițând utilizatorilor să mențină proprietatea deplină a datelor și controlul asupra infrastructurii lor digitale. Directorul este structurat printr-o taxonomie ierarhică ce organizează o colecție vastă de aplicații în categorii logice, variind de la gestionarea media și analiza datelor la comunicare privată și instrumente de productivitate în echipă. Se distinge printr-un proces colaborativ de peer-review, unde membrii comunității validează calitatea și relevanța fiecărei trimiteri pentru a se asigura că directorul rămâne precis și fiabil. Proiectul acoperă o suprafață largă de capabilități, inclusiv automatizarea infrastructurii, implementarea serviciilor bazate pe containere și gestionarea configurației declarative. Aceste instrumente ajută utilizatorii să mențină medii de server reproductibile și să gestioneze dependențele complexe ale serviciilor pe hardware privat. Directorul este menținut ca un repository controlat prin versiuni, asigurându-se că toate actualizările și modificările conduse de comunitate sunt urmărite și transparente.

    Extracts and formats text from online articles to provide a distraction-free reading experience.

    awesomeawesome-listcloud
    Vezi pe GitHub↗299,516
  • obra/superpowersAvatar obra

    obra/superpowers

    229,538Vezi pe GitHub↗

    Superpowers este un motor de dezvoltare a jocurilor bazat pe browser și un mediu de dezvoltare integrat colaborativ. Oferă un spațiu de lucru unificat pentru construirea de experiențe interactive bidimensionale, permițând utilizatorilor să gestioneze codul, activele și logica scenelor direct într-un browser web, fără a fi nevoie de compilatoare locale sau software desktop greoi. Platforma se distinge printr-o arhitectură de scripting modulară, bazată pe componente, unde obiectele de joc sunt definite prin logică atașată și proprietăți vizuale. Suportă sincronizarea în timp real, permițând mai multor dezvoltatori să lucreze simultan la același proiect. Acest mediu este conceput pentru a funcționa ca un instrument educațional, predând concepte de programare prin crearea integrată de grafică, audio și logică. Sistemul include un pipeline de build cuprinzător care gestionează compilarea markdown a site-urilor statice și rutarea bazată pe sistemul de fișiere. Automatizează fluxul de lucru de dezvoltare prin rezolvarea dependențelor la momentul build-ului, injectând componente UI reutilizabile și gestionând pipeline-urile de active pentru a asigura livrarea eficientă a resurselor.

    Transforms lightweight markup into structured documentation through an integrated build-time pipeline.

    Shell
    Vezi pe GitHub↗229,538
  • vuejs/vueAvatar vuejs

    vuejs/vue

    209,900Vezi pe GitHub↗

    Vue este un framework JavaScript progresiv, bazat pe componente, conceput pentru construirea de interfețe utilizator reactive și aplicații single-page. Se concentrează pe un sistem de template-uri declarativ care transformă HTML-ul în funcții de randare eficiente, permițând dezvoltatorilor să organizeze interfețe complexe în unități izolate, reutilizabile, care se sincronizează automat cu starea aplicației. Framework-ul se distinge printr-un sistem de reactivitate bazat pe urmărirea dependențelor care monitorizează accesul la date în timpul randării pentru a declanșa actualizări precise. Oferă o arhitectură flexibilă care suportă atât adoptarea incrementală ca bibliotecă ușoară, cât și dezvoltarea de aplicații la scară largă. Dezvoltatorii pot utiliza un model de extensibilitate robust, bazat pe plugin-uri, pentru a injecta logică globală, în timp ce reconcilierea virtuală a DOM-ului framework-ului asigură actualizări eficiente ale interfeței prin calcularea mutațiilor minime. Dincolo de capabilitățile sale de randare de bază, proiectul include o suită cuprinzătoare de instrumente pentru gestionarea stării aplicației, rutarea bazată pe URL și randarea pe partea de server. Oferă suport extins pentru compunerea componentelor, distribuția conținutului și gestionarea animațiilor, alături de măsuri de securitate încorporate, cum ar fi escaparea automată a conținutului pentru a preveni vulnerabilitățile comune. Framework-ul este distribuit cu declarații oficiale de tip pentru a susține analiza statică și poate fi instalat prin manageri de pachete standard sau integrat direct în mediile de browser prin tag-uri script.

    Transforms markdown text into formatted HTML within the user interface.

    TypeScriptframeworkfrontendjavascript
    Vezi pe GitHub↗209,900
  • avelino/awesome-goAvatar avelino

    avelino/awesome-go

    175,576Vezi pe GitHub↗

    This project serves as a comprehensive language ecosystem index, functioning as a centralized, community-curated directory for the Go programming language. It organizes a vast landscape of software components, libraries, and development tools into a structured, navigable hierarchy, enabling developers to efficiently discover resources tailored to specific functional domains. The repository distinguishes itself through a decentralized contribution model, where community-driven updates ensure the index remains current with the rapidly evolving software landscape. Beyond simple resource listing,

    Exposes a variety of packages for generating, reading, and manipulating spreadsheet and presentation file formats.

    Goawesomeawesome-listgo
    Vezi pe GitHub↗175,576
  • microsoft/markitdownAvatar microsoft

    microsoft/markitdown

    154,485Vezi pe GitHub↗

    This project is an AI-powered document processing engine designed to transform diverse file formats into structured Markdown. By leveraging multimodal language models, it performs complex layout analysis and semantic text extraction, allowing for the conversion of both unstructured files and scanned images into machine-readable content. The toolkit distinguishes itself through a modular, plugin-based architecture that orchestrates multi-stage extraction pipelines. Users can steer the parsing behavior by injecting custom instructions, enabling the system to adapt to domain-specific document st

    Applies machine learning to perform layout analysis and extract structured data from complex, multi-format files.

    Pythonautogenautogen-extensionlangchain
    Vezi pe GitHub↗154,485
  • langgenius/difyAvatar langgenius

    langgenius/dify

    145,458Vezi pe GitHub↗

    Dify is an open-source platform for building, orchestrating, and deploying generative AI applications and autonomous agents. It provides a visual development environment that allows users to design complex, multi-step logic chains and conversational flows, which can then be published as APIs, web interfaces, or embedded widgets. The platform acts as a centralized infrastructure layer, managing model connections, prompt templates, and knowledge retrieval to support scalable AI-powered services. What distinguishes the platform is its focus on stateful application design and workflow orchestrati

    The platform automates the translation of text keys using language models that preserve formatting and synchronize updates across all supported languages through automated version control.

    TypeScriptagentagentic-aiagentic-framework
    Vezi pe GitHub↗145,458
  • chalarangelo/30-seconds-of-codeAvatar Chalarangelo

    Chalarangelo/30-seconds-of-code

    128,121Vezi pe GitHub↗

    30-seconds-of-code is a comprehensive knowledge base and programming snippet library designed to support software engineering education and professional development. It provides a curated collection of reusable code units and technical guides that help developers master core language mechanics, design patterns, and architectural philosophies. The project distinguishes itself by offering a wide-ranging library of algorithmic solutions and web development patterns that are organized into modular, independently testable units. It emphasizes functional programming paradigms and declarative logic,

    Provides utilities for reading and parsing text files into structured data.

    JavaScriptastroawesome-listcss
    Vezi pe GitHub↗128,121
  • ripienaar/free-for-devAvatar ripienaar

    ripienaar/free-for-dev

    123,154Vezi pe GitHub↗

    This project is a community-maintained directory of technical resources, tools, and services that offer free tiers for developers. It serves as a centralized reference point for discovering infrastructure, software, and educational materials, helping individuals and teams minimize operational costs while building and scaling applications. The directory distinguishes itself through a collaborative, community-driven curation model that aggregates metadata about third-party services. By utilizing a hierarchical taxonomy and storing all content in version-controlled, plain-text files, the project

    Localize application content using translation management systems that automate workflows for adapting software to global markets.

    HTMLawesome-listfree-for-developers
    Vezi pe GitHub↗123,154
  • garrytan/gstackAvatar garrytan

    garrytan/gstack

    110,596Vezi pe GitHub↗

    gstack is an AI agent framework and development workflow system designed to automate the software development lifecycle. It coordinates specialized AI personas to manage tasks across product design, engineering management, and quality assurance, transforming product intent into technical specifications and final releases. The project is distinguished by its deep integration of headless browser automation and semantic code memory. It utilizes a persistent Chromium daemon for web scraping and visual auditing, and implements a searchable knowledge base that logs architectural decisions and repos

    Renders markdown files into publication-quality PDFs featuring professional typography and vector diagrams.

    TypeScript
    Vezi pe GitHub↗110,596
  • jaywcjlove/awesome-macAvatar jaywcjlove

    jaywcjlove/awesome-mac

    105,841Vezi pe GitHub↗

    This project is a comprehensive, curated collection of software resources designed for the macOS ecosystem. It serves as a centralized directory for discovering applications across a wide range of functional domains, including professional development, system management, and personal productivity. The directory distinguishes itself by offering a highly granular classification of tools that cater to specific technical and creative workflows. It highlights specialized software for software engineering, such as terminal emulators, version control clients, and API development tools, alongside a b

    Find editors and previewing tools designed for writing and formatting documents using lightweight markup syntax.

    Swiftappappleapplication
    Vezi pe GitHub↗105,841
  • googlechrome/puppeteerAvatar GoogleChrome

    GoogleChrome/puppeteer

    94,974Vezi pe GitHub↗

    Puppeteer is a JavaScript library for programmatically controlling Chrome and Firefox through the Chrome DevTools Protocol or the WebDriver BiDi protocol. It launches and manages browser instances—typically without a visible user interface—to automate interactions with web pages, enabling navigation, clicking, typing, and data extraction entirely through code. The library distinguishes itself through deep integration with the Chromium embedding layer, allowing fine-grained process configuration with custom flags, permissions, and sandbox policies. It maintains multiple concurrent command stre

    Renders web pages and HTML content into static PDF documents or image snapshots with precise viewport control.

    TypeScript
    Vezi pe GitHub↗94,974
  • oven-sh/bunAvatar oven-sh

    oven-sh/bun

    93,257Vezi pe GitHub↗

    Bun is a high-performance runtime environment designed to execute JavaScript and TypeScript applications with minimal latency and high throughput. Built on a native core implemented in Zig, it provides a unified execution engine that leverages JavaScriptCore for efficient memory management and low-latency startup. The project functions as an all-in-one toolchain, integrating a native bundler, transpiler, package manager, and test runner into a single command-line interface. What distinguishes Bun is its focus on native system integration and developer productivity. It features a high-performa

    Parses and modifies HTML content using CSS selectors to dynamically update documents during request or response handling.

    Rustbunbundlerjavascript
    Vezi pe GitHub↗93,257
  • gohugoio/hugoAvatar gohugoio

    gohugoio/hugo

    88,701Vezi pe GitHub↗

    Hugo is a high-performance static site generator that transforms source content and templates into optimized web assets. Built with a focus on speed and scalability, it provides a comprehensive framework for managing large-scale documentation and editorial projects through structured content organization, taxonomies, and a flexible template-driven rendering engine. The project distinguishes itself through a sophisticated build system that utilizes incremental caching to minimize redundant processing during site updates. It supports complex content requirements by enabling multidimensional mod

    Handles project localization by transforming and formatting dates, currencies, numbers, and translated strings.

    Goblog-enginecmscontent-management-system
    Vezi pe GitHub↗88,701
  • infiniflow/ragflowAvatar infiniflow

    infiniflow/ragflow

    82,922Vezi pe GitHub↗

    This project is a comprehensive retrieval-augmented generation platform designed for building, managing, and deploying knowledge-based AI applications. It provides a unified environment for organizing datasets, configuring conversational chat assistants, and developing autonomous agents that execute multi-step reasoning workflows. By integrating document intelligence with advanced retrieval pipelines, the platform enables the creation of grounded, verifiable responses supported by traceable citations. The platform distinguishes itself through deep document understanding and sophisticated know

    Offers programmatic methods to asynchronously parse and extract content from various document types for further processing.

    Pythonagentagenticagentic-ai
    Vezi pe GitHub↗82,922
  • frooodle/stirling-pdfAvatar Frooodle

    Frooodle/Stirling-PDF

    81,168Vezi pe GitHub↗

    Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents. It functions as a self-hosted document manager, providing a centralized interface for users to manipulate files on a private server. The system features a workflow automation engine that allows for the creation of processing pipelines to handle large volumes of documents without writing custom code. It also includes an optical character recognition tool to convert scanned PDFs into searchable and editable text. Access is managed through single sign-on integration and OIDC comp

    Offers a unified interface for merging, splitting, editing, and converting PDF files.

    Java
    Vezi pe GitHub↗81,168
  • stirling-tools/stirling-pdfAvatar Stirling-Tools

    Stirling-Tools/Stirling-PDF

    81,109Vezi pe GitHub↗

    Stirling-PDF is a self-hosted document processing suite designed for secure, private file management. It functions as a comprehensive transformation engine that executes complex operations—such as merging, splitting, converting, and redacting documents—directly on the host machine. The platform provides both a browser-based interface for interactive editing and a programmatic, API-first architecture that allows for the automation of document workflows through standard HTTP requests. The project distinguishes itself through its focus on private, infrastructure-agnostic deployment and granular

    Executes complex document transformations and rendering tasks locally to ensure data privacy.

    TypeScriptdockerhacktoberfestjava
    Vezi pe GitHub↗81,109
  • awesomedata/awesome-public-datasetsAvatar awesomedata

    awesomedata/awesome-public-datasets

    75,979Vezi pe GitHub↗

    This project is a community-maintained, open-access directory of high-quality public datasets. It serves as a centralized reference point for researchers, developers, and data scientists to locate reliable information sources across a wide spectrum of industries and scientific fields. By providing a structured index, the repository facilitates the discovery of data necessary for exploratory analysis, machine learning model training, and the development of data-intensive applications. The directory distinguishes itself through a lightweight, platform-agnostic approach to resource indexing that

    Employs human-readable text files to simplify community contributions and version-controlled updates.

    aaron-swartzawesome-public-datasetsdatasets
    Vezi pe GitHub↗75,979
  • tesseract-ocr/tesseractAvatar tesseract-ocr

    tesseract-ocr/tesseract

    74,751Vezi pe GitHub↗

    Tesseract is a neural network-based optical character recognition engine designed to convert scanned images and digital documents into machine-readable, searchable text. It functions as both a command-line utility for automating large-scale digitization workflows and a cross-platform library that can be embedded into desktop, mobile, or server-side applications. By utilizing long short-term memory networks, the engine provides robust text extraction across more than one hundred languages and dozens of scripts. The project distinguishes itself through a sophisticated document layout analysis f

    Streamlines large-scale document digitization workflows by automating image processing, text extraction, and structured output generation.

    C++hacktoberfestlstmmachine-learning
    Vezi pe GitHub↗74,751
  • pewdiepie-archdaemon/odysseusAvatar pewdiepie-archdaemon

    pewdiepie-archdaemon/odysseus

    72,184Vezi pe GitHub↗

    Odysseus is a self-hosted AI workspace and autonomous agent framework designed for deploying and managing large language models. It serves as a centralized platform for orchestrating agentic tasks, utilizing a model context protocol server to connect AI models to external system utilities, browser automation, and local hardware. The system distinguishes itself through a combination of retrieval-augmented generation and a RAG knowledge base, using vector stores and local embeddings to provide persistent semantic memory. It further integrates AI-driven communication management to triage email i

    Renders PDF pages in a viewer panel and provides tools for document interaction.

    Python
    Vezi pe GitHub↗72,184
Înapoi123456…36Înainte
  1. Home
  2. Content Management & Publishing
  3. Content Processing and Transformation

Explorează sub-etichetele

  • Content Extraction Engines4 sub-tag-uriTools that parse and normalize raw web content from disparate sources into structured, usable data. **Distinguishing note:** Focuses on the extraction and normalization of data from external services rather than general content management.
  • Content Parsers1 sub-tagTools for intercepting and transforming specific syntax markers in source files. **Distinguishing note:** Focuses on injecting dynamic content via custom tag syntax.
  • Content Processing11 sub-tag-uriUtilities that manipulate, parse, or transform text and data content into different formats or structures.
  • Document Filter InterfacesInterfaces for external programs to manipulate document trees via serialized data. **Distinguishing note:** Focuses on JSON-based interop for external filter programs.
  • Document Processing and Conversion11 sub-tag-uriEngines and APIs that automate the conversion and processing of documents between various file formats.
  • Document Transformation Pipelines3 sub-tag-uriFrameworks for programmatically manipulating document structures through filters and scripts. **Distinguishing note:** Focuses on programmatic document manipulation rather than simple format conversion.
  • Internationalization & LocalizationPlatforms for managing multilingual assets and automated translation workflows.
  • Internationalized Web Content1 sub-tagFrameworks for translating and localizing technical documentation into multiple languages.
  • Markdown and Markup Tools3 sub-tag-uriTools that parse, render, and format Markdown and other lightweight markup languages into structured output.
  • PDF CompressionAlgorithms and tools to reduce PDF file size while preserving document integrity.
  • PDF Generation Libraries3 sub-tag-uriFrameworks for creating and formatting professional PDF documents. **Distinguishing note:** Focuses on high-level document construction.
  • PDF Layout OperationsTools for modifying document structure, including imposition, scaling, and overlaying.
  • PDF Processing Libraries1 sub-tagLibraries for generating, reading, and manipulating PDF documents. **Distinguishing note:** None available; minting under content management umbrella.