awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
adithya-s-k avatar

adithya-s-k/omniparse

0
View on GitHub↗
7,618 estrellas·650 forks·Python·GPL-3.0·7 vistasomniparse.cognitivelab.in↗

Omniparse

Omniparse is a multimodal content parser and generative AI ingestion engine designed to convert documents, images, and multimedia into a uniform format. It functions as a data preprocessing pipeline that transforms diverse raw data sources into structured markdown to improve the performance of large language model workflows.

The system extracts text and structural data from PDFs, images, audio, and video files. It includes a web crawler that converts dynamic website content into clean markdown and a multimodal transformation process that maps disparate input formats into a unified data schema.

The tool's capabilities cover layout-aware document parsing for PDFs and slides, visual element extraction from images, and speech-to-text transcription for multimedia recordings. These processes enable the extraction of tables, objects, and spoken content for use in generative AI frameworks.

Features

  • Data Preprocessing Pipelines - Provides a comprehensive pipeline to clean and format diverse raw data sources specifically for large language model workflows.
  • Document to Markdown Converters - Transforms diverse document and media inputs into a standardized markdown format for LLM consumption.
  • Multimedia Content Analyzers - Analyzes image and document files to identify text and objects for metadata extraction.
  • LLM Data Ingestion Engines - Provides a comprehensive pipeline that optimizes diverse raw data sources for generative AI frameworks.
  • Speech to Text Transcription - Provides automated transcription of audio recordings into analyzable text using speech recognition engines.
  • Automated Video Transcribers - Converts spoken audio from video and audio files into time-synced text transcripts.
  • Visual-to-Text Generation - Detects objects and text within images to translate visual data into searchable text strings.
  • Data Preprocessing - Prepares diverse documents and media by extracting and structuring them for LLM ingestion.
  • Document Parsing and Extraction - Extracts text and tables from PDF, PowerPoint, and Word files to produce LLM-ready formats.
  • Web Crawling - Retrieves content from interactive websites and extracts clean raw information for further processing.
  • Multimodal Parsers - Extracts structural data and text from PDFs, images, audio, and video files into a uniform format.
  • Web Content Scrapers - Crawls dynamic websites and converts retrieved HTML content into structured markdown.
  • Ingestion Pipelines - Builds standardized data flows to convert raw files into structured formats for AI knowledge bases.
  • Multimodal Unified Schemas - Maps disparate input formats into a single structural representation for consistent AI framework compatibility.
  • Image - Converts visual text within image files into digital strings for analysis.
  • Layout-Aware Extraction - Identifies tables and structural elements in PDFs and slides to preserve spatial relationships during text extraction.
  • Web Page Markdown Converters - Crawls dynamic websites and converts rendered web page content into clean markdown for RAG and LLM training.

Historial de estrellas

Gráfico del historial de estrellas de adithya-s-k/omniparseGráfico del historial de estrellas de adithya-s-k/omniparse

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Preguntas frecuentes

¿Qué hace adithya-s-k/omniparse?

Omniparse is a multimodal content parser and generative AI ingestion engine designed to convert documents, images, and multimedia into a uniform format. It functions as a data preprocessing pipeline that transforms diverse raw data sources into structured markdown to improve the performance of large language model workflows.

¿Cuáles son las características principales de adithya-s-k/omniparse?

Las características principales de adithya-s-k/omniparse son: Data Preprocessing Pipelines, Document to Markdown Converters, Multimedia Content Analyzers, LLM Data Ingestion Engines, Speech to Text Transcription, Automated Video Transcribers, Visual-to-Text Generation, Data Preprocessing.

¿Qué alternativas de código abierto existen para adithya-s-k/omniparse?

Las alternativas de código abierto para adithya-s-k/omniparse incluyen: run-llama/liteparse — A fast, helpful, and open-source document parser. samuraigpt/ai-youtube-shorts-generator — This project is an AI-driven suite of tools designed to repurpose long-form video content into short-form clips. It… mli/autocut — Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video… linyqh/narratoai — NarratoAI is an automated video production pipeline that uses large language models to generate scripts, voiceovers,… mendableai/firecrawl-mcp-server — This project is a Model Context Protocol server that connects large language models to web scraping and crawling… llmware-ai/llmware — llmware is a Python framework for AI agent orchestration and model management, designed to coordinate multi-model…

Alternativas open-source a Omniparse

Proyectos open-source similares, clasificados según cuántas características comparten con Omniparse.
  • run-llama/liteparseAvatar de run-llama

    run-llama/liteparse

    10,782Ver en GitHub↗

    A fast, helpful, and open-source document parser

    Rustdocument-ocrdocument-processingocr
    Ver en GitHub↗10,782
  • mli/autocutAvatar de mli

    mli/autocut

    7,579Ver en GitHub↗

    Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video clips by modifying a text transcript instead of using a traditional timeline. The system operates as an FFmpeg video processor and subtitle manipulation utility. It converts spoken audio into text and compacts subtitle files into simplified formats, enabling the removal of unwanted video segments by deleting corresponding sentences from a transcription file. The project covers automated video transcription, non-linear video cutting, and subtitle file management. It supports hard

    Python
    Ver en GitHub↗7,579
  • linyqh/narratoaiAvatar de linyqh

    linyqh/NarratoAI

    8,091Ver en GitHub↗

    NarratoAI is an automated video production pipeline that uses large language models to generate scripts, voiceovers, and edited video commentary. It functions as a combined scriptwriter, voiceover generator, and video editor to streamline the creation of movie and television commentary content. The system automates the production workflow by converting input data into structured narrative scripts, synthesizing artificial speech for narration, and programmatically assembling video clips based on script timestamps. It also converts spoken audio from video files into written text for subtitles a

    Pythonaiagentaiopsgemini-api
    Ver en GitHub↗8,091
  • samuraigpt/ai-youtube-shorts-generatorAvatar de SamurAIGPT

    SamurAIGPT/AI-Youtube-Shorts-Generator

    3,037Ver en GitHub↗

    This project is an AI-driven suite of tools designed to repurpose long-form video content into short-form clips. It integrates a speech-to-text engine for automated transcription, a highlighting system that ranks engaging segments based on emotional hooks, and a video processor that converts horizontal footage into vertical formats. The system distinguishes itself through intelligent video cropping that utilizes face tracking and motion smoothing to keep subjects centered. It also employs an analysis system to extract viral highlights by scoring segments for engagement and practical value. T

    Pythonai-video-generatorartificial-intelligenceimage-to-video
    Ver en GitHub↗3,037
  • Ver las 30 alternativas a Omniparse→