5 repositorios
Automated workflows that apply machine learning models to extract metadata or identify objects within media.
Explore 5 awesome GitHub repositories matching graphics & multimedia · Computer Vision Pipelines. Refine with filters or upvote what's useful.
Immich is a self-hosted media management platform designed to provide a centralized, private repository for photos and videos. It functions as a comprehensive system for organizing, backing up, and viewing personal media collections across mobile devices, web browsers, and external storage locations. By maintaining full control over data ownership and storage infrastructure, the platform ensures that users retain sovereignty over their digital assets. The system distinguishes itself through a distributed architecture that coordinates background media synchronization, real-time filesystem moni
Automates facial recognition, object detection, and metadata extraction using integrated machine learning models.
Tesseract is a neural network-based optical character recognition engine designed to convert scanned images and digital documents into machine-readable, searchable text. It functions as both a command-line utility for automating large-scale digitization workflows and a cross-platform library that can be embedded into desktop, mobile, or server-side applications. By utilizing long short-term memory networks, the engine provides robust text extraction across more than one hundred languages and dozens of scripts. The project distinguishes itself through a sophisticated document layout analysis f
Decomposes visual documents into hierarchical structures, including text blocks, lines, and individual characters.
Supervision is a computer vision toolset for normalizing model outputs, managing datasets, and visualizing annotations. It provides a framework to convert predictions from various classification and detection models into a standardized data format to ensure interoperability across different computer vision pipelines. The library features a post-processor for filtering, counting, and tracking detected objects across image frames and video streams. It includes capabilities for large image tiling to improve the detection of small objects and tools for assigning persistent identities to objects t
Standardizes the processing and visualization of detection results within computer vision pipelines.
Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion. The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks
Identifies and isolates distinct sections within documents to improve data extraction accuracy.
SimSwap es un framework de aprendizaje profundo para el intercambio de rostros y un procesador de medios de visión artificial construido con PyTorch. Funciona como una herramienta de síntesis de imágenes diseñada para reemplazar la identidad de una persona en imágenes y videos con un rostro objetivo utilizando un único modelo entrenado. El sistema opera como una herramienta de reemplazo de identidad de video que intercambia identidades entre fotogramas mientras preserva las expresiones y la iluminación originales de los medios fuente. Permite la manipulación de identidad digital y la producción de medios sintéticos mediante el mapeo automatizado de características faciales. El framework admite tanto la aplicación de modelos entrenados para intercambiar rostros en medios como la capacidad de entrenar modelos personalizados de intercambio de rostros utilizando conjuntos de datos de imágenes específicos.
Ships computer vision pipelines that apply machine learning models for automated facial feature mapping.