399 dépôts
Specialized tools and frameworks for processing visual data, including object tracking, face analysis, and image segmentation.
Explore 399 awesome GitHub repositories matching artificial intelligence & ml · Computer Vision Systems. Refine with filters or upvote what's useful.
Ce projet est un répertoire complet, organisé par la communauté, qui structure un vaste paysage de bibliothèques, frameworks et outils logiciels Python. Il sert de base de connaissances centralisée conçue pour faciliter la navigation dans l'écosystème et accélérer la découverte par les développeurs tout au long du cycle de vie du développement logiciel. Le répertoire se distingue en fournissant un index structuré de ressources classées par domaine technique, allant des utilitaires de développement fondamentaux aux domaines d'ingénierie spécialisés. Il couvre des capacités de haut niveau, notamment l'intelligence artificielle, la science des données, le développement web et la gestion d'infrastructure, permettant aux développeurs d'identifier des solutions éprouvées pour des défis techniques spécifiques. Le projet englobe une large surface de capacités, notamment des outils pour la gestion des dépendances, l'analyse de code statique et les tests automatisés. Il catalogue également des ressources pour le stockage de données persistantes, l'orchestration d'infrastructure cloud et le développement d'interfaces, fournissant une référence unifiée pour la construction et la maintenance de systèmes logiciels complexes.
Identifies resources for applying machine learning techniques to visual data analysis and image recognition.
Ce projet est un répertoire de logiciels open source organisé par la communauté, conçu pour être déployé dans des environnements de serveurs privés et des laboratoires domestiques. Il sert de ressource complète pour découvrir des alternatives indépendantes et auto-hébergées aux services cloud grand public, permettant aux utilisateurs de conserver la pleine propriété des données et le contrôle de leur infrastructure numérique. Le répertoire est structuré par une taxonomie hiérarchique qui organise une vaste collection d'applications en catégories logiques, allant de la gestion multimédia et de l'analyse de données à la communication privée et aux outils de productivité d'équipe. Il se distingue par un processus de revue par les pairs collaboratif, où les membres de la communauté valident la qualité et la pertinence de chaque soumission pour garantir que le répertoire reste précis et fiable. Le projet couvre une large surface de capacités, notamment l'automatisation de l'infrastructure, le déploiement de services basés sur des conteneurs et la gestion de configuration déclarative. Ces outils aident les utilisateurs à maintenir des environnements de serveur reproductibles et à gérer des dépendances de services complexes sur du matériel privé. Le répertoire est maintenu en tant que dépôt contrôlé par version, garantissant que toutes les mises à jour et les changements pilotés par la communauté sont suivis et transparents.
Analyzes video streams in real time to identify movement or specific objects and trigger alerts.
Ce projet est un dépôt centralisé de tutoriels pratiques piloté par la communauté, conçu pour faciliter l'acquisition de compétences par la construction pratique d'applications logicielles réelles. Il sert de répertoire complet qui agrège la documentation externe et les supports pédagogiques, fournissant un chemin structuré pour que les développeurs maîtrisent des langages de programmation et des domaines techniques spécifiques. Le dépôt se distingue en organisant des ressources techniques disparates dans une structure hiérarchique basée sur une taxonomie qui permet aux développeurs de découvrir et de naviguer dans diverses disciplines du génie logiciel. En regroupant des projets individuels en séquences logiques, il fournit une roadmap qui aide les apprenants à progresser des concepts fondamentaux à la mise en œuvre avancée. Le contenu est maintenu par des contributions collaboratives, garantissant que la collection reste une ressource actuelle et expansive pour la communauté des développeurs. Le projet couvre une large surface de capacités, couvrant des domaines tels que le développement web full-stack, l'ingénierie d'applications mobiles et le développement de jeux interactifs. Il inclut des ressources pour un large éventail de langages de programmation, allant des langages système comme C, C++ et Rust aux langages de haut niveau et fonctionnels tels que Python, Ruby, Haskell et Clojure. Ces supports soutiennent une maîtrise technique spécialisée dans des domaines incluant l'apprentissage automatique, la science des données et la programmation réseau. Le répertoire est structuré pour permettre une découverte efficace par langage de programmation et domaine technique, avec une table des matières claire pour aider les utilisateurs à localiser des informations spécifiques. Il fonctionne comme un index persistant de liens externes, connectant les développeurs à la documentation et aux tutoriels tiers pour approfondir leur compréhension des concepts techniques.
Apply mathematical transformations to visual data streams and static files to perform real-time image analysis, object detection, and feature tracking.
Ce projet est un dépôt complet d'implémentations computationnelles vérifiées conçu pour servir de ressource éducative pour l'informatique et la résolution de problèmes algorithmiques. Il fournit une collection structurée d'exemples de code qui couvrent les structures de données fondamentales, les opérations mathématiques et les concepts de programmation de base, permettant aux utilisateurs d'étudier la logique et la complexité derrière diverses méthodes computationnelles. Le dépôt se distingue par un modèle d'implémentation modulaire basé sur des références qui organise le code dans des espaces de noms logiques. Cette approche facilite l'exécution indépendante et la clarté éducative, permettant aux utilisateurs d'explorer l'évolution des stratégies computationnelles, des approches naïves par force brute aux solutions optimisées haute performance. En découplant les abstractions de structures de données des opérations algorithmiques, le projet garantit que les implémentations restent interchangeables et faciles à analyser. La surface de capacités couvre un large éventail de domaines techniques, notamment l'apprentissage automatique, la cryptographie, le calcul scientifique et la vision par ordinateur. Il inclut des implémentations pour la modélisation prédictive, les réseaux de neurones et l'analyse statistique, aux côtés d'outils pour le traitement du signal numérique, la gestion des flux réseau et la modélisation financière. La collection répond également à des besoins mathématiques spécialisés, tels que l'algèbre linéaire, les calculs géométriques et la manipulation de bits, fournissant une base large pour la recherche et les applications d'ingénierie.
Interpret visual data from digital media to detect objects, features, and patterns through automated processing routines.
Immich is a self-hosted media management platform designed to provide a centralized, private repository for photos and videos. It functions as a comprehensive system for organizing, backing up, and viewing personal media collections across mobile devices, web browsers, and external storage locations. By maintaining full control over data ownership and storage infrastructure, the platform ensures that users retain sovereignty over their digital assets. The system distinguishes itself through a distributed architecture that coordinates background media synchronization, real-time filesystem moni
Analyzes facial features through configurable parameters like recognition distance to improve biometric accuracy within large collections.
Deep-Live-Cam is a generative video transformation tool designed for real-time facial manipulation and cinematic enhancement. It functions as a local-first AI runtime, performing all media processing directly on the user's hardware to ensure complete data privacy without external network dependencies. By utilizing a high-performance processing pipeline, the application enables live face swapping and interactive video modifications during active streaming sessions or on pre-recorded media. The system distinguishes itself through a hardware-abstraction execution layer that dynamically routes co
Swaps faces while maintaining consistent lighting, expressions, and movement.
OpenCV is an open-source computer vision library and visual analysis toolkit. It provides a framework for processing static images and dynamic video frames to analyze visual data and extract information using deep learning. The project functions as a real-time image processing framework, enabling the execution of vision algorithms on live video streams for immediate analysis and data processing. The toolkit covers a broad range of capabilities including image pattern recognition, real-time video analysis, and visual data extraction. It also supports automated visual inspection for detecting
Serves as a comprehensive software library for image recognition and camera stream processing.
OpenCV is a comprehensive computer vision library designed for real-time performance and cross-platform deployment. It provides a native execution environment that leverages multi-threaded operations and automated memory management to handle intensive computational tasks, including image processing and machine learning model inference. The library distinguishes itself through a data-oriented matrix framework that utilizes proxy-based array abstractions to provide a consistent interface for multidimensional data. By employing factory-pattern algorithm interfaces and runtime type dispatching, i
Identifies, localizes, and maintains the trajectory of objects within static imagery or live video streams.
This project is a community-driven educational repository that serves as a comprehensive directory of university-level computer science video lectures. It provides a structured learning path for students and professionals, aggregating high-quality academic resources to facilitate self-paced study across a wide range of technical disciplines. The repository distinguishes itself through a collaborative maintenance model, utilizing version control workflows to allow contributors to expand and update the collection. Content is organized within a single, version-controlled document that leverages
Groups academic video resources that explore computer vision techniques and image processing methodologies.
This project is an open-source, interactive educational platform designed to teach deep learning through a comprehensive, code-first curriculum. It provides a structured learning path that covers foundational mathematics, modern neural network architectures, and practical optimization techniques, enabling practitioners to master complex artificial intelligence concepts through hands-on experimentation. The platform distinguishes itself by integrating technical explanations with executable Jupyter notebooks. This design allows readers to modify code and hyperparameters in real-time, facilitati
Details modern algorithmic approaches for identifying and tracking objects within complex visual environments.
This repository serves as a centralized collection of state-of-the-art deep learning architectures and reference implementations designed for research and application development. It provides a comprehensive toolkit for computer vision and natural language processing, offering pre-built models and training pipelines for tasks ranging from image classification and object detection to complex sequence modeling. The project distinguishes itself by providing a flexible execution harness that manages the entire training lifecycle, including data ingestion and backpropagation. It supports scalable
Bundles specialized pipelines and benchmarking utilities for developing and managing complex computer vision workflows.
Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I
Creates structured visual patterns by iteratively refining noise through a specialized generative machine learning pipeline.
This project is a comprehensive, community-driven directory of machine learning resources, software libraries, and educational materials. It serves as a centralized knowledge base for developers and researchers, organizing tools and frameworks by their primary programming language and technical domain to simplify discovery across the artificial intelligence ecosystem. The collection distinguishes itself by providing a cross-language development index that spans diverse programming environments, including C, C++, Rust, Clojure, and Python. It covers a wide range of specialized capabilities, fr
Lists specialized software utilities for image recognition and the processing of camera streams.
This project is a structured educational resource and technical guide for designing and implementing autonomous systems using large language models. It provides a comprehensive curriculum and code samples focused on agentic design patterns, autonomous development, and the creation of systems capable of planning and executing multi-step tasks. The resource details the implementation of agentic retrieval-augmented generation, where models autonomously plan and refine data searches. It covers a wide array of orchestrators and design patterns, including metacognitive reflection for self-correctin
Implements computer vision to verify interface elements and page states via visual screenshot analysis.
Openpilot est un système d'assistance à la conduite open source qui s'intègre aux unités de contrôle des véhicules pour fournir une direction, une accélération et un freinage automatisés. Il fonctionne comme un middleware de robotique automobile, utilisant un environnement d'exécution spécialisé pour traiter les données des capteurs et exécuter des commandes de contrôle en temps réel qui gèrent la dynamique du véhicule. La plateforme se distingue par une interface agnostique au matériel qui traduit les commandes de conduite standardisées dans les protocoles propriétaires requis par un large éventail de marques et de modèles de véhicules. Elle utilise une planification de trajectoire basée sur des réseaux de neurones pour prédire les trajectoires à partir de données visuelles et historiques, tandis qu'une boucle de contrôle déterministe assure des ajustements à haute fréquence pour la stabilité du véhicule. Pour maintenir la sécurité opérationnelle, le système intègre un processus de surveillance indépendant qui surveille les performances et déclenche un désengagement immédiat en cas de détection d'anomalies. L'architecture logicielle repose sur la fusion de capteurs en temps réel pour synchroniser les entrées des caméras et des radars dans une représentation environnementale unifiée. Les composants du système communiquent via un bus basé sur les messages pour faciliter l'échange de données à faible latence entre les capteurs et les actionneurs, soutenu par une couche de traduction modulaire qui permet l'intégration avec divers protocoles de communication automobile.
Executes real-time logic that bridges high-level driving intelligence with low-level vehicle control units.
This project is a community-curated directory of resources, libraries, and tools designed to support developers working with the Flutter framework. It functions as a centralized knowledge base, organizing high-quality external references into a structured, human-readable format to assist in the discovery of technical materials for cross-platform application development. The directory distinguishes itself through a comprehensive index of the global Flutter ecosystem, including local user groups, meetups, and communication channels that connect developers to international support networks. It m
Connects developers with vision-focused libraries capable of processing live camera feeds for object, face, and barcode recognition.
Ultralytics is a comprehensive computer vision framework designed for training, validating, and deploying deep learning models across a wide range of visual recognition tasks. It provides a unified interface for core operations including object detection, instance segmentation, pose estimation, and image classification. By utilizing a modular architecture, the platform allows users to swap model components to balance inference speed and accuracy requirements for diverse applications. The framework distinguishes itself through its support for real-time processing and flexible deployment. It in
Analyzes spatial orientation and movement by tracking keypoint coordinates across video sequences.
YOLOv5 is a comprehensive computer vision framework designed for end-to-end deep learning, specializing in real-time object detection, image classification, and instance segmentation. It provides a unified toolkit that manages the entire lifecycle of a model, from initial dataset configuration and hyperparameter tuning to high-speed inference and deployment. The framework utilizes a modular neural architecture, allowing users to swap backbone and head components to tailor models for specific visual tasks. What distinguishes this project is its focus on production-ready deployment and model ef
Analyzes live video streams to detect and track entities for immediate automated decision-making.
This is a Python facial recognition library designed to detect, encode, and identify human faces in images and video. It functions as a biometric identification tool that converts facial features into numerical encodings to compare and match identities. The library provides a computer vision command line interface for batch processing face detection and recognition tasks across image directories. It also supports a GPU accelerated vision API that utilizes CUDA and NVIDIA hardware to increase the speed of facial analysis and identification. Its capabilities cover human face detection and faci
Locates human faces by analyzing gradients of image intensity using Histogram of Oriented Gradients.
Faceswap is a comprehensive framework for automated media manipulation and neural face synthesis. It provides a modular pipeline that manages the entire lifecycle of facial feature extraction, deep learning model training, and image conversion. By coordinating complex computer vision workflows, the system enables users to map facial identities between source and destination datasets while maintaining structural alignment and lighting consistency across video frames. The project distinguishes itself through a highly extensible plugin-based architecture that handles hardware-accelerated process
Identifies and locates faces within image frames using rotation and scaling detection models.