40 مستودعات
Tools for image augmentation, 3D reconstruction, and visual tracking.
Explore 40 awesome GitHub repositories matching part of an awesome list · Computer Vision and Image Processing. Refine with filters or upvote what's useful.
YOLOv5 is a comprehensive computer vision framework designed for end-to-end deep learning, specializing in real-time object detection, image classification, and instance segmentation. It provides a unified toolkit that manages the entire lifecycle of a model, from initial dataset configuration and hyperparameter tuning to high-speed inference and deployment. The framework utilizes a modular neural architecture, allowing users to swap backbone and head components to tailor models for specific visual tasks. What distinguishes this project is its focus on production-ready deployment and model ef
Real-time object detection and image recognition framework.
Detectron2 is a PyTorch computer vision framework and visual recognition platform designed for training and deploying models for object detection, image segmentation, and visual recognition. It provides a research-oriented environment for training complex vision models with multi-GPU acceleration. The project includes a specialized object detection library for identifying and locating multiple objects via bounding boxes, as well as an image segmentation toolkit for creating pixel-level masks through instance, semantic, and panoptic segmentation. Additionally, it features a human pose estimati
Research platform for object detection and segmentation.
EasyOCR is a deep learning-based computer vision library designed to perform optical character recognition on images and video frames. It functions as a comprehensive pipeline that automates the transformation of visual text into machine-readable strings, enabling the digitization of physical documents, forms, and receipts into searchable data. The engine distinguishes itself through a multi-stage processing workflow that combines convolutional neural networks for spatial feature extraction with sequence-based decoding mechanisms. This architecture allows the system to identify and interpret
Multi-language optical character recognition.
SHAP is an explainable AI toolkit that provides a game theoretic framework for interpreting machine learning model predictions. It functions as a feature attribution engine, decomposing model outputs into the sum of individual feature effects to clarify how specific input variables influence a final decision. By assigning importance values to these inputs, the library enables users to understand the logic behind complex predictive models. The project distinguishes itself through its versatility and specialized calculation methods. It operates as a model-agnostic diagnostic library, capable of
Game-theoretic approach for explaining machine learning model outputs.
imgaug is a Python library for machine learning data augmentation and computer vision dataset expansion. It provides tools to increase the volume and variety of training sets by applying random geometric, color, and noise transformations to images. The library ensures spatial consistency by synchronizing transformations across images and their associated annotations, such as bounding boxes, keypoints, and segmentation maps. It uses a compositional pipeline pattern to chain multiple augmentations into sequences and employs deterministic seed management to reproduce specific data samples. The
Image augmentation library for machine learning.
Meshroom is a node-based photogrammetry software designed to transform collections of two-dimensional images into three-dimensional models and scene geometry. It provides a visual interface for constructing and managing modular data pipelines, allowing users to automate complex computer vision tasks such as feature extraction, depth map estimation, and mesh generation. The software distinguishes itself through a distributed computational framework that dispatches resource-intensive tasks across local hardware or remote render farms. By utilizing a directed acyclic graph execution model, it en
3D reconstruction software based on photogrammetry.
Libvips is a C-based image processing library designed to manipulate large visual assets through a low-memory, parallel processing pipeline. It functions as a streaming image processor that avoids loading entire files into system memory, enabling the handling of massive images in resource-constrained environments. The library distinguishes itself through a demand-driven architecture that constructs a deferred execution plan, computing only the necessary pixels for a final output. By utilizing a cache-friendly tiled processing model and memory-mapped file access, it minimizes latency and redun
High-performance image processing library with low memory usage.
Techniques for deep learning with satellite & aerial imagery
Resources for deep learning on aerial imagery.
Fawkes هو مولد صور عدائي وأداة لإخفاء التعرف على الوجه مصممة لحماية الخصوصية عن طريق تعتيم ملامح الوجه في الصور. يعمل كمخفي لخصوصية الصور يضيف اضطرابات بكسل غير مرئية إلى الصور، مما يمنع نماذج التعرف على الوجه من تحديد هوية الشخص بدقة مع الحفاظ على وضوح الصورة بصريًا للبشر. يستخدم النظام تعيين الاضطراب العدائي وتعتيم مساحة الميزات لتضليل مصنفات التعلم الآلي. ومن خلال استخدام حلقة تحسين تكرارية وتوليد ضوضاء مستقل عن النموذج، فإنه يعدل تمثيلات الوجه لمنع أنظمة التعرف من استخراج هوية متسقة عبر بنيات مختلفة.
Privacy tool for obfuscating faces from recognition.
Yolact هو إطار عمل للرؤية الحاسوبية ونموذج تجزئة مثيل في الوقت الفعلي. يستخدم شبكة عصبية تلافيفية بالكامل لاكتشاف الكائنات وإنشاء أقنعة على مستوى البكسل للصور وتدفقات الفيديو. يستخدم النظام توليد أقنعة نموذجية لإنشاء نماذج أقنعة عالمية يتم دمجها خطياً للحصول على نتائج خاصة بالمثيل. يدمج طبقات تلافيفية قابلة للتشوه وتجميع مناطق الاهتمام القابلة للتشوه لتكييف أخذ العينات المكانية مع الأشكال غير المنتظمة للكائنات. يغطي إطار العمل دورة حياة تطوير النموذج بالكامل، بما في ذلك التدريب على مجموعات بيانات مخصصة، وتقييم الدقة باستخدام متوسط الدقة (mAP)، واستخدام التدريب الموزع متعدد وحدات معالجة الرسومات لتوسيع سرعة المعالجة. كما يوفر أدوات معالجة الوسائط لتطبيق أقنعة التجزئة على الصور وتصدير ملفات الفيديو المشروحة. يتضمن المشروع أدوات استمرارية الحالة لإدارة نقاط التحقق واستئناف التدريب، إلى جانب التسجيل لتسجيل المقاييس وقيم الخسارة.
Real-time instance segmentation model.
This is an interactive notebook-based course that teaches machine learning from Python fundamentals through deep learning and natural language processing. It uses real datasets and multiple frameworks within a structured, hands-on curriculum that combines concise explanations with executable code cells, built-in datasets, and embedded exercise checkpoints. Learning progresses through data preparation and exploration, classical machine learning workflows, computer vision with convolutional neural networks, and natural language processing with deep learning, all delivered as a cohesive progressi
Implements convolutional neural networks and image augmentation to recognize and classify visual patterns.
MathUtilities هي مجموعة من مجموعات الأدوات المتخصصة التي توفر محركات للهندسة، والرؤية الحاسوبية، والرياضيات، ومحاكاة الفيزياء، ومعالجة الإشارات. تعمل كمكتبة شاملة للرياضيات والفيزياء تركز على الجبر الخطي، والتحسين العددي، والحسابات الهندسية للتطبيقات التقنية. يتميز المشروع بمجموعة أدوات لمحاكاة الفيزياء ومحرك هندسي ثلاثي الأبعاد. توفر هذه الأدوات قدرات لتكامل Verlet، وحلالات الحركية العكسية التكرارية، وعرض حقل المسافة عبر تتبع الأشعة الحجمي، وتشوه هندسة الشبكة. كما يتضمن أداة للرؤية الحاسوبية لتقدير حركة الكاميرا النسبية وإنشاء إسقاطات عين السمكة. تغطي المكتبة مجالات قدرات واسعة بما في ذلك أنظمة اكتشاف التصادم باستخدام فروق Minkowski والتجزئة المكانية، وتخطيط حركة الروبوتات، ومعالجة الإشارات لتقليل الضوضاء باستخدام مرشحات Kalman. تشمل الوظائف الإضافية تحسين البيانات العددية، وعمليات الجبر الخطي لملاءمة مجموعة النقاط، وتسلسل JSON القائم على الانعكاس لتسلسلات الكائنات.
Provides utilities for estimating relative camera motion and generating wide fisheye projections from image sources.
هذا المشروع هو مورد تعليمي للتعلم العميق ومجموعة مشاريع للشبكات العصبية. يوفر مجموعة من تطبيقات TensorFlow العملية ومشاريع البرمجة المصممة لإظهار تطبيق بنيات الشبكات العصبية المختلفة على بيانات العالم الحقيقي. يتضمن المشروع عينات محددة للشبكات التنافسية التوليدية (GANs)، مع التركيز على توليد الصور الاصطناعية وترجمة الأنماط. كما يوفر أمثلة لبناء نماذج التعلم العميق عبر نماذج تعلم مختلفة. يغطي الكود المصدري مجموعة واسعة من القدرات، بما في ذلك رؤية الكمبيوتر للتعرف على أنماط الصور، ومعالجة اللغة الطبيعية، وتحليل البيانات للتنبؤ بالسلاسل الزمنية. كما يشمل التعلم التعزيزي لتدريب الوكلاء المستقلين واستخدام المعالجة التسلسلية المتكررة.
Provides capabilities for classifying visual data into categories using neural networks trained on labeled images.
pysot is a computer vision framework designed for single object tracking. It provides a platform for implementing and evaluating algorithms that locate and follow specific target objects across sequences of video frames. The project includes implementations of the SiamRPN architecture for region proposal network based localization and the SiamMask model, which combines tracking with binary mask generation to provide pixel-level segmentation of objects. The framework also contains a visual tracking evaluation toolkit used to measure the accuracy and reliability of tracking algorithms against
High-performance codebase for visual tracking research.
Segment Geospatial هي مجموعة أدوات Python لعزل الميزات الجغرافية في صور الاستشعار عن بُعد باستخدام نموذج Segment Anything. تعمل كمعالج صور استشعار عن بُعد يحول بلاطات الخرائط إلى تنسيقات مرجعية جغرافياً لتوليد أقنعة التجزئة من بيانات الأقمار الصناعية. يتيح النظام استخراج الكائنات الجغرافية من خلال توليد القناع التلقائي أو المطالبات اليدوية، مثل الأوصاف النصية، ومربعات الإحاطة، والعلامات التفاعلية. ويدعم تجزئة صور السلاسل الزمنية لتتبع أو تحديد الكائنات عبر تسلسلات من الصور على تواريخ مختلفة ويوفر مصور قناع جغرافي مكاني لعرض النتائج على خرائط تفاعلية. يغطي المشروع مجموعة واسعة من العمليات المكانية، بما في ذلك الحصول على بلاطات الخرائط، وتحويل النقطي إلى متجه، وإعادة بناء حافة الميزة لتحسين حدود الكائنات. كما يتضمن REST API يكشف عن وظائف التجزئة ومعالجة البيانات هذه للتطبيقات البعيدة. تدعم إمكانيات التصدير الصور النقطية المرجعية جغرافياً وتنسيقات المتجهات القياسية بما في ذلك GeoJSON، و Shapefile، و GeoPackage.
Geospatial data segmentation using foundation models.
Visual tracking library based on PyTorch.
Framework for visual object tracking and segmentation.
3D Computer Vision Framework
Photogrammetric framework for 3D reconstruction.
Pytorch implementation of FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
Optical flow estimation using deep networks.
A Simple and Versatile Framework for Object Detection and Instance Recognition
Framework for object detection and instance recognition.
C++ image processing and machine learning library with using of SIMD: SSE, AVX, AVX-512, AMX for x86/x64, NEON, SVE for ARM, HVX for Hexagon
SIMD-accelerated image processing and machine learning library.