awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

15 مستودعات

Awesome GitHub RepositoriesComputer Vision Benchmarks

Standardized evaluation suites for measuring the accuracy and generalization of visual recognition systems.

Distinguishing note: Specifically targets vision-based model evaluation rather than general-purpose ML benchmarking.

Explore 15 awesome GitHub repositories matching artificial intelligence & ml · Computer Vision Benchmarks. Refine with filters or upvote what's useful.

Awesome Computer Vision Benchmarks GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • openai/clipالصورة الرمزية لـ openai

    openai/CLIP

    33,779عرض على GitHub↗

    CLIP is a neural network architecture designed to map visual and textual data into a shared latent vector space. By utilizing transformer-based feature extraction and multi-modal tokenization, the system aligns images and natural language strings, enabling cross-modal similarity analysis and semantic classification. The project functions as a zero-shot classification engine, identifying image content by calculating the cosine similarity between visual features and arbitrary text labels without requiring task-specific retraining. Beyond inference, it serves as a research toolkit for evaluating

    Evaluating how well visual recognition systems generalize across diverse datasets and identifying performance gaps in real-world application scenarios.

    Jupyter Notebookdeep-learningmachine-learning
    عرض على GitHub↗33,779
  • jbhuang0604/awesome-computer-visionالصورة الرمزية لـ jbhuang0604

    jbhuang0604/awesome-computer-vision

    23,074عرض على GitHub↗

    This project is a comprehensive, community-driven repository that serves as a centralized catalog for computer vision research and development. It functions as a structured index of academic papers, open-source software libraries, public datasets, and educational tutorials, providing a navigation point for the complex landscape of modern vision technology. The repository distinguishes itself through a taxonomy-based indexing system that maps the relationships between foundational research, influential academic figures, and their corresponding software implementations. By utilizing a lightweig

    Acts as a comprehensive research catalog for influential figures, algorithms, and benchmarking suites.

    عرض على GitHub↗23,074
  • zalandoresearch/fashion-mnistالصورة الرمزية لـ zalandoresearch

    zalandoresearch/fashion-mnist

    12,754عرض على GitHub↗

    This project is a computer vision benchmark and image classification dataset used to measure and compare the accuracy of machine learning models. It provides a standardized collection of labeled fashion product images and training data formatted to be compatible with the MNIST dataset structure. The dataset consists of fixed-dimension grayscale images and label-based category mappings, stored in a binary format. It includes pre-split training and testing sets and a static distribution to ensure consistent cross-model benchmarking. The repository supports image classification benchmarking and

    Serves as a reference dataset to measure and compare the accuracy of image classifiers.

    Pythonbenchmarkcomputer-visionconvolutional-neural-networks
    عرض على GitHub↗12,754
  • xpixelgroup/basicsrالصورة الرمزية لـ XPixelGroup

    XPixelGroup/BasicSR

    8,297عرض على GitHub↗

    BasicSR is a PyTorch-based image restoration toolbox and framework designed for training and deploying deep learning models to upscale, denoise, and deblur images and videos. It serves as a comprehensive system for image super-resolution and video quality restoration, providing the necessary infrastructure to recover fine visual details and increase pixel density. The project distinguishes itself through specialized toolkits for facial image enhancement and high-fidelity face synthesis, as well as a dedicated video quality restoration suite that utilizes deformable convolutions and generative

    Computes standard restoration benchmarks including PSNR, SSIM, LPIPS, NIQE, and FID.

    Pythonbasicsrbasicvsrdfdnet
    عرض على GitHub↗8,297
  • mikel-brostrom/boxmotالصورة الرمزية لـ mikel-brostrom

    mikel-brostrom/boxmot

    8,212عرض على GitHub↗

    Boxmot is a multi-object tracking framework designed to follow multiple objects across video frames using motion and appearance algorithms to maintain consistent identities. It functions as a system for tracking objects with specific orientations using rotated bounding boxes and corresponding intersection-over-union computations. The project includes a re-identification model optimizer that converts neural networks into formats for hardware-accelerated execution. It also features an evolutionary hyperparameter tuner that iteratively mutates tracker settings to maximize accuracy for specific d

    Uses standardized evaluation suites to measure the accuracy and consistency of visual tracking systems.

    Pythonboosttrackbotsortbytetrack
    عرض على GitHub↗8,212
  • rafaelpadilla/object-detection-metricsالصورة الرمزية لـ rafaelpadilla

    rafaelpadilla/Object-Detection-Metrics

    5,098عرض على GitHub↗

    This project is an object detection evaluation library and benchmarking tool designed to calculate precision, recall, and average precision for computer vision models. It provides a suite of utilities for parsing bounding box coordinates from text files and calculating spatial overlap to determine detection accuracy. The toolkit features a command line interface for comparing ground truth files against model predictions. It includes a precision-recall curve generator to visualize the relationship between precision and recall across different confidence thresholds and an intersection over unio

    Provides a standardized evaluation suite for measuring the accuracy and generalization of object detection models.

    Pythonaverage-precisionbounding-boxesmean-average-precision
    عرض على GitHub↗5,098
  • kaiyangzhou/deep-person-reidالصورة الرمزية لـ KaiyangZhou

    KaiyangZhou/deep-person-reid

    4,849عرض على GitHub↗

    هذا المشروع عبارة عن إطار عمل PyTorch لإعادة تحديد هوية الأشخاص مصمم لتدريب وتقييم النماذج التي تحدد الأفراد عبر مشاهد كاميرا مختلفة. يوفر خط أنابيب تدريب نموذج كامل، ومستخرج ميزات تعلم عميق لتحويل الصور إلى متجهات رقمية، ومجموعة من أدوات قياس الرؤية الحاسوبية لقياس دقة استرجاع الهوية. يتضمن إطار العمل مجموعة أدوات تعلم نقل متخصصة تدعم تجميد الطبقات، وتحسين معدل التعلم المرحلي، ومعدلات تعلم تفاضلية لضبط النماذج المدربة مسبقاً. يتميز بمحرك قابل للتوسيع يسمح بتطوير منطق تدريب مخصص وتنفيذ أهداف تحسين محددة مثل تعدين خسارة الثلاثي للعينة الصعبة وتنعيم التسميات. يغطي النظام إدارة شاملة لمجموعات البيانات، بما في ذلك دعم المعايير القياسية، وأخذ عينات الدفعات المتوازنة، وتعزيز الصور. يوفر أدوات تقييم لحساب رتب الاسترجاع ومسافات الميزات، بالإضافة إلى أدوات تصور لتوليد خرائط حرارة التنشيط ومعارض الاسترجاع المصنفة. تم تنفيذ المشروع بلغة Python ويستفيد من PyTorch لعمليات التعلم العميق الخاصة به.

    Provides a suite for evaluating identity retrieval accuracy using standard re-identification benchmarks.

    Pythoncomputer-visioncross-domaindeep-learning
    عرض على GitHub↗4,849
  • openimages/datasetالصورة الرمزية لـ openimages

    openimages/dataset

    4,366عرض على GitHub↗

    هذا المشروع عبارة عن مستودع لمجموعة بيانات الرؤية الحاسوبية وتعليقات الصور التوضيحية مصمم لتدريب وتقييم نماذج التعلم الآلي. يوفر مجموعة كبيرة من الصور المصنفة، ويعمل كمعيار لاكتشاف الكائنات ومصدر لبيانات التجزئة على مستوى البكسل. يتميز المستودع كمجموعة بيانات بصرية متعددة الوسائط من خلال إقران الصور بصوت ونصوص ومسارات ماوس متزامنة لدعم فهم السرد. كما يتيح تحليل عدالة النموذج من خلال تضمين السمات الديموغرافية والتعليقات التوضيحية الشاملة. تغطي مجموعة البيانات نطاقاً واسعاً من إمكانيات الرؤية الحاسوبية، بما في ذلك اكتشاف الكائنات عبر صناديق التحديد، وتجزئة مثيل الصورة باستخدام أقنعة البكسل، ورسم خرائط العلاقات البصرية من خلال ثلاثيات الكائن-السمة. كما تدعم التصنيف على مستوى النقطة، والتعرف الهرمي على النصوص، واسترجاع مجموعات فرعية من البيانات المنسقة بناءً على تصفية الفئة أو السمة.

    Serves as a standardized benchmark for computing precision and recall in object detection and classification models.

    Python
    عرض على GitHub↗4,366
  • facebookresearch/deitالصورة الرمزية لـ facebookresearch

    facebookresearch/deit

    4,348عرض على GitHub↗

    DeiT هو إطار عمل محول رؤية (vision transformer) لـ PyTorch مصمم لتصنيف الصور. ينفذ معمارية قائمة على المحولات تعالج الصور كتسلسلات من الرقع المسطحة باستخدام طبقات الانتباه الذاتي ونمذجة التسلسل الواعية بالموقع بدلاً من المرشحات التلافيفية. يركز المشروع على التدريب الفعال للبيانات من خلال إطار عمل لتقطير المعرفة. يسمح هذا النظام لنموذج الطالب بتقليد التسميات اللينة لنموذج معلم عالي الأداء لتحسين الدقة والتعميم، خاصة عند التدريب على مجموعات بيانات أصغر. تغطي المكتبة دورة حياة التطوير الكاملة، بما في ذلك تدريب تصنيف الصور، وتحسين فقدان الإنتروبيا المتقاطعة، ونشر الأوزان المدربة مسبقاً للاستنتاج. كما تتضمن أداة قياس لتقييم أداء النموذج ودقته مقابل مجموعات البيانات القياسية.

    Includes tools for evaluating model accuracy against standard computer vision benchmarking datasets.

    Python
    عرض على GitHub↗4,348
  • richzhang/perceptualsimilarityالصورة الرمزية لـ richzhang

    richzhang/PerceptualSimilarity

    4,244عرض على GitHub↗

    PerceptualSimilarity هو إطار عمل للتعلم العميق مصمم لقياس وتقييم المسافة الإدراكية بين الصور. يوفر نظاماً لقياس مدى تشابه صورتين أو رقعتين من الصور مع الرؤية البشرية باستخدام تمثيلات الميزات العميقة بدلاً من الاختلافات على مستوى البكسل. ينفذ المشروع مقياس مسافة قابل للتمايز يعمل كدالة خسارة، مما يسمح بتحسين بكسلات الصورة عبر الانتشار العكسي للوصول إلى مظهر بصري مستهدف. يتضمن طبقة خطية قابلة للتدريب يمكن تطبيقها على الميزات العميقة المجمدة لتعلم مقاييس المسافة المرجحة المتوافقة مع الإدراك البشري. يغطي إطار العمل قدرات واسعة في تقييم جودة الصورة، وتدريب مقياس التشابه، وقياس أداء رؤية الكمبيوتر. يتم تقييم دقة النموذج من خلال مقارنة درجات المسافة المتوقعة مقابل مجموعات بيانات الحكم البشري باستخدام أطر عمل مثل اختبارات الاختيار القسري البديلين.

    Tests the accuracy of visual similarity models against standardized human judgment datasets.

    Python
    عرض على GitHub↗4,244
  • princeton-vl/raftالصورة الرمزية لـ princeton-vl

    princeton-vl/RAFT

    4,057عرض على GitHub↗

    RAFT is a PyTorch computer vision framework and deep learning system designed for optical flow estimation. It functions as a GPU-accelerated motion estimator that calculates per-pixel motion vectors between video frames to determine object movement. The implementation utilizes recurrent all-pairs field transforms and custom CUDA kernels to optimize the memory and compute overhead associated with high-dimensional correlation calculations. This hardware-level acceleration reduces GPU memory usage during the forward pass. The toolkit covers supervised flow learning and model training using mixe

    Evaluates the accuracy of motion estimation models against standardized computer vision datasets.

    Python
    عرض على GitHub↗4,057
  • hustvl/vimالصورة الرمزية لـ hustvl

    hustvl/Vim

    3,882عرض على GitHub↗

    Vim is a state space model vision framework designed for image classification and visual representation learning. It functions as a computer vision research tool that converts two-dimensional image grids into one-dimensional sequences to extract spatial features. The system implements a linear-scaling image classifier that replaces quadratic attention mechanisms with state space operations. This approach utilizes bidirectional sequence modeling and selective gating mechanisms to process visual data. The framework covers computer vision benchmarking and image classification research, providin

    Measures vision model accuracy and performance against standard industry datasets.

    Python
    عرض على GitHub↗3,882
  • open-mmlab/mmtrackingالصورة الرمزية لـ open-mmlab

    open-mmlab/mmtracking

    3,881عرض على GitHub↗

    mmtracking is a PyTorch video perception framework designed for training and deploying computer vision models that analyze sequential image data. It provides specialized tools for multi-object tracking, video instance segmentation, and a configuration-driven system for managing deep learning models. The project utilizes a deep learning model registry and a configuration-driven pipeline to swap model backbones and detectors without modifying the core codebase. This modular approach allows for the development of custom perception architectures by combining various components and configurations.

    Provides a standardized evaluation suite for measuring the tracking precision of visual recognition systems.

    Pythonmulti-object-trackingsingle-object-trackingtracking
    عرض على GitHub↗3,881
  • roboflow/trackersالصورة الرمزية لـ roboflow

    roboflow/trackers

    2,565عرض على GitHub↗

    This project is a multi-object tracking library and computer vision toolkit designed to maintain consistent identity IDs for objects across video frames. It provides a motion-based object tracking system that converts raw detections into stable temporal tracks, enabling the analysis of object movement and behavior over time. The toolkit distinguishes itself through advanced identity maintenance, utilizing Kalman filters for linear motion tracking and sparse optical flow for camera motion estimation. It features multi-stage object association to recover occluded objects and non-linear motion t

    Evaluates the accuracy and precision of tracking algorithms against ground-truth datasets using standardized metrics.

    Pythonbytetrackmulti-object-trackingoc-sort
    عرض على GitHub↗2,565
  • vincentqyw/image-matching-webuiالصورة الرمزية لـ Vincentqyw

    Vincentqyw/image-matching-webui

    1,283عرض على GitHub↗

    This project is a web-based platform designed for benchmarking, visualizing, and evaluating computer vision algorithms focused on image feature extraction and matching. It provides a unified interface to compare the performance and accuracy of different models by processing image pairs or live video streams. The system distinguishes itself through a modular architecture that allows users to define custom processing pipelines and register external algorithms via configuration files. It incorporates geometric verification techniques to refine visual data and improve the precision of detected co

    Provides a platform for comparing the accuracy and performance of various feature extraction and matching algorithms.

    Pythonaspanformerdeep-learningfeature-matching
    عرض على GitHub↗1,283
  1. Home
  2. Artificial Intelligence & ML
  3. Computer Vision Benchmarks

استكشف الوسوم الفرعية

  • Restoration BenchmarksStandardized evaluation suites for comparing the performance of image and video restoration models. **Distinct from Computer Vision Benchmarks:** Distinct from Computer Vision Benchmarks: focuses on restoration-specific metrics (LPIPS, NIQE, SSIM) rather than general recognition accuracy.