7 रिपॉजिटरी
Generating numerical vectors that describe keypoints to allow image matching.
Distinct from Computer Vision Features: Focuses on the creation of the descriptor vector rather than just the extraction of the feature point.
Explore 7 awesome GitHub repositories matching artificial intelligence & ml · Feature Descriptor Computation. Refine with filters or upvote what's useful.
This repository contains programming assignments and lecture notes from Andrew Ng's foundational deep learning course specialization on Coursera. The materials cover core neural network training techniques including optimization algorithms, normalization methods, regularization approaches, parameter initialization strategies, and learning rate scheduling to improve model convergence and generalization. The coursework explores design principles where successive neural network layers learn progressively more abstract feature representations from input data. It provides guidance on selecting ope
Build deeper layers that compute more complex input features than earlier layers in a neural network.
GoCV is a computer vision library and Go language binding for OpenCV. It serves as an image processing toolkit and deep learning inference engine, providing programmatic access to a wide range of algorithms for image manipulation, object detection, and video analysis. The project differentiates itself through high-performance native bindings and hardware acceleration. It utilizes a foreign function interface to map Go calls to C++ functions and includes a hardware-agnostic backend dispatch to route neural network tasks to computation engines such as CUDA and OpenVINO. The library covers a br
Generates numerical representations of keypoints to enable comparison and matching of different images.
DeepSORT एक रियल-टाइम मल्टी-ऑब्जेक्ट ट्रैकिंग फ़्रेमवर्क है जिसे वीडियो फ़्रेम में कई ऑब्जेक्ट्स की सुसंगत पहचान बनाए रखने के लिए डिज़ाइन किया गया है। यह वीडियो डेटा के अनुक्रम के माध्यम से ऑब्जेक्ट्स को ट्रैक करने के लिए मोशन डिस्क्रिप्टर्स के साथ डीप लर्निंग अपीयरेंस सुविधाओं को इंटीग्रेट करता है। सिस्टम व्यक्ति की पुन: पहचान (re-identification) के लिए उच्च-आयामी विज़ुअल डिस्क्रिप्टर्स उत्पन्न करने के लिए एक डीप कन्वेन्शनल न्यूरल नेटवर्क का उपयोग करता है। इन अपीयरेंस सुविधाओं को Kalman फ़िल्टरिंग के माध्यम से मोशन एस्टिमेशन के साथ जोड़ा जाता है और मौजूदा ट्रैक्स के साथ डिटेक्शन्स को बेहतर ढंग से जोड़ने के लिए हंगेरियन एल्गोरिदम का उपयोग करके हल किया जाता है। फ़्रेमवर्क में ऑब्जेक्ट लाइफसाइकिल को संभालने के लिए गेटिंग-आधारित एसोसिएशन फ़िल्टरिंग और स्टेट-आधारित ट्रैक प्रबंधन के लिए क्षमताएं शामिल हैं। यह वीडियो फ़्रेम पर ट्रैकिंग रिज़ल्ट्स को रेंडर करने और स्थापित बेंचमार्क के खिलाफ ट्रैकिंग प्रदर्शन का मूल्यांकन करने के लिए टूल्स भी प्रदान करता है।
Generates numerical feature descriptors for bounding boxes to enable similarity comparison.
OpenCVSharp is a .NET library that wraps native OpenCV functions, providing C# developers with access to OpenCV's computer vision capabilities through an API that mirrors the native C/C++ style. It serves as a managed wrapper for image processing, feature detection, object detection, and image manipulation tasks, while also handling automatic disposal of unmanaged OpenCV resources like Mat objects to prevent memory leaks in .NET applications. The library enables keypoint detection and descriptor extraction using algorithms such as AKAZE, BRISK, or FAST, with brute-force or FLANN-based matchin
Chains keypoint detection, descriptor extraction, and brute-force or FLANN-based matching.
ArrayFire एक हार्डवेयर-अज्ञेयवादी (hardware-agnostic) कंप्यूट फ्रेमवर्क और JIT-कंपाइल किया गया टेंसर इंजन है जिसे उच्च-प्रदर्शन संख्यात्मक कंप्यूटिंग के लिए डिज़ाइन किया गया है। यह एक GPU न्यूमेरिकल कंप्यूटिंग लाइब्रेरी और पैरेलल सिग्नल प्रोसेसिंग टूलकिट के रूप में कार्य करता है जो हार्डवेयर बैकएंड को एब्स्ट्रैक्ट करता है, जिससे एक ही कोडबेस विभिन्न GPU आर्किटेक्चर और CPUs पर निष्पादित हो सकता है। यह प्रोजेक्ट एक JIT इंजन के माध्यम से खुद को अलग करता है जो ऑपरेशन्स को फ्यूज करने और मेमोरी ओवरहेड को कम करने के लिए एक्सप्रेशन कंपाइलेशन का उपयोग करता है। यह कंप्यूटेशन चेन को ऑप्टिमाइज़ करने के लिए एक डिफर्ड एक्जीक्यूशन ग्राफ का उपयोग करता है और CUDA तथा OpenCL जैसे बाहरी कंप्यूट प्लेटफॉर्म के साथ डेटा और निष्पादन संदर्भ साझा करने के लिए इंटरऑपरेबिलिटी प्रिमिटिव्स प्रदान करता है। यह लाइब्रेरी पैरेलल लीनियर अलजेब्रा, डिजिटल सिग्नल प्रोसेसिंग, और त्वरित कंप्यूटर विज़न सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करती है। यह मशीन लर्निंग इम्प्लीमेंटेशन, वित्तीय मॉडलिंग सिमुलेशन, और भौतिक प्रणाली सिमुलेशन के लिए आंशिक अंतर समीकरणों (partial differential equations) को हल करने के लिए उपकरण प्रदान करती है। इसका टेंसर मैनेजमेंट सिस्टम मल्टी-डायमेंशनल ऐरे एलोकेशन, स्लाइसिंग, और होस्ट-डिवाइस डेटा ट्रांसफर को संभालता है।
Generates numerical representations of image regions to enable efficient comparison between different images.
This project is a Python bio-imaging toolkit and analysis suite designed for processing and analyzing microscopy and medical images. It provides a collection of tools for image quantification, medical image segmentation, and general bio-imaging workflows. The suite includes specialized capabilities for quantifying biological data, such as measuring neuron branching complexity via Sholl analysis, calculating particle size distributions, and tracking wound area in scratch assays. It also features a medical image segmentation library that implements U-Net architectures for isolating anatomical s
Generates numerical descriptors for keypoints that capture scale and orientation for feature matching.
Vim is a state space model vision framework designed for image classification and visual representation learning. It functions as a computer vision research tool that converts two-dimensional image grids into one-dimensional sequences to extract spatial features. The system implements a linear-scaling image classifier that replaces quadratic attention mechanisms with state space operations. This approach utilizes bidirectional sequence modeling and selective gating mechanisms to process visual data. The framework covers computer vision benchmarking and image classification research, providin
Organizes visual data across multiple abstraction levels to capture local and global context.