20 repository-uri
Systems designed to maintain the persistent identity of multiple objects across continuous video streams and live feeds.
Explore 20 awesome GitHub repositories matching artificial intelligence & ml · Object Tracking Systems. Refine with filters or upvote what's useful.
Ultralytics is a comprehensive computer vision framework designed for training, validating, and deploying deep learning models across a wide range of visual recognition tasks. It provides a unified interface for core operations including object detection, instance segmentation, pose estimation, and image classification. By utilizing a modular architecture, the platform allows users to swap model components to balance inference speed and accuracy requirements for diverse applications. The framework distinguishes itself through its support for real-time processing and flexible deployment. It in
Maintains persistent identity across continuous video feeds for multiple detected objects.
PaddleDetection is an object detection framework designed for the end-to-end development, training, and deployment of computer vision models. It provides a comprehensive library of modular neural network architectures and pipelines that support object detection, instance segmentation, and multi-object tracking tasks. The project distinguishes itself through a configuration-driven approach that decouples model components like backbones and heads, allowing for the flexible assembly of custom vision workflows. It incorporates advanced techniques such as anchor-free detection logic, joint detecti
Monitors moving objects across single or multiple camera feeds to analyze traffic flow and pedestrian movement patterns in real-time.
jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti
Maintains unique object identities across a network of multiple cameras to handle occlusions.
This project is a comprehensive collection of educational examples and reference implementations for building vision and language models using PyTorch. It serves as a deep learning tutorial covering the end-to-end process of developing neural networks, from initial architecture definition to final production deployment. The repository provides detailed guides on implementing a wide range of domain-specific models, including convolutional neural networks for object detection and segmentation, as well as transformer and recurrent architectures for natural language processing. It emphasizes gene
Implements assignment algorithms to match detected object boxes with existing tracking identities.
This project is a computer vision system for object segmentation and tracking across images and videos. It employs models capable of identifying and masking objects using text prompts, bounding boxes, click points, or image exemplars. The system differentiates itself through memory-based video tracking and shared-memory architectures that maintain consistent object identities over time. It supports multi-object processing in single computation passes to increase frame throughput and utilizes iterative refinement to correct segmentation boundaries through sequential prompts. The software also
Tracks multiple objects simultaneously using a shared-memory approach to maximize frame throughput.
ByteTrack is a multi-object tracking framework that implements the ByteTrack algorithm, an ECCV 2022 method designed to recover occluded objects and reduce trajectory fragmentation. The core innovation of the project is its association algorithm, which processes every detection box—including low-confidence ones—by using separate high and low score thresholds, Kalman filter motion prediction, and Hungarian algorithm matching to produce consistent object identities across video frames. The project distinguishes itself by its comprehensive approach to handling occlusions and fragmented trajector
Implements the ByteTrack association algorithm that matches every detection box to existing track IDs.
Follows shoppers through a store by stitching together video feeds from multiple cameras to analyze movement patterns.
DeepSORT este un framework de urmărire multi-obiect în timp real conceput pentru a menține identități consistente ale mai multor obiecte pe parcursul cadrelor video. Integrează caracteristici de aspect deep learning cu descriptori de mișcare pentru a urmări obiectele printr-o secvență de date video. Sistemul utilizează o rețea neuronală convoluțională profundă pentru a genera descriptori vizuali de înaltă dimensiune pentru re-identificarea persoanelor. Aceste caracteristici de aspect sunt combinate cu estimarea mișcării prin filtrare Kalman și rezolvate folosind algoritmul maghiar pentru a asocia optim detecțiile cu pistele existente. Framework-ul include capabilități pentru filtrarea asocierii bazată pe gating și gestionarea pistelor bazată pe stare pentru a gestiona ciclurile de viață ale obiectelor. Oferă, de asemenea, instrumente pentru redarea rezultatelor urmăririi pe cadrele video și evaluarea performanței urmăririi față de benchmark-urile stabilite.
Maintains consistent identities of multiple objects across a sequence of video frames.
Gluon-CV este o bibliotecă de computer vision pentru MXNet care oferă o colecție cuprinzătoare de arhitecturi de viziune pre-implementate și pipeline-uri de antrenament. Servește drept toolkit de cercetare în deep learning și o grădină zoologică de modele (model zoo) care conține ponderi pre-antrenate de ultimă generație pentru analiza imaginilor și a videoclipurilor. Proiectul include o bibliotecă specializată de estimare a posturii umane și un toolkit de compresie a modelelor. Aceste instrumente permit tăierea (pruning) și cuantizarea modelelor de deep learning pentru a crește viteza de inferență și a facilita implementarea pe hardware edge cu resurse limitate. Biblioteca acoperă o gamă largă de capabilități de viziune, inclusiv clasificarea imaginilor, detectarea obiectelor și segmentarea semantică și de instanță. De asemenea, oferă instrumente pentru analiza video, cum ar fi recunoașterea acțiunilor, urmărirea obiectelor și estimarea adâncimii monoculare. Antrenamentul este susținut prin pipeline-uri automatizate și sarcini de lucru distribuite multi-GPU pentru a accelera convergența modelului.
Matches and identifies specific individuals across different camera scenes using visual features.
Acest proiect este un framework PyTorch de re-identificare a persoanelor, conceput pentru antrenarea și evaluarea modelelor care identifică indivizi prin diferite unghiuri ale camerelor video. Oferă un pipeline complet de antrenare a modelelor, un extractor de caracteristici deep learning pentru convertirea imaginilor în vectori numerici și o suită de instrumente de benchmarking pentru viziunea artificială pentru a măsura acuratețea regăsirii identității. Framework-ul include un toolkit specializat de transfer learning care suportă înghețarea straturilor, optimizarea etapizată a ratei de învățare și rate de învățare diferențiale pentru fine-tuning-ul modelelor preantrenate. Se distinge printr-un motor extensibil care permite dezvoltarea de logică de antrenare personalizată și implementarea unor obiective de optimizare specifice, cum ar fi hard-sample triplet loss mining și label smoothing. Sistemul acoperă gestionarea cuprinzătoare a seturilor de date, inclusiv suport pentru benchmark-uri standard, eșantionare echilibrată a batch-urilor și augmentarea imaginilor. Oferă utilitare de evaluare pentru calcularea rangurilor de regăsire și a distanțelor dintre caracteristici, precum și instrumente de vizualizare pentru generarea de hărți de activare (heatmaps) și galerii de regăsire clasificate. Proiectul este implementat în Python și utilizează PyTorch pentru operațiunile sale de deep learning.
Computes specialized accuracy, rank, and distance measures to quantify the effectiveness of identity matching across camera views.
Roboflow Sports is a sports video analysis system that combines object detection and tracking with bird's-eye field visualization. Its core pipeline detects and tracks players, referees, and balls across video frames, then maps those tracked positions onto a radar-style overhead view of the playing field. The system goes beyond basic detection by localizing field boundaries and key landmarks such as pitch lines and corners, enabling spatial mapping of player positions relative to the field geometry. It classifies detected players by team affiliation through visual feature extraction and clust
Associates detections across frames using Kalman filters for motion prediction and appearance features for re-identifying occluded objects.
Acest proiect este o resursă educațională cuprinzătoare și un curs pentru construirea de rețele neuronale folosind PyTorch. Acoperă elementele fundamentale ale deep learning-ului, inclusiv manipularea tensorilor, diferențierea automată și construcția componentelor modulare de rețele neuronale. Repository-ul servește drept ghid tehnic pentru mai multe domenii specializate. Oferă detalii de implementare pentru sarcini de computer vision, cum ar fi clasificarea imaginilor, detecția obiectelor și segmentarea semantică, precum și fluxuri de lucru de procesare a limbajului natural (NLP) care implică transformatoare, rețele recurente și modele generative. În plus, include o referință pentru AI generativ, concentrându-se în mod specific pe sinteza de imagini prin modele de difuzie și rețele adversariale. Materialul se extinde către optimizarea modelelor și pipeline-uri de deployment. Acoperă tehnici pentru reducerea dimensiunii modelelor și creșterea vitezei de inferență prin cuantizare și exportul modelelor în formate precum ONNX și TensorRT. Alte domenii de capabilitate includ ingineria datelor pentru încărcarea paralelă, evaluarea modelelor folosind metrici personalizate și deployment-ul modelelor de limbaj mari (LLM) open-source. Proiectul este livrat în principal sub formă de serie de Jupyter Notebooks.
Associates new detections with existing tracking IDs based on the intersection over union of bounding boxes.
This project is a PyTorch-based deep learning framework and supervised learning baseline for person and vehicle re-identification. It provides a complete pipeline for training and evaluating models designed to extract identity-based feature embeddings and match the same entity across different camera views. The framework distinguishes itself with support for cross-modality identity matching, enabling the retrieval of identities across different imaging sensors such as RGB and infrared. It also includes advanced retrieval refinement through re-ranking techniques, utilizing reciprocal encoding
Implements a complete PyTorch framework for training and evaluating person re-identification models.
Acest proiect este un framework de urmărire multi-obiect conceput pentru a atribui identități persistente casetelor de delimitare detectate în cadre video consecutive. Acesta funcționează ca un algoritm de urmărire prin viziune computerizată care monitorizează mai multe ținte în mișcare în timp real, asociind detecțiile cu etichete consistente. Sistemul utilizează o abordare de estimare a stării centrată pe un filtru Kalman pentru a prezice pozițiile viitoare ale obiectelor și a menține identitatea în timpul pauzelor de detecție. Utilizează algoritmul maghiar pentru asocierea optimă a datelor și calculează intersecția peste reuniune (IoU) pentru a potrivi locațiile de urmărire prezise cu detecțiile reale. Pipeline-ul de procesare gestionează un registru de urmăriri active folosind un model liniar de viteză constantă pentru a simplifica tranzițiile de stare. Acesta efectuează procesarea recursivă cadru cu cadru pentru a actualiza starea tuturor obiectelor urmărite pe măsură ce sunt analizate imagini noi.
Provides a comprehensive system for assigning persistent identities to detected objects across video streams.
FairMOT este un framework de tracking multi-obiect și un model de deep learning conceput pentru a identifica și urmări entități multiple în cadre video. Implementează un pipeline unificat care integrează detecția obiectelor și re-identificarea identității într-o rețea comună single-stage. Sistemul utilizează o metodă de detecție fără ancore pentru a prezice centrele obiectelor și dimensiunile bounding box-urilor. Menține consistența identității între cadre consecutive prin generarea de vectori de embedding de înaltă dimensiune pentru re-identificare și utilizarea unui filtru Kalman pentru predicția stării de mișcare. Framework-ul acoperă o gamă largă de capabilități de computer vision, inclusiv detecția obiectelor în timp real și utilizarea algoritmului Hungarian pentru atribuirea tracklet-urilor. Include, de asemenea, utilitare pentru antrenarea modelelor pe seturi de date personalizate și generarea de vizualizări video cu bounding box-uri suprapuse și identificatori persistenți.
Provides a complete system for maintaining the persistent identity of multiple objects across continuous video streams.
fast-reid is a PyTorch-based computer vision framework designed for building, training, and deploying deep learning models for identity-based vision tasks. It provides a specialized toolbox for person re-identification and vehicle re-identification, enabling the matching of individuals and vehicles across non-overlapping camera views. The project includes tools for person attribute recognition to identify specific physical characteristics and traits. It features a modular model zoo that allows for the swapping and benchmarking of different re-identification architectures. The framework cover
Matches individuals across non-overlapping camera views using deep learning for identity tracking.
Human is a TensorFlow.js computer vision library used for face, body, and hand tracking within the browser or Node.js. It provides a framework for human pose and gesture tracking, facial recognition, and biometric liveness detection to verify a live human presence. The project distinguishes itself through a full suite of identity and motion tools, including a facial recognition framework that generates embeddings for similarity matching and a background segmenter for separating humans from their environment. It incorporates a liveness detector to prevent spoofing during facial analysis. The
Implements logic to associate detected body parts and features with specific individuals for consistent tracking.
This project is a multi-object tracking library and computer vision toolkit designed to maintain consistent identity IDs for objects across video frames. It provides a motion-based object tracking system that converts raw detections into stable temporal tracks, enabling the analysis of object movement and behavior over time. The toolkit distinguishes itself through advanced identity maintenance, utilizing Kalman filters for linear motion tracking and sparse optical flow for camera motion estimation. It features multi-stage object association to recover occluded objects and non-linear motion t
Maintains consistent identity IDs for multiple objects across video frames to analyze movement and behavior.
Acest proiect este un pipeline de computer vision care integrează detectarea și urmărirea obiectelor pentru a monitoriza obiectele în mișcare în fluxurile video. Funcționează ca un instrument de analiză end-to-end care procesează cadrele video pentru a identifica, clasifica și menține identitatea unică a obiectelor pe măsură ce se mișcă printr-o scenă. Sistemul utilizează o combinație de inferență deep learning pentru detectare și estimarea mișcării pentru a asigura continuitatea temporală. Prin împerecherea descriptorilor de aspect vizual cu modelarea predictivă a mișcării, menține identitățile obiectelor chiar și în timpul ocluziilor temporale sau când suprapunerea spațială este insuficientă. Framework-ul folosește procesarea secvențială pentru a sincroniza rezultatele detectării cu logica de urmărire, permițând monitorizarea consistentă a tiparelor de mișcare. Dincolo de urmărirea de bază, software-ul include capabilități pentru cuantificarea activității într-un flux video. Suportă calcularea numărului total de obiecte sau vehicule pe măsură ce traversează linii desemnate sau intră în zone specifice. Implementarea este structurată ca un framework de dezvoltare pentru construirea de aplicații de viziune personalizate care interpretează și extrag date din medii dinamice.
Implements a computer vision pipeline that detects and tracks objects across video frames using deep learning models.
This project is a computer vision framework designed for the detection, identification, and tracking of human subjects within video streams. It provides an integrated system for locating individuals, generating biometric models from image datasets, and maintaining identity labels across consecutive video frames. The system distinguishes itself through its ability to maintain identity persistence across multiple camera feeds. By utilizing deep learning inference to extract feature vector embeddings and applying motion prediction algorithms, it links unique identity signatures across disparate
Links unique identity signatures across disparate camera feeds to maintain consistent tracking in complex environments.