10 रिपॉजिटरी
Utilities for tracking GPU memory and throughput during model inference.
Explore 10 awesome GitHub repositories matching testing & quality assurance · GPU Performance Profilers. Refine with filters or upvote what's useful.
Faceswap is a comprehensive framework for automated media manipulation and neural face synthesis. It provides a modular pipeline that manages the entire lifecycle of facial feature extraction, deep learning model training, and image conversion. By coordinating complex computer vision workflows, the system enables users to map facial identities between source and destination datasets while maintaining structural alignment and lighting consistency across video frames. The project distinguishes itself through a highly extensible plugin-based architecture that handles hardware-accelerated process
Benchmarks graphics hardware by tracking memory usage and throughput across varying batch sizes to refine pipeline performance.
nvtop एक टर्मिनल-आधारित डैशबोर्ड है जिसका उपयोग कई ग्राफिक्स प्रोसेसर और हार्डवेयर एक्सेलेरेटर के प्रदर्शन, मेमोरी उपयोग और तापमान की निगरानी के लिए किया जाता है। यह एक ही सिस्टम पर कई उपकरणों के स्वास्थ्य और कंप्यूट लोड को ट्रैक करने के लिए एक केंद्रीकृत प्रशासन टूल के रूप में कार्य करता है। यह टूल सिस्टम प्रोसेस IDs को हार्डवेयर संसाधन खपत के साथ सहसंबंधित करके खुद को अलग करता है, जिससे उपयोगकर्ताओं को GPU संसाधनों का उपभोग करने वाले विशिष्ट अनुप्रयोगों की पहचान करने की अनुमति मिलती है। यह एक ही इंटरफ़ेस के भीतर कई अलग-अलग निर्माताओं के हार्डवेयर का समर्थन करने के लिए एक वेंडर-अज्ञेयवादी एब्स्ट्रैक्शन लेयर का उपयोग करता है। सॉफ्टवेयर टेक्स्ट-आधारित इंटरफ़ेस का उपयोग करके रीयल-टाइम प्रदर्शन मेट्रिक्स और प्रति-प्रक्रिया संसाधन एट्रिब्यूशन प्रदान करता है। उपयोगकर्ता इंटरफ़ेस लेआउट का प्रबंधन कर सकते हैं और सत्रों में सेटिंग्स बनाए रखने के लिए स्थानीय कॉन्फ़िगरेशन फ़ाइल के माध्यम से प्रदर्शन प्राथमिकताओं को सहेज सकते हैं।
Analyzes individual process resource consumption on GPUs to identify bottlenecks and memory leaks.
jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti
Analyzes and debugs GPU-accelerated workloads to optimize AI, graphics, and compute performance.
This project is a comprehensive collection of educational examples and reference implementations for building vision and language models using PyTorch. It serves as a deep learning tutorial covering the end-to-end process of developing neural networks, from initial architecture definition to final production deployment. The repository provides detailed guides on implementing a wide range of domain-specific models, including convolutional neural networks for object detection and segmentation, as well as transformer and recurrent architectures for natural language processing. It emphasizes gene
Tracks CPU and GPU utilization and data throughput to identify system-level application bottlenecks.
Lists running GPU processes with PID, user, and memory, and allows termination from the interface.
Measures GPU throughput, utilization, cache hit rates, and memory throughput to identify optimization opportunities.
capa is a binary capability scanner that identifies high-level behaviors and actions an executable can perform, such as network communication or file manipulation. It functions as a malware behavior analysis tool and a MITRE ATT&CK mapping framework, scanning PE, ELF, .NET, and shellcode files through both static analysis and dynamic sandbox report processing. The tool distinguishes itself through a YAML-based detection rule engine that defines detection logic in human-readable files, with conditions expressed as feature combinations and logical operators. It integrates with IDA Pro, Ghidra,
Limits analysis to specific processes by PID when processing dynamic sandbox reports.
btrace is a JVM dynamic tracing tool and performance profiler used for injecting safe instrumentation scripts into a running Java Virtual Machine without requiring a process restart. It functions as a Java agent framework and a Model Context Protocol server, exposing JVM diagnostic operations and tracing tools to large language models and AI assistants. The project distinguishes itself by enabling real-time code injection and bytecode-level instrumentation via a secure binary protocol. It ensures production stability through a static safety analysis engine that blocks unstable code patterns,
Includes utilities for tracking GPU memory and throughput during deep learning model inference.
Perfetto is a platform for system-level performance tracing and analysis on Linux and Android. It combines a high-throughput trace recorder, a SQL-based query engine, and a browser-based visualizer into a single toolchain. The platform covers CPU scheduling and call-stack profiling, native and Java heap memory allocation tracking, GPU and graphics events, and system-wide counters such as CPU frequency and power consumption. The architecture decouples trace recording from offline analysis, using a compact protobuf format for event encoding and columnar storage for efficient SQL queries. The we
Organizes GPU traces into timelines by device and process, displaying per-kernel metric tables and details.
regl is a declarative WebGL library that manages graphics state and GPU resources through functional commands instead of manual binding and state tracking. It provides a command-based drawing abstraction where shaders, attributes, and render state are encapsulated into reusable, compiled functions that can be executed efficiently. What sets regl apart is its scoped state inheritance system, which allows nested drawing commands to inherit and override render state from parent scopes for organized rendering. The library automatically recovers from GPU context loss by restoring buffer and textur
Regl collects CPU and GPU timing data and draw call counts per command for rendering performance diagnostics.