21 个仓库
Deep learning frameworks built on PyTorch for building and training neural network models with GPU acceleration.
Distinct from Deep Learning Frameworks: Distinct from general Deep Learning Frameworks: specifies PyTorch as the underlying framework, not framework-agnostic.
Explore 21 awesome GitHub repositories matching artificial intelligence & ml · PyTorch-Based Frameworks. Refine with filters or upvote what's useful.
This library provides a comprehensive collection of modular building blocks and research-backed architectures for implementing vision transformers within the PyTorch framework. It serves as a centralized repository for constructing, training, and analyzing attention-based models, offering a wide array of specialized variants designed for image classification and visual representation learning. The project distinguishes itself through a focus on architectural efficiency and flexibility, supporting diverse input formats including non-square images and volumetric data like video. It incorporates
Provides specialized transformer variants and tools for image classification, visual representation learning, and model observability.
This project is a deep learning curriculum and a collection of PyTorch tutorials designed for deep learning education. It provides a structured set of technical documents and runnable notebooks that translate theoretical machine learning concepts into executable code. The repository includes implementation guides for various neural network architectures, specifically covering convolutional, recurrent, and transformer-based models. It provides practical examples for building computer vision pipelines for object detection and semantic segmentation, as well as natural language processing tools f
Implements neural network models using the PyTorch framework for tensor operations and automatic differentiation.
This PyTorch-based deep learning library provides a framework for analyzing and forecasting temporal data. It implements specialized architectures for time series forecasting, anomaly detection, data imputation, and classification. The project distinguishes itself through the inclusion of zero-shot inference capabilities, allowing large-scale temporal models to be evaluated on unseen datasets without requiring task-specific fine-tuning. The framework covers a broad range of analytical capabilities, including the recovery of missing values in incomplete datasets, the identification of irregul
Provides a PyTorch-based deep learning framework specifically for temporal data analysis and forecasting.
AllenNLP is a PyTorch-based research library and deep learning language toolkit designed for developing and training neural network architectures for linguistic tasks. It provides a distributed training system that coordinates data and gradients across multiple GPUs and a framework for integrating pretrained transformer architectures. The system distinguishes itself with a dedicated algorithmic bias mitigation tool used to identify and reduce bias in linguistic model predictions. It also includes model influence analysis to interpret predictions by calculating the influence of specific traini
Built as a deep learning framework on top of PyTorch for linguistic research and development.
BigDL 是一个 PyTorch 加速框架和分布式推理引擎,专为大语言模型设计。它提供了一个在 Intel 硬件上运行模型的工具包,集成了量化工具和用于参数高效微调的库。 该项目通过使用流水线并行将模型工作负载分布在多个硬件加速器上而脱颖而出。它利用低位整数量化和推测解码来减少内存占用并降低文本生成延迟。 该系统涵盖了模型优化的广泛功能,包括权重压缩和量化模型加载。它还支持硬件加速的训练例程,以使预训练模型适应特定任务。
Provides a toolkit for optimizing and executing PyTorch models on hardware accelerators via weight compression and parallelism.
DocTR is a deep learning OCR library built on PyTorch that detects and transcribes text in document images using a two-stage detection-recognition pipeline. It provides a complete framework for building and deploying OCR pipelines with pretrained models available through the Hugging Face Hub, and supports exporting trained models to ONNX format for cross-runtime deployment. The library offers end-to-end OCR pipelines that combine text detection and recognition to extract all text from document images or PDFs, with support for rotated page handling and varied text orientations. It includes cap
Provides a complete PyTorch-based framework for building and deploying OCR pipelines with pretrained models.
该项目是 EfficientDet 架构的 PyTorch 实现,专为实时目标检测而设计。它提供了一个神经网络和推理引擎,能够识别并定位图像或视频流中的多个对象。 该实现包括带有优化权重的预训练计算机视觉模型,无需从头开始训练即可进行即时推理和微调。 该项目涵盖了计算机视觉模型优化的完整流水线,包括自定义目标检测训练和模型权重优化。它结合了双向特征融合、复合缩放神经网络架构和基于锚点的区域建议等结构组件,以平衡推理速度和检测准确性。
Builds upon the PyTorch framework to implement a real-time object detection system with GPU acceleration.
SimSwap 是一个基于 PyTorch 构建的深度学习换脸框架与计算机视觉媒体处理器。它作为一种图像合成工具,旨在利用单个训练模型将图像与视频中的人物身份替换为目标人脸。 该系统作为视频身份替换工具运行,在保持源媒体原始表情与光照的同时,跨帧交换身份。它通过自动化的面部特征映射,实现了数字身份操纵与合成媒体的制作。 该框架既支持应用训练好的模型在媒体中进行换脸,也支持使用特定图像数据集训练自定义换脸模型。
Built as a deep learning framework leveraging PyTorch for generating synthetic facial imagery.
moco 是一个用于自监督视觉表示学习的动量对比(Momentum Contrast)PyTorch 实现。它作为一个基于研究的框架,通过最大化同一图像不同视图之间的相似性,从无标签数据集中提取高层图像特征。 该系统利用非对称编码器架构,由快速学习的在线编码器和缓慢演进的动量编码器组成,以稳定训练过程。它采用基于字典的方法,将查询图像与动态负样本队列进行对比,从而在无需人工标注的情况下学习具有区分度的视觉特征。 该框架涵盖了端到端的对比学习工作流,包括无监督视觉表示学习和无标签图像分析。它利用 GPU 加速的张量运算进行高维向量相似度计算和模型训练。
Ships a PyTorch-based framework for momentum contrast to learn visual representations from unlabeled data.
Kaolin 是一个 PyTorch 3D 深度学习库,提供了一套全面的工具,用于 3D 几何处理、物理模拟、数据可视化和用于计算机视觉的梯度渲染。 该库包括一个可微分的 3D 渲染器和一个用于转换和变换 3D 表示(如网格和点云)的几何处理工具包。它还具有一个 3D 物理模拟引擎,用于计算三维物体和场景之间的物理交互和碰撞。 该工具包提供用于 3D 数据可视化的实用工具,包括创建交互式视图和转盘动画。其他功能涵盖 3D 数据集管理、数据预处理和 3D 表示渲染。
Acts as a PyTorch-based framework for accelerating research in 3D computer vision and deep learning.
该项目是一个 PyTorch 人员重识别框架,专为训练和评估识别不同摄像机视角下个人的模型而设计。它提供了一个完整的模型训练管线、用于将图像转换为数字向量的深度学习特征提取器,以及一套用于衡量身份检索准确性的计算机视觉基准测试工具。 该框架包括一个专门的迁移学习工具包,支持层冻结、分阶段学习率优化和用于微调预训练模型的差异化学习率。它通过一个可扩展的引擎脱颖而出,该引擎允许开发自定义训练逻辑,并实现特定的优化目标,如困难样本三元组损失挖掘(hard-sample triplet loss mining)和标签平滑。 该系统涵盖了全面的数据集管理,包括对标准基准、平衡批次采样和图像增强的支持。它提供用于计算检索排名和特征距离的评估实用程序,以及用于生成激活热力图和排名检索库的可视化工具。 该项目使用 Python 实现,并利用 PyTorch 进行深度学习操作。
A comprehensive PyTorch-based framework for training and evaluating person re-identification models.
mmocr 是一个基于 PyTorch 的光学字符识别(OCR)框架,旨在训练和部署文本检测、识别和关键信息提取模型。它作为一个全面的场景文本检测和识别工具箱,提供用于定位文本区域并将视觉文本转换为机器编码字符串的专用库。 该项目的独特之处在于用于关键信息提取的研究框架和高级文本定位功能。这些包括使用 Transformer 的基于点的定位,以及使用参数化贝塞尔曲线来识别和转录任意形状的文本。 该框架涵盖了广泛的计算机视觉功能,包括用于增强和标准化多样化 OCR 数据集的流水线管理、具有分布式扩展的模型训练,以及使用标准 OCR 指标的性能评估。它还提供用于几何多边形操作和结果可视化的实用程序,以便根据真实标注审计预测。 该系统使用 Python 实现,并支持通过 Docker 环境打包进行安装。
Serves as a PyTorch-based toolbox for training and deploying text detection, recognition, and key information extraction models.
此项目是一个自监督对比学习框架,旨在训练深度学习模型从图像中学习视觉表示,而无需使用人类提供的标签。它提供了一个系统,用于开发可适应下游计算机视觉任务的预训练视觉表示模型。 该框架包括用于半监督图像分类的工具,它结合了大型未标记数据集和小型标记集以提高准确性。它还具有线性探测评估工具,通过在冻结的表示之上训练简单的线性分类器来评估学习到的图像特征的质量。 代码库涵盖了分布式深度学习训练和硬件加速以处理大批量数据,以及优化原语,如余弦衰减学习率调度和权重衰减正则化。它还提供了模型管理实用程序,包括在不同深度学习框架格式之间转换预训练检查点,以及用于模型部署的工具。 该实现以 Jupyter Notebooks 集合的形式提供。
Provides a specialized framework for training models to learn visual representations using contrastive objectives.
CV-Backbones 是一个计算机视觉骨干网络库和模型库,提供了一系列预定义的神经网络架构,用于提取视觉特征和处理图像数据。它作为一个可重用的深度学习组件的 PyTorch 视觉框架,专为图像分析和视觉表征学习而设计。 该库专注于高效的神经网络架构,以在保持特征提取性能的同时降低计算开销。这是通过实现 GhostNet 和 MLP 等轻量级模型设计来实现的。 该项目涵盖了广泛的模型架构,包括卷积神经网络和 Transformer。它包含一个用于切换骨干网络实现的模块化系统,以及一个用于加载预训练权重以加速收敛的机制。
Provides a set of reusable deep learning components built on the PyTorch framework for image analysis.
Tensor-Puzzles 是一套教育练习和数值计算教程,旨在掌握 PyTorch 中的张量运算和广播规则。它作为一个实现训练器,用户通过重新实现深度学习数学原语来练习将数学公式转换为代码。 该项目利用一套基于约束的练习,限制可用的库调用以强制使用特定的张量原语。这些挑战被结构化为顺序谜题,要求用户使用模块化实现模式解决任务,其中复杂函数被分解为更简单的依赖运算。 通过集成的 PyTorch 执行环境确保正确性,该环境使用参考实现验证和数值容差检查。系统验证用户定义的输出是否与参考结果匹配,并遵守标准的多元数组广播规则。
Provides an execution environment that runs user code within a live PyTorch session for validation.
Imaginaire 是一个 PyTorch 图像合成库和神经图像翻译框架,旨在生成高分辨率的合成视觉内容。它作为一个深度学习视觉生成器,使用监督和无监督方法将语义图像和视频映射为照片级真实版本。 该项目包括一个用于渲染 3D 环境的专用工具,它将基于块的世界表示转换为照片级真实场景,同时保持长期的视觉一致性。它进一步支持利用参考图像确保序列间时间稳定性的照片级真实视频翻译。 该框架涵盖了广泛的视觉合成功能,包括多域风格迁移、图像到图像映射,以及基于深度学习研究的合成图像和视频生成。
Built as a deep learning framework leveraging PyTorch for high-dimensional tensor computation and neural network training.
该项目是用于视频动作识别的 3D 残差网络(3D Residual Networks)的 PyTorch 实现。它提供了一种时空架构,通过分析空间帧和时间运动来对视频片段中的人类活动进行分类。 该系统包含一个分布式模型训练框架,以加速跨多个计算节点的学习过程。它支持预训练模型权重的部署与微调,允许将现有网络适配到特定的新数据集。 代码库涵盖了时空学习的全流程,包括用于将原始文件转换为图像序列的视频数据集预处理工具、动作推理功能以及用于计算识别准确率的指标。
Built as a PyTorch-based framework utilizing GPUs for deep learning model training and deployment.
这是一个基于 PyTorch 的场景文本识别框架和工具包。它提供了一个深度学习流水线,用于从自然环境的图像中提取字符和单词,涵盖了从训练数据准备到模型验证的完整过程。 该框架作为衡量文本识别模型准确性和推理速度的标准化基准。它包括用于计算识别准确率和测量每张图像 GPU 处理时间的工具,以评估模型在一致数据集上的性能。 该系统结合了视觉和序列处理阶段,利用卷积特征提取和循环序列建模。它包括用于文本和索引转换的数据工程工具,以及用于在训练期间管理数据集分布的批处理级数据平衡功能。
Provides a PyTorch-based deep learning framework for extracting text from images using visual and sequential stages.
Instructor-embedding 是一个自然语言处理框架,旨在将非结构化文本转换为高维数值向量。通过利用基于 Transformer 的编码器架构,该系统促进了大规模数据集上的语义检索、数据分类和相似度分析。 该框架通过指令条件向量投影脱颖而出,它将自然语言指令直接纳入嵌入过程,从而在无需额外训练的情况下提高特定任务的性能。它作为一个对比学习库,允许用户在自定义数据集上微调预训练语言模型,为特定领域创建专业化的嵌入。 该项目提供了一套全面的向量表示管理工具,包括针对标准化指标对模型准确性进行基准测试,以及为快速相似度搜索建立嵌入索引的功能。为了支持在资源受限环境中的部署,该框架包含了混合精度模型量化等优化功能,以减少内存使用并加速推理速度。
Provides a toolkit for fine-tuning pretrained language models on custom datasets to create specialized embeddings for niche domains and specific tasks.
Graph Nets is a graph neural network library and educational toolkit implemented in PyTorch, providing implementations of popular graph representation learning algorithms and research papers. The project covers core graph machine learning tasks including semi-supervised node classification, inductive and unsupervised node embedding generation, and neighborhood feature aggregation. The library supports diverse algorithmic approaches for processing network structures, ranging from shared-parameter graph convolutions and attention-weighted neighborhood aggregation to spectral Chebyshev filtering
Provides a collection of deep learning models and representation algorithms built on top of PyTorch.