9 个仓库
Techniques and activation functions that maintain gradient flow across deep network layers.
Distinct from Gradient Optimization Techniques: Distinct from general gradient optimization: focuses on activation-based stability for deep networks.
Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Gradient Flow Stabilizers. Refine with filters or upvote what's useful.
This project is an educational platform and research toolkit designed to teach deep learning through a combination of mathematical theory, visual diagrams, and executable code. It provides a comprehensive environment for building, training, and evaluating neural networks, grounding complex concepts in interactive computational notebooks that allow for hands-on experimentation. The framework distinguishes itself by interleaving theoretical foundations—including linear algebra, calculus, and probability—with practical implementations across multiple industry-standard libraries. It supports flex
Selects activation functions that maintain gradient flow to prevent vanishing gradients in deep networks.
This project is a comprehensive educational resource and curriculum designed to teach the mathematical foundations and practical implementation of neural networks. It provides a structured path for understanding how computers learn from data, covering core concepts such as gradient descent, backpropagation, and the biological inspiration behind artificial neurons. The platform distinguishes itself by combining theoretical proofs with hands-on implementation exercises. It demonstrates the universal approximation theorem through visual explanations and guides users in building various architect
Analyzes gradient stability across layers to identify and resolve slow learning in deep architectures.
This project is a PyTorch-based generative framework and implementation template for building Generative Adversarial Networks. It provides a collection of foundational toolkits and architectural patterns designed to synthesize high-quality artificial data while focusing on the stability of adversarial neural networks. The framework distinguishes itself through a specialized toolkit for conditional image generation, which integrates discrete labels and auxiliary classification into the training process. It utilizes specific mechanisms to guide the generative process toward target classes by co
Uses leaky activations and avoids max-pooling to maintain stable gradient flow across deep network layers.
This project is a static educational website and comprehensive curriculum focused on computer vision and deep learning. It serves as a public repository of instructional materials, lecture notes, and technical guides specifically detailing convolutional neural networks and visual recognition. The site is developed using static-site generation to host course documentation and student project directories. It provides structured academic resources that guide learners through image classification, generative modeling, and the implementation of various neural network architectures. The curriculum
Provides theoretical guidance on how operations like addition and multiplication influence gradient flow.
The PyTorch Tutorials repository is a collection of educational resources that provides step-by-step guidance on building, training, and deploying neural networks using the PyTorch framework. It covers the complete machine learning workflow, from data loading and model definition through optimization loops and model persistence, with dedicated guides for distributed training, model fine-tuning, and deployment. The tutorials offer practical demonstrations of adapting pre-trained models to new tasks through transfer learning, scaling training across multiple GPUs or machines using PyTorch's dis
Visualizes gradient propagation through network layers to identify vanishing or exploding gradients.
This repository collects illustrated single-page cheat sheets that compress the core topics of Stanford's CS 230 deep learning course into visual reference summaries. The collection covers convolutional neural networks, recurrent neural networks, and practical training techniques, pairing schematic diagrams with mathematical notation to bridge intuition and formal understanding. The cheat sheets are organized by subject area and link related concepts across topics, such as connecting vanishing gradients to LSTM gates, to reinforce the full deep learning workflow. Practical training advice on
Covers GRU and LSTM gated architectures that prevent vanishing and exploding gradients.
Composer 是一个 PyTorch 分布式训练框架,旨在实现大规模模型在多节点 GPU 集群上的扩展。它兼具大语言模型训练器、分布式模型优化器和训练生命周期管理器的功能。 该项目作为深度学习正则化库脱颖而出,提供诸如 Sharpness Aware Minimization、MixUp 和 CutMix 等专业优化技术,以提升模型的泛化能力。它还通过序列长度预热、渐进式层冻结以及用于大规模模型恢复的分片状态检查点技术,优化了训练流程。 该框架涵盖了广泛的功能领域,包括分布式训练编排、混合精度硬件管理和云原生数据流。它还为 GPU 内存诊断、训练发散检测和吞吐量跟踪提供了丰富的监控与可观测性工具。 该项目包含一个命令行启动器,可自动执行跨节点的分布式多 GPU 训练任务。
Clips gradients and manages layer freezing to stabilize and accelerate the training process.
本项目是一个 TensorFlow 元学习框架和研究工具包,旨在实现和训练学习到的优化器。它提供了一套用于开发学习如何优化其他模型的神经网络的工具,取代了传统的基于梯度的优化算法。 该框架包括一个问题集成管理器,允许将多个不同的优化任务组合成单个加权损失函数进行同步训练。它使用工厂模式进行网络实例化,并支持定义自定义目标函数和损失图作为学习算法的目标。 该工具包涵盖了广泛的功能,包括基于梯度的元优化、模型基准测试以及具有可配置展开长度的训练循环执行。它还提供了用于梯度预处理、序列化状态持久化以及报告实验统计数据(如平均最终误差和 epoch 持续时间)的工具。
Applies logarithmic scaling and sign extraction to gradients to ensure numerical stability during training.
This project is a collection of structured study notes and notebooks serving as an educational resource for deep learning and neural network fundamentals. It provides a technical reference for implementing machine learning theory, covering everything from basic network design to the construction of advanced architectures. The material specifically focuses on the implementation of convolutional neural networks for computer vision and sequence models for natural language processing. It includes detailed guidance on building object detection systems, face recognition, and speech transcription mo
Covers weight initialization strategies like He and Xavier to prevent vanishing and exploding gradients.