22 रिपॉजिटरी
Logarithmic error functions used to optimize classification performance and prevent learning slowdowns.
Distinct from Binary Cross-Entropy Calculators: Distinct from Binary Cross-Entropy Calculators: covers general cross-entropy implementations for multi-class classification, not just binary tasks.
Explore 22 awesome GitHub repositories matching artificial intelligence & ml · Cross-Entropy Loss Functions. Refine with filters or upvote what's useful.
This project is a collection of educational examples and code for implementing deep learning architectures using the PyTorch framework. It serves as a tutorial and implementation guide for building various neural network architectures for machine learning tasks. The project provides practical implementations for computer vision, including image classification and neural style transfer, as well as natural language processing examples for building sequence models and language predictors. It also covers generative models using adversarial and variational networks to synthesize or transform visua
Implements cross-entropy loss functions to guide the training of classification models.
This project is a comprehensive educational resource and curriculum designed to teach the mathematical foundations and practical implementation of neural networks. It provides a structured path for understanding how computers learn from data, covering core concepts such as gradient descent, backpropagation, and the biological inspiration behind artificial neurons. The platform distinguishes itself by combining theoretical proofs with hands-on implementation exercises. It demonstrates the universal approximation theorem through visual explanations and guides users in building various architect
Uses cross-entropy cost functions to optimize network training and prevent learning saturation.
dalle-mini is a text-to-image model and generative AI system designed to transform natural language descriptions into synthetic images. It functions as an image generation training toolkit and a generative model capable of creating visual representations from text prompts. The project provides a containerized deployment for consistent execution across different computing environments. It includes the necessary scripts and configuration files to train custom generative models from datasets. The system utilizes an autoregressive transformer architecture that treats visual data as discrete toke
Utilizes cross-entropy loss functions to optimize the prediction of image tokens during model training.
This project is a static educational website and comprehensive curriculum focused on computer vision and deep learning. It serves as a public repository of instructional materials, lecture notes, and technical guides specifically detailing convolutional neural networks and visual recognition. The site is developed using static-site generation to host course documentation and student project directories. It provides structured academic resources that guide learners through image classification, generative modeling, and the implementation of various neural network architectures. The curriculum
Instructs on implementing cross-entropy loss functions and regularization to guide the optimization of classifiers.
This project is a comprehensive collection of educational examples and reference implementations for building vision and language models using PyTorch. It serves as a deep learning tutorial covering the end-to-end process of developing neural networks, from initial architecture definition to final production deployment. The repository provides detailed guides on implementing a wide range of domain-specific models, including convolutional neural networks for object detection and segmentation, as well as transformer and recurrent architectures for natural language processing. It emphasizes gene
Implements general cross-entropy loss functions for multi-class classification tasks in PyTorch.
Liger-Kernel is a collection of pre-built fused Triton kernels and patching utilities designed to accelerate large language model training. It provides drop-in kernel replacements for common LLM operations such as RMSNorm, cross-entropy loss, and attention, enabling increased throughput and reduced memory usage while preserving bitwise-exact gradients. The project serves as a toolkit for composing custom model architectures from individual optimized kernels and for patching pre-existing models with minimal code changes. The project distinguishes itself through its ability to perform runtime m
Ships an optimized fused cross-entropy loss kernel for large-vocabulary classification tasks.
tiny-dnn is a header-only C++14 deep learning framework for building, training, and running inference on neural networks. It constructs static computational graphs at compile time using template-based layer composition, with a gradient-based backpropagation engine and minibatch stochastic gradient descent for training, all without external dependencies beyond the C++14 standard library. The framework supports importing pre-trained models from the Caffe framework directly, parsing its binary serialization format without requiring external protocol buffer libraries. It provides CPU-optimized te
Measures the difference between predicted and target values using cross-entropy, mean squared error, or mean absolute error.
Torchtune is a PyTorch-native library for fine-tuning, aligning, and quantizing large language models. It provides a config-driven system for instantiating components, orchestrating distributed training, and managing parameter-efficient fine-tuning with quantization support, all through YAML-based configurations and command-line overrides. The library distinguishes itself through its comprehensive post-training workflow orchestration, combining supervised fine-tuning, preference optimization (DPO, PPO, GRPO), knowledge distillation, and quantization-aware training in a single configurable pip
Provides selectable DPO and RSO loss functions for controlling how models penalize un-preferred responses.
Torchtune is a PyTorch-native library for fine-tuning, aligning, and quantizing large language models. It provides a configurable training pipeline orchestrated through YAML recipes, with CLI overrides and component swapping, distributed training via FSDP2, memory optimizations, and parameter-efficient fine-tuning methods like LoRA, DoRA, and QLoRA. The library distinguishes itself through its YAML-driven configuration system that defines all training parameters and instantiates components from config files, with full CLI override capability for any field or component at launch time. It suppo
Supports switching between DPO and RSO loss variants via a configuration flag to control alignment strategy.
Flashlight एक स्टैंडअलोन C++ मशीन लर्निंग लाइब्रेरी और टेंसर लाइब्रेरी है जिसका उपयोग न्यूरल नेटवर्क बनाने और ट्रेन करने के लिए किया जाता है। यह एक व्यापक न्यूरल नेटवर्क फ्रेमवर्क और ऑटोमैटिक डिफरेंशिएशन इंजन के रूप में कार्य करता है, जो कम्प्यूटेशन ग्राफ बनाने और बैकप्रोपैगेशन के माध्यम से ग्रेडिएंट्स की गणना करने के लिए उपकरण प्रदान करता है। यह प्रोजेक्ट एक वितरित ट्रेनिंग फ्रेमवर्क के रूप में कार्य करता है, जो कई कंप्यूट नोड्स और डिवाइसेस पर ग्रेडिएंट्स और पैरामीटर्स को सिंक्रोनाइज़ करने के लिए ऑल-रिड्यूस ऑपरेशन्स का उपयोग करता है। यह उच्च-प्रदर्शन टेंसर मैनिपुलेशन, नेटिव डिवाइस मेमोरी इंटरऑपरेबिलिटी और बड़े पैमाने पर मॉडल ट्रेनिंग को गति देने के लिए वितरित वर्कर्स में वेट्स को सिंक्रोनाइज़ करने के सिस्टम के गहरे एकीकरण के माध्यम से खुद को अलग करता है। यह फ्रेमवर्क रेजिडुअल ब्लॉक्स और रिकरेंट सेल्स जैसे जटिल आर्किटेक्चर को डिज़ाइन करने के लिए मॉड्यूलर लेयर कंपोज़िशन सहित डीप लर्निंग क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह मॉडल स्टेट्स को बनाए रखने के लिए सीरियलाइजेशन सिस्टम के साथ-साथ इनजेशन और प्रीफेचिंग के लिए व्यापक डेटा प्रबंधन यूटिलिटीज प्रदान करता है। इसके अतिरिक्त, इसमें ट्रेनिंग मेट्रिक्स को ट्रैक करने और सीक्वेंस एरर्स को मापने के लिए मॉनिटरिंग और ऑब्जर्वेबिलिटी टूल्स का एक सूट शामिल है। यह लाइब्रेरी C++ में इम्प्लीमेंट की गई है।
Calculates errors between predictions and targets using standard loss functions like Mean Squared Error and Cross Entropy.
Neuraltalk is an automated image captioning system that generates natural language descriptions for images. It utilizes a deep learning model that integrates a pretrained convolutional neural network for visual feature extraction with a recurrent neural network decoder to produce text sequences. The project provides a full workflow for training and evaluating captioning models, including weight optimization via backpropagation and gradient descent. It includes tools for measuring caption accuracy by comparing generated text against reference descriptions. The system covers data preprocessing
Utilizes cross-entropy loss functions to measure the difference between predicted word distributions and ground-truth labels.
xtuner बड़े भाषा मॉडल के लिए एक व्यापक प्रशिक्षण इंजन है, जो प्री-ट्रेनिंग, सुपरवाइज्ड फाइन-ट्यूनिंग और विज़न-लैंग्वेज मल्टीमॉडल मॉडल के अनुकूलन के लिए एक टूलकिट प्रदान करता है। यह एक वितरित प्रशिक्षण त्वरक और Mixture-of-Experts मॉडल को स्केल करने और मानव फीडबैक से सुदृढीकरण शिक्षण के माध्यम से मॉडल व्यवहार को संरेखित करने के लिए एक विशेष फ्रेमवर्क के रूप में कार्य करता है। प्रोजेक्ट उन्नत मेमोरी और कंप्यूट अनुकूलन के माध्यम से खुद को अलग करता है, जैसे अल्ट्रा-लॉन्ग कॉन्टेक्स्ट विंडो के लिए सीक्वेंस पैरेललिज्म और GPU आइडल समय को कम करने के लिए इंटरलीव्ड पाइपलाइन पैरेललिज्म। यह प्राथमिकता अनुकूलन के लिए एक समर्पित सूट प्रदान करता है, जो मॉडल नीतियों और इनाम प्रणालियों को परिष्कृत करने के लिए Group Relative Policy Optimization और Direct Preference Optimization जैसी तकनीकों को लागू करता है। व्यापक क्षमता क्षेत्र कई नोड्स में वितरित मॉडल प्रशिक्षण, मल्टीमॉडल डेटासेट तैयारी और एडाप्टर-आधारित फाइन-ट्यूनिंग के प्रबंधन को कवर करते हैं। इंजन में मॉडल मूल्यांकन, वेट मर्जिंग और प्रशिक्षित मापदंडों को इन्फरेंस इंजन में निर्यात करने के लिए टूल भी शामिल हैं। प्रशिक्षण का प्रबंधन मानकीकृत कॉन्फ़िगरेशन फाइलों और वितरित लॉन्चरों के माध्यम से किया जाता है ताकि कंप्यूटिंग क्लस्टर में सुसंगत परिणाम सुनिश्चित किए जा सकें।
Compute objective functions for cross-entropy or reinforcement learning to guide model optimization.
LightFM एक Python रिकमेंडेशन लाइब्रेरी और मशीन लर्निंग फ़्रेमवर्क है जिसे यूजर की प्राथमिकताओं की भविष्यवाणी करने के लिए डिज़ाइन किया गया है। यह एक हाइब्रिड रिकमेंडेशन इंजन लागू करता है जो यूजर-आइटम इंटरैक्शन डेटा को वर्णनात्मक मेटाडेटा के साथ एकीकृत करके सहयोगी फ़िल्टरिंग (collaborative filtering) को कंटेंट फ़िल्टरिंग के साथ जोड़ता है। यह सिस्टम उपयोगकर्ताओं और आइटम्स के लेटेंट रिप्रेजेंटेशन को सीखने के लिए हाइब्रिड मैट्रिक्स फ़ैक्टराइज़ेशन का उपयोग करता है। इसे विशेष रूप से इम्प्लिसिट फ़ीडबैक को संभालने के लिए डिज़ाइन किया गया है, जो नकारात्मक रेटिंग की कमी वाले डेटासेट के लिए आइटम प्राथमिकताओं को ऑप्टिमाइज़ करने के लिए Weighted Approximate Rank Pairwise और Bayesian Personalized Ranking जैसे विशेष लॉस फ़ंक्शंस का उपयोग करता है। यह लाइब्रेरी स्टोकेस्टिक ग्रेडिएंट डिसेंट के माध्यम से मॉडल्स को ट्रेन करने, आइटम प्राथमिकता भविष्यवाणियों की गणना करने और मॉडल परिशुद्धता का मूल्यांकन करने के लिए टूल्स प्रदान करती है। यह इंटरैक्शन मैट्रिसेस को फ़ीचर एम्बेडिंग्स के साथ सिंथेसाइज़ करके व्यक्तिगत आइटम रैंकिंग और यूजर बिहेवियर प्रेडिक्शन को सपोर्ट करती है।
Optimizes item preferences using WARP and BPR loss functions for datasets without negative ratings.
यह रिपॉजिटरी एक व्यापक शैक्षिक कार्यक्रम और डीप लर्निंग फ्रेमवर्क है, जिसे नोटबुक और कोड उदाहरणों के माध्यम से PyTorch का उपयोग करके व्यावहारिक डीप लर्निंग सिखाने के लिए डिज़ाइन किया गया है। यह न्यूरल नेटवर्क बनाने, प्रशिक्षित करने और डिप्लॉय करने के लिए एक हाई-लेवल लाइब्रेरी के रूप में कार्य करता है। यह प्रोजेक्ट कंप्यूटर विज़न, नेचुरल लैंग्वेज प्रोसेसिंग और टैबुलर डेटा प्रीप्रोसेसिंग के लिए विशेष टूलकिट प्रदान करता है। यह डिस्क्रिमिनेटिव लर्निंग रेट्स, ट्रेनिंग लॉजिक को कस्टमाइज़ करने के लिए टू-वे कॉलबैक सिस्टम और हाई-लेवल लर्नर एब्स्ट्रैक्शन जैसे उन्नत ट्रेनिंग कंट्रोल्स के माध्यम से खुद को अलग करता है। यह प्रोजेक्ट Jupyter Notebooks की एक श्रृंखला के रूप में उपलब्ध है।
Computes model loss using a variety of algorithms including Cross Entropy and Mean Squared Error.
यह प्रोजेक्ट एक PyTorch पर्सन री-आइडेंटिफिकेशन फ्रेमवर्क है जिसे विभिन्न कैमरा व्यूज में व्यक्तियों की पहचान करने वाले मॉडल्स को प्रशिक्षित और मूल्यांकन करने के लिए डिज़ाइन किया गया है। यह एक पूर्ण मॉडल ट्रेनिंग पाइपलाइन, छवियों को संख्यात्मक वैक्टर में बदलने के लिए एक डीप लर्निंग फीचर एक्सट्रैक्टर और पहचान पुनर्प्राप्ति सटीकता को मापने के लिए कंप्यूटर विज़न बेंचमार्किंग टूल्स का एक सुइट प्रदान करता है। फ्रेमवर्क में एक विशेष ट्रांसफर लर्निंग टूलकिट शामिल है जो प्रीट्रेन्ड मॉडल्स को फाइन-ट्यून करने के लिए लेयर फ्रीज़िंग, स्टेज्ड लर्निंग रेट ऑप्टिमाइज़ेशन और डिफरेंशियल लर्निंग रेट्स का समर्थन करता है। यह एक एक्स्टेंसिबल इंजन के माध्यम से खुद को अलग करता है जो कस्टम ट्रेनिंग लॉजिक के विकास और हार्ड-सैंपल ट्रिपलेट लॉस माइनिंग और लेबल स्मूथिंग जैसे विशिष्ट ऑप्टिमाइज़ेशन उद्देश्यों के इम्प्लीमेंटेशन की अनुमति देता है। सिस्टम व्यापक डेटासेट प्रबंधन को कवर करता है, जिसमें स्टैंडर्ड बेंचमार्क, बैलेंस्ड बैच सैंपलिंग और इमेज ऑगमेंटेशन के लिए समर्थन शामिल है। यह पुनर्प्राप्ति रैंक और फीचर दूरियों की गणना के लिए मूल्यांकन उपयोगिताएं, साथ ही एक्टिवेशन हीटमैप और रैंक की गई पुनर्प्राप्ति गैलरी उत्पन्न करने के लिए विज़ुअलाइज़ेशन टूल्स प्रदान करता है। यह प्रोजेक्ट Python में इम्प्लीमेंट किया गया है और अपने डीप लर्निंग ऑपरेशंस के लिए PyTorch का लाभ उठाता है।
Implements cross-entropy loss with optional label smoothing to regularize the training of classification models.
This is a PyTorch image classification framework designed for training and evaluating convolutional neural networks. It provides a comprehensive library of pre-defined architectures and a training pipeline specifically implemented for the CIFAR-100 benchmark dataset. The framework includes a variety of convolutional neural network implementations, ranging from standard research models to lightweight versions optimized for mobile devices. It features a modular model registry to initialize specific architectures and a benchmarking system to compare the effectiveness of different network designs
Implements multi-class cross-entropy loss functions to guide the training of image classification models.
This is an educational implementation that builds a generative pre-trained transformer (GPT) language model from scratch using PyTorch. The project is structured as a step-by-step tutorial, walking through the construction of a decoder-only transformer architecture and its training loop with clean git commits and an accompanying video lecture for a hands-on learning experience. What sets this implementation apart is its focus on practical reproduction: it provides a workflow to train a 124-million-parameter model from scratch in about one hour on cloud GPU hardware, costing under ten dollars.
Uses cross-entropy loss as the objective function for next-token prediction during language model training.
This project is a comprehensive instructional resource and course for building neural networks using PyTorch. It covers the fundamental building blocks of deep learning, including tensor manipulation, automatic differentiation, and the construction of modular neural network components. The repository serves as a technical guide for several specialized domains. It provides implementation details for computer vision tasks such as image classification, object detection, and semantic segmentation, as well as natural language processing workflows involving transformers, recurrent networks, and gen
Implements general cross-entropy loss functions used to optimize classification performance.
This project is a PyTorch-based Chinese text classification framework. It provides a transformer-based pipeline designed to categorize Chinese language sequences into predefined labels using deep learning models. The implementation supports both BERT and ERNIE language models for processing and tagging complex Chinese text. These models are used to perform tasks such as sentiment analysis and general text categorization. The system utilizes transformer-based text encoding and attention-weighted sequence pooling to convert raw characters into document vectors. It employs pre-trained model fin
Employs cross-entropy loss optimization to measure prediction error and update model weights during training.
This project is a high-performance C++ and CUDA neural network library designed for fast training and inference of small networks on NVIDIA GPUs. It serves as a specialized backend for neural radiance fields and coordinate-based networks, providing a fused GPU kernel library and a hash grid encoder for transforming raw input dimensions into high-dimensional representations. The library distinguishes itself through the use of C++ template metaprogramming and fused-kernel execution, which merge neural network layers into single GPU device functions to eliminate memory bottlenecks. It leverages
Computes standard cross entropy loss for probability density function predictions.