3 repository-uri
Mechanisms for importing and preprocessing raw data for efficient consumption by neural network models.
Distinct from Data Ingestion: Distinct from database ingestion; focuses specifically on preparing data for ML model training pipelines.
Explore 3 awesome GitHub repositories matching artificial intelligence & ml · Training Data Ingestion. Refine with filters or upvote what's useful.
Flashlight este o bibliotecă C++ standalone de machine learning și tensori, utilizată pentru construirea și antrenarea rețelelor neuronale. Aceasta funcționează ca un framework cuprinzător de rețele neuronale și motor de diferențiere automată, oferind instrumentele necesare pentru a construi grafuri de calcul și a calcula gradienții prin backpropagation. Proiectul servește drept framework de antrenare distribuită, utilizând operațiuni all-reduce pentru a sincroniza gradienții și parametrii pe mai multe noduri de calcul și dispozitive. Se distinge prin integrarea profundă a manipulării de înaltă performanță a tensorilor, interoperabilitatea nativă a memoriei dispozitivului și un sistem pentru sincronizarea ponderilor între workerii distribuiți pentru a accelera antrenarea modelelor la scară largă. Framework-ul acoperă o gamă largă de capabilități de deep learning, inclusiv compoziția modulară a straturilor pentru proiectarea arhitecturilor complexe precum blocuri reziduale și celule recurente. Oferă utilitare extinse de gestionare a datelor pentru ingestie și prefetching, alături de sisteme de serializare pentru persistența stărilor modelelor. În plus, include o suită de instrumente de monitorizare și observabilitate pentru urmărirea metricilor de antrenare și măsurarea erorilor de secvență. Biblioteca este implementată în C++.
Provides built-in utilities to ingest and preprocess data for efficient delivery to neural network models.
mmtracking is a PyTorch video perception framework designed for training and deploying computer vision models that analyze sequential image data. It provides specialized tools for multi-object tracking, video instance segmentation, and a configuration-driven system for managing deep learning models. The project utilizes a deep learning model registry and a configuration-driven pipeline to swap model backbones and detectors without modifying the core codebase. This modular approach allows for the development of custom perception architectures by combining various components and configurations.
Provides mechanisms for importing and streaming large-scale video datasets into neural network training pipelines.
TensorFlowOnSpark is a distributed framework for running TensorFlow machine learning workloads and model training across Apache Spark clusters. It functions as a cluster computing orchestrator that manages worker processes and resource allocation to scale deep learning tasks across multiple computing nodes. The platform enables distributed deep learning training and large-scale model inference, allowing users to execute tasks across a cluster of servers to handle datasets that exceed the memory of a single machine. It integrates deep learning workloads with Spark data processing to create end
Provides mechanisms for importing and preprocessing raw training data from HDFS or Spark for neural network consumption.