5 रिपॉजिटरी
Systems for defining and executing complex data dependencies across distributed compute clusters.
Distinct from Distributed ML Trainers: Manages the overall pipeline workflow rather than just the specific distributed training process.
Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Distributed ML Pipeline Managers. Refine with filters or upvote what's useful.
Flyte is a distributed machine learning pipeline manager and MLOps workflow engine. It functions as a Kubernetes-native orchestrator used to coordinate data, models, and compute resources for executing machine learning pipelines and autonomous agents at scale. The platform provides specialized infrastructure for the full machine learning lifecycle, including a dedicated model serving platform to deploy trained models as scalable production-ready inference services. It also enables the coordination and state management of autonomous AI agents. The system manages scalable pipeline execution th
Defines and executes complex data dependencies and compute tasks across distributed clusters.
SynapseML एक Apache Spark मशीन लर्निंग लाइब्रेरी है जिसे वितरित क्लस्टर में मशीन लर्निंग वर्कफ़्लो और डेटा पाइपलाइनों के निर्माण और स्केलिंग के लिए डिज़ाइन किया गया है। यह बड़े पैमाने पर डेटासेट पर हार्डवेयर-त्वरित भविष्यवाणियों और डीप लर्निंग कार्यों को निष्पादित करने के लिए एक वितरित मशीन लर्निंग पाइपलाइन फ्रेमवर्क और एक वितरित अनुमान इंजन के रूप में कार्य करता है। यह प्रोजेक्ट एक क्लाउड AI एकीकरण परत के रूप में कार्य करता है, जो उपयोगकर्ताओं को वितरित पाइपलाइनों के भीतर टेक्स्ट, विज़न और स्पीच के लिए पूर्व-प्रशिक्षित कृत्रिम बुद्धिमत्ता सेवाओं को लागू करने की अनुमति देता है। इसमें उच्च-आयामी डेटा में मल्टीवेरिएट और टाइम-सीरीज़ आउटलेर्स की पहचान करने के लिए वितरित विसंगति पहचान के लिए उपकरणों का एक समर्पित सूट भी शामिल है। लाइब्रेरी क्षमताओं की एक विस्तृत श्रृंखला को कवर करती है, जिसमें चेहरा और छवि विश्लेषण के लिए वितरित कंप्यूटर विज़न, टेक्स्ट एनालिटिक्स और अनुवाद के लिए स्केलेबल नेचुरल लैंग्वेज प्रोसेसिंग, और ग्रेडिएंट बूस्टेड डिसीजन ट्री का प्रशिक्षण शामिल है। यह k-निकटतम पड़ोसी मॉडलिंग के माध्यम से समानता खोज, फ़ीचर एट्रिब्यूशन के माध्यम से मॉडल व्याख्यात्मकता और रीइन्फोर्समेंट लर्निंग वर्कफ़्लो के ऑर्केस्ट्रेशन के लिए उपकरण प्रदान करती है। सिस्टम एक कंपोज़ेबल पाइपलाइन आर्किटेक्चर का उपयोग करता है और क्रॉस-प्लेटफ़ॉर्म संगतता के लिए ONNX-आधारित मॉडल अनुमान का समर्थन करता है।
Provides a composable framework for sequencing data featurization and model training across distributed compute clusters.
Mmlspark Apache Spark क्लस्टर में मशीन लर्निंग मॉडल, डेटा परिवर्तन और AI सेवा एकीकरण को निष्पादित करने के लिए एक वितरित फ्रेमवर्क है। यह एक वितरित मशीन लर्निंग लाइब्रेरी और पाइपलाइन ऑर्केस्ट्रेटर के रूप में कार्य करता है, जो उपयोगकर्ताओं को बड़े पैमाने पर बैच और स्ट्रीमिंग वर्कफ़्लो में पूर्व-प्रशिक्षित संज्ञानात्मक सेवाओं और कस्टम मॉडल को एकीकृत करने की अनुमति देता है। यह प्रोजेक्ट टेक्स्ट और विज़न विश्लेषण के लिए बिग डेटा पाइपलाइनों में बाहरी AI सेवाओं और वेब API को सीधे शामिल करने की अपनी क्षमता से प्रतिष्ठित है। यह एक स्केलेबल मॉडल प्रशिक्षण फ्रेमवर्क प्रदान करता है जो ग्रेडिएंट बूस्टिंग और वर्गीकरण कार्यों को लचीले ढंग से रिसाइज़ करने योग्य कंप्यूट क्लस्टर में समन्वयित करता है, वितरित मॉडल अनुमान के लिए हार्डवेयर त्वरण का उपयोग करता है। टूलसेट छवि, भाषण और टेक्स्ट के लिए मल्टीमॉडल सामग्री विश्लेषण, साथ ही टाइम-सीरीज़ और मल्टीवेरिएट डेटा के लिए उन्नत विसंगति पहचान सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। इसमें डेटा फ़ीचराइजेशन, ONNX मॉडल का निष्पादन, और योगात्मक योगदान मूल्यों का उपयोग करके मॉडल निष्पक्षता ऑडिटिंग और भविष्यवाणी व्याख्या के लिए जिम्मेदार AI उपकरण शामिल हैं। फ्रेमवर्क विभिन्न डेटाबेस और क्लाउड स्टोरेज सिस्टम में पढ़ने और लिखने के लिए एक एकीकृत डेटा एक्सेस इंटरफ़ेस भी प्रदान करता है।
Orchestrates the integration of pre-trained AI services and custom models into large-scale batch and streaming workflows.
Deep Java Library is a Java deep learning framework and JVM model inference engine. It provides a high-level API for building and deploying deep learning models within the Java ecosystem, acting as a cross-platform runtime for executing models across CPUs, GPUs, and mobile devices. The library is engine-agnostic, allowing users to switch between different deep learning engines such as PyTorch, TensorFlow, and MXNet while maintaining a single unified API. This enables the deployment of the same model across different backends without changing the application code. The framework supports the f
Enables the integration of deep learning inference into large-scale distributed big data processing pipelines.
Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li
Manages complex data dependencies and hyperparameters across distributed compute clusters for audio experiments.