9 مستودعات
Systems for downloading and switching between different versions of diffusion models.
Distinct from Diffusion Models: Focuses on the operational management and switching of model files rather than the model architecture itself.
Explore 9 awesome GitHub repositories matching artificial intelligence & ml · Model Version Management. Refine with filters or upvote what's useful.
AISystem is a comprehensive AI full-stack infrastructure project covering the entire pipeline from AI chip architecture to high-level training frameworks. It encompasses the development of AI compiler frameworks, inference engines, and distributed training orchestrators designed to coordinate workloads across a heterogeneous compute stack of CPUs, GPUs, and NPUs. The project focuses on the deep integration of software and hardware, employing software-hardware co-design to align tensor layouts with physical memory structures. It provides specialized capabilities for accelerating Transformer mo
Tracks experiment metadata and performance metrics to ensure reproducibility and enable version rollback.
This project is a structured educational program and comprehensive training curriculum designed to teach the end-to-end lifecycle of machine learning models. It serves as a resource for engineers to master the transition of data science projects from development into reliable, production-ready systems. The curriculum focuses on the practical application of engineering best practices, emphasizing the orchestration of complex data processing and training sequences. It provides instruction on building repeatable workflows, managing experiment metadata, and implementing infrastructure automation
Logs model parameters and performance metrics to maintain a reproducible history of training iterations.
DiffusionBee is a Stable Diffusion desktop client for macOS that functions as an AI image generator and editor. It allows for the local generation of images from text prompts and the management of diffusion models without requiring external cloud services or technical setup. The application includes a local diffusion model manager for importing and switching between custom trained model files to achieve specific artistic styles. It also features a system for tracking generation history and uploading assets to a public gallery. The software covers several image synthesis and manipulation work
Manages the downloading and switching of diffusion model versions to alter output characteristics.
ClearML is a comprehensive MLOps platform designed to manage the end-to-end machine learning lifecycle, from initial experimentation to production deployment. It provides a suite of integrated tools including a pipeline orchestrator for automating workflows, an experiment tracking tool for logging hyperparameters and metrics, and a metadata-driven data versioning system for managing large-scale datasets and model artifacts. The platform is distinguished by its advanced compute management and serving capabilities. It features a GPU compute manager that supports fractional resource slicing and
Automatically records hyperparameters, performance metrics, and plots to ensure AI experiments are reproducible and comparable.
ZenML is an orchestration platform designed for building, deploying, and monitoring reproducible machine learning pipelines and agentic workflows. It provides a unified framework that manages the entire lifecycle of machine learning assets, from data processing and model training to the deployment of persistent inference services. By decoupling pipeline logic from underlying compute and storage, the platform enables teams to transition workflows seamlessly from local development environments to production-grade cloud infrastructure. The platform distinguishes itself through a service-oriented
Automatically captures, versions, and stores intermediate data objects and configuration parameters produced by pipeline steps to ensure reproducibility.
lakeFS هو نظام إصدارات لبحيرات البيانات يوفر تفرعاً (branching) والتزامات (commits) تشبه Git لمجموعات البيانات الكبيرة المخزنة في تخزين الكائنات. يعمل كطبقة تحكم في الإصدار، مما يتيح إنشاء لقطات غير قابلة للتغيير، والتزامات ذرية، وتفرعاً بدون نسخ (zero-copy) لإنشاء بيئات معزولة لتجارب البيانات دون تكرار الملفات الفيزيائية. يعمل النظام كبوابة تخزين متوافقة مع S3 وفهرس Iceberg REST، مما يسمح لبروتوكولات التخزين السحابي القياسية والعملاء المتوافقين بإدارة الجداول ذات الإصدارات. يعمل كحارس لجودة البيانات باستخدام نظام خطافات (hooks) قائم على الأحداث للتحقق من مجموعات البيانات مقابل سياسات الحوكمة قبل دمج التغييرات في الإنتاج. تغطي المنصة قدرات واسعة لحوكمة البيانات، بما في ذلك التعاون عبر طلبات السحب (pull requests)، والتحكم في الوصول القائم على الأدوار، وتتبع أصل البيانات. يوفر تكاملاً لتنسيق سير العمل، وخطوط أنابيب التعلم الآلي، ومحركات حوسبة البيانات الضخمة المختلفة، ويدعم اتصال التخزين متعدد السحابة ومزامنة الهوية عبر SSO وSCIM. يمكن تثبيت البرنامج باستخدام ملفات ثنائية، أو حاويات، أو Helm charts للنشر على Kubernetes.
Records model performance and data provenance by attaching custom metrics to specific commits.
Sacred هي أداة لإدارة التجارب وإطار عمل لإعادة الإنتاج مصمم لتنظيم عمليات تشغيل متعددة لعملية ما بتكوينات مختلفة. تعمل كمتتبع لتجارب التعلم الآلي ومدير لتكوين المعلمات الفائقة (hyperparameters)، حيث تسجل المعلمات الفائقة والمقاييس والبيانات الوصفية في قاعدة بيانات لضمان بقاء عمليات التنفيذ التجريبية قابلة للتتبع. يركز المشروع على إعادة إنتاج النتائج العلمية من خلال إدارة البذور العشوائية (random seeds) وتتبع تبعيات النظام تلقائياً. ويسمح بتنفيذ متغيرات التجربة من خلال تجاوز معلمات سطر الأوامر وحقن المعلمات الديناميكي، مما يتيح تعديل الإعدادات دون تغيير الكود المصدري الأساسي. يوفر إطار العمل إمكانيات لتسجيل البيانات الوصفية المدعومة بقاعدة بيانات، والتقاط تفاصيل الأجهزة وإصدارات البرامج للحفاظ على سجل قابل للبحث لكل عملية تشغيل. كما يدعم تسلسل حالة التنفيذ لتمكين النسخ الدقيق للنتائج التجريبية.
Saves configuration settings, system dependencies, and hardware details to a database for future analysis.
Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as a control plane for managing distributed deep learning workloads, automated machine learning pipelines, and experiment tracking. The platform distinguishes itself through specialized services for distributed training management, including MPI-based coordination for PyTorch and TensorFlow. It provides an automated hyperparameter optimization service utilizing Bayesian, random, and grid search algorithms, alongside managed interactive AI workspaces for launching Jupyter notebook
Provides a system for recording hyperparameters, performance metrics, and version history to ensure scientific reproducibility of AI experiments.
يعمل هذا المشروع كمنصة متخصصة لأبحاث التصوير الطبي السريري، حيث يوفر مجموعة من دفاتر الملاحظات التعليمية والأدوات القياسية للتعلم العميق. يعمل كإطار عمل لبناء وتدريب الشبكات العصبية المصممة خصيصاً للخصائص الهندسية وكثافة بيانات الصور الطبية، ويدعم مهام مثل التجزئة (segmentation)، والتصنيف، والتسجيل. تتميز المنصة بتركيزها على سير عمل الأبحاث من البداية إلى النهاية، حيث تقدم قوالب معيارية توحد معالجة البيانات مسبقاً، وتدريب النماذج، والاستدلال. تتضمن قدرات للنمذجة التوليدية، مثل الانتشار الكامن (latent diffusion) والشبكات التنافسية، لإنشاء صور اصطناعية أو إجراء ترجمة من صورة إلى صورة. علاوة على ذلك، توفر أدوات آلية لتعليق وتجزئة الصور الطبية لتقليل الجهد اليدوي في إعداد مجموعات البيانات. يدعم إطار العمل الأبحاث عالية الأداء من خلال دمج تنسيق الحوسبة الموزعة، والتدريب بالدقة المختلطة، وخطوط أنابيب البيانات القائمة على الموترات (tensor-based) للتعامل مع مجموعات البيانات ثلاثية الأبعاد واسعة النطاق. كما يتضمن ميزات لإدارة بيانات تعريف التجارب لضمان القابلية للتكرار، ويوفر مسارات لتغليف النماذج المدربة في خدمات جاهزة للإنتاج لدعم القرارات السريرية. تم تنظيم المستودع كسلسلة من دفاتر Jupyter التفاعلية التي توضح سير العمل هذا، مع خيارات لتنفيذ المهام في بيئات سحابية مهيأة مسبقاً.
Logs training metrics and tracks experiment configurations to ensure reproducibility in clinical research.