awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
SchedMD avatar

SchedMD/slurm

0
View on GitHub↗
4,059 stars·855 forks·C·3 vuesslurm.schedmd.com↗

Slurm

Slurm est un gestionnaire de charge de travail de cluster et un planificateur de tâches conçu pour les environnements de calcul haute performance. Il fonctionne comme un orchestrateur de calcul distribué qui met en file d'attente et exécute des tâches computationnelles à grande échelle sur plusieurs nœuds de calcul dans un cluster.

Le système agit comme un arbitre de ressources, distribuant les nœuds matériels et les processeurs entre les utilisateurs simultanés pour éviter les conflits de ressources et maximiser l'efficacité. Il coordonne le lancement simultané de multiples processus sur différents serveurs physiques pour exécuter des jobs parallèles et des charges de travail scientifiques.

La plateforme couvre de vastes domaines de capacités, notamment la planification de jobs par lots, l'allocation de ressources de calcul et l'exécution de charges de travail parallèles. Il gère le timing et l'exécution des jobs en fonction de la disponibilité des ressources et de la priorité.

Features

  • HPC Resource Allocation - Manages hardware resource allocation and job priority specifically for high-performance computing environments.
  • High-Performance Computing - Executes computationally intensive tasks across distributed hardware architectures using a cluster manager.
  • Managed Cluster Batch Execution - Queues and executes large-scale computational batch workloads across multiple compute nodes in a cluster.
  • Process Coordination - Synchronizes the simultaneous launch of multiple processes across different physical servers for parallel jobs.
  • HPC Cluster Orchestrations - Provides orchestration and resource management of large-scale computational workloads across high-performance computing clusters.
  • Cluster Workload Managers - Controls the execution and timing of high performance computing jobs to ensure optimal hardware utilization.
  • Priority-Based Job Schedulers - Queues and executes batch computational workloads based on priority and resource availability.
  • Batch Schedulers - Queues and executes computational batch tasks based on priority and available hardware.
  • Job Execution Controllers - Provides runtime control over the start, stop, and monitoring of work across allocated compute nodes.
  • Remote Job Dispatchers - Dispatches control signals from a central manager to specific nodes to trigger task execution.
  • Resource Allocation - Distributes hardware nodes and processors across a cluster to ensure efficient utilization.
  • Queue Priority Scheduling - Orders pending workloads in a logical queue and dispatches them according to user priorities.
  • Parallel Job Orchestrators - Coordinates the simultaneous launch and execution of parallel workloads across multiple processes and distributed compute nodes.
  • Distributed Workload Execution - Runs single computational tasks across multiple compute nodes simultaneously to reduce processing time.
  • Resource Access Arbitrators - Resolves conflicting requests for available system resources through a centralized coordinator.
  • Distributed Node Agents - Uses lightweight agents on every compute node to manage local job execution and report health status.
  • Distributed Job Executors - Executes and monitors computational workloads by integrating with distributed compute nodes.
  • Scientific Workload Schedulers - Schedules the timing and execution of complex scientific workloads across a network of distributed servers.
  • State-Based Resource Tracking - Maintains a real-time database of node availability and allocations to prevent hardware oversubscription.
  • Infrastructure and Deployment - Scalable workload manager for high-performance computing.
  • Gestion d'infrastructure - Gestionnaire de charge de travail évolutif pour le calcul haute performance (HPC).
  • Resource Scheduling - Highly scalable workload manager for high-performance computing.

Historique des stars

Graphique de l'historique des stars pour schedmd/slurmGraphique de l'historique des stars pour schedmd/slurm

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Slurm

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Slurm.
  • polyaxon/polyaxonAvatar de polyaxon

    polyaxon/polyaxon

    3,707Voir sur GitHub↗

    Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as a control plane for managing distributed deep learning workloads, automated machine learning pipelines, and experiment tracking. The platform distinguishes itself through specialized services for distributed training management, including MPI-based coordination for PyTorch and TensorFlow. It provides an automated hyperparameter optimization service utilizing Bayesian, random, and grid search algorithms, alongside managed interactive AI workspaces for launching Jupyter notebook

    MDX
    Voir sur GitHub↗3,707
  • microsoftdocs/azure-docsAvatar de MicrosoftDocs

    MicrosoftDocs/azure-docs

    10,894Voir sur GitHub↗

    Azure Docs is the official technical documentation repository for Microsoft Azure, the cloud computing platform. It provides comprehensive guidance on the full spectrum of Azure services, covering everything from core infrastructure components like virtual machines, Kubernetes clusters, and serverless computing to platform services for AI, machine learning, data analytics, and storage. The documentation details how to provision, manage, and govern cloud resources at scale, including policy enforcement, identity management, and cost optimization. The documentation distinguishes Azure through i

    Markdownskilling
    Voir sur GitHub↗10,894
  • zhaochenyang20/awesome-ml-sys-tutorialAvatar de zhaochenyang20

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371Voir sur GitHub↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Python
    Voir sur GitHub↗5,371
  • inngest/inngestAvatar de inngest

    inngest/inngest

    5,499Voir sur GitHub↗

    Inngest is a durable execution framework and event-driven automation engine designed to orchestrate background workflows. It enables developers to build resilient, stateful processes by memoizing function steps, ensuring that long-running tasks can automatically resume from the last successful operation after failures, timeouts, or infrastructure restarts. The platform distinguishes itself through its event-driven architecture, which uses a schema-validated bus to trigger functions and coordinate complex, multi-step logic. It employs an onion-model middleware approach for cross-cutting concer

    Go
    Voir sur GitHub↗5,499
Voir les 30 alternatives à Slurm→

Questions fréquentes

Que fait schedmd/slurm ?

Slurm est un gestionnaire de charge de travail de cluster et un planificateur de tâches conçu pour les environnements de calcul haute performance. Il fonctionne comme un orchestrateur de calcul distribué qui met en file d'attente et exécute des tâches computationnelles à grande échelle sur plusieurs nœuds de calcul dans un cluster.

Quelles sont les fonctionnalités principales de schedmd/slurm ?

Les fonctionnalités principales de schedmd/slurm sont : HPC Resource Allocation, High-Performance Computing, Managed Cluster Batch Execution, Process Coordination, HPC Cluster Orchestrations, Cluster Workload Managers, Priority-Based Job Schedulers, Batch Schedulers.

Quelles sont les alternatives open-source à schedmd/slurm ?

Les alternatives open-source à schedmd/slurm incluent : polyaxon/polyaxon — Polyaxon is a Kubernetes-native machine learning orchestration platform and MLOps pipeline orchestrator. It serves as… microsoftdocs/azure-docs — Azure Docs is the official technical documentation repository for Microsoft Azure, the cloud computing platform. It… zhaochenyang20/awesome-ml-sys-tutorial — This project provides a comprehensive technical guide and framework for engineering large-scale machine learning… inngest/inngest — Inngest is a durable execution framework and event-driven automation engine designed to orchestrate background… dask/dask — Dask is a parallel computing framework and distributed task scheduler designed to scale Python data science workflows… ansible/ansible — Ansible is an agentless infrastructure automation engine designed to manage remote servers and network devices. It…