awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 dépôts

Awesome GitHub RepositoriesSystem Monitoring and Scaling

Techniques for distributed training and monitoring data distribution shifts in production.

Distinct from Distributed and Scaling Strategies: Combines training scale with production monitoring, whereas the parent is focused on scaling strategies.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · System Monitoring and Scaling. Refine with filters or upvote what's useful.

Awesome System Monitoring and Scaling GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • alirezadir/machine-learning-interviewsAvatar de alirezadir

    alirezadir/Machine-Learning-Interviews

    8,455Voir sur GitHub↗

    This project is a comprehensive machine learning interview guide and technical study resource designed for individuals preparing for machine learning and AI engineering roles. It provides a collection of materials and practice problems covering core algorithms, theoretical fundamentals, and the implementation of neural network architectures. The resource serves as a technical reference for generative AI development, focusing on the design and optimization of large language models and diffusion systems. It includes frameworks for system design, covering the architecture of production machine l

    Includes guidelines on implementing distributed training and detecting data distribution shifts.

    Jupyter Notebookagenticaiai-agents
    Voir sur GitHub↗8,455
  • subhashchy/the-accidental-ctoAvatar de subhashchy

    subhashchy/The-Accidental-CTO

    3,168Voir sur GitHub↗

    The Accidental CTO is a comprehensive collection of guides and frameworks focused on distributed systems architecture, resilience engineering, and system observability. It provides strategies for scaling applications from thousands to millions of users while maintaining high availability. The project offers specific methodologies for managing data volume through replication, sharding, and caching. It includes a framework for analyzing cloud infrastructure spending and evaluating transitions to self-hosted environments to reduce operational expenses. The resource covers the implementation of

    Provides methodologies for growing user capacity using replication, sharding, and caching strategies.

    TypeScriptscalingsystem-design
    Voir sur GitHub↗3,168
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Training & Tuning
  6. Distributed and Scaling Strategies
  7. System Monitoring and Scaling

Explorer les sous-tags

  • Scaling StrategiesArchitectural methods for increasing system capacity through data and resource distribution. **Distinct from System Monitoring and Scaling:** Focuses on the high-level strategy of scaling distributed systems rather than specific training-set monitoring.