For a tool for monitoring model drift in production, the strongest matches are evidentlyai/evidently (Evidently is an AI observability platform that monitors deployed), arize-ai/phoenix (Phoenix is an LLM observability platform that monitors AI) and deepchecks/deepchecks (Deepchecks provides continuous validation of ML models and data). nannyml/nannyml and shap/shap round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Tools for tracking model performance, detecting data drift, and identifying degradation in deployed machine learning systems.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Evidently is an AI observability platform that monitors deployed ML models for data drift, performance degradation, and quality issues, with real-time dashboards and evaluation tools, directly matching the model monitoring intent.
Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and
Phoenix is an LLM observability platform that monitors AI applications for data drift through embedding visualization and includes evaluation and performance tracking, so it fits the machine-learning model monitoring category but is specialised for large language models rather than general ML models.
Deepchecks is a machine learning model validation framework and MLOps testing library. It serves as an AI data quality suite and performance evaluator designed to verify the integrity and performance of models and datasets from research through production. The project functions as a model monitoring tool for tracking data drift and performance degradation in production environments. It allows for the creation of custom validation suites and utilizes a pluggable check architecture to automate quality checks within continuous integration pipelines. The framework covers a broad range of capabil
Deepchecks provides continuous validation of ML models and data with drift detection and performance checks from research to production, which aligns with the need to monitor deployed models for drift and degradation, though it may require integration for real-time dashboards and alerts.
nannyml: post-deployment data science in python
NannyML is a Python library purpose-built for post-deployment ML model monitoring, offering drift detection and performance metrics out of the box, making it a direct fit for monitoring deployed models even though it may require additional setup for dashboards and real-time alerts.
SHAP is an explainable AI toolkit that provides a game theoretic framework for interpreting machine learning model predictions. It functions as a feature attribution engine, decomposing model outputs into the sum of individual feature effects to clarify how specific input variables influence a final decision. By assigning importance values to these inputs, the library enables users to understand the logic behind complex predictive models. The project distinguishes itself through its versatility and specialized calculation methods. It operates as a model-agnostic diagnostic library, capable of
SHAP is a model interpretability library for explaining predictions via feature attribution, not a production monitoring system — it lacks drift detection, performance alerts, and real-time dashboards for deployed models.
RagaAI-Catalyst is a suite of software implementation tools providing an SDK, dashboard, and platform for monitoring, debugging, red-teaming, and evaluating agentic AI workflows. It serves as an observability framework for tracing the execution paths of large language models and multi-agent systems. The project distinguishes itself through a security suite for automated red-teaming and vulnerability scanning to detect biases, alongside a centralized prompt registry that decouples templates from application code. It further provides an evaluation platform that combines synthetic data generatio
RagaAI-Catalyst is an observability and debugging platform for agentic AI workflows and LLMs, not a general ML model monitoring tool for drift and degradation in production—it lacks support for traditional model drift detection and performance metrics.
Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo
Giskard is an evaluation and testing framework for LLMs, focused on pre-deployment robustness and security rather than continuous production monitoring of drift and degradation across all ML models.
ClearML is a comprehensive MLOps platform designed to manage the entire machine learning lifecycle. It functions as an experiment tracking tool, a data versioning system, and a pipeline orchestrator, while providing infrastructure for GPU cluster management and model serving. The platform is distinguished by its ability to handle hybrid-cloud compute scheduling and fractional GPU allocation, allowing multiple workloads to share a single hardware accelerator. It employs a metadata-based approach to data versioning, using virtual views to track large datasets and artifacts without duplicating r
ClearML is an MLOps platform that includes model performance monitoring, but its core focus is on experiment tracking and pipeline orchestration rather than dedicated drift detection and root-cause analysis for production models.
Lit is a machine learning interpretability framework and model debugging tool designed to analyze model behavior and performance. It serves as an interpretability dashboard for large language models and a general performance analyzer for text, image, and tabular datasets. The project distinguishes itself through a comprehensive suite of interpretability tools, including salience map generation for feature attribution, the creation of synthetic and counterfactual examples to test robustness, and the projection of high-dimensional embeddings into visual spaces via UMAP or PCA. It further enable
LIT is a model interpretability and debugging dashboard, not a production monitoring tool for drift detection or real-time degradation alerts, so it addresses a related but different need.
SigNoz is a full-stack observability platform designed to collect, store, and visualize metrics, logs, and distributed traces in a unified environment. It leverages OpenTelemetry-based data collection to ingest telemetry from diverse sources using vendor-neutral protocols, ensuring interoperability across complex microservices architectures. The platform utilizes a high-performance columnar storage engine to enable rapid aggregation and filtering, providing a centralized backend for monitoring application health and performance. What distinguishes the platform is its focus on automated instru
SigNoz is a general-purpose observability platform for application metrics, logs, and traces, but it lacks built-in support for ML model monitoring—features like drift detection, model performance metrics, and root cause analysis for model degradation are not part of its scope, so it is not the specialized tool this search targets.
KServe is a Kubernetes-native platform for deploying and serving machine learning models as scalable inference services. It supports both generative AI models, including large language models, and traditional predictive models from frameworks such as TensorFlow, PyTorch, Scikit-Learn, XGBoost, and ONNX. The platform manages the full lifecycle of model deployments, including revision tracking, canary rollouts, A/B testing, and automatic rollbacks, and provides serverless scale-to-zero capabilities for cost-efficient resource management. KServe distinguishes itself through a standardized infere
KServe is a model serving and deployment platform, not a monitoring tool — it focuses on inference serving with canary rollouts and A/B testing, but does not provide built-in drift detection, performance degradation alerts, or real-time observability of model behavior in production.
Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo
Oumi is a development and evaluation platform for LLMs, not a tool for monitoring deployed models; it focuses on training and synthesis rather than production drift detection or observability.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| evidentlyai/evidently | 7.1K | Jupyter Notebook | apache-2.0 | |
| arize-ai/phoenix | 8.6K | Jupyter Notebook | other | |
| deepchecks/deepchecks | 4K | Python | NOASSERTION | |
| nannyml/nannyml | 2.1K | Python | Apache-2.0 | |
| shap/shap | 25K | Jupyter Notebook | mit | |
| raga-ai-hub/ragaai-catalyst | 16.2K | Python | Apache-2.0 | |
| giskard-ai/giskard | 5.4K | Python | Apache-2.0 | |
| allegroai/clearml | 6.7K | Python | Apache-2.0 | |
| pair-code/lit | 3.6K | TypeScript | apache-2.0 | |
| signoz/signoz | 27.4K | TypeScript | NOASSERTION |