awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 repository-uri

Awesome GitHub RepositoriesModel Experiment Execution

Running a set of tasks against a dataset and applying evaluators to compare results across versions.

Distinct from Automated Dataset Evaluation: Focuses on comparative experimentation rather than just the execution of a single automated evaluation.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Model Experiment Execution. Refine with filters or upvote what's useful.

Awesome Model Experiment Execution GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • arize-ai/phoenixAvatar Arize-ai

    Arize-ai/phoenix

    8,605Vezi pe GitHub↗

    Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

    Executes tasks against datasets and applies evaluators to compare performance across model or prompt iterations.

    Jupyter Notebookagentsai-monitoringai-observability
    Vezi pe GitHub↗8,605
  • llm-attacks/llm-attacksAvatar llm-attacks

    llm-attacks/llm-attacks

    4,509Vezi pe GitHub↗

    This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.

    Implements a system for running harmful prompts across multiple models to compare safety robustness.

    Python
    Vezi pe GitHub↗4,509
  1. Home
  2. Artificial Intelligence & ML
  3. Dataset Management
  4. Evaluation Datasets
  5. Model Experiment Execution