awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 مستودعات

Awesome GitHub RepositoriesModel Experiment Execution

Running a set of tasks against a dataset and applying evaluators to compare results across versions.

Distinct from Automated Dataset Evaluation: Focuses on comparative experimentation rather than just the execution of a single automated evaluation.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Model Experiment Execution. Refine with filters or upvote what's useful.

Awesome Model Experiment Execution GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • arize-ai/phoenixالصورة الرمزية لـ Arize-ai

    Arize-ai/phoenix

    8,605عرض على GitHub↗

    Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

    Executes tasks against datasets and applies evaluators to compare performance across model or prompt iterations.

    Jupyter Notebookagentsai-monitoringai-observability
    عرض على GitHub↗8,605
  • llm-attacks/llm-attacksالصورة الرمزية لـ llm-attacks

    llm-attacks/llm-attacks

    4,509عرض على GitHub↗

    This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.

    Implements a system for running harmful prompts across multiple models to compare safety robustness.

    Python
    عرض على GitHub↗4,509
  1. Home
  2. Artificial Intelligence & ML
  3. Dataset Management
  4. Evaluation Datasets
  5. Model Experiment Execution