awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
typpo avatar

typpo/promptfoo

0
View on GitHub↗
22,295 Stars·1,992 Forks·TypeScript·MIT·8 Aufrufepromptfoo.dev↗

Promptfoo

promptfoo is an evaluation framework for measuring the performance of large language model prompts, agents, and retrieval augmented generation pipelines. It provides a suite of tools for conducting comparative benchmarking and executing automated quality and security regressions.

The system features a benchmarking suite for running identical prompts across different model providers to compare output quality side-by-side. It also includes a dedicated red teaming tool for identifying security vulnerabilities and prompt injection risks through automated penetration testing.

The framework supports declarative evaluation pipelines and metric-based scoring to quantify model reliability. These capabilities are designed for integration into continuous integration and deployment workflows to prevent regressions in model behavior. Results can be visualized in shared reports to facilitate team reviews of performance data and security findings.

Features

  • Prompt Evaluation Tools - Offers a comprehensive toolkit for comparing output quality across different prompt variations and models to identify the most effective instructions.
  • LLM Evaluation - Provides a comprehensive framework for measuring LLM output quality using custom metrics and automated judges.
  • AI Model Benchmarking - Provides frameworks for running standardized tests to compare the performance and reliability of different LLM providers.
  • Scoring Pipelines - Implements scoring pipelines that apply algorithmic checks to quantify model quality and detect inaccuracies.
  • Model Comparison Interfaces - Enables side-by-side visual and analytical comparison of outputs from different LLM providers.
  • Model Benchmarking Suites - Conducts comparative analysis of model accuracy and reasoning using standardized datasets across providers.
  • Provider-Agnostic Model Interfaces - Standardizes inputs and outputs across different large language models to enable side-by-side performance comparisons.
  • RAG Evaluation Frameworks - Offers specialized frameworks for assessing RAG-specific metrics like groundedness and retrieval relevance.
  • AI Red Teaming - Evaluates and probes vulnerabilities in language models through automated red teaming and penetration testing.
  • Automated Prompt Testing - Provides a framework for integrating prompt evaluation and data-driven quality checks into continuous integration pipelines.
  • Adversarial Red Teaming Toolkits - Provides specialized toolkits for generating adversarial prompts to test for security bypasses and injections.
  • Automated Agent Quality Assurance - Integrates automated model behavior checks into CI/CD pipelines to ensure quality and prevent regressions.
  • Automated Assertion Validators - Provides a framework for validating LLM outputs against programmatic assertions and predefined quality metrics.
  • CI/CD Pipeline Integrations - Integrates evaluation runs into CI/CD pipelines to block deployments when model performance fails thresholds.
  • Automation Pipelines - Automates the execution of quality checks and security scans within the software delivery pipeline.
  • Vulnerability Scanning Utilities - Includes tools for performing automated vulnerability assessments and red teaming on AI pipelines.
  • Application Development - Tool for testing, evaluating, and comparing LLM outputs.

Star-Verlauf

Star-Verlauf für typpo/promptfooStar-Verlauf für typpo/promptfoo

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Promptfoo

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Promptfoo.
  • promptfoo/promptfooAvatar von promptfoo

    promptfoo/promptfoo

    10,529Auf GitHub ansehen↗

    Promptfoo is an evaluation framework designed for testing, benchmarking, and red-teaming language models and agentic workflows. It provides a unified environment to run prompts against multiple providers, allowing developers to systematically validate model outputs against objective assertions, semantic similarity metrics, and custom grading rubrics. The platform distinguishes itself through a provider-agnostic execution layer and a stateful orchestrator capable of simulating multi-turn conversations and complex tool-use trajectories. It includes a dedicated adversarial mutation pipeline that

    TypeScriptcici-cdcicd
    Auf GitHub ansehen↗10,529
  • vibrantlabsai/ragasAvatar von vibrantlabsai

    vibrantlabsai/ragas

    12,659Auf GitHub ansehen↗

    Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin

    Pythonevaluationllmllmops
    Auf GitHub ansehen↗12,659
  • confident-ai/deepevalAvatar von confident-ai

    confident-ai/deepeval

    13,733Auf GitHub ansehen↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    Auf GitHub ansehen↗13,733
  • evidentlyai/evidentlyAvatar von evidentlyai

    evidentlyai/evidently

    7,137Auf GitHub ansehen↗

    Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of

    Jupyter Notebookdata-driftdata-qualitydata-science
    Auf GitHub ansehen↗7,137
Alle 30 Alternativen zu Promptfoo anzeigen→

Häufig gestellte Fragen

Was macht typpo/promptfoo?

promptfoo is an evaluation framework for measuring the performance of large language model prompts, agents, and retrieval augmented generation pipelines. It provides a suite of tools for conducting comparative benchmarking and executing automated quality and security regressions.

Was sind die Hauptfunktionen von typpo/promptfoo?

Die Hauptfunktionen von typpo/promptfoo sind: Prompt Evaluation Tools, LLM Evaluation, AI Model Benchmarking, Scoring Pipelines, Model Comparison Interfaces, Model Benchmarking Suites, Provider-Agnostic Model Interfaces, RAG Evaluation Frameworks.

Welche Open-Source-Alternativen gibt es zu typpo/promptfoo?

Open-Source-Alternativen zu typpo/promptfoo sind unter anderem: promptfoo/promptfoo — Promptfoo is an evaluation framework designed for testing, benchmarking, and red-teaming language models and agentic… vibrantlabsai/ragas — Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and… confident-ai/deepeval — Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for… evidentlyai/evidently — Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine… internlm/opencompass — OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to… microsoft/promptbase — Promptbase is a prompt engineering framework designed for designing, testing, and optimizing prompts for large…