awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
uptrain-ai avatar

uptrain-ai/uptrain

0
View on GitHub↗
2,352 Stars·202 Forks·Python·Apache-2.0·2 Aufrufeuptrain.ai↗

Uptrain

UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.

Features

  • Evaluation and Observability - Tool to evaluate and improve LLM applications.
  • Evaluation Frameworks - Unified platform for evaluating and improving generative AI applications.

Star-Verlauf

Star-Verlauf für uptrain-ai/uptrainStar-Verlauf für uptrain-ai/uptrain

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Uptrain

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Uptrain.
  • explodinggradients/ragasAvatar von explodinggradients

    explodinggradients/ragas

    14,400Auf GitHub ansehen↗

    Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s

    Python
    Auf GitHub ansehen↗14,400
  • truera/trulensAvatar von truera

    truera/trulens

    3,384Auf GitHub ansehen↗

    Evaluation and Tracking for LLM Experiments and AI Agents

    Python
    Auf GitHub ansehen↗3,384
  • confident-ai/deepevalAvatar von confident-ai

    confident-ai/deepeval

    13,733Auf GitHub ansehen↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    Auf GitHub ansehen↗13,733
  • bigcode-project/bigcode-evaluation-harnessAvatar von bigcode-project

    bigcode-project/bigcode-evaluation-harness

    1,049Auf GitHub ansehen↗

    A framework for the evaluation of autoregressive code generation language models.

    Python
    Auf GitHub ansehen↗1,049
Alle 16 Alternativen zu Uptrain anzeigen→

Häufig gestellte Fragen

Was macht uptrain-ai/uptrain?

UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.

Was sind die Hauptfunktionen von uptrain-ai/uptrain?

Die Hauptfunktionen von uptrain-ai/uptrain sind: Evaluation and Observability, Evaluation Frameworks.

Welche Open-Source-Alternativen gibt es zu uptrain-ai/uptrain?

Open-Source-Alternativen zu uptrain-ai/uptrain sind unter anderem: explodinggradients/ragas — Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented… confident-ai/deepeval — Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for… truera/trulens — Evaluation and Tracking for LLM Experiments and AI Agents. bigcode-project/bigcode-evaluation-harness — A framework for the evaluation of autoregressive code generation language models. comet-ml/opik — Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It… arize-ai/phoenix — Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and…