awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to pkunlp-icler/pca-eval

Projects sharing features with PCA EVAL

30 open-source projects similar to pkunlp-icler/pca-eval, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • facebookresearch/habitat-labfacebookresearch avatar

    facebookresearch/habitat-lab

    2,848View on GitHub↗

    Habitat-Lab is an open-source platform for training and evaluating embodied AI agents in photorealistic 3D indoor environments. It functions as a high-performance 3D indoor environment simulator that supports physics-based interaction, enabling research into navigation and manipulation tasks. The platform provides a modular task-environment abstraction that separates task logic from environment simulation, using configuration-driven pipeline assembly to compose simulation and training pipelines. It includes a hierarchical sensor-actuator architecture for mixing and matching perception and act

    Pythonaicomputer-visiondeep-learning
    View on GitHub↗2,848
  • comet-ml/opikcomet-ml avatar

    comet-ml/opik

    17,787View on GitHub↗

    Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn

    Pythonevaluationhacktoberfesthacktoberfest2025
    View on GitHub↗17,787
  • confident-ai/deepevalconfident-ai avatar

    confident-ai/deepeval

    13,733View on GitHub↗

    Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs

    Pythonevaluation-frameworkevaluation-metricsllm-evaluation
    View on GitHub↗13,733

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • craftjarvis/jarvis-1CraftJarvis avatar

    CraftJarvis/JARVIS-1

    398View on GitHub↗

    [Website](http://craftjarvis-jarvis1.github.io/) [Paper](https://arxiv.org/abs/2311.05997) [Twitter](https://twitter.com/jeasinema/status/1723900032653643796)

    Java
    View on GitHub↗398
  • cvs-health/uqlmcvs-health avatar

    cvs-health/uqlm

    1,169View on GitHub↗

    UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

    Python
    View on GitHub↗1,169
  • declare-lab/instruct-evaldeclare-lab avatar

    declare-lab/instruct-eval

    553View on GitHub↗

    This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on held-out tasks.

    Python
    View on GitHub↗553
  • eleutherai/lm-evaluation-harnessEleutherAI avatar

    EleutherAI/lm-evaluation-harness

    11,460View on GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    View on GitHub↗11,460
  • embodied-generalist/embodied-generalistembodied-generalist avatar

    embodied-generalist/embodied-generalist

    485View on GitHub↗

    An Embodied Generalist Agent in 3D World

    Python
    View on GitHub↗485
  • evalplus/evalplusevalplus avatar

    evalplus/evalplus

    1,765View on GitHub↗

    Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

    Python
    View on GitHub↗1,765
  • explodinggradients/ragasexplodinggradients avatar

    explodinggradients/ragas

    14,400View on GitHub↗

    Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s

    Python
    View on GitHub↗14,400
  • freshllms/freshqaF

    freshllms/freshqa

    0View on GitHub↗
    View on GitHub↗0
  • future-agi/ai-evaluationF

    future-agi/ai-evaluation

    0View on GitHub↗

    Assess, Guard, and Monitor Your LLM Applications Built by Future AGI | Docs | Platform

    View on GitHub↗0
  • giskard-ai/giskardGiskard-AI avatar

    Giskard-AI/giskard

    5,434View on GitHub↗

    Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo

    Python
    View on GitHub↗5,434
  • huangwl18/language-plannerH

    huangwl18/language-planner

    0View on GitHub↗
    View on GitHub↗0
  • huangwl18/voxposerhuangwl18 avatar

    huangwl18/VoxPoser

    816View on GitHub↗

    Wenlong Huang 1 , Chen Wang 1 , Ruohan Zhang 1 , Yunzhu Li 1,2 , Jiajun Wu 1 , Li Fei-Fei 1

    Python
    View on GitHub↗816
  • huggingface/evaluatehuggingface avatar

    huggingface/evaluate

    2,455View on GitHub↗

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    Python
    View on GitHub↗2,455
  • huggingface/lightevalhuggingface avatar

    huggingface/lighteval

    2,453View on GitHub↗

    Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction

    Pythonevaluationevaluation-frameworkevaluation-metrics
    View on GitHub↗2,453
  • ianarawjo/chainforgeianarawjo avatar

    ianarawjo/ChainForge

    2,997View on GitHub↗

    An open-source visual programming environment for battle-testing prompts to LLMs.

    TypeScript
    View on GitHub↗2,997
  • johnsnowlabs/langtestJohnSnowLabs avatar

    JohnSnowLabs/langtest

    561View on GitHub↗

    Deliver safe & effective language models

    Python
    View on GitHub↗561
  • langchain-ai/agentevalslangchain-ai avatar

    langchain-ai/agentevals

    629View on GitHub↗

    Agentic applications give an LLM freedom over control flow in order to solve problems. While this freedom can be extremely powerful, the black box nature of LLMs can make it difficult to understand how changes in one part of your agent will affect others downstream. This makes evaluating your…

    Python
    View on GitHub↗629
  • langfuse/langfuselangfuse avatar

    langfuse/langfuse

    29,190View on GitHub↗

    Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model

    TypeScriptanalyticsautogenevaluation
    View on GitHub↗29,190
  • lhao499/instructrlL

    lhao499/instructrl

    0View on GitHub↗
    View on GitHub↗0
  • lm-sys/fastchatlm-sys avatar

    lm-sys/FastChat

    39,472View on GitHub↗

    FastChat is a training and serving platform for large language models that provides an integrated toolkit for fine-tuning, hosting, and benchmarking chatbots. It functions as an inference server capable of hosting multiple models and exposing them via a standardized API for chat applications. The platform distinguishes itself through a distributed model controller that manages worker nodes and routes requests across a hardware-agnostic inference layer supporting various accelerators. It includes a dedicated evaluation framework for assessing model quality using automated judges, multi-turn di

    Python
    View on GitHub↗39,472
  • microsoft/promptbenchmicrosoft avatar

    microsoft/promptbench

    2,808View on GitHub↗

    A unified evaluation framework for large language models

    Python
    View on GitHub↗2,808
  • minedojo/voyagerMineDojo avatar

    MineDojo/Voyager

    6,987View on GitHub↗

    Voyager is an autonomous embodied agent and lifelong learning framework that uses a large language model to explore virtual environments. It functions as a code-based action controller, translating natural language instructions into executable scripts to interact with its surroundings. The system features an automatic curriculum generator that creates sequences of exploration goals to discover new items and behaviors without human intervention. It maintains a skill library manager that stores learned behaviors as reusable code fragments, which can be composed to execute complex tasks. The fr

    JavaScriptembodied-learninglarge-language-modelsminecraft
    View on GitHub↗6,987
  • mlgroupjlu/llm-eval-surveyMLGroupJLU avatar

    MLGroupJLU/LLM-eval-survey

    1,600View on GitHub↗

    The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

    benchmarkevaluationlarge-language-models
    View on GitHub↗1,600
  • noahshinn024/reflexionN

    noahshinn024/reflexion

    0View on GitHub↗
    View on GitHub↗0
  • notmahi/clip-fieldsnotmahi avatar

    notmahi/clip-fields

    189View on GitHub↗

    [Paper](https://arxiv.org/abs/2210.05663) [Website](https://mahis.life/clip-fields/) [Code](https://github.com/notmahi/clip-fields) [Data](https://osf.io/famgv) [Video](https://youtu.be/bKu7GvRiSQU)

    Python
    View on GitHub↗189
  • notmahi/dobb-enotmahi avatar

    notmahi/dobb-e

    623View on GitHub↗

    Project webpage · Documentation (gitbooks) · Paper

    G-code
    View on GitHub↗623
  • onestardao/wfgyonestardao avatar

    onestardao/WFGY

    1,489View on GitHub↗
    Jupyter Notebookai-interpretabilityalignmentembedding
    View on GitHub↗1,489