awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

134 repository-uri

Awesome GitHub RepositoriesEvaluation Benchmarks

Tools and datasets for measuring model performance and reasoning capabilities.

Explore 134 awesome GitHub repositories matching part of an awesome list · Evaluation Benchmarks. Refine with filters or upvote what's useful.

Awesome Evaluation Benchmarks GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • microsoft/jarvisAvatar microsoft

    microsoft/JARVIS

    24,854Vezi pe GitHub↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Implements a benchmark suite to measure model success in automating complex machine learning tasks using standardized datasets.

    Python
    Vezi pe GitHub↗24,854
  • datawhalechina/prompt-engineering-for-developersAvatar datawhalechina

    datawhalechina/prompt-engineering-for-developers

    24,267Vezi pe GitHub↗

    This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap

    Implements systematic methods for measuring model performance and reasoning capabilities using benchmark datasets.

    Jupyter Notebook
    Vezi pe GitHub↗24,267
  • pyspur-dev/pyspurAvatar PySpur-Dev

    PySpur-Dev/pyspur

    5,677Vezi pe GitHub↗

    Runs pre-built benchmarks from academic datasets to measure an AI workflow's reasoning and knowledge.

    TypeScriptagentagentsai
    Vezi pe GitHub↗5,677
  • llm-attacks/llm-attacksAvatar llm-attacks

    llm-attacks/llm-attacks

    4,509Vezi pe GitHub↗

    This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.

    Universal and transferable adversarial attacks on aligned models.

    Python
    Vezi pe GitHub↗4,509
  • microsoft/malmoAvatar Microsoft

    Microsoft/malmo

    4,265Vezi pe GitHub↗

    Malmo este o platformă de simulare bazată pe voxeli, concepută pentru cercetarea în domeniul inteligenței artificiale și studiul comportamentelor agenților autonomi. Construit ca un mediu sandbox folosind Minecraft, acesta servește drept framework pentru simularea multi-agent și cercetarea în învățarea prin consolidare (reinforcement learning) într-o grilă 3D de blocuri. Proiectul se distinge printr-un framework de simulare multi-agent care coordonează și sincronizează mai mulți agenți autonomi pentru a îndeplini misiuni colaborative. Oferă o interfață standardizată care respectă specificațiile de învățare prin consolidare, permițându-i să funcționeze ca un mediu pentru antrenarea agenților prin încercare și eroare. Platforma acoperă o gamă largă de capabilități, inclusiv generarea de medii de cercetare cu definiții de sarcini reproductibile și integrarea backend-urilor de joc externe. Suportă agenți scriși în mai multe limbaje de programare printr-un strat de comunicare agnostică față de limbaj și oferă instrumente pentru vizualizarea stării la distanță a simulărilor. Motorul de simulare și dependențele serverului sunt furnizate ca implementări containerizate pentru a asigura o instalare consistentă pe diferite sisteme.

    Uses structured configuration files to define reproducible goals and environmental constraints for AI agent tasks.

    Java
    Vezi pe GitHub↗4,265
  • verazuo/jailbreak_llmsAvatar verazuo

    verazuo/jailbreak_llms

    3,563Vezi pe GitHub↗

    This project is a comprehensive ecosystem of frameworks, toolkits, and datasets designed to evaluate model vulnerabilities and analyze jailbreak patterns. It serves as an adversarial testing framework and research toolkit for measuring the effectiveness of safety guardrails in large language models. The system includes a library of real-world prompt injection datasets harvested from social media to study bypass strategies. It provides specialized tools for semantic attack analysis and prompt visualization, allowing for the mapping of relationships between adversarial prompts to discover commo

    Quantifies model vulnerability by comparing generated outputs against predefined forbidden criteria.

    Jupyter Notebookchatgptjailbreakjailbreaking
    Vezi pe GitHub↗3,563
  • memodb-io/acontextAvatar memodb-io

    memodb-io/Acontext

    3,035Vezi pe GitHub↗

    Acontext is an LLM orchestration backend and agent memory framework designed to manage session state and knowledge for AI agents. It functions as a context manager and orchestration layer that integrates model providers with a secure code sandbox and a zero-knowledge data store. The project is distinguished by its approach to knowledge distillation, capturing agent learnings as reusable Markdown skills and structured memory files. It provides a secure execution environment where shell commands and scripts run in isolated containers with the ability to mount these persistent skill files direct

    The product sets custom standards for evaluating whether an extracted task was successfully completed or failed.

    TypeScriptagentagent-development-kitagent-observability
    Vezi pe GitHub↗3,035
  • zjunlp/easyeditAvatar zjunlp

    zjunlp/EasyEdit

    2,718Vezi pe GitHub↗

    EasyEdit is a framework and toolkit designed for updating, inserting, or erasing specific factual information within large language models without requiring full retraining. It functions as a parameter modifier and knowledge editing system capable of performing targeted weight updates across diverse model architectures. The project distinguishes itself by supporting both text-based and multimodal model editing, allowing for knowledge updates across image and text modalities. It provides utilities for model steering to adjust personality and reasoning patterns in real time via activation inter

    Detoxifies models using knowledge editing techniques.

    Jupyter Notebookartificial-intelligencebaichuanchatgpt
    Vezi pe GitHub↗2,718
  • xlang-ai/osworldAvatar xlang-ai

    xlang-ai/OSWorld

    2,584Vezi pe GitHub↗

    OSWorld is an evaluation framework and multimodal agent benchmark designed to test the ability of large language models to complete complex tasks within virtualized operating system environments. It provides a virtualized desktop sandbox and a virtual machine orchestrator to deploy, snapshot, and reset cloud-based desktops, ensuring reproducible test states for AI agent interactions. The system distinguishes itself by providing an OS-level action space that translates model decisions into mouse clicks, keyboard inputs, and system commands. It employs a standardized interface to integrate vari

    Computes success scores for multi-step tasks by comparing final system states against ground truth rules.

    Pythonagentartificial-intelligencebenchmark
    Vezi pe GitHub↗2,584
  • rlinf/rlinfAvatar RLinf

    RLinf/RLinf

    2,502Vezi pe GitHub↗

    RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the infrastructure to train vision-language-action models and robotic policies using a combination of reinforcement learning and supervised fine-tuning. The system is designed for scaling workloads across GPU clusters, managing the placement of actors, rollout workers, and environment components. It features a specialized robotics data collection pipeline for gathering teleoperated demonstrations and simulation trajectories into standardized replay buffers, alongside a hardware interface

    Verifies model checkpoints in simulation to measure task success signals and rewards.

    Pythonagentic-aiembodied-aireinforcement-learning
    Vezi pe GitHub↗2,502
  • vchitect/vbenchAvatar Vchitect

    Vchitect/VBench

    1,656Vezi pe GitHub↗

    CVPR2024 Highlight VBench - We Evaluate Video Generation

    Comprehensive suite for evaluating video generation models.

    Python
    Vezi pe GitHub↗1,656
  • stormzhang/ai-coding-guideAvatar stormzhang

    stormzhang/ai-coding-guide

    1,164Vezi pe GitHub↗

    Acest proiect oferă un ghid și un framework complet pentru implementarea asistenților de codare AI autonomi în medii de dezvoltare locale. Se concentrează pe orchestrarea echipelor de agenți care pot planifica, executa și verifica sarcini complexe de inginerie software, cum ar fi refactorizarea, rezolvarea bug-urilor și generarea de teste, menținând în același timp o conștientizare profundă a contextului și memoriei specifice proiectului. Sistemul se distinge printr-o arhitectură robustă, axată pe securitate, care impune controale de acces granulare, izolarea execuției și aprobări obligatorii de tip human-in-the-loop pentru toate modificările de fișiere și apelurile către instrumente externe. Suportă automatizarea sofisticată a fluxului de lucru permițând dezvoltatorilor să definească abilități personalizate, reutilizabile și instrucțiuni ierarhice care persistă între sesiuni, asigurând un comportament consistent și retenția cunoștințelor pe tot parcursul ciclului de viață al dezvoltării software. Dincolo de automatizarea de bază, platforma oferă instrumente extinse de observabilitate și gestionare, inclusiv monitorizarea utilizării token-urilor în timp real, vizualizarea interactivă a diff-urilor de cod și monitorizarea sesiunilor în fundal. Se integrează direct în fluxurile de lucru bazate pe terminal și suportă diverși furnizori de inteligență artificială, permițând utilizatorilor să optimizeze performanța și costurile operaționale prin selecția modelului și ajustări ale raționamentului specifice sarcinii. Repository-ul servește atât ca resursă educațională pentru stăpânirea dezvoltării integrate cu AI, cât și ca toolkit funcțional pentru implementarea agenților autonomi care operează în limite de securitate definite.

    Enables users to define success conditions through test cases to facilitate model self-validation and iteration.

    agentai-codinganthropic
    Vezi pe GitHub↗1,164
  • thu-coai/safety-promptsAvatar thu-coai

    thu-coai/Safety-Prompts

    1,176Vezi pe GitHub↗

    Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。

    Assesses safety and alignment in Chinese language models.

    attack-defensechatgptchinese-language
    Vezi pe GitHub↗1,176
  • centerforaisafety/harmbenchAvatar centerforaisafety

    centerforaisafety/HarmBench

    991Vezi pe GitHub↗

    📰 Latest News 📰 - 🗡️ What is HarmBench 🛡️ - 🌐 Overview 🌐 - ☕ Quick Start ☕ - ⚙️ Installation - 🛠️ Running the Evaluation Pipeline - ➕ Using your own models in HarmBench - ➕ Using your own red teaming methods in HarmBench - 🤗 Classifiers - ⚓ Documentation ⚓ - 🌱 HarmBench's Roadmap 🌱 -…

    Standardized framework for automated red teaming and refusal.

    Jupyter Notebook
    Vezi pe GitHub↗991
  • easyjailbreak/easyjailbreakAvatar EasyJailbreak

    EasyJailbreak/EasyJailbreak

    870Vezi pe GitHub↗

    An easy-to-use Python framework to generate adversarial jailbreak prompts.

    Unified framework for executing and studying jailbreak attacks.

    Python
    Vezi pe GitHub↗870
  • yuliang-liu/multimodalocrAvatar Yuliang-Liu

    Yuliang-Liu/MultimodalOCR

    852Vezi pe GitHub↗

    On the Hidden Mystery of OCR in Large Multimodal Models (OCRBench)

    Investigating hidden OCR capabilities in multimodal models.

    Python
    Vezi pe GitHub↗852
  • bradyfu/video-mmeAvatar BradyFU

    BradyFU/Video-MME

    779Vezi pe GitHub↗

    ✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    Comprehensive benchmark for video analysis capabilities.

    Vezi pe GitHub↗779
  • openrobotlab/embodiedscanAvatar OpenRobotLab

    OpenRobotLab/EmbodiedScan

    669Vezi pe GitHub↗

    EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

    Comprehensive perception suite for 3D scene understanding.

    Python
    Vezi pe GitHub↗669
  • vision-x-nyu/thinking-in-spaceAvatar vision-x-nyu

    vision-x-nyu/thinking-in-space

    674Vezi pe GitHub↗

    Evaluating spatial reasoning and recall in multimodal models.

    Python
    Vezi pe GitHub↗674
  • jailbreakbench/jailbreakbenchAvatar JailbreakBench

    JailbreakBench/jailbreakbench

    615Vezi pe GitHub↗

    JailbreakBench -->

    Open robustness benchmark for testing jailbreak attacks.

    Python
    Vezi pe GitHub↗615
Înapoi123456…7Înainte
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Evaluation Benchmarks

Explorează sub-etichetele

  • Automation Success Metrics2 sub-tag-uriStandardized measures for evaluating how effectively models can automate complex multi-step tasks. **Distinct from Evaluation Benchmarks:** Focuses specifically on task automation success rates rather than general model reasoning or domain performance
  • Task Scenario DefinitionsStructured data formats used to define specific evaluation scenarios and goals for AI agents. **Distinct from Evaluation Benchmarks:** Focuses on the definition of the scenario/task rather than the general toolkit for measurement