134 repository-uri
Tools and datasets for measuring model performance and reasoning capabilities.
Explore 134 awesome GitHub repositories matching part of an awesome list · Evaluation Benchmarks. Refine with filters or upvote what's useful.
JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat
Implements a benchmark suite to measure model success in automating complex machine learning tasks using standardized datasets.
This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap
Implements systematic methods for measuring model performance and reasoning capabilities using benchmark datasets.
Runs pre-built benchmarks from academic datasets to measure an AI workflow's reasoning and knowledge.
This repository provides tools and methodologies for studying adversarial attacks on large language models. It focuses on understanding how carefully crafted inputs can manipulate or bypass the safety mechanisms of LLMs, enabling researchers to probe model vulnerabilities and improve their robustness. The project covers techniques for generating adversarial prompts, evaluating model responses under attack conditions, and analyzing the effectiveness of different attack strategies.
Universal and transferable adversarial attacks on aligned models.
Malmo este o platformă de simulare bazată pe voxeli, concepută pentru cercetarea în domeniul inteligenței artificiale și studiul comportamentelor agenților autonomi. Construit ca un mediu sandbox folosind Minecraft, acesta servește drept framework pentru simularea multi-agent și cercetarea în învățarea prin consolidare (reinforcement learning) într-o grilă 3D de blocuri. Proiectul se distinge printr-un framework de simulare multi-agent care coordonează și sincronizează mai mulți agenți autonomi pentru a îndeplini misiuni colaborative. Oferă o interfață standardizată care respectă specificațiile de învățare prin consolidare, permițându-i să funcționeze ca un mediu pentru antrenarea agenților prin încercare și eroare. Platforma acoperă o gamă largă de capabilități, inclusiv generarea de medii de cercetare cu definiții de sarcini reproductibile și integrarea backend-urilor de joc externe. Suportă agenți scriși în mai multe limbaje de programare printr-un strat de comunicare agnostică față de limbaj și oferă instrumente pentru vizualizarea stării la distanță a simulărilor. Motorul de simulare și dependențele serverului sunt furnizate ca implementări containerizate pentru a asigura o instalare consistentă pe diferite sisteme.
Uses structured configuration files to define reproducible goals and environmental constraints for AI agent tasks.
This project is a comprehensive ecosystem of frameworks, toolkits, and datasets designed to evaluate model vulnerabilities and analyze jailbreak patterns. It serves as an adversarial testing framework and research toolkit for measuring the effectiveness of safety guardrails in large language models. The system includes a library of real-world prompt injection datasets harvested from social media to study bypass strategies. It provides specialized tools for semantic attack analysis and prompt visualization, allowing for the mapping of relationships between adversarial prompts to discover commo
Quantifies model vulnerability by comparing generated outputs against predefined forbidden criteria.
Acontext is an LLM orchestration backend and agent memory framework designed to manage session state and knowledge for AI agents. It functions as a context manager and orchestration layer that integrates model providers with a secure code sandbox and a zero-knowledge data store. The project is distinguished by its approach to knowledge distillation, capturing agent learnings as reusable Markdown skills and structured memory files. It provides a secure execution environment where shell commands and scripts run in isolated containers with the ability to mount these persistent skill files direct
The product sets custom standards for evaluating whether an extracted task was successfully completed or failed.
EasyEdit is a framework and toolkit designed for updating, inserting, or erasing specific factual information within large language models without requiring full retraining. It functions as a parameter modifier and knowledge editing system capable of performing targeted weight updates across diverse model architectures. The project distinguishes itself by supporting both text-based and multimodal model editing, allowing for knowledge updates across image and text modalities. It provides utilities for model steering to adjust personality and reasoning patterns in real time via activation inter
Detoxifies models using knowledge editing techniques.
OSWorld is an evaluation framework and multimodal agent benchmark designed to test the ability of large language models to complete complex tasks within virtualized operating system environments. It provides a virtualized desktop sandbox and a virtual machine orchestrator to deploy, snapshot, and reset cloud-based desktops, ensuring reproducible test states for AI agent interactions. The system distinguishes itself by providing an OS-level action space that translates model decisions into mouse clicks, keyboard inputs, and system commands. It employs a standardized interface to integrate vari
Computes success scores for multi-step tasks by comparing final system states against ground truth rules.
RLinf is a distributed reinforcement learning orchestrator and embodied AI training framework. It provides the infrastructure to train vision-language-action models and robotic policies using a combination of reinforcement learning and supervised fine-tuning. The system is designed for scaling workloads across GPU clusters, managing the placement of actors, rollout workers, and environment components. It features a specialized robotics data collection pipeline for gathering teleoperated demonstrations and simulation trajectories into standardized replay buffers, alongside a hardware interface
Verifies model checkpoints in simulation to measure task success signals and rewards.
CVPR2024 Highlight VBench - We Evaluate Video Generation
Comprehensive suite for evaluating video generation models.
Acest proiect oferă un ghid și un framework complet pentru implementarea asistenților de codare AI autonomi în medii de dezvoltare locale. Se concentrează pe orchestrarea echipelor de agenți care pot planifica, executa și verifica sarcini complexe de inginerie software, cum ar fi refactorizarea, rezolvarea bug-urilor și generarea de teste, menținând în același timp o conștientizare profundă a contextului și memoriei specifice proiectului. Sistemul se distinge printr-o arhitectură robustă, axată pe securitate, care impune controale de acces granulare, izolarea execuției și aprobări obligatorii de tip human-in-the-loop pentru toate modificările de fișiere și apelurile către instrumente externe. Suportă automatizarea sofisticată a fluxului de lucru permițând dezvoltatorilor să definească abilități personalizate, reutilizabile și instrucțiuni ierarhice care persistă între sesiuni, asigurând un comportament consistent și retenția cunoștințelor pe tot parcursul ciclului de viață al dezvoltării software. Dincolo de automatizarea de bază, platforma oferă instrumente extinse de observabilitate și gestionare, inclusiv monitorizarea utilizării token-urilor în timp real, vizualizarea interactivă a diff-urilor de cod și monitorizarea sesiunilor în fundal. Se integrează direct în fluxurile de lucru bazate pe terminal și suportă diverși furnizori de inteligență artificială, permițând utilizatorilor să optimizeze performanța și costurile operaționale prin selecția modelului și ajustări ale raționamentului specifice sarcinii. Repository-ul servește atât ca resursă educațională pentru stăpânirea dezvoltării integrate cu AI, cât și ca toolkit funcțional pentru implementarea agenților autonomi care operează în limite de securitate definite.
Enables users to define success conditions through test cases to facilitate model self-validation and iteration.
Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。
Assesses safety and alignment in Chinese language models.
📰 Latest News 📰 - 🗡️ What is HarmBench 🛡️ - 🌐 Overview 🌐 - ☕ Quick Start ☕ - ⚙️ Installation - 🛠️ Running the Evaluation Pipeline - ➕ Using your own models in HarmBench - ➕ Using your own red teaming methods in HarmBench - 🤗 Classifiers - ⚓ Documentation ⚓ - 🌱 HarmBench's Roadmap 🌱 -…
Standardized framework for automated red teaming and refusal.
An easy-to-use Python framework to generate adversarial jailbreak prompts.
Unified framework for executing and studying jailbreak attacks.
On the Hidden Mystery of OCR in Large Multimodal Models (OCRBench)
Investigating hidden OCR capabilities in multimodal models.
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Comprehensive benchmark for video analysis capabilities.
EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
Comprehensive perception suite for 3D scene understanding.
Evaluating spatial reasoning and recall in multimodal models.
JailbreakBench -->
Open robustness benchmark for testing jailbreak attacks.