awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

43 Repos

Awesome GitHub RepositoriesModel Capability Assessment

Tools for benchmarking and selecting models based on specific requirements.

Explore 43 awesome GitHub repositories matching artificial intelligence & ml · Model Capability Assessment. Refine with filters or upvote what's useful.

Awesome Model Capability Assessment GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • langchain-ai/langchainAvatar von langchain-ai

    langchain-ai/langchain

    139,458Auf GitHub ansehen↗

    LangChain is an orchestration framework designed for building, managing, and deploying applications powered by large language models. It provides a unified integration layer that normalizes disparate model provider APIs into a consistent set of primitives, enabling developers to build complex, multi-step AI workflows that manage state, memory, and tool execution. The project distinguishes itself through a durable execution runtime that maintains persistent state across long-running processes by checkpointing progress to external storage. It models agent workflows as directed graphs, allowing

    Benchmarks model performance to assist in selecting appropriate providers for specific requirements.

    Pythonagentsaiai-agents
    Auf GitHub ansehen↗139,458
  • yidadaa/chatgpt-next-webAvatar von Yidadaa

    Yidadaa/ChatGPT-Next-Web

    88,263Auf GitHub ansehen↗

    ChatGPT-Next-Web is a web-based chat interface for interacting with large language models via API or self-hosted model runners. It functions as a prompt management tool and a cross-platform application available for web, mobile, and desktop environments. The project distinguishes itself through a plugin integration gateway that extends model capabilities with external tools like network search and calculators. It includes a self-hosted administrative dashboard for controlling model lists, member permissions, and access passwords on private infrastructure. The application covers prompt engine

    Allows administrators to edit available models, rename entries, and designate vision capabilities.

    TypeScript
    Auf GitHub ansehen↗88,263
  • voltagent/awesome-claude-code-subagentsAvatar von VoltAgent

    VoltAgent/awesome-claude-code-subagents

    21,906Auf GitHub ansehen↗

    This project provides a framework for managing multi-agent systems, designed to automate complex software development, infrastructure, and business workflows. It functions as a multi-agent workflow orchestrator that routes tasks to domain-specific workers while maintaining state persistence and infrastructure automation. By leveraging large language models, the system decomposes high-level objectives into actionable plans, ensuring that complex operations are executed with consistency and reliability. The framework distinguishes itself through its hierarchical agent registry and policy-driven

    Automatically assigns subagents to specific models based on task complexity to balance reasoning depth against execution speed and cost.

    Shellai-agent-frameworkai-agent-toolsai-agents
    Auf GitHub ansehen↗21,906
  • letta-ai/lettaAvatar von letta-ai

    letta-ai/letta

    21,168Auf GitHub ansehen↗

    Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com

    Uses language models to judge open-ended agent responses based on nuanced or creative criteria.

    Pythonaiai-agentsllm
    Auf GitHub ansehen↗21,168
  • vxcontrol/pentagiAvatar von vxcontrol

    vxcontrol/pentagi

    17,766Auf GitHub ansehen↗

    Pentagi is an autonomous security testing framework and agent orchestrator designed to plan and execute end-to-end security assessments. It utilizes a coordination engine to decompose complex goals into actionable subtasks, performing automated penetration testing and vulnerability research within isolated container environments. The system distinguishes itself through a temporal knowledge graph that tracks semantic relationships between entities and vulnerabilities to reuse intelligence across projects. It includes a web intelligence reconnaissance tool for automated data gathering and agent

    Provides diagnostic utilities to benchmark AI provider configurations and embedding functions.

    Goai-agentsai-security-toolanthropic
    Auf GitHub ansehen↗17,766
  • decolua/9routerAvatar von decolua

    decolua/9router

    17,690Auf GitHub ansehen↗

    9router is an AI model gateway designed to route requests from AI coding tools to multiple model providers through a single unified API. It provides administration for self-hosted AI proxy deployments, allowing users to manage API keys and model access on local servers or edge networks. The system differentiates itself through multi-provider API normalization, which translates incompatible request and response formats to ensure compatibility across different AI models. It features AI provider failover management to automatically switch between providers or accounts when quotas are exhausted o

    Ensures specialized inputs are sent to the most capable providers by reordering target models per request.

    JavaScriptai-agentsai-gatewayanthropic
    Auf GitHub ansehen↗17,690
  • vercel/vercelAvatar von vercel

    vercel/vercel

    15,738Auf GitHub ansehen↗

    Vercel is a cloud platform for building, deploying, and scaling web applications. It provides a unified infrastructure that automates the build process by detecting project frameworks and distributing static and dynamic content through a global content delivery network. The platform executes application logic using serverless functions that scale automatically based on real-time traffic demand. The platform distinguishes itself through a centralized AI gateway that proxies requests to multiple model providers, enabling standardized authentication, observability, and cost tracking. It supports

    Allows developers to programmatically retrieve available models and their configurations.

    TypeScriptclicloudcommand
    Auf GitHub ansehen↗15,738
  • kilo-org/kilocodeAvatar von Kilo-Org

    Kilo-Org/kilocode

    15,616Auf GitHub ansehen↗

    Kilocode is an autonomous engineering platform designed to orchestrate AI agents for complex software development tasks. It functions as a comprehensive system for automating coding, testing, and repository management by integrating directly with your codebase and terminal. The platform provides a unified gateway for model orchestration, allowing for the management of agentic workflows, event-driven automation, and persistent session state across distributed development environments. The platform distinguishes itself through its federated task management and policy-based access control, which

    Provides a public endpoint to retrieve supported language models, their specifications, and pricing details.

    TypeScriptaiai-ageai-coding
    Auf GitHub ansehen↗15,616
  • chiphuyen/aie-bookAvatar von chiphuyen

    chiphuyen/aie-book

    13,779Auf GitHub ansehen↗

    This project serves as a comprehensive educational resource and technical handbook for engineers building applications powered by large language models. It provides a structured framework for mastering the principles of artificial intelligence engineering, covering the full lifecycle of model development from initial design to production deployment. The repository distinguishes itself by offering a deep dive into the practical implementation of advanced design patterns, including retrieval-augmented generation, agentic tool orchestration, and parameter-efficient model adaptation. It emphasize

    Provides tools for benchmarking and selecting models based on specific application requirements.

    Jupyter Notebook
    Auf GitHub ansehen↗13,779
  • owainlewis/awesome-artificial-intelligenceAvatar von owainlewis

    owainlewis/awesome-artificial-intelligence

    12,960Auf GitHub ansehen↗

    This project is a comprehensive repository and curated index of resources, research papers, and development frameworks designed to support the construction and deployment of intelligent systems. It serves as a centralized knowledge base for developers seeking to navigate the technical landscape of artificial intelligence, ranging from foundational educational materials to specialized implementation guides. The repository distinguishes itself by providing structured directories for comparing generative artificial intelligence providers, including aggregated performance metrics, pricing data, a

    Aggregates benchmarks and pricing data to assess and compare the capabilities of different generation providers.

    aiartificial-intelligencedeep-learning
    Auf GitHub ansehen↗12,960
  • jacobgil/pytorch-grad-camAvatar von jacobgil

    jacobgil/pytorch-grad-cam

    12,893Auf GitHub ansehen↗

    Dieses Projekt ist eine Computer-Vision-Bibliothek für erklärbare KI und ein Framework für PyTorch, das eine Suite von Tools zur Visualisierung und Prüfung der internen Entscheidungsprozesse tiefer neuronaler Netze bereitstellt. Es dient als Attributions-Tool für neuronale Netze und Debugging-Dienstprogramm, um zu identifizieren, welche Bildregionen Modellvorhersagen steuern. Die Bibliothek zeichnet sich durch ihre Unterstützung sowohl für gradientenbasierte als auch für gradientenfreie Attributionsmethoden aus, was die Generierung visueller Heatmaps und Attributionskarten ermöglicht, ohne dass Änderungen am ursprünglichen Modellquellcode erforderlich sind. Sie differenziert sich zudem durch die Entdeckung visueller Konzepte, wobei Matrixfaktorisierung verwendet wird, um interne Aktivierungen in interpretierbare Muster zu zerlegen und latente Einbettungen auf Pixelwichtigkeit abzubilden. Das Framework deckt ein breites Spektrum an Fähigkeiten ab, einschließlich Heatmap-Generierung und -Verfeinerung, räumlicher Transformation für Architekturen wie Vision-Transformer und Anpassungen für multimodale Vision-Ziele wie Objekterkennung und semantische Segmentierung. Es enthält zudem eine Suite zur Bewertung der Modelltreue, die Störungsanalysen, Ablationsstudien und Lokalisierungsmessungen verwendet, um die Genauigkeit generierter Erklärungen zu quantifizieren. Das Projekt bietet Mechanismen für dynamisches Aktivierungs-Hooking, benutzerdefinierte Architektur-Anpassung und zielorientierte Zielkonfiguration, um Erklärbarkeits-Tools mit verschiedenen Modellausgaben zu verbinden.

    Verifies if an explanation accurately reflects the model's process by measuring confidence drops after removing relevant pixels.

    Python
    Auf GitHub ansehen↗12,893
  • microsoft/vscode-copilot-chatAvatar von microsoft

    microsoft/vscode-copilot-chat

    9,493Auf GitHub ansehen↗

    This project is an AI-powered IDE extension and LLM coding assistant that provides a conversational interface for generating, refactoring, and debugging code. It functions as an AI agent framework and a Model Context Protocol client, connecting AI models to external data sources and tools to automate complex development tasks. The system is distinguished by its use of autonomous AI agents capable of multi-step task execution, including the ability to read files, modify code, and run terminal commands iteratively. It supports recursive agent orchestration through subagent delegation and employ

    Manages lists of available models, their capabilities, and context sizes for selection.

    TypeScript
    Auf GitHub ansehen↗9,493
  • getpaseo/paseoAvatar von getpaseo

    getpaseo/paseo

    9,118Auf GitHub ansehen↗

    Paseo is an LLM coding agent orchestrator and multi-agent workflow manager designed to coordinate multiple AI agents across isolated git worktrees. It provides a unified control interface for managing these agents and their associated environments to execute complex programming tasks. The system distinguishes itself through a remote agent daemon that enables secure access to local coding agents via encrypted relays. It employs a git worktree environment manager to isolate parallel tasks into dedicated directories and branch-based server URLs, preventing file collisions and network port confli

    Provides administrative tools for adding, relabeling, and refining the list of available AI models.

    TypeScriptadeagentsclaude-code
    Auf GitHub ansehen↗9,118
  • oumi-ai/oumiAvatar von oumi-ai

    oumi-ai/oumi

    8,858Auf GitHub ansehen↗

    Oumi is a comprehensive large language model development platform designed for synthesizing data, fine-tuning models, and running performance evaluations. It serves as a unified environment for the entire model lifecycle, encompassing a training and fine-tuning suite, an evaluation framework, and tools for synthetic data generation and model distillation. The platform is distinguished by its iterative, failure-driven synthesis approach, which analyzes model weaknesses during evaluation to generate targeted training data. It utilizes an LLM-based judge framework to programmatically score respo

    Assesses model outputs for instruction following, safety, and truthfulness using general-purpose evaluation dimensions.

    Pythondpoevaluationfine-tuning
    Auf GitHub ansehen↗8,858
  • airbnb/epoxyAvatar von airbnb

    airbnb/epoxy

    8,556Auf GitHub ansehen↗

    Epoxy is an Android library for building complex RecyclerView screens using a model-driven approach. It generates RecyclerView adapter models at compile time from annotated custom views, data binding layouts, or view holders, eliminating the manual boilerplate typically associated with view holders and adapters. The library provides a diffing engine that automatically compares model lists and applies minimal updates with animations for insertions, removals, and moves. The library distinguishes itself through its controller-based model building, where a controller class with a buildModels meth

    Manages lists of models that define items and their order in a RecyclerView with change notifications.

    Java
    Auf GitHub ansehen↗8,556
  • internlm/opencompassAvatar von InternLM

    InternLM/opencompass

    7,096Auf GitHub ansehen↗

    OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta

    Tests model stability and security by applying various attack methods and evaluating tool-use capabilities.

    Python
    Auf GitHub ansehen↗7,096
  • open-compass/opencompassAvatar von open-compass

    open-compass/opencompass

    6,678Auf GitHub ansehen↗

    OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API

    Measures model performance across examination, knowledge, reasoning, understanding, language, and safety dimensions.

    Pythonbenchmarkchatgptevaluation
    Auf GitHub ansehen↗6,678
  • cleverhans-lab/cleverhansAvatar von cleverhans-lab

    cleverhans-lab/cleverhans

    6,443Auf GitHub ansehen↗

    Cleverhans is an adversarial machine learning library and toolkit designed to generate adversarial examples, incorporate them into training loops, and benchmark the resilience of machine learning models. It provides a gradient-based attack framework for constructing both white-box and black-box attacks to identify model misclassifications. The project includes capabilities for model robustness benchmarking, allowing users to evaluate and verify how models resist evasion attacks and malicious input perturbations. It also facilitates adversarial training to increase a model's resistance to pert

    Provides tools to measure and verify the resilience of machine learning models against adversarial attacks across multiple frameworks.

    Jupyter Notebookbenchmarkingmachine-learningsecurity
    Auf GitHub ansehen↗6,443
  • tensorflow/cleverhansAvatar von tensorflow

    tensorflow/cleverhans

    6,443Auf GitHub ansehen↗

    Cleverhans ist eine TensorFlow-Bibliothek für Adversarial Machine Learning, die als Angriffs-Framework, Robustheits-Benchmark und Verteidigungsbibliothek dient. Sie bietet eine Sammlung von Tools zur Generierung von Adversarial Examples, zum Testen der Sicherheit neuronaler Netze und zur Implementierung von Schutzmechanismen, um die Widerstandsfähigkeit von Modellen gegen bösartige Eingaben zu erhöhen. Das Projekt konzentriert sich auf die Erstellung von gestörten Eingaben, die darauf ausgelegt sind, Machine-Learning-Modelle zu falschen Vorhersagen zu verleiten. Es ermöglicht die Bewertung der Stabilität und Genauigkeit von Deep-Learning-Modellen unter Adversarial-Noise und bietet Referenzimplementierungen bekannter Angriffsmethoden zur Identifizierung von Sicherheitslücken. Das Toolkit deckt die Generierung von Adversarial Examples, die Verteidigung von Machine-Learning-Modellen und das Benchmarking der Robustheit neuronaler Netze ab. Es nutzt eine modellagnostische Schnittstelle und differenzierbare Angriffsimplementierungen, um gradientenbasierte Störungen und iterative Optimierungsschleifen auszuführen.

    Measures model stability and accuracy by subjecting neural networks to simulated adversarial attacks.

    Jupyter Notebook
    Auf GitHub ansehen↗6,443
  • diegosouzapw/omnirouteAvatar von diegosouzapw

    diegosouzapw/OmniRoute

    6,391Auf GitHub ansehen↗

    OmniRoute is a unified LLM API gateway that connects multiple AI providers to a single endpoint. Its primary purpose is to simplify the integration of various AI models into tools and agents by translating different provider formats into a standardized API. The project distinguishes itself through a multi-strategy request routing system that optimizes for cost, speed, and availability, including automatic model fallbacks and a circuit-breaker resilience model to isolate provider failures. It employs a local-first security posture, using AES-256-GCM encryption to store API keys and conversatio

    Maintains a registry of AI model specifications, capabilities, and pricing synced from an external source.

    TypeScript
    Auf GitHub ansehen↗6,391
Vorherige123Nächste
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Infrastructure
  5. Evaluation & Validation
  6. Model Capability Assessment

Unter-Tags erkunden

  • Adversarial Robustness Testing2 Sub-TagsEvaluation of model stability and security through adversarial attack methods. **Distinct from Model Capability Assessment:** Specifically targets adversarial attack simulation rather than general capability benchmarking.
  • Model Capability Queries3 Sub-TagsMechanisms for retrieving lists of available models and their configuration details. **Distinct from Model Capability Assessment:** Focuses on programmatic discovery of model endpoints and capabilities, distinct from benchmarking or assessment.
  • Model Routing StrategiesMechanisms for assigning tasks to specific models based on complexity and cost. **Distinct from Model Capability Assessment:** Distinct from Model Capability Assessment: focuses on the routing logic for task dispatching rather than benchmarking.
  • Project-Based Model AssessmentsEvaluation methods that validate learner mastery through the implementation of models on real-world datasets. **Distinct from Model Capability Assessment:** Focuses on assessing a student's ability to build a model rather than benchmarking a model's internal capabilities.