awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 repositorios

Awesome GitHub RepositoriesModel Benchmarks

Comparative performance metrics, pricing data, and evaluation tools for generative artificial intelligence providers.

Distinct from Large Language Models: Distinct from Large Language Models: focuses on the comparative evaluation and benchmarking of providers rather than the models themselves.

Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Model Benchmarks. Refine with filters or upvote what's useful.

Awesome Model Benchmarks GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • microsoft/jarvisAvatar de microsoft

    microsoft/JARVIS

    24,854Ver en GitHub↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Evaluates the capability of large language models to automate complex tasks using standardized benchmarking datasets.

    Python
    Ver en GitHub↗24,854
  • owainlewis/awesome-artificial-intelligenceAvatar de owainlewis

    owainlewis/awesome-artificial-intelligence

    12,960Ver en GitHub↗

    This project is a comprehensive repository and curated index of resources, research papers, and development frameworks designed to support the construction and deployment of intelligent systems. It serves as a centralized knowledge base for developers seeking to navigate the technical landscape of artificial intelligence, ranging from foundational educational materials to specialized implementation guides. The repository distinguishes itself by providing structured directories for comparing generative artificial intelligence providers, including aggregated performance metrics, pricing data, a

    Aggregates performance metrics, pricing data, and evaluation tools to facilitate objective comparison of generative artificial intelligence providers.

    aiartificial-intelligencedeep-learning
    Ver en GitHub↗12,960
  • ray-project/llm-numbersAvatar de ray-project

    ray-project/llm-numbers

    4,310Ver en GitHub↗

    llm-numbers es un conjunto de herramientas de cálculo y benchmarks utilizados para predecir requisitos de hardware, uso de tokens y costos operativos en varios niveles de modelos. Proporciona una calculadora de costos y recursos basada en fórmulas y benchmarks para estimar tokens, memoria de GPU y gastos operativos para modelos de lenguaje grandes. El proyecto incluye un planificador de requisitos de hardware para calcular la VRAM y la memoria de GPU necesarias para alojar modelos basados en recuentos de parámetros. También cuenta con un estimador de tokens que convierte recuentos de palabras en estimaciones de tokens para predecir la facturación de la API y el uso de la ventana de contexto, junto con benchmarks de precios que comparan costos y compensaciones de rendimiento entre diferentes métodos de alojamiento. El conjunto de herramientas cubre el benchmarking de modelos de IA y la previsión de costos, planificación de recursos de GPU y análisis de rendimiento para medir las ganancias de rendimiento del procesamiento por lotes. Utiliza fórmulas deterministas y datasets de benchmark estáticos para mapear parámetros a memoria y calcular las relaciones costo-beneficio entre modelos base y ajuste fino (fine-tuning).

    Provides comparative pricing and throughput benchmarks for different generative AI model tiers and hosting methods.

    Ver en GitHub↗4,310
  • openai/simple-evalsAvatar de openai

    openai/simple-evals

    4,354Ver en GitHub↗

    This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and performance of models across diverse datasets. It provides a system for implementing model-based graders, running standardized tests for mathematical reasoning, coding, and factuality, and calculating quantified performance metrics such as precision, recall, F1 scores, and pass-at-k. The framework utilizes model-based grading and rubrics to validate response quality against expert-defined criteria. It includes a multi-model benchmarking loop and a model-agnostic API interface to co

    Runs a suite of standardized benchmarks to measure language model accuracy on reasoning, math, and coding.

    Python
    Ver en GitHub↗4,354
  • orchestra-research/ai-research-skillsAvatar de Orchestra-Research

    Orchestra-Research/AI-Research-SKILLs

    3,641Ver en GitHub↗

    This project is an LLM research orchestrator and autonomous AI agent framework designed to automate the scientific lifecycle. It functions as an end-to-end research pipeline and model training toolkit, managing everything from initial literature reviews and hypothesis testing to the final drafting of academic papers. The system is distinguished by its ability to convert unstructured academic PDFs into machine-executable knowledge layers, allowing agents to reproduce and extend research findings. It employs a two-loop orchestration architecture and a specialized research engineering skill libr

    Evaluates the ability of AI systems to autonomously design and analyze scientific experiments with rigor.

    TeXaiai-researchclaude
    Ver en GitHub↗3,641
  • mahonzhan/awesome-coding-planAvatar de mahonzhan

    mahonzhan/awesome-coding-plan

    1,641Ver en GitHub↗

    Awesome Coding Plan is a community-driven knowledge repository that provides a comparative analysis of subscription-based coding environments and artificial intelligence development tools. It functions as a tracker for developer tool costs, aggregating data on pricing structures, usage quotas, and token limits to assist in the selection of cloud-based coding services. The project utilizes a standardized framework to evaluate the performance and economic efficiency of various language models. By organizing technical metrics into a unified format, it allows for the objective assessment of proce

    Benchmarks processing speeds and token costs across different language models used for code generation.

    Ver en GitHub↗1,641
  1. Home
  2. Artificial Intelligence & ML
  3. Large Language Models
  4. Model Benchmarks

Explorar subetiquetas

  • Automation Capability Benchmarks1 sub-etiquetaStandardized benchmarks specifically designed to measure the automation efficiency of AI models. **Distinct from Model Benchmarks:** Focuses on the ability to automate complex tasks rather than static model performance or pricing
  • Multilingual Accuracy Evaluations1 sub-etiquetaAssessments that measure model performance and accuracy across different natural languages using translated datasets. **Distinct from Model Benchmarks:** Focuses on linguistic accuracy and translation consistency across languages, whereas the parent covers general provider benchmarks.