awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 个仓库

Awesome GitHub RepositoriesModel Benchmarks

Comparative performance metrics, pricing data, and evaluation tools for generative artificial intelligence providers.

Distinct from Large Language Models: Distinct from Large Language Models: focuses on the comparative evaluation and benchmarking of providers rather than the models themselves.

Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Model Benchmarks. Refine with filters or upvote what's useful.

Awesome Model Benchmarks GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • microsoft/jarvismicrosoft 的头像

    microsoft/JARVIS

    24,854在 GitHub 上查看↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Evaluates the capability of large language models to automate complex tasks using standardized benchmarking datasets.

    Python
    在 GitHub 上查看↗24,854
  • owainlewis/awesome-artificial-intelligenceowainlewis 的头像

    owainlewis/awesome-artificial-intelligence

    12,960在 GitHub 上查看↗

    This project is a comprehensive repository and curated index of resources, research papers, and development frameworks designed to support the construction and deployment of intelligent systems. It serves as a centralized knowledge base for developers seeking to navigate the technical landscape of artificial intelligence, ranging from foundational educational materials to specialized implementation guides. The repository distinguishes itself by providing structured directories for comparing generative artificial intelligence providers, including aggregated performance metrics, pricing data, a

    Aggregates performance metrics, pricing data, and evaluation tools to facilitate objective comparison of generative artificial intelligence providers.

    aiartificial-intelligencedeep-learning
    在 GitHub 上查看↗12,960
  • ray-project/llm-numbersray-project 的头像

    ray-project/llm-numbers

    4,310在 GitHub 上查看↗

    llm-numbers 是一套计算工具和基准测试,用于预测各种模型层级的硬件要求、令牌使用量和运营成本。它提供了一个基于公式和基准测试的成本和资源计算器,用于估算大语言模型的令牌、GPU 内存和运营费用。 该项目包括一个硬件需求规划器,用于根据参数数量计算托管模型所需的 VRAM 和 GPU 内存。它还具有一个令牌估算器,将字数转换为令牌估算值以预测 API 账单和上下文窗口使用情况,以及比较不同托管方法之间成本和吞吐量权衡的定价基准。 该工具集涵盖了 AI 模型基准测试和成本预测、GPU 资源规划以及用于衡量批处理吞吐量增益的性能分析。它利用确定性公式和静态基准数据集将参数映射到内存,并计算基础模型与微调模型之间的成本效益比。

    Provides comparative pricing and throughput benchmarks for different generative AI model tiers and hosting methods.

    在 GitHub 上查看↗4,310
  • openai/simple-evalsopenai 的头像

    openai/simple-evals

    4,354在 GitHub 上查看↗

    This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and performance of models across diverse datasets. It provides a system for implementing model-based graders, running standardized tests for mathematical reasoning, coding, and factuality, and calculating quantified performance metrics such as precision, recall, F1 scores, and pass-at-k. The framework utilizes model-based grading and rubrics to validate response quality against expert-defined criteria. It includes a multi-model benchmarking loop and a model-agnostic API interface to co

    Runs a suite of standardized benchmarks to measure language model accuracy on reasoning, math, and coding.

    Python
    在 GitHub 上查看↗4,354
  • orchestra-research/ai-research-skillsOrchestra-Research 的头像

    Orchestra-Research/AI-Research-SKILLs

    3,641在 GitHub 上查看↗

    This project is an LLM research orchestrator and autonomous AI agent framework designed to automate the scientific lifecycle. It functions as an end-to-end research pipeline and model training toolkit, managing everything from initial literature reviews and hypothesis testing to the final drafting of academic papers. The system is distinguished by its ability to convert unstructured academic PDFs into machine-executable knowledge layers, allowing agents to reproduce and extend research findings. It employs a two-loop orchestration architecture and a specialized research engineering skill libr

    Evaluates the ability of AI systems to autonomously design and analyze scientific experiments with rigor.

    TeXaiai-researchclaude
    在 GitHub 上查看↗3,641
  • mahonzhan/awesome-coding-planmahonzhan 的头像

    mahonzhan/awesome-coding-plan

    1,641在 GitHub 上查看↗

    Awesome Coding Plan is a community-driven knowledge repository that provides a comparative analysis of subscription-based coding environments and artificial intelligence development tools. It functions as a tracker for developer tool costs, aggregating data on pricing structures, usage quotas, and token limits to assist in the selection of cloud-based coding services. The project utilizes a standardized framework to evaluate the performance and economic efficiency of various language models. By organizing technical metrics into a unified format, it allows for the objective assessment of proce

    Benchmarks processing speeds and token costs across different language models used for code generation.

    在 GitHub 上查看↗1,641
  1. Home
  2. Artificial Intelligence & ML
  3. Large Language Models
  4. Model Benchmarks

探索子标签

  • Automation Capability Benchmarks1 个子标签Standardized benchmarks specifically designed to measure the automation efficiency of AI models. **Distinct from Model Benchmarks:** Focuses on the ability to automate complex tasks rather than static model performance or pricing
  • Multilingual Accuracy Evaluations1 个子标签Assessments that measure model performance and accuracy across different natural languages using translated datasets. **Distinct from Model Benchmarks:** Focuses on linguistic accuracy and translation consistency across languages, whereas the parent covers general provider benchmarks.