awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 个仓库

Awesome GitHub RepositoriesAutomation Capability Benchmarks

Standardized benchmarks specifically designed to measure the automation efficiency of AI models.

Distinct from Model Benchmarks: Focuses on the ability to automate complex tasks rather than static model performance or pricing

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Automation Capability Benchmarks. Refine with filters or upvote what's useful.

Awesome Automation Capability Benchmarks GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • microsoft/jarvismicrosoft 的头像

    microsoft/JARVIS

    24,854在 GitHub 上查看↗

    JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat

    Evaluates the capability of large language models to automate complex tasks using standardized benchmarking datasets.

    Python
    在 GitHub 上查看↗24,854
  • orchestra-research/ai-research-skillsOrchestra-Research 的头像

    Orchestra-Research/AI-Research-SKILLs

    3,641在 GitHub 上查看↗

    This project is an LLM research orchestrator and autonomous AI agent framework designed to automate the scientific lifecycle. It functions as an end-to-end research pipeline and model training toolkit, managing everything from initial literature reviews and hypothesis testing to the final drafting of academic papers. The system is distinguished by its ability to convert unstructured academic PDFs into machine-executable knowledge layers, allowing agents to reproduce and extend research findings. It employs a two-loop orchestration architecture and a specialized research engineering skill libr

    Evaluates the ability of AI systems to autonomously design and analyze scientific experiments with rigor.

    TeXaiai-researchclaude
    在 GitHub 上查看↗3,641
  1. Home
  2. Artificial Intelligence & ML
  3. Large Language Models
  4. Model Benchmarks
  5. Automation Capability Benchmarks

探索子标签

  • Scientific Experimentation BenchmarksStandardized evaluations of an AI's ability to design and analyze scientific experiments. **Distinct from Automation Capability Benchmarks:** Specializes automation benchmarks to the scientific method and research rigor