2 个仓库
Standardized benchmarks specifically designed to measure the automation efficiency of AI models.
Distinct from Model Benchmarks: Focuses on the ability to automate complex tasks rather than static model performance or pricing
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Automation Capability Benchmarks. Refine with filters or upvote what's useful.
JARVIS is a system for large language model task orchestration, deployment management, and automation benchmarking. It utilizes a task orchestrator to decompose complex requests into actionable steps and coordinates various expert models to synthesize final responses. The project includes an AI model deployment manager to handle the local deployment of expert models across different hardware scales. It further provides an AI workflow API consisting of web endpoints used to trigger automated task workflows and retrieve results from model selection stages. The framework incorporates an automat
Evaluates the capability of large language models to automate complex tasks using standardized benchmarking datasets.
This project is an LLM research orchestrator and autonomous AI agent framework designed to automate the scientific lifecycle. It functions as an end-to-end research pipeline and model training toolkit, managing everything from initial literature reviews and hypothesis testing to the final drafting of academic papers. The system is distinguished by its ability to convert unstructured academic PDFs into machine-executable knowledge layers, allowing agents to reproduce and extend research findings. It employs a two-loop orchestration architecture and a specialized research engineering skill libr
Evaluates the ability of AI systems to autonomously design and analyze scientific experiments with rigor.