awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

1 repository

Awesome GitHub RepositoriesLanguage Model Math Evaluations

Measures language model accuracy on mathematical reasoning tasks by prompting step-by-step solving and comparing answers against ground truth.

Distinct from Mathematical Problem Solving Toolkits: Distinct from general Mathematical Problem Solving Toolkits: focuses on evaluating LLM math ability rather than providing interactive solvers.

Explore 1 awesome GitHub repository matching scientific & mathematical computing · Language Model Math Evaluations. Refine with filters or upvote what's useful.

Awesome Language Model Math Evaluations GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • openai/simple-evalsAvatar openai

    openai/simple-evals

    4,354Vezi pe GitHub↗

    This project is a language model evaluation framework and benchmarking tool designed to measure the accuracy and performance of models across diverse datasets. It provides a system for implementing model-based graders, running standardized tests for mathematical reasoning, coding, and factuality, and calculating quantified performance metrics such as precision, recall, F1 scores, and pass-at-k. The framework utilizes model-based grading and rubrics to validate response quality against expert-defined criteria. It includes a multi-model benchmarking loop and a model-agnostic API interface to co

    Measures language model accuracy on mathematical reasoning tasks by prompting step-by-step solving and comparing answers against ground truth.

    Python
    Vezi pe GitHub↗4,354
  1. Home
  2. Scientific & Mathematical Computing
  3. Mathematical Problem Solving Toolkits
  4. Language Model Math Evaluations