awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to mikegu721/xiezhibenchmark

Open-source alternatives to XiezhiBenchmark

30 open-source projects similar to mikegu721/xiezhibenchmark, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best XiezhiBenchmark alternative.

  • sjtu-lit/cevalالصورة الرمزية لـ SJTU-LIT

    SJTU-LIT/ceval

    1,854عرض على GitHub↗

    Official github repo for C-Eval, a Chinese evaluation suite for foundation models NeurIPS 2023

    Python
    عرض على GitHub↗1,854
  • open-compass/opencompassالصورة الرمزية لـ open-compass

    open-compass/opencompass

    6,678عرض على GitHub↗

    OpenCompass is an open-source framework for standardized benchmarking of large language models. It provides a configurable evaluation pipeline that supports both objective and subjective assessment, using a dual-engine architecture to handle closed-form answer comparison and open-ended response rating. The framework is designed as a modular platform where datasets, models, and metrics are composed through declarative YAML configuration files. The framework distinguishes itself through its extensible model integration layer, which supports custom models, HuggingFace models, and third-party API

    Pythonbenchmarkchatgptevaluation
    عرض على GitHub↗6,678
  • haonan-li/cmmluالصورة الرمزية لـ haonan-li

    haonan-li/CMMLU

    821عرض على GitHub↗

    CMMLU: Measuring massive multitask language understanding in Chinese

    Python
    عرض على GitHub↗821
  • ruixiangcui/agievalالصورة الرمزية لـ ruixiangcui

    ruixiangcui/AGIEval

    775عرض على GitHub↗

    This repository contains information about AGIEval, data, code and output of baseline systems for the benchmark.

    Python
    عرض على GitHub↗775
  • flagopen/flagevalالصورة الرمزية لـ FlagOpen

    FlagOpen/FlagEval

    13عرض على GitHub↗

    FlagEval, launched by BAAI in 2023, is a comprehensive large model evaluation system that encompasses over 800 open-source and closed-source models from around the globe. It features more than 40 capability dimensions, including reasoning, mathematical skills, and task-solving abilities, along…

    عرض على GitHub↗13

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • thu-coai/safety-promptsالصورة الرمزية لـ thu-coai

    thu-coai/Safety-Prompts

    1,176عرض على GitHub↗

    Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。

    attack-defensechatgptchinese-language
    عرض على GitHub↗1,176
  • michael-wzhu/promptcblueالصورة الرمزية لـ michael-wzhu

    michael-wzhu/PromptCBLUE

    394عرض على GitHub↗

    PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese

    Python
    عرض على GitHub↗394
  • cluebenchmark/supercluelybالصورة الرمزية لـ CLUEbenchmark

    CLUEbenchmark/SuperCLUElyb

    144عرض على GitHub↗

    SuperCLUE琅琊榜:中文通用大模型匿名对战评价基准

    عرض على GitHub↗144
  • sjwhitworth/golearnالصورة الرمزية لـ sjwhitworth

    sjwhitworth/golearn

    9,438عرض على GitHub↗

    GoLearn is a machine learning library for the Go programming language. It provides a supervised learning framework and a toolkit for building, training, and evaluating predictive models through a standardized interface. The project implements a data frame system that loads CSV files into structured grids for matrix operations. It includes a preprocessing library for discretizing continuous variables and a model evaluation toolkit that utilizes confusion matrices and cross-validation to measure precision and recall. The library covers data engineering and management, including the ability to

    Go
    عرض على GitHub↗9,438
  • nvidia/isaac-gr00tالصورة الرمزية لـ NVIDIA

    NVIDIA/Isaac-GR00T

    6,222عرض على GitHub↗
    Jupyter Notebook
    عرض على GitHub↗6,222
  • ageron/handson-mlالصورة الرمزية لـ ageron

    ageron/handson-ml

    25,608عرض على GitHub↗

    This is a machine learning educational repository consisting of a collection of notebooks and code examples. It provides practical implementations of diverse machine learning algorithms and workflows, ranging from traditional scientific computing to deep learning. The project features specific implementations of Scikit-Learn models, such as decision trees, random forests, and support vector machines, as well as TensorFlow examples for building neural networks, convolutional layers, and recurrent architectures. It also includes tutorials on reinforcement learning development and the creation o

    Jupyter Notebook
    عرض على GitHub↗25,608
  • eleutherai/lm-evaluation-harnessالصورة الرمزية لـ EleutherAI

    EleutherAI/lm-evaluation-harness

    11,460عرض على GitHub↗

    This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc

    Pythonevaluation-frameworklanguage-modeltransformer
    عرض على GitHub↗11,460
  • eth-sri/matharenaE

    eth-sri/matharena

    0عرض على GitHub↗
    عرض على GitHub↗0
  • felixgithub2017/cg-evalالصورة الرمزية لـ Felixgithub2017

    Felixgithub2017/CG-Eval

    13عرض على GitHub↗

    Chinese Generation Evaluation

    عرض على GitHub↗13
  • felixgithub2017/mmcuالصورة الرمزية لـ Felixgithub2017

    Felixgithub2017/MMCU

    90عرض على GitHub↗

    MEASURING MASSIVE MULTITASK CHINESE UNDERSTANDING

    Python
    عرض على GitHub↗90
  • hazyresearch/legalbenchالصورة الرمزية لـ HazyResearch

    HazyResearch/legalbench

    597عرض على GitHub↗

    An open science effort to benchmark legal reasoning in foundation models

    Python
    عرض على GitHub↗597
  • hc-guo/owlالصورة الرمزية لـ HC-Guo

    HC-Guo/Owl

    237عرض على GitHub↗

    A Large Language Model for IT Operations

    Python
    عرض على GitHub↗237
  • huggingface/evaluateالصورة الرمزية لـ huggingface

    huggingface/evaluate

    2,455عرض على GitHub↗

    🤗 Evaluate: A library for easily evaluating machine learning models and datasets.

    Python
    عرض على GitHub↗2,455
  • huggingface/evaluation-guidebookالصورة الرمزية لـ huggingface

    huggingface/evaluation-guidebook

    2,125عرض على GitHub↗

    Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

    Jupyter Notebookevaluationevaluation-metricsguidebook
    عرض على GitHub↗2,125
  • huggingface/lightevalالصورة الرمزية لـ huggingface

    huggingface/lighteval

    2,453عرض على GitHub↗

    Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction

    Pythonevaluationevaluation-frameworkevaluation-metrics
    عرض على GitHub↗2,453
  • huggingface/yourbenchH

    huggingface/yourbench

    0عرض على GitHub↗
    عرض على GitHub↗0
  • ibm/aif360الصورة الرمزية لـ IBM

    IBM/AIF360

    2,827عرض على GitHub↗

    A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.

    Python
    عرض على GitHub↗2,827
  • internlm/opencompassالصورة الرمزية لـ InternLM

    InternLM/opencompass

    7,096عرض على GitHub↗

    OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta

    Python
    عرض على GitHub↗7,096
  • joelniklaus/lextremeالصورة الرمزية لـ JoelNiklaus

    JoelNiklaus/LEXTREME

    25عرض على GitHub↗

    This repository provides scripts for evaluating NLP models on the LEXTREME benchmark, a set of diverse multilingual tasks in legal NLP

    Python
    عرض على GitHub↗25
  • langchain-ai/auto-evaluatorالصورة الرمزية لـ langchain-ai

    langchain-ai/auto-evaluator

    780عرض على GitHub↗

    Context

    TypeScript
    عرض على GitHub↗780
  • mlfoundations/evalchemyالصورة الرمزية لـ mlfoundations

    mlfoundations/evalchemy

    597عرض على GitHub↗

    Automatic evals for LLMs

    HTML
    عرض على GitHub↗597
  • modelscope/evalscopeالصورة الرمزية لـ modelscope

    modelscope/evalscope

    2,955عرض على GitHub↗

    A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

    Python
    عرض على GitHub↗2,955
  • modelscope/openjudgeM

    modelscope/OpenJudge

    0عرض على GitHub↗
    عرض على GitHub↗0
  • noudald/pyrocN

    noudald/pyroc

    0عرض على GitHub↗
    عرض على GitHub↗0
  • open-compass/lawbenchالصورة الرمزية لـ open-compass

    open-compass/LawBench

    434عرض على GitHub↗

    Benchmarking Legal Knowledge of Large Language Models

    Python
    عرض على GitHub↗434