awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépÎts open-source grùce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre mĂ©thodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
THUDM avatar

THUDM/DataSciBench

0
View on GitHub↗
62 stars·7 forks·Python·3 vues

DataSciBench

📃 [DataSciBench] [GitHub] [Evaluation Data] [Website]

Features

  • Coding Benchmarks - Benchmark for data science-focused LLM agents.

Historique des stars

Graphique de l'historique des stars pour thudm/datascibenchGraphique de l'historique des stars pour thudm/datascibench

Recherche par IA

Explorez plus de dépÎts awesome

DĂ©crivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sĂ©lectionnĂ©s par pertinence.

Start searching with AI

Questions fréquentes

Que fait thudm/datascibench ?

📃 [DataSciBench] [GitHub] [Evaluation Data] [Website]

Quelles sont les fonctionnalités principales de thudm/datascibench ?

Les fonctionnalités principales de thudm/datascibench sont : Coding Benchmarks.

Quelles sont les alternatives open-source Ă  thudm/datascibench ?

Les alternatives open-source à thudm/datascibench incluent : swe-bench/swe-bench — SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering


Alternatives open source Ă  DataSciBench

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec DataSciBench.
  • swe-bench/swe-benchAvatar de SWE-bench

    SWE-bench/SWE-bench

    4,321Voir sur GitHub↗

    SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering tasks. It measures how effectively models can generate and apply code patches that resolve actual GitHub issues, using a standardized dataset and scoring system built around Docker-based patch verification against original project test suites. The framework provides curated benchmark datasets spanning comprehensive, fast, verified, multilingual, and multimodal evaluation splits, allowing targeted assessment of model capabilities across different programming languages and issue ty

    Pythonbenchmarklanguage-modelsoftware-engineering
    Voir sur GitHub↗4,321