📃 [DataSciBench] [GitHub] [Evaluation Data] [Website]
Die Hauptfunktionen von thudm/datascibench sind: Coding Benchmarks.
Open-Source-Alternativen zu thudm/datascibench sind unter anderem: swe-bench/swe-bench — SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering…
SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering tasks. It measures how effectively models can generate and apply code patches that resolve actual GitHub issues, using a standardized dataset and scoring system built around Docker-based patch verification against original project test suites. The framework provides curated benchmark datasets spanning comprehensive, fast, verified, multilingual, and multimodal evaluation splits, allowing targeted assessment of model capabilities across different programming languages and issue ty