SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering tasks. It measures how effectively models can generate and apply code patches that resolve actual GitHub issues, using a standardized dataset and scoring system built around Docker-based patch verification against original project test suites. The framework provides curated benchmark datasets spanning comprehensive, fast, verified, multilingual, and multimodal evaluation splits, allowing targeted assessment of model capabilities across different programming languages and issue ty
📃 [DataSciBench] [GitHub] [Evaluation Data] [Website]
Las características principales de thudm/datascibench son: Coding Benchmarks.
Las alternativas de código abierto para thudm/datascibench incluyen: swe-bench/swe-bench — SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering…