1 रिपॉजिटरी
Creating new evaluation instances from user-provided repositories for benchmark expansion.
Distinct from Benchmark Data Generators: Distinct from Benchmark Data Generators: generates software engineering task instances from real repositories, not synthetic data.
Explore 1 awesome GitHub repository matching testing & quality assurance · Software Issue Generators. Refine with filters or upvote what's useful.
SWE-bench is an automated evaluation framework that tests large language models on real-world software engineering tasks. It measures how effectively models can generate and apply code patches that resolve actual GitHub issues, using a standardized dataset and scoring system built around Docker-based patch verification against original project test suites. The framework provides curated benchmark datasets spanning comprehensive, fast, verified, multilingual, and multimodal evaluation splits, allowing targeted assessment of model capabilities across different programming languages and issue ty
Runs a data collection procedure on user-provided repositories to generate new evaluation instances.