awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
hendrycks avatar

hendrycks/apps

0
View on GitHub↗
532 stele·70 fork-uri·Python·MIT·1 vizualizare

Apps

This is the repository for Measuring Coding Challenge Competence With APPS by Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt.

Features

  • Benchmarks and Datasets - Benchmark for measuring coding challenge competence.

Istoric stele

Graficul istoricului de stele pentru hendrycks/appsGraficul istoricului de stele pentru hendrycks/apps

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru Apps

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu Apps.
  • rllm-org/rllmAvatar rllm-org

    rllm-org/rllm

    5,641Vezi pe GitHub↗

    rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline that runs the same agent code for both evaluation and training, automatically capturing traces for gradient computation. The framework supports distributed reinforcement learning across multiple GPUs and nodes using pluggable backends, and executes agents in isolated sandboxes—either locally or in the cloud—for safe and scalable rollout collection. It trains agents built with LangGraph, SmolAgents, OpenAI Agents SDK, or custom frameworks without requiring core logic changes. T

    Pythonagent-frameworkagentic-workflowcoding-agent
    Vezi pe GitHub↗5,641
  • evalplus/evalplusAvatar evalplus

    evalplus/evalplus

    1,765Vezi pe GitHub↗

    Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

    Python
    Vezi pe GitHub↗1,765
  • gersteinlab/biocoderAvatar gersteinlab

    gersteinlab/BioCoder

    58Vezi pe GitHub↗

    BioCoder is a challenging bioinformatics code generation benchmark for examining the capabilities of state-of-the-art large language models (LLMs).

    Jupyter Notebook
    Vezi pe GitHub↗58
  • bytedance/fullstackbenchAvatar bytedance

    bytedance/FullStackBench

    121Vezi pe GitHub↗

    FullStack Bench: Evaluating LLMs as Full Stack Coders

    Python
    Vezi pe GitHub↗121
Vezi toate cele 15 alternative pentru Apps→

Întrebări frecvente

Ce face hendrycks/apps?

This is the repository for Measuring Coding Challenge Competence With APPS by Dan Hendrycks\, Steven Basart\, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt.

Care sunt principalele funcționalități ale hendrycks/apps?

Principalele funcționalități ale hendrycks/apps sunt: Benchmarks and Datasets.

Care sunt câteva alternative open-source pentru hendrycks/apps?

Alternativele open-source pentru hendrycks/apps includ: rllm-org/rllm — rllm is an asynchronous reinforcement learning framework for training language agents. It provides a unified pipeline… evalplus/evalplus — Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024. gersteinlab/biocoder — BioCoder is a challenging bioinformatics code generation benchmark for examining the capabilities of state-of-the-art… google-research/google-research — This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum… leolty/repobench — ✨ RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems - ICLR 2024. bytedance/fullstackbench — FullStack Bench: Evaluating LLMs as Full Stack Coders.