For generative pre trained transformer lists, the strongest matches are dair-ai/prompt-engineering-guide (This repository is a curated educational guide providing comprehensive), microsoft/generative-ai-for-beginners (This repository provides a structured educational curriculum and learning) and rasbt/llm-architecture-gallery (This repository provides structured source data on various LLM). rohitg00/awesome-ai-apps and eleutherai/lm-evaluation-harness round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Hand-picked generative pre-trained transformer repositories, ranked by GitHub stars and activity. Compare the top LLM resources and pick the right one.
This project is a comprehensive educational resource and technical guide focused on the development, optimization, and application of large language models. It provides a structured curriculum for mastering prompt engineering, ranging from foundational principles of instruction design to advanced techniques for improving model reasoning, accuracy, and reliability. The guide distinguishes itself by offering deep technical insights into agentic workflows and autonomous system design. It covers the implementation of multi-step reasoning chains, tool integration through function calling, and stat
This repository is a curated educational guide providing comprehensive resources on prompt engineering, agentic workflows, and large language model applications, matching your search for curated LLM materials.
This project is a comprehensive, open-source educational curriculum designed to guide developers through the mastery of generative artificial intelligence. It provides a structured learning path that covers foundational concepts, prompt engineering, and the practical application of large language models. The repository serves as a central hub for skill acquisition, offering sequential modules that progress from basic model mechanics to advanced architectural patterns. The curriculum distinguishes itself by focusing on the end-to-end lifecycle of intelligent software, including the implementat
This repository provides a structured educational curriculum and learning path for generative AI, featuring prompt engineering and application guides, though it functions as a learning resource rather than a raw list of papers or models.
The llm-architecture-gallery is a structured catalog and reference platform for exploring machine learning model architectures. It maintains a curated directory of language model specifications, release dates, parameter counts, and technical fact sheets. The platform includes a dedicated comparison tool that aligns model attributes side by side to evaluate structural differences, design tradeoffs, and attention mechanisms. Users can browse the reference collection and analyze aggregated metadata across various architectures. Underlying data is maintained in structured YAML files and process
This repository provides structured source data on various LLM architectures, which aligns with the request for curated resources on model architectures, though it lacks some of the broader guides and tools in this category.
This project serves as a curated directory and resource hub for developers working with generative artificial intelligence. It provides a comprehensive index of open-source software solutions, frameworks, and project examples designed to help users discover and implement advanced AI systems. The repository focuses on practical implementations of agentic, multimodal, and retrieval-augmented generation architectures. It highlights tools for building conversational assistants, voice-enabled agents, and automated workflows that leverage large language models. By showcasing diverse technical domai
This project serves as a curated directory of generative AI applications and frameworks, making it a relevant resource hub for the requested topic, though it emphasizes application examples over academic papers and foundational models.
This project is a standardized framework for benchmarking large language models across a wide range of academic and reasoning datasets. It provides a platform for executing automated evaluation tasks to measure model accuracy and performance, ensuring consistent assessment through a structured configuration schema. The framework distinguishes itself by incorporating a dedicated utility for data decontamination, which identifies and removes overlapping training samples from evaluation sets to prevent data leakage. It also features a flexible task builder that allows users to define custom benc
This repository is a specialized benchmarking framework for language models rather than a curated resource list, making it a valuable tool but the wrong format for this search.
LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.
LLM Comparator is an interactive evaluation tool rather than a comprehensive resource directory, making it a specialized component rather than the broad collection the visitor is looking for.
The official evaluation suite and dynamic data release for MixEval.
MixEval is an evaluation suite and benchmark tool for language models rather than a comprehensive curated resource list of papers, models, and guides.
Evaluation and Tracking for LLM Experiments and AI Agents
TruLens is an evaluation and tracking framework for LLMs rather than a curated resource list, serving as a specialised component for assessing applications rather than a collection of references and papers.
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
This library provides tools for evaluating machine learning models and datasets, which covers the evaluation aspect, but it is a software library rather than a curated list of resources.
OpenCompass is a comprehensive evaluation platform, benchmarking suite, and distributed model evaluator designed to measure the performance and accuracy of large language models. It provides a framework for benchmarking both open-source and API-based models against diverse datasets using standardized metrics and reproducible pipelines. The project features an automated judging framework that uses language models as judges to score and verify the quality of generated text. It includes a performance leaderboard system for comparing the relative capabilities of various models across industry-sta
This repository is a dedicated model evaluation and benchmarking platform rather than a curated resource list of papers and models.
Lighteval is an open-source framework for running standardized benchmarks and custom evaluation tasks against language models. It provides a system for defining new evaluation tasks with custom prompts, metrics, and scoring in YAML configuration files, and integrates with the Hugging Face Hub for storing and comparing results. The framework supports evaluating models across multiple inference backends, including transformers, vllm, and custom APIs, through a unified generation and log-probability interface. It includes a pluggable metric registry for built-in and custom scoring, a prediction
Lighteval is a model evaluation and benchmarking framework rather than a curated resource list of papers and models, making it a specialized tool component rather than the requested directory.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| dair-ai/prompt-engineering-guide | 75.7K | MDX | MIT | |
| microsoft/generative-ai-for-beginners | 112K | Jupyter Notebook | MIT | |
| rasbt/llm-architecture-gallery | 1.3K | — | Apache-2.0 | |
| rohitg00/awesome-ai-apps | 723 | HTML | apache-2.0 | |
| eleutherai/lm-evaluation-harness | 11.5K | Python | mit | |
| pair-code/llm-comparator | 528 | JavaScript | Apache-2.0 | |
| psycoy/mixeval | 255 | Python | — | |
| truera/trulens | 3.4K | Python | MIT | |
| huggingface/evaluate | 2.5K | Python | Apache-2.0 | |
| internlm/opencompass | 7.1K | Python | Apache-2.0 |