awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to phoebussi/alpaca-cot

Projects sharing features with Alpaca CoT

20 open-source projects similar to phoebussi/alpaca-cot, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • lianjiatech/belleLianjiaTech avatar

    LianjiaTech/BELLE

    8,273View on GitHub↗

    BELLE is a specialized implementation of Chinese conversational large language models, encompassing a full instruction tuning framework. It provides a pipeline for training, evaluating, and deploying models optimized for natural language understanding and dialogue tasks in the Chinese language. The project is distinguished by its integrated approach to model refinement, combining the curation of multi-million entry instruction datasets with a distributed training pipeline. This pipeline supports both full fine-tuning and low-rank adaptation to optimize conversational performance. The system

    HTMLbloomchinese-nlpgpt-evaluation
    View on GitHub↗8,273
  • da-southampton/redgptDA-southampton avatar

    DA-southampton/RedGPT

    70View on GitHub↗

    English Version

    View on GitHub↗70
  • facebookresearch/codellamafacebookresearch avatar

    facebookresearch/codellama

    16,307View on GitHub↗

    Code Llama is a large language model based on Llama 2 trained specifically for programming tasks and software development. It provides specialized model types optimized for general code generation, instruction following, and context-aware infilling. The project includes an instruction-tuned programming model for executing technical tasks via natural language prompts and a code infilling model that predicts missing sections based on surrounding source context. A large context code model is also provided to analyze extensive blocks of source code for improved coherence. The system covers capab

    Python
    View on GitHub↗16,307
  • freedomintelligence/huatuo-26mFreedomIntelligence avatar

    FreedomIntelligence/Huatuo-26M

    335View on GitHub↗

    The Largest-scale Chinese Medical QA Dataset: with 26,000,000 question answer pairs.

    View on GitHub↗335

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • ibm/dromedaryIBM avatar

    IBM/Dromedary

    1,138View on GitHub↗

    Dromedary: towards helpful, ethical and reliable LLMs.

    Python
    View on GitHub↗1,138
  • instruction-tuning-with-gpt-4/gpt-4-llmInstruction-Tuning-with-GPT-4 avatar

    Instruction-Tuning-with-GPT-4/GPT-4-LLM

    4,335View on GitHub↗

    This project is an instruction tuning framework and synthetic data generator that uses high-capacity teacher models to produce instruction-following pairs for training smaller student models. It provides datasets and tools for supervised instruction tuning and reinforcement learning from human feedback. The framework specializes in cross-lingual tuning, offering high-quality instruction-following examples in English and Chinese to improve model generalization across different scripts. It includes a reward modeling tool for creating preference datasets and comparative ratings used to train rew

    HTMLalpacachatgptgpt-4
    View on GitHub↗4,335
  • mbzuai-nlp/lamini-lmmbzuai-nlp avatar

    mbzuai-nlp/LaMini-LM

    822View on GitHub↗

    LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions

    View on GitHub↗822
  • nlpxucan/wizardlmnlpxucan avatar

    nlpxucan/WizardLM

    9,486View on GitHub↗

    WizardLM is a large language model and instruction-tuning framework designed to execute sophisticated coding, mathematical, and conversational tasks. It functions as an AI system for mathematical reasoning and code generation, as well as a synthetic dataset generator used to train other language models. The project is distinguished by its evolutionary instruction tuning, which uses a method to rewrite simple instructions into complex tasks. This process expands training dataset difficulty and produces a high volume of open-domain tasks across various difficulty levels. The system covers capa

    Python
    View on GitHub↗9,486
  • peterwestai2/symbolic-knowledge-distillationpeterwestai2 avatar

    peterwestai2/symbolic-knowledge-distillation

    142View on GitHub↗

    This is the repository for the project Symbolic Knowledge Distillation: from General Language Models to Commonsense Models

    Python
    View on GitHub↗142
  • plexpt/chatgpt-corpusPlexPt avatar

    PlexPt/chatgpt-corpus

    964View on GitHub↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    View on GitHub↗964
  • qiuhuachuan/smileqiuhuachuan avatar

    qiuhuachuan/smile

    530View on GitHub↗

    EMNLP 2024 中文领域心理健康对话大模型MeChat

    Python
    View on GitHub↗530
  • sahil280114/codealpacasahil280114 avatar

    sahil280114/codealpaca

    1,512View on GitHub↗

    This is the repo for the Code Alpaca project, which aims to build and share an instruction-following LLaMA model for code generation. This repo is fully based on Stanford Alpaca ,and only changes the data used for training. Training approach is the same.

    Python
    View on GitHub↗1,512
  • servicenow/promptmix-emnlp-2023ServiceNow avatar

    ServiceNow/PromptMix-EMNLP-2023

    12View on GitHub↗

    This is the repository for the EMNLP 2023 paper PromptMix: A Class Boundary Augmentation Method for Large Language Model Distillation.

    Python
    View on GitHub↗12
  • toyhom/chinese-medical-dialogue-dataToyhom avatar

    Toyhom/Chinese-medical-dialogue-data

    1,720View on GitHub↗

    Chinese medical dialogue data 中文医疗对话数据集

    Python
    View on GitHub↗1,720
  • xuefuzhao/instructionwildXueFuzhao avatar

    XueFuzhao/InstructionWild

    462View on GitHub↗

    We release InstructWild v2 under data v2 dir, which includes over 110K high-quailty user-based instructions. We did not use self-instruct to generate any instructions. We also label a subset of these instructions with instruction type and speical tag. Please see README for details.

    View on GitHub↗462
  • ydli-ai/cslydli-ai avatar

    ydli-ai/CSL

    671View on GitHub↗

    COLING 2022 CSL: A Large-scale Chinese Scientific Literature Dataset 中文科学文献数据集

    Pythonchinese-nlpdatasetmachine-learning
    View on GitHub↗671
  • yhydhx/auggptyhydhx avatar

    yhydhx/AugGPT

    54View on GitHub↗

    \Accepted by TBD\ Code for: AugGPT: Leveraging ChatGPT for Text Data Augmentation

    Python
    View on GitHub↗54
  • yizhongw/self-instructyizhongw avatar

    yizhongw/self-instruct

    4,602View on GitHub↗

    Self-instruct is a framework for generating synthetic instruction datasets and fine-tuning large language models to improve their instruction-following capabilities. It provides a pipeline for aligning pretrained models with human intentions through a supervised fine-tuning workflow. The system utilizes a synthetic data generator that uses a seed set of tasks to prompt a model to create new instructional data. It includes an instruction dataset curator to remove redundant or low-quality entries, maintaining dataset diversity through a filtered task pool. The framework covers the full alignme

    Python
    View on GitHub↗4,602
  • cluebenchmark/pclueCLUEbenchmark avatar

    CLUEbenchmark/pCLUE

    505View on GitHub↗
    Jupyter Notebookchinesecluedatasets
    View on GitHub↗505
  • zexuehe/tdgZexueHe avatar

    ZexueHe/TDG

    0View on GitHub↗
    View on GitHub↗0