awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to bigcode-project/octopack

Open-source alternatives to Octopack

23 open-source projects similar to bigcode-project/octopack, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Octopack alternative.

  • bigscience-workshop/promptsourceالصورة الرمزية لـ bigscience-workshop

    bigscience-workshop/promptsource

    3,025عرض على GitHub↗

    Toolkit for creating, sharing and using natural language prompts.

    Python
    عرض على GitHub↗3,025
  • bigscience-workshop/xmtfالصورة الرمزية لـ bigscience-workshop

    bigscience-workshop/xmtf

    536عرض على GitHub↗

    This repository provides an overview of all components used for the creation of BLOOMZ & mT0 and xP3 introduced in the paper Crosslingual Generalization through Multitask Finetuning. Link to 25min video on the paper by Samuel Albanie; Link to 4min video on the paper by Niklas Muennighoff.

    Jupyter Notebook
    عرض على GitHub↗536
  • facebookresearch/metaseqالصورة الرمزية لـ facebookresearch

    facebookresearch/metaseq

    6,546عرض على GitHub↗

    Metaseq is a transformer sequence modeling toolkit designed for training, fine-tuning, and deploying sequence-to-sequence models using open pre-trained weights. It provides a comprehensive framework for large language model training, including dedicated tools for sequence dataset processing and a standalone inference server for generating text via API requests. The project features specialized utilities for model quantization to reduce parameter precision to eight bits, which lowers memory usage and increases inference speed. It also includes a checkpoint conversion pipeline to transform mode

    Python
    عرض على GitHub↗6,546
  • google-research/flanالصورة الرمزية لـ google-research

    google-research/FLAN

    1,566عرض على GitHub↗

    Original Flan (2021) | The Flan Collection (2022) | Flan 2021 Citation | License

    Python
    عرض على GitHub↗1,566

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • google-research/text-to-text-transfer-transformerالصورة الرمزية لـ google-research

    google-research/text-to-text-transfer-transformer

    6,528عرض على GitHub↗

    This is a machine learning framework for treating diverse natural language processing tasks as a unified text-to-text problem. It provides a toolkit for pre-training and fine-tuning large-scale transformer models, utilizing a system where both inputs and outputs are formatted as raw text sequences. The framework is distinguished by its distributed training system, which uses mesh-based strategies to scale model weights and training batches across multiple TPU cores. It supports multi-task learning by combining diverse datasets into a single training stream using configurable mixture rates, al

    Python
    عرض على GitHub↗6,528
  • hkust-nlp/deitaالصورة الرمزية لـ hkust-nlp

    hkust-nlp/deita

    597عرض على GitHub↗

    🤗 HF Repo 📄 Paper 📚 6K Data 📚 10K Data

    Python
    عرض على GitHub↗597
  • imoneoi/openchatالصورة الرمزية لـ imoneoi

    imoneoi/openchat

    5,481عرض على GitHub↗

    OpenChat is a framework for the training, fine-tuning, and deployment of large language models optimized for conversational and mathematical reasoning tasks. It provides a comprehensive lifecycle for these models, ranging from training pipelines and deployment stacks to a web-based chat interface. The project focuses on enabling high-performance model execution on consumer-grade hardware without the need for enterprise-grade accelerators. It includes a production-ready inference server that implements the OpenAI chat completion protocol and utilizes dynamic request batching to optimize hardwa

    Python
    عرض على GitHub↗5,481
  • ise-uiuc/magicoderI

    ise-uiuc/magicoder

    0عرض على GitHub↗

    🎩 Models | 📚 Dataset | 🚀 Quick Start | 👀 Demo | 📝 Citation | 🙏 Acknowledgements

    عرض على GitHub↗0
  • luohongyin/sailالصورة الرمزية لـ luohongyin

    luohongyin/SAIL

    161عرض على GitHub↗

    Towards Robust Grounded Language Modeling [DEMO](https://huggingface.co/spaces/luohy/SAIL-7B) | [WEB](https://openlsr.org/sail-7b)

    Python
    عرض على GitHub↗161
  • namisan/mt-dnnالصورة الرمزية لـ namisan

    namisan/mt-dnn

    2,259عرض على GitHub↗

    New Release We released Adversarial training for both LM pre-training/finetuning and f-divergence.

    Python
    عرض على GitHub↗2,259
  • nardien/kardالصورة الرمزية لـ Nardien

    Nardien/KARD

    44عرض على GitHub↗

    Official Code Repository for the paper Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-intensive Tasks (NeurIPS 2023).

    Python
    عرض على GitHub↗44
  • nlpxucan/wizardlmالصورة الرمزية لـ nlpxucan

    nlpxucan/WizardLM

    9,486عرض على GitHub↗

    WizardLM is a large language model and instruction-tuning framework designed to execute sophisticated coding, mathematical, and conversational tasks. It functions as an AI system for mathematical reasoning and code generation, as well as a synthetic dataset generator used to train other language models. The project is distinguished by its evolutionary instruction tuning, which uses a method to rewrite simple instructions into complex tasks. This process expands training dataset difficulty and produces a high volume of open-domain tasks across various difficulty levels. The system covers capa

    Python
    عرض على GitHub↗9,486
  • ofa-sys/expertllamaالصورة الرمزية لـ OFA-Sys

    OFA-Sys/ExpertLLaMA

    299عرض على GitHub↗

    This repo introduces ExpertLLaMA, a solution to produce high-quality, elaborate, expert-like responses by augmenting vanilla instructions with specialized Expert Identity description. This repo contains: - Brief introduction on the method. - 52k Instruction-Following Expert Data generated by…

    Python
    عرض على GitHub↗299
  • orhonovich/unnatural-instructionsالصورة الرمزية لـ orhonovich

    orhonovich/unnatural-instructions

    181عرض على GitHub↗

    This repository contains the Unnatural Instructions dataset. Unnatural Instructions is a dataset of instructions automatically generated by a Large Language model. See full details in the paper: "Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor"

    عرض على GitHub↗181
  • renzelou/muffinالصورة الرمزية لـ RenzeLou

    RenzeLou/Muffin

    16عرض على GitHub↗

    This repository contains the source code for reproducing the data curation of MUFFIN (Multi-faceted Instructions).

    Python
    عرض على GitHub↗16
  • sahil280114/codealpacaالصورة الرمزية لـ sahil280114

    sahil280114/codealpaca

    1,512عرض على GitHub↗

    This is the repo for the Code Alpaca project, which aims to build and share an instruction-following LLaMA model for code generation. This repo is fully based on Stanford Alpaca ,and only changes the data used for training. Training approach is the same.

    Python
    عرض على GitHub↗1,512
  • thunlp/ultrachatالصورة الرمزية لـ thunlp

    thunlp/UltraChat

    2,786عرض على GitHub↗

    UltraChat is a collection of large-scale conversational datasets and instruction-tuning data designed for training and evaluating generative AI models. It provides structured JSON data consisting of complex, multi-round dialogue sequences intended to refine the performance of large language models in chat tasks. The project focuses on improving reasoning and response quality through a diverse set of interactions across multiple sectors. These datasets are used for supervised fine-tuning and instruction tuning workflows to improve how models follow complex directions and maintain context acros

    Pythonchatbotchatgptdeep-learning
    عرض على GitHub↗2,786
  • tianyi-lab/debatuneالصورة الرمزية لـ tianyi-lab

    tianyi-lab/DEBATunE

    25عرض على GitHub↗

    Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements (ACL'24) Chinese Version: [知乎](https://zhuanlan.zhihu.com/p/720237237)

    Python
    عرض على GitHub↗25
  • universal-ner/universal-nerالصورة الرمزية لـ universal-ner

    universal-ner/universal-ner

    375عرض على GitHub↗

    UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition Wenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen, Hoifung Poon (*Equal Contribution)

    Python
    عرض على GitHub↗375
  • urchade/glinerالصورة الرمزية لـ urchade

    urchade/GLiNER

    3,333عرض على GitHub↗

    Generalist and Lightweight Model for Named Entity Recognition (Extract any entity types from texts)

    Python
    عرض على GitHub↗3,333
  • yizhongw/self-instructالصورة الرمزية لـ yizhongw

    yizhongw/self-instruct

    4,602عرض على GitHub↗

    Self-instruct is a framework for generating synthetic instruction datasets and fine-tuning large language models to improve their instruction-following capabilities. It provides a pipeline for aligning pretrained models with human intentions through a supervised fine-tuning workflow. The system utilizes a synthetic data generator that uses a seed set of tasks to prompt a model to create new instructional data. It includes an instruction dataset curator to remove redundant or low-quality entries, maintaining dataset diversity through a filtered task pool. The framework covers the full alignme

    Python
    عرض على GitHub↗4,602
  • yyding1/gnerالصورة الرمزية لـ yyDing1

    yyDing1/GNER

    60عرض على GitHub↗

    Rethinking Negative Instances for Generative Named Entity Recognition

    Python
    عرض على GitHub↗60
  • zjukg/knowpatالصورة الرمزية لـ zjukg

    zjukg/KnowPAT

    193عرض على GitHub↗

    Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering

    Python
    عرض على GitHub↗193