awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bigcode-project avatar

bigcode-project/octopack

0
View on GitHub↗
478 stars·28 forks·Jupyter Notebook·MIT·5 viewsarxiv.org/abs/2308.07124↗

Octopack

This repository provides an overview of all components from the paper OctoPack: Instruction Tuning Code Large Language Models. Link to 5-min video on the paper presented by Niklas Muennighoff.

Features

  • Instruction Tuning - Techniques for instruction tuning code models on diverse datasets.

Star history

Star history chart for bigcode-project/octopackStar history chart for bigcode-project/octopack

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Octopack

Similar open-source projects, ranked by how many features they share with Octopack.
  • bigscience-workshop/xmtfbigscience-workshop avatar

    bigscience-workshop/xmtf

    536View on GitHub↗

    This repository provides an overview of all components used for the creation of BLOOMZ & mT0 and xP3 introduced in the paper Crosslingual Generalization through Multitask Finetuning. Link to 25min video on the paper by Samuel Albanie; Link to 4min video on the paper by Niklas Muennighoff.

    Jupyter Notebook
    View on GitHub↗536
  • facebookresearch/metaseqfacebookresearch avatar

    facebookresearch/metaseq

    6,546View on GitHub↗

    Metaseq is a transformer sequence modeling toolkit designed for training, fine-tuning, and deploying sequence-to-sequence models using open pre-trained weights. It provides a comprehensive framework for large language model training, including dedicated tools for sequence dataset processing and a standalone inference server for generating text via API requests. The project features specialized utilities for model quantization to reduce parameter precision to eight bits, which lowers memory usage and increases inference speed. It also includes a checkpoint conversion pipeline to transform mode

    Python
    View on GitHub↗6,546
  • google-research/flangoogle-research avatar

    google-research/FLAN

    1,566View on GitHub↗

    Original Flan (2021) | The Flan Collection (2022) | Flan 2021 Citation | License

    Python
    View on GitHub↗1,566
  • bigscience-workshop/promptsourcebigscience-workshop avatar

    bigscience-workshop/promptsource

    3,025View on GitHub↗

    Toolkit for creating, sharing and using natural language prompts.

    Python
    View on GitHub↗3,025
See all 23 alternatives to Octopack→

Frequently asked questions

What does bigcode-project/octopack do?

This repository provides an overview of all components from the paper OctoPack: Instruction Tuning Code Large Language Models. Link to 5-min video on the paper presented by Niklas Muennighoff.

What are the main features of bigcode-project/octopack?

The main features of bigcode-project/octopack are: Instruction Tuning.

What are some open-source alternatives to bigcode-project/octopack?

Open-source alternatives to bigcode-project/octopack include: bigscience-workshop/xmtf — This repository provides an overview of all components used for the creation of BLOOMZ & mT0 and xP3 introduced in the… facebookresearch/metaseq — Metaseq is a transformer sequence modeling toolkit designed for training, fine-tuning, and deploying… google-research/flan — Original Flan (2021) | The Flan Collection (2022) | Flan 2021 Citation | License. google-research/text-to-text-transfer-transformer — This is a machine learning framework for treating diverse natural language processing tasks as a unified text-to-text… hkust-nlp/deita — 🤗 HF Repo     📄 Paper     📚 6K Data     📚 10K Data. bigscience-workshop/promptsource — Toolkit for creating, sharing and using natural language prompts.