awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

Open-Source Coding Models

Ranking updated Jul 13, 2026

For an open source model for code generation, the first results are meta-llama/codellama (CodeLlama is a specialized family of models built specifically for code generation, offering instruction-tuned variants, multi-language support, and large context windows that directly address the requirements for software development tasks), facebookresearch/codellama (Code Llama is a foundational model specifically trained for software development tasks, offering instruction-tuned variants, support for large context windows, and specialized capabilities for code infilling and generation) and deepseek-ai/deepseek-coder-v2 (This is a state-of-the-art mixture-of-experts model specifically trained for code generation and software development, featuring a 128K context window, instruction tuning, and support for fill-in-the-middle objectives). qwenlm/qwen2.5 and baichuan-inc/baichuan2 round out the shortlist. Compare the match explanations and check the project documentation against your requirements.

We curate open-source GitHub repositories matching “best open source coding models”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.

Open-Source Coding Models

Find the best repos with AI.We'll search the best matching repositories with AI.
  • meta-llama/codellamameta-llama avatar

    meta-llama/codellama

    16,307View on GitHub↗

    CodeLlama is a family of large language models derived from the Llama 2 architecture and specialized for producing, completing, and refactoring source code across multiple programming languages. It functions as a code generation model capable of synthesizing source code from natural language descriptions. The project includes specific model variants designed for different programming tasks. This includes instruction-tuned models trained to follow complex natural language directions and code infilling models that predict and insert missing code segments into existing files by analyzing surroun

    CodeLlama is a specialized family of models built specifically for code generation, offering instruction-tuned variants, multi-language support, and large context windows that directly address the requirements for software development tasks.

    PythonInstruction Fine-tuningInstruction-Tuned Language ModelsInstruction-Following Models
    View on GitHub↗16,307
  • facebookresearch/codellamafacebookresearch avatar

    facebookresearch/codellama

    16,307View on GitHub↗

    Code Llama is a large language model based on Llama 2 trained specifically for programming tasks and software development. It provides specialized model types optimized for general code generation, instruction following, and context-aware infilling. The project includes an instruction-tuned programming model for executing technical tasks via natural language prompts and a code infilling model that predicts missing sections based on surrounding source context. A large context code model is also provided to analyze extensive blocks of source code for improved coherence. The system covers capab

    Code Llama is a foundational model specifically trained for software development tasks, offering instruction-tuned variants, support for large context windows, and specialized capabilities for code infilling and generation.

    PythonInstruction Fine-tuningInstruction-Tuned Language ModelsInstruction-Following Models
    View on GitHub↗16,307
  • deepseek-ai/deepseek-coder-v2deepseek-ai avatar

    deepseek-ai/DeepSeek-Coder-V2

    6,462View on GitHub↗

    This is a state-of-the-art mixture-of-experts model specifically trained for code generation and software development, featuring a 128K context window, instruction tuning, and support for fill-in-the-middle objectives.

    128K-Token Context WindowsLong Context Processing
    View on GitHub↗6,462
  • qwenlm/qwen2.5QwenLM avatar

    QwenLM/Qwen2.5

    27,307View on GitHub↗

    Qwen2.5 is a suite of large language model foundation models designed for natural language generation, code production, and complex mathematical reasoning. The project encompasses a multilingual language model capable of processing dozens of languages and a specialized code generation model for technical problem solving and debugging. The framework is distinguished by its long context capabilities, enabling the analysis of massive inputs ranging from 256K up to 1 million tokens. It further functions as an agentic framework, utilizing standardized templates and parsers to execute autonomous wo

    Qwen2.5 is a comprehensive suite of foundation models specifically optimized for code generation and technical reasoning, featuring extensive multi-language support, instruction tuning, and industry-leading context window sizes.

    PythonLong Context ProcessingLong-Context ModelsInstruction-Following Models
    View on GitHub↗27,307
  • baichuan-inc/baichuan2baichuan-inc avatar

    baichuan-inc/Baichuan2

    4,098View on GitHub↗

    Baichuan2 is a collection of pre-trained large language models, including base and chat variants, designed for natural language generation and multi-turn conversational AI. It provides an inference engine and a fine-tuning framework to adapt these models to custom datasets and specialized domains. The project features a quantization toolkit and an inference engine that enable model execution across diverse hardware, including graphics processors, central processors, and specialized accelerators. These tools support low-bit weight quantization to reduce memory usage and increase inference spee

    This is a general-purpose foundation model that can be fine-tuned for code generation, though it lacks the specialized pre-training for software development found in dedicated coding models.

    PythonModel QuantizationModel QuantizationSupervised Instruction Fine-Tuning
    View on GitHub↗4,098
  • facico/chinese-vicunaFacico avatar

    Facico/Chinese-Vicuna

    4,121View on GitHub↗

    Chinese-Vicuna is a Chinese large language model and instruction-following AI based on the LLaMA architecture. It is specifically designed for natural language understanding and generation in the Chinese language, utilizing an instruction-tuned model to follow complex user prompts across conversations. The project provides a LoRA fine-tuning framework and quantization systems to enable model adaptation and inference on consumer hardware. It implements quantized inference to reduce memory usage on both CPUs and GPUs, supported by a low-level C++ implementation to minimize system resource requi

    This is a general-purpose Chinese-language instruction-tuned model that includes support for code-related tasks and low-level inference optimization, though it is not exclusively specialized for software development.

    CInstruction Fine-tuningInstruction TuningInstruction-Tuned Language Models
    View on GitHub↗4,121
  • bigcode-project/starcoderbigcode-project avatar

    bigcode-project/starcoder

    7,508View on GitHub↗

    Starcoder is a large language model and associated framework designed to generate, complete, and evaluate source code across multiple programming languages. It functions as a source code model that can produce complete function implementations and predict subsequent characters in a line of code based on provided prompts. The project provides a specialized toolkit for adapting base models to specific coding tasks and instruction-following behaviors. This includes a conversational code assistant framework for training models to generate code via natural language chat, as well as a parameter-eff

    StarCoder is a specialized large language model explicitly designed for code generation and completion across many programming languages, offering the instruction tuning and framework support needed for software development tasks.

    PythonInstruction-Tuned Language Models
    View on GitHub↗7,508
  • salesforce/codegensalesforce avatar

    salesforce/CodeGen

    5,175View on GitHub↗

    CodeGen is a trained large language model and program synthesis model designed to generate functional source code. It utilizes a neural network architecture to synthesize executable code from natural language descriptions or partial code snippets. The model enables automated program synthesis and AI-assisted coding by predicting and filling in missing sections of code within a program. It transforms natural language descriptions into functional programming logic to automate the creation of boilerplate and logic.

    CodeGen is a specialized large language model designed specifically for program synthesis and code generation, providing the core capability of transforming natural language into functional code as requested.

    PythonNatural Language to Code GeneratorsAI Coding AssistantsCausal Language Modeling
    View on GitHub↗5,175
  • salesforce/codet5salesforce avatar

    salesforce/CodeT5

    3,098View on GitHub↗

    Home of CodeT5: Open Code LLMs for Code Understanding and Generation

    CodeT5 is a specialized family of open-source models explicitly designed for code understanding and generation tasks, offering the instruction tuning and multi-language support required for software development workflows.

    PythonAI Coding AssistantsCode Generation ModelsLarge Language Models
    View on GitHub↗3,098
  • bigcode-project/starcoder2bigcode-project avatar

    bigcode-project/starcoder2

    2,075View on GitHub↗

    StarCoder2 is a family of code generation models (3B, 7B, and 15B), trained on 600+ programming languages from The Stack v2 and some natural language text such as Wikipedia, Arxiv, and GitHub issues. The models use Grouped Query Attention, a context window of 16,384 tokens, with sliding window…

    StarCoder2 is a family of open-source models specifically trained on a massive corpus of code and technical documentation, offering the instruction tuning, multi-language support, and large context window required for high-performance code generation.

    PythonLarge Language ModelsLarge Language Models (LLMs)Pre-training Research
    View on GitHub↗2,075
  • ibm-granite/granite-code-modelsibm-granite avatar

    ibm-granite/granite-code-models

    1,250View on GitHub↗

    Granite Code Models is a family of transformer-based foundational models designed for software engineering and logical reasoning tasks. These models are trained on high-quality programming datasets to interpret natural language prompts and generate functional source code, explain complex logic, repair code defects, and produce technical documentation. The project distinguishes itself through specialized training methodologies that align model behavior with complex programming instructions and mathematical problem-solving. By utilizing chain-of-thought reasoning and instruction-tuned parameter

    This repository provides a family of foundation models specifically trained for code intelligence tasks, offering instruction-tuned variants that support multiple programming languages and are designed for integration into development workflows.

    Instruction TuningInstruction-Tuned Language Models
    View on GitHub↗1,250
  • zai-org/glm-4zai-org avatar

    zai-org/GLM-4

    7,058View on GitHub↗

    GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod

    GLM-4 is a powerful, instruction-tuned large language model that explicitly supports code generation and long-context processing, making it a strong candidate for software development tasks despite being a general-purpose model rather than one exclusively dedicated to coding.

    PythonLong Context ProcessingLong-Context Models
    View on GitHub↗7,058
  • salesforce/codegen2salesforce avatar

    salesforce/CodeGen2

    270View on GitHub↗

    Official research release for the CodeGen2 models (1B, 3B, 7B, 16B) for Program Synthesis as presented in ICLR 2023:

    This repository provides a series of pre-trained models specifically designed for program synthesis and code generation, serving as a foundational tool for developers looking to implement or fine-tune code-focused LLMs.

    PythonCode Generation Models
    View on GitHub↗270
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
meta-llama/codellama16.3KPythonNOASSERTIONAug 12, 2024
facebookresearch/codellama16.3KPythonNOASSERTIONAug 12, 2024
deepseek-ai/deepseek-coder-v2
6.5K
—
mit
Nov 11, 2025
qwenlm/qwen2.527.3KPython—Jan 9, 2026
baichuan-inc/baichuan24.1KPythonApache-2.0Nov 8, 2024
facico/chinese-vicuna4.1KCApache-2.0Apr 18, 2025
bigcode-project/starcoder7.5KPythonApache-2.0Feb 27, 2024
salesforce/codegen5.2KPythonApache-2.0Jun 2, 2026
salesforce/codet53.1KPythonBSD-3-ClauseJun 2, 2026
bigcode-project/starcoder22.1KPythonApache-2.0Mar 21, 2024

Related searches

  • an open-source Copilot alternative
  • an open source model for image generation
  • an open source model for local deployment
  • an open source cli for ai coding
  • an open source model for local inference
  • an open source model for video generation
  • an open source tool for automated code review
  • an open source platform for local LLMs