awesome-repositories.com
Blog
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Plataforma de modelos de lenguaje (LLM)

Clasificación actualizada el 30 jun 2026

For una plataforma de código abierto para LLMs locales, the strongest matches are 01-ai/yi (Yi is a publicly available bilingual language model with), google/gemma_pytorch (This repository provides the official PyTorch implementation of Google's) and qwenlm/qwen2.5 (Qwen2). zai-org/chatglm2-6b and internlm/internlm round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Curamos repositorios de código abierto en GitHub que coinciden con “open source llm”. Los resultados están clasificados por relevancia según tu búsqueda; usa los filtros de abajo para acotar o refina con IA.

Plataforma de modelos de lenguaje (LLM)

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • 01-ai/yiAvatar de 01-ai

    01-ai/Yi

    7,822Ver en GitHub↗

    Yi is a bilingual language model and foundation model designed for natural language processing, reasoning, and reading comprehension in both English and Chinese. It is built as a transformer-based architecture capable of general purpose text generation and conversational tasks. The model is distinguished by its ability to function as a long context system, processing and analyzing extended input sequences up to 200k tokens. It also supports quantized versions that use low-bit precision to reduce memory footprints, enabling execution on consumer-grade hardware. The project covers a broad rang

    Yi is a publicly available bilingual language model with up to 200k-token context, quantization support for consumer hardware, and fine-tuning capabilities, directly meeting the request for a self-hostable LLM for text generation with several of the sought-after features.

    Jupyter NotebookLong-Context ModelsLong Context ProcessingModel Quantization
    Ver en GitHub↗7,822
  • google/gemma_pytorchAvatar de google

    google/gemma_pytorch

    5,697Ver en GitHub↗

    The official PyTorch implementation of Google's Gemma models

    This repository provides the official PyTorch implementation of Google's open-weight Gemma models, letting you self-host, fine-tune via LoRA, and run text generation with a permissive license and Hugging Face integration.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗5,697
  • qwenlm/qwen2.5Avatar de QwenLM

    QwenLM/Qwen2.5

    27,307Ver en GitHub↗

    Qwen2.5 is a suite of large language model foundation models designed for natural language generation, code production, and complex mathematical reasoning. The project encompasses a multilingual language model capable of processing dozens of languages and a specialized code generation model for technical problem solving and debugging. The framework is distinguished by its long context capabilities, enabling the analysis of massive inputs ranging from 256K up to 1 million tokens. It further functions as an agentic framework, utilizing standardized templates and parsers to execute autonomous wo

    Qwen2.5 offers open-weight decoder models with long-context handling (256K–1M tokens), multilingual support, and standard fine-tuning compatibility, squarely meeting the need for a self-hosted, publicly available LLM for text generation.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗27,307
  • zai-org/chatglm2-6bAvatar de zai-org

    zai-org/ChatGLM2-6B

    15,564Ver en GitHub↗

    ChatGLM2-6B is a bilingual chat large language model designed for natural conversation and text generation in both English and Chinese. It functions as a fine-tunable language model that supports updating weights via specialized scripts to adapt to specific datasets and tasks. The project serves as a quantized inference engine and multi-GPU model orchestrator, enabling the execution of large models on consumer-grade hardware. It is capable of processing long context sequences up to 32K tokens to maintain understanding across extended documents. The system covers capabilities for multilingual

    ChatGLM2-6B is an open-source bilingual LLM with publicly available weights that can be self-hosted and fine-tuned, featuring inference optimization via quantization and multi-GPU support, a long context window of 32K tokens, and integration with the Hugging Face ecosystem, making it a comprehensive fit for your text generation and fine-tuning needs.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗15,564
  • internlm/internlmAvatar de InternLM

    InternLM/InternLM

    7,224Ver en GitHub↗

    InternLM is a large language model and a comprehensive suite of weights designed for text generation and complex reasoning. It functions as an inference engine for serving responses, a fine-tuning framework for adjusting model weights, and a platform for building autonomous AI agents. The system is capable of processing long-context input sequences up to one million tokens for document analysis. It employs chain-of-thought reasoning to solve knowledge-intensive tasks by generating intermediate logic steps before producing a final answer. The project covers model weight optimization through s

    InternLM is a publicly available, self-hostable large language model with support for long-context processing (up to 1M tokens), fine-tuning, and inference optimization, making it a comprehensive fit for text generation and advanced reasoning tasks.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗7,224
  • zai-org/glm-4Avatar de zai-org

    zai-org/GLM-4

    7,058Ver en GitHub↗

    GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod

    GLM-4 is an openly available large language model with a built-in fine-tuning framework that supports parameter-efficient adapters (like LoRA), inference optimization, and a context window of up to one million tokens, making it a comprehensive choice for self-hosted text generation and customization.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗7,058
  • meta-llama/llama-modelsAvatar de meta-llama

    meta-llama/llama-models

    7,643Ver en GitHub↗

    This project provides a foundational framework and reference implementation for executing causal language modeling and multimodal reasoning on local systems. It includes a set of core components for managing model assets, a fine-tuning framework, and structural definitions required to instantiate transformer-based architectures. The system is distinguished by its ability to process combined text and image inputs through multimodal transformer models for visual reasoning and document analysis. It also supports the deployment of quantized models, reducing memory footprints through low-precision

    This is the official Meta repository for the Llama family of LLMs, providing publicly available weights, a fine-tuning framework, inference with quantization, and long-context support — a flagship open-source model that exactly matches the request for a self-hostable, fine-tunable text generation model.

    PythonLong Context ProcessingModel Quantization
    Ver en GitHub↗7,643
  • thudm/chatglm2-6bAvatar de THUDM

    THUDM/ChatGLM2-6B

    15,565Ver en GitHub↗

    ChatGLM2-6B is an open-weight large language model designed for natural language conversations and text generation in both English and Chinese. It functions as a bilingual chat model capable of processing and maintaining coherence across text sequences up to 32K tokens. The model is optimized for local deployment through precision quantization, which reduces memory requirements to allow execution on consumer-grade hardware. It supports distributing model weights across multiple graphics cards to handle parameters that exceed the memory of a single device. The project covers capabilities for

    ChatGLM2-6B is an open-weight bilingual LLM with up to 32K token context and quantization for local deployment, making it a solid fit for self-hosted text generation and fine-tuning, though the description does not explicitly confirm a permissive license or Hugging Face ecosystem integration.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗15,565
  • zihangdai/xlnetAvatar de zihangdai

    zihangdai/xlnet

    6,182Ver en GitHub↗

    This project is a natural language processing framework focused on a generalized autoregressive pretrainer designed for unsupervised language representation. It implements a language model that combines permutation-based training with a Transformer-XL backbone to function as a long-context text processor. The system is distinguished by its ability to handle text sequences that exceed standard length limits through the use of segment-level recurrence and relative positional encoding. It scales high-performance pretraining across multiple GPUs and TPU clusters using distributed training impleme

    XLNet is a transformer-based language model with publicly available pretrained weights that you can self-host and fine-tune, and its Transformer-XL backbone gives it a long context window—but its parameter scale is smaller than modern LLMs and it does not prominently cover inference optimization or LoRA specifics.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗6,182
  • ymcui/chinese-llama-alpaca-2Avatar de ymcui

    ymcui/Chinese-LLaMA-Alpaca-2

    7,136Ver en GitHub↗

    This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce

    This repository provides a Chinese-optimized LLaMA-2 large language model with publicly available weights, supporting LoRA fine-tuning, quantization, and 64K-token contexts, so it fits the search for a self-hostable, fine-tunable LLM for text generation.

    PythonLong-Context ModelsLong Context ProcessingSpeculative Decoding Strategies
    Ver en GitHub↗7,136
  • openbmb/minicpmAvatar de OpenBMB

    OpenBMB/MiniCPM

    9,464Ver en GitHub↗

    MiniCPM is a collection of small language models designed for local, on-device deployment in resource-constrained environments. The project focuses on running dense Transformer models on consumer hardware, including GPUs, CPUs, and Apple Silicon, without requiring custom code forks. The project distinguishes itself through heavy optimization for edge hardware, utilizing quantized weight compression in GGUF and MLX formats to reduce memory overhead. It implements advanced inference techniques such as speculative sampling and radix-tree prefix caching to accelerate generation speed and throughp

    MiniCPM is a collection of small, self-hostable LLMs with publicly available weights and heavy inference optimization for consumer hardware, making it a direct fit for local text generation — fine-tuning details are less emphasised but the category is right.

    Jupyter NotebookInference OptimizationsInference OptimizationsLong Context Processing
    Ver en GitHub↗9,464
  • allenai/olmoAvatar de allenai

    allenai/OLMo

    6,313Ver en GitHub↗

    OLMo is an open-source large language model from AI2 with publicly available weights, supports fine-tuning including LoRA, integrates with Hugging Face, and offers various sizes and quantization options, making it a comprehensive answer for self-hosted text generation.

    PythonLarge Language Model Training Frameworks8-Bit Inference Quantizers8-Bit Load-Time Quantizers
    Ver en GitHub↗6,313
  • huggingface/smollmAvatar de huggingface

    huggingface/smollm

    3,624Ver en GitHub↗

    SmolLM is a project dedicated to the development of small language models. It focuses on training and fine-tuning compact models that maintain high performance while utilizing fewer parameters. The project emphasizes efficient AI inference and on-device text generation, aiming to enable the deployment of lightweight models on edge devices with limited memory and processing power. It utilizes synthetic data generation to produce artificial datasets that improve the reasoning and training of these AI systems. The system supports a variety of optimization and training capabilities, including we

    SmolLM from Hugging Face provides openly available small language model weights that can be self-hosted, fine-tuned, and used for text generation, fitting your request for an open-source LLM with accessible weights, though its compact scale is narrower than a flagship large model.

    PythonInference OptimizationLong-Context Models
    Ver en GitHub↗3,624
  • qwenlm/qwen3Avatar de QwenLM

    QwenLM/Qwen3

    27,324Ver en GitHub↗

    Qwen3 is a transformer-based large language model designed as a generative AI foundation for understanding, reasoning, and generating human language. It functions as a comprehensive ecosystem for model training, fine-tuning, and production-ready inference, providing the underlying architecture and weights necessary to build diverse artificial intelligence applications. The project distinguishes itself through extensive support for model quantization and distributed inference, enabling efficient execution across a wide range of hardware from consumer-grade devices to scalable cloud infrastruct

    Qwen3 is an open-source transformer-based large language model with publicly available weights that supports fine-tuning, quantization, and distributed inference, making it a solid fit for self-hosted text generation — though the provided description does not confirm permissive licensing or full feature details.

    PythonModel Quantization
    Ver en GitHub↗27,324
  • zai-org/chatglm-6bAvatar de zai-org

    zai-org/ChatGLM-6B

    41,039Ver en GitHub↗

    ChatGLM-6B is a generative AI inference engine designed for local execution of transformer-based language models. It provides a comprehensive runtime environment that allows users to load and run pre-trained neural network weights directly on their own hardware, ensuring data privacy and independence from external cloud services. The project distinguishes itself through a hardware-agnostic execution backend that supports deployment across diverse environments, including standard processors, Apple Silicon, and multi-GPU configurations. It incorporates advanced optimization techniques such as w

    ChatGLM-6B is an open-source large language model with 6 billion parameters, publicly available weights, and a repository that provides inference and fine-tuning tooling, fitting your need for a self-hostable, fine-tunable text generation model.

    PythonModel Quantization
    Ver en GitHub↗41,039
  • thudm/chatglm-6bAvatar de THUDM

    THUDM/ChatGLM-6B

    41,040Ver en GitHub↗

    ChatGLM-6B is an open-source bilingual large language model designed for natural dialogue and text generation in both English and Chinese. It is structured as a dialogue model capable of tasks such as role-playing and information extraction. The project provides implementations for quantized language models, using low-precision weights to reduce GPU memory requirements for local inference. It also supports parameter-efficient fine-tuning, allowing model behavior to be optimized for specific tasks without requiring full retraining. The model includes capabilities for local execution on GPUs a

    ChatGLM-6B is an open-source bilingual LLM that can be self-hosted and fine-tuned for text generation, with support for quantized inference and parameter-efficient fine-tuning, making it a solid fit for this search—though its license is not fully permissive and its context window is moderate.

    PythonTransformer ModelsBilingual Language ModelsDecoder Architectures
    Ver en GitHub↗41,040
  • deepseek-ai/deepseek-v3Avatar de deepseek-ai

    deepseek-ai/DeepSeek-V3

    103,753Ver en GitHub↗

    DeepSeek-V3 is a large language model that provides comprehensive resources for model utilization, including technical specifications, pre-trained weights, and evaluation benchmarks. The project details the core transformer architecture, including parameter counts and multi-token prediction modules, while supporting native 8-bit floating-point quantization. The repository offers extensive support for local and distributed inference through integration with multiple frameworks and engines. It includes documentation for deploying the model across various hardware configurations, such as GPUs an

    DeepSeek-V3 is a large language model with publicly available pre-trained weights and support for local inference and deployment, making it a strong fit for self-hosted text generation and fine-tuning, though the provided description does not explicitly mention LoRA or Hugging Face integration.

    PythonModel WeightsInference FrameworksFrontier Models
    Ver en GitHub↗103,753
  • cstankonrad/long_llamaAvatar de CStanKonrad

    CStanKonrad/long_llama

    1,465Ver en GitHub↗

    Long Llama is a transformer-based language model and fine-tuning framework designed to process and maintain logical coherence across input sequences that significantly exceed standard length limits. By utilizing a focused transformer architecture, the project enables models to handle massive documents or entire books by training attention layers to track distant tokens. The framework distinguishes itself through specialized attention mechanisms that allow for the processing of hundreds of thousands of tokens. It incorporates memory-efficient inference techniques, such as key-value caching and

    LongLLaMA is an open-source large language model specifically fine-tuned for handling long contexts, which directly meets the core requirement for a self-hostable text generation model with publicly available weights, though it does not explicitly cover LoRA fine-tuning or inference optimization.

    PythonLong-Context ModelsLong Context Processing
    Ver en GitHub↗1,465
Compara los 10 mejores de un vistazo
RepositorioEstrellasLenguajeLicenciaÚltimo push
01-ai/yi7.8KJupyter NotebookApache-2.027 nov 2024
google/gemma_pytorch5.7KPythonApache-2.030 may 2025
qwenlm/qwen2.527.3KPython—9 ene 2026
zai-org/chatglm2-6b15.6KPythonNOASSERTION27 jun 2024
internlm/internlm7.2KPythonApache-2.030 oct 2025
zai-org/glm-47.1KPythonapache-2.04 jul 2025
meta-llama/llama-models7.6KPythonNOASSERTION11 feb 2026
thudm/chatglm2-6b15.6KPythonNOASSERTION27 jun 2024
zihangdai/xlnet6.2KPythonApache-2.028 may 2023
ymcui/chinese-llama-alpaca-27.1KPythonApache-2.019 abr 2026

Related searches

  • plataforma open source para LLMs locales
  • un modelo de código abierto para inferencia local
  • un framework de código abierto para LLMs locales
  • un modelo de código abierto para despliegue local
  • an open source engine for local LLMs
  • plataforma open source para hosting de LLMs
  • un motor de inferencia para ejecutar LLMs locales
  • un framework de código abierto para aplicaciones de LLM