5 Repos
Techniques for forcing model outputs into specific structured formats via logit manipulation.
Distinct from Sequence Decoders: Distinct from general sequence decoding by focusing on the constraint/forcing of specific output formats.
Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Constrained Decoding. Refine with filters or upvote what's useful.
Ludwig is a multimodal machine learning platform and low-code framework designed for building, training, and deploying neural networks. It enables the construction of models that process text, images, audio, and tabular data through a unified interface using declarative configuration files rather than custom code. The system features a specialized low-code framework for large language models, supporting supervised fine-tuning, preference alignment, and a constrained decoding tool to force structured data output via logit extraction. It also includes an automated model architecture search to i
Forces large language models to produce structured data using logit extraction and constrained decoding.
mistral.rs is an inference engine for large language models that runs locally and exposes models behind OpenAI and Anthropic-compatible APIs. It serves as a multi-model serving platform, capable of loading several models in a single server process with per-request routing and on-demand loading and unloading. The engine supports multimodal inference, processing text alongside images, video, audio, and speech inputs, and includes a quantized model deployment runtime that reduces memory use and speeds up inference on consumer hardware. The project distinguishes itself through an agentic tool exe
Enforces JSON Schema on tool call arguments during decoding to prevent malformed output.
LiteRT-LM ist ein Hochleistungs-Inferenz-Framework, das darauf ausgelegt ist, Large Language Models lokal auf Mobil-, Desktop- und IoT-Hardware auszuführen. Es dient als On-Device-Modell-Laufzeitumgebung, die CPU-, GPU- und NPU-Beschleunigung nutzt, um eine Verarbeitung mit geringer Latenz zu ermöglichen. Das Framework zeichnet sich durch die Fähigkeit aus, Text-, Bild- und Audioeingaben über eine einzige multimodale Inferenz-Engine zu verarbeiten. Es verfügt über einen lokalen HTTP-Server, der OpenAI-kompatible API-Endpunkte emuliert, sowie eine WebGPU-basierte Laufzeitumgebung zur Ausführung von Modellen direkt im Webbrowser. Um die Zuverlässigkeit der Ausgabe zu gewährleisten, enthält es einen eingeschränkten Textgenerator, der JSON-Schemas oder Grammatikregeln für Modellantworten erzwingt. Das Projekt bietet umfassende Funktionen für zustandsbehaftetes Konversationsmanagement, spekulative Dekodierung für höhere Token-Generierungsgeschwindigkeiten und eine Tool-Calling-Schnittstelle, die Modellanfragen auf externe Funktionen abbildet. Es beinhaltet zudem eine spezialisierte Integration für das Apple-Ökosystem und ein dediziertes Plugin für die Modellausführung in Flutter. Benutzer können Modelle über eine Befehlszeilenschnittstelle ausführen oder sie über native APIs in Anwendungen integrieren.
Provides constrained decoding to ensure model outputs follow specific structured formats via logit manipulation.
This project is a retrieval augmented generation framework designed to build pipelines that connect unstructured data and knowledge graphs with large language models. It functions as a vector database orchestrator for indexing text and multimodal content, as well as a system for translating natural language queries into structured database commands. The framework integrates a hybrid retrieval engine that combines dense vector search with sparse keyword matching to increase the precision of retrieved contexts. It further enhances reasoning and relationship mapping through a graph-augmented ret
Uses constrained decoding and validation to force model outputs into predefined structured formats.
LightLLM is a high-performance serving framework for deploying and executing large language models. It functions as a multi-GPU inference engine and server capable of handling dense architectures, mixture-of-experts designs, and multimodal models that process both text and images. The system is distinguished by its specialized support for Mixture-of-Experts models using expert parallelism and fused kernels. It implements structured text generation through deterministic state machines and pushdown automata to enforce precise output formats. To optimize throughput, the framework employs specula
Enforces structured text generation using deterministic state machines to ensure responses follow precise formats.