20 Repos
Utilities for executing single model predictions directly from a terminal interface.
Distinct from Command Line: Candidates focus on shell completions or general task management; this is specifically for running ML model inference via CLI.
Explore 20 awesome GitHub repositories matching development tools & productivity · Command Line Model Inferences. Refine with filters or upvote what's useful.
TensorFlow.js is a JavaScript machine learning library used for training and deploying models in web browsers and server-side environments. It functions as a browser-based model trainer, a WebAssembly inference engine, and a WebGPU accelerated tensor library for low-level linear algebra. The project also includes a model converter to transform Python-based models into optimized formats for JavaScript execution. The library distinguishes itself through a pluggable backend architecture that allows mathematical operations to be executed via CPU, WebGL, or WebGPU. It supports the conversion of Py
Provides command-line utilities to process input tensors and perform model inference.
The simplest way to run LLaMA on your local machine
Manages language models through terminal commands for local use.
Mistral Inference is a library for running Mistral large language models on a GPU, generating text from prompts with token streaming. It loads pretrained model weights from local disk or a remote registry into GPU memory, then produces output tokens one by one for real-time display in interactive applications. The library supports multimodal prompts that accept image URLs alongside text, enabling visual description and reasoning. It includes content safety guardrails that scan generated text against predefined policies to block or flag policy violations. For structured interactions, it provid
Starts a command-line session that accepts user prompts and streams model responses conversationally.
Dieses Projekt ist eine Inference-Bibliothek und ein Framework für Large Language Models, das darauf ausgelegt ist, Modelle für Textgenerierung, Problemlösung und Coding-Assistenz auszuführen. Es enthält ein multimodales Framework für die Verarbeitung kombinierter Bild- und Texteingaben sowie eine Tool-Use-Implementierung, die die Ausführung externer Funktionen basierend auf Modell-Reasoning ermöglicht. Das System verfügt über eine verteilte GPU-Inference-Engine, die große Modell-Workloads auf mehrere Grafikprozessoren verteilt, um die Verarbeitungsgeschwindigkeit zu erhöhen und Speicheranforderungen zu erfüllen. Es bietet zudem containerisiertes Modell-Deployment durch vorverpackte Images und Abhängigkeiten für das Serving von Inference-Engines in isolierten Umgebungen. Die Bibliothek deckt eine Reihe von Funktionen ab, einschließlich multimodaler Eingabeanalyse, Integration von Function-Calling und Fill-in-the-Middle-Coding zur Vorhersage fehlender Code-Segmente. Zudem unterstützt sie interaktiven Modell-Chat via Command-Line-Interface für die Aufrechterhaltung von Konversationssitzungen.
Provides a command-line interface for maintaining interactive conversational sessions with models.
WhisperLiveKit is a real-time speech-to-text server that transcribes streaming audio into text with ultra-low latency using Whisper models. It serves transcription capabilities through REST endpoints and WebSocket connections, enabling external applications to send audio and receive transcriptions as words are spoken, making it suitable for live captioning or voice interfaces. The project distinguishes itself by combining real-time transcription with speaker diarization, assigning transcribed words to individual speakers during live audio streams for meeting or interview transcripts. It also
Manages model lifecycle through CLI commands for listing, downloading, and deleting speech recognition models.
Cog is a machine learning packaging tool and containerized model wrapper that bundles models and their dependencies into standardized Docker containers. It functions as an environment manager and inference server, ensuring consistent model execution across different hardware systems by resolving GPU drivers, system libraries, and Python dependencies. The project distinguishes itself by automatically generating RESTful HTTP servers and OpenAPI schemas based on defined model input and output types. It manages large model weights as external fixtures to optimize image size and utilizes a slot-ba
Provides a command-line interface to execute a single prediction through a containerized model.
PrusaSlicer is a G-code generator that converts 3D models into machine instructions for FFF and mSLA printers, handling slicing, infill, and support generation. It provides a command-line slicing interface for processing models and profiles via terminal commands without a graphical user interface, and includes a G-code customization engine that inserts user-defined macros, variables, and post-processing scripts into generated G-code for tailored machine control. The software also manages multi-material prints by coordinating multiple extruders and filament colors, assigning materials to model
Processes 3D models and profiles via terminal commands to produce printable G-code without a graphical user interface.
Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP
Provides a command-line interface for executing model inferences with configurable sampling parameters.
Vowpal Wabbit is an open-source machine learning system designed for online learning, where models update incrementally from streaming data without requiring full retraining. It provides a reduction-based learning framework that composes complex tasks from simpler algorithms, and includes a feature hashing trick that maps unbounded feature names into a fixed-size vector space to keep memory usage constant regardless of dataset size. The system supports distributed training across a cluster using an allreduce protocol for synchronized updates, and offers an active learning query strategy that s
Trains and evaluates machine learning models directly from the terminal using compact argument syntax.
Metaseq ist ein Transformer-Sequenzmodellierungs-Toolkit, das für das Training, Fine-Tuning und die Bereitstellung von Sequence-to-Sequence-Modellen unter Verwendung offener, vortrainierter Gewichte entwickelt wurde. Es bietet ein umfassendes Framework für das Training von Large Language Models, einschließlich dedizierter Tools für die Verarbeitung von Sequenz-Datensätzen und einen eigenständigen Inference-Server zur Textgenerierung via API-Anfragen. Das Projekt bietet spezialisierte Dienstprogramme für Modell-Quantisierung, um die Parameterpräzision auf acht Bit zu reduzieren, was den Speicherverbrauch senkt und die Inferenzgeschwindigkeit erhöht. Es enthält zudem eine Checkpoint-Konvertierungs-Pipeline, um Modellgewichte in Strukturen umzuwandeln, die für leistungsstarke Inferenz-Engines optimiert sind. Das Framework unterstützt groß angelegtes Training über GPU-Cluster hinweg durch den Einsatz von Tensor-Parallelität und Sharded-Data-Parallelität. Zusätzliche Funktionen decken die Vorbereitung von NLP-Datensätzen, das Laden vortrainierter Gewichte für Transfer Learning und die Nachverfolgung von Trainingsmetriken zur Fortschrittsvisualisierung ab.
Supports interactive command-line sessions for loading models and generating text with configurable sampling parameters.
The TensorFlow Cookbook is a collection of code examples and recipes for building, training, and deploying machine learning models using TensorFlow. It covers the full model lifecycle, from constructing neural networks and training them with configurable parameters to packaging trained models for production deployment with unit tests and multi-device support. The project also integrates TensorBoard for logging and visualizing computational graphs, scalar summaries, and histograms during training. The cookbook demonstrates a wide range of machine learning techniques, including convolutional ne
Manages model training and inference through explicit session creation, variable initialization, and cleanup.
Dieses Projekt ist eine Bibliothek und ein Kommandozeilen-Interface für die Inferenz lokaler Large Language Models. Es ermöglicht die Generierung von Textvervollständigungen und Chat-Antworten verschiedener Modellarchitekturen. Das Projekt bietet Tools für Weight-Quantization, um den Speicherbedarf zu reduzieren, und integriert Hardwarebeschleunigung durch GPU-Offloading, um die Berechnungsgeschwindigkeit zu erhöhen. Zudem enthält es Utilities für die Modellevaluierung durch Messung der Perplexity auf spezifischen Datensätzen. Die Funktionen decken den gesamten Inferenz-Lifecycle ab, einschließlich binärem Modell-Laden, Template-basierter Prompt-Strukturierung und Sitzungspersistenz zur Beibehaltung des Konversationskontexts. Es unterstützt zudem Aufgaben-Orchestrierung, was es ermöglicht, mehrere Modellaufrufe in Pipelines für mehrstufige Operationen zu sequenzieren.
Allows saving and loading the state of an interaction to maintain context across sessions.
SAHI ist ein Sliced-Inference-Framework und eine Computer-Vision-Pipeline, die entwickelt wurde, um kleine Objekte in hochauflösenden Bildern zu erkennen. Es bietet ein System zur Unterteilung großer Bilder in überlappende Patches, um den Detailverlust zu verhindern, der typischerweise bei der Standard-Modell-Herunterskalierung auftritt, sowie ein Bild-Tiling-Dienstprogramm und ein COCO-Datensatz-Toolkit. Das Projekt zeichnet sich durch einen modellagnostischen Vorhersage-Wrapper aus, der verschiedene Machine-Learning-Frameworks in eine einheitliche Schnittstelle standardisiert. Dies ermöglicht die Implementierung von Sliced Inference und Objekterkennung über verschiedene Modell-Backends hinweg bei gleichzeitiger Beibehaltung eines konsistenten Ausgabeformats. Über die Inferenz hinaus deckt das Framework das Datensatzmanagement für COCO- und YOLO-Formate ab, einschließlich Tools für annotiertes Bild-Slicing, Kategorien-Remapping und Datensatz-Zusammenführung. Es enthält zudem eine Suite zur Bewertung und Überwachung der Modellleistung, mit Metrikberechnung für Präzision und Recall, Erkennungsfehleranalyse und Ergebnisvisualisierung. Das Toolset ist über eine Befehlszeilenschnittstelle zugänglich, um Inferenz-Workflows über Bildverzeichnisse und Videostreams hinweg zu automatisieren.
Provides a command-line interface for executing object detection predictions and dataset operations.
Dieses Projekt ist ein containerisierter Entwicklungs-Stack und ein Anwendungs-Framework für den Aufbau von RAG-Systemen (Retrieval-Augmented Generation). Es bietet eine dockerisierte KI-Sandbox, die lokale Modell-Runtimes, Wissensgraphen und Vektorspeicher integriert, um die Erstellung kontextbezogener Chatbots zu ermöglichen. Der Stack zeichnet sich durch seinen graphenbasierten Vektorspeicher aus, der strukturierte Wissensgraphen mit Vektorindizes für semantisches und strukturelles Daten-Retrieval kombiniert. Er ermöglicht das lokale Hosten von Modellen mit CPU- oder GPU-Beschleunigung, wodurch generative Aufgaben ohne Abhängigkeit von externen Cloud-APIs möglich sind. Das Framework deckt ein breites Spektrum an Funktionen ab, einschließlich der Verarbeitung und Indizierung von PDF-Dokumenten, der Orchestrierung containerbasierter KI-Dienste und der Implementierung von Grounded-Response-Generierung. Es enthält eine webbasierte Chat-Oberfläche mit inkrementellem Response-Streaming sowie eine standardisierte Schnittstelle zum Wechseln zwischen verschiedenen Sprachmodell-Anbietern. Die Umgebung wird mittels Container-Orchestrierung gebootstrapt, um einen vorkonfigurierten Stack aus Modellen und Datenbanken schnell bereitzustellen.
Automates the download and installation of local language model runtimes and system services.
xtuner ist eine umfassende Trainings-Engine für Large Language Models und bietet ein Toolkit für Pre-Training, Supervised Fine-Tuning und die Optimierung von vision-sprachlichen multimodalen Modellen. Sie dient als verteilter Trainingsbeschleuniger und spezialisiertes Framework zur Skalierung von Mixture-of-Experts-Modellen sowie zur Ausrichtung von Modellverhalten durch Reinforcement Learning from Human Feedback. Das Projekt zeichnet sich durch fortgeschrittene Speicher- und Rechenoptimierungen aus, wie Sequence-Parallelism für ultra-lange Kontextfenster und Interleaved-Pipeline-Parallelism zur Reduzierung von GPU-Idle-Zeiten. Es bietet eine dedizierte Suite für Preference-Optimization und implementiert Techniken wie Group Relative Policy Optimization und Direct Preference Optimization, um Modell-Policies und Belohnungssysteme zu verfeinern. Breite Funktionsbereiche decken verteiltes Modelltraining über mehrere Knoten hinweg, multimodale Datensatzvorbereitung und die Verwaltung von Adapter-basiertem Fine-Tuning ab. Die Engine enthält zudem Tools für Modellevaluation, Weight-Merging und den Export trainierter Parameter in Inferenz-Engines. Das Training wird über standardisierte Konfigurationsdateien und verteilte Launcher verwaltet, um konsistente Ergebnisse über Rechencluster hinweg sicherzustellen.
Executes interactive chat sessions using specific prompt templates and optional adapter weights.
MedicalGPT is an open-source framework for fine-tuning large language models, with a dedicated focus on adapting general models to the medical domain. It provides a complete pipeline that covers continued pretraining on domain-specific corpora, supervised instruction tuning, tokenizer vocabulary extension with medical terminology, and alignment to clinician preferences through direct preference optimization, reinforcement learning, or knowledge distillation. The framework also supports training models to invoke external tools and functions in multi-turn clinical conversations. The platform di
Loads the fine‑tuned model weights and supports interactive chat or batch text generation from a command‑line session within the framework
Aigcpanel is a visual workflow automation tool and model lifecycle manager designed for generative AI media pipelines. It provides a unified interface to install, launch, and configure both local and remote AI model endpoints, acting as an orchestration platform for large language models and AI tools. The system features a drag-and-drop node editor for chaining AI models and scripts into automated processing pipelines. It distinguishes itself with a breakpoint-aware execution model that allows users to pause and resume long media tasks from specific points in the workflow. Additionally, it in
Provides a command line interface for executing model functions and querying available models for script integration.
ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,
ExecuTorch loads and executes a language model on-device, wrapping the runtime for text generation tasks.
Distributed-llama is a distributed inference engine and command line tool for running large language models across multiple networked machines. It functions as a compute cluster manager that coordinates worker nodes to share the computational load of a single model. The system utilizes tensor parallelism to shard model weights across different hosts, allowing the execution of models that exceed the memory capacity of a single piece of hardware. It includes a dedicated format converter to transform standard model files into a compatible binary layout optimized for distributed loading. The eng
Provides a command-line interface and server for interactive chat sessions and batch text generation.
Foundry-Local ist ein Machine-Learning-Entwicklungstool, das für private Inferenz auf dem Gerät und Modellmanagement entwickelt wurde. Es bietet eine lokale Serverumgebung, die Machine-Learning-Modelle direkt auf der Hardware des Benutzers hostet und sicherstellt, dass die gesamte Datenverarbeitung, einschließlich Prompt-Handhabung und Audiotranskription, innerhalb der lokalen Umgebung verbleibt, ohne dass eine externe Cloud-Konnektivität erforderlich ist. Das Projekt zeichnet sich durch die Automatisierung des gesamten Modell-Lebenszyklus aus, einschließlich der Entdeckung, des Herunterladens und der Versionierung von Assets, um die Kompatibilität mit der Host-Hardware aufrechtzuerhalten. Es verfügt über eine Hardware-Abstraktionsschicht, die automatisch den effizientesten verfügbaren Prozessor für rechenintensive Aufgaben erkennt und auswählt, was eine hardwarebeschleunigte Ausführung ohne manuelle Konfiguration ermöglicht. Über die Kern-Inferenz hinaus enthält das Tool eine Befehlszeilenschnittstelle für interaktive Modell-Exploration und Leistungsüberprüfung. Es bietet zudem standardisiertes API-Proxying, das eingehende Anfragen unter Verwendung branchenüblicher Protokolle auf lokale Modell-Endpunkte abbildet, um die Integration mit externen Software-Frameworks zu unterstützen.
Provides an interactive command-line interface for developers to test inference performance and verify model outputs directly.