20 dépôts
Utilities for executing single model predictions directly from a terminal interface.
Distinct from Command Line: Candidates focus on shell completions or general task management; this is specifically for running ML model inference via CLI.
Explore 20 awesome GitHub repositories matching development tools & productivity · Command Line Model Inferences. Refine with filters or upvote what's useful.
TensorFlow.js is a JavaScript machine learning library used for training and deploying models in web browsers and server-side environments. It functions as a browser-based model trainer, a WebAssembly inference engine, and a WebGPU accelerated tensor library for low-level linear algebra. The project also includes a model converter to transform Python-based models into optimized formats for JavaScript execution. The library distinguishes itself through a pluggable backend architecture that allows mathematical operations to be executed via CPU, WebGL, or WebGPU. It supports the conversion of Py
Provides command-line utilities to process input tensors and perform model inference.
The simplest way to run LLaMA on your local machine
Manages language models through terminal commands for local use.
Mistral Inference is a library for running Mistral large language models on a GPU, generating text from prompts with token streaming. It loads pretrained model weights from local disk or a remote registry into GPU memory, then produces output tokens one by one for real-time display in interactive applications. The library supports multimodal prompts that accept image URLs alongside text, enabling visual description and reasoning. It includes content safety guardrails that scan generated text against predefined policies to block or flag policy violations. For structured interactions, it provid
Starts a command-line session that accepts user prompts and streams model responses conversationally.
This project is a large language model inference library and framework designed to run models for text generation, problem solving, and coding assistance. It includes a multimodal framework for processing combined image and text inputs and a tool-use implementation that enables the execution of external functions based on model reasoning. The system features a distributed GPU inference engine that spreads large model workloads across multiple graphics processors to increase processing speed and meet memory requirements. It also provides containerized model deployment through pre-packaged imag
Provides a command-line interface for maintaining interactive conversational sessions with models.
WhisperLiveKit is a real-time speech-to-text server that transcribes streaming audio into text with ultra-low latency using Whisper models. It serves transcription capabilities through REST endpoints and WebSocket connections, enabling external applications to send audio and receive transcriptions as words are spoken, making it suitable for live captioning or voice interfaces. The project distinguishes itself by combining real-time transcription with speaker diarization, assigning transcribed words to individual speakers during live audio streams for meeting or interview transcripts. It also
Manages model lifecycle through CLI commands for listing, downloading, and deleting speech recognition models.
Cog is a machine learning packaging tool and containerized model wrapper that bundles models and their dependencies into standardized Docker containers. It functions as an environment manager and inference server, ensuring consistent model execution across different hardware systems by resolving GPU drivers, system libraries, and Python dependencies. The project distinguishes itself by automatically generating RESTful HTTP servers and OpenAPI schemas based on defined model input and output types. It manages large model weights as external fixtures to optimize image size and utilizes a slot-ba
Provides a command-line interface to execute a single prediction through a containerized model.
PrusaSlicer is a G-code generator that converts 3D models into machine instructions for FFF and mSLA printers, handling slicing, infill, and support generation. It provides a command-line slicing interface for processing models and profiles via terminal commands without a graphical user interface, and includes a G-code customization engine that inserts user-defined macros, variables, and post-processing scripts into generated G-code for tailored machine control. The software also manages multi-material prints by coordinating multiple extruders and filament colors, assigning materials to model
Processes 3D models and profiles via terminal commands to produce printable G-code without a graphical user interface.
Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP
Provides a command-line interface for executing model inferences with configurable sampling parameters.
Vowpal Wabbit is an open-source machine learning system designed for online learning, where models update incrementally from streaming data without requiring full retraining. It provides a reduction-based learning framework that composes complex tasks from simpler algorithms, and includes a feature hashing trick that maps unbounded feature names into a fixed-size vector space to keep memory usage constant regardless of dataset size. The system supports distributed training across a cluster using an allreduce protocol for synchronized updates, and offers an active learning query strategy that s
Trains and evaluates machine learning models directly from the terminal using compact argument syntax.
Metaseq est une boîte à outils de modélisation de séquences transformer conçue pour entraîner, affiner et déployer des modèles séquence-à-séquence en utilisant des poids pré-entraînés ouverts. Il fournit un framework complet pour l'entraînement de grands modèles de langage, incluant des outils dédiés pour le traitement de datasets de séquences et un serveur d'inférence autonome pour générer du texte via des requêtes API. Le projet présente des utilitaires spécialisés pour la quantification de modèles afin de réduire la précision des paramètres à huit bits, ce qui diminue l'utilisation de la mémoire et augmente la vitesse d'inférence. Il inclut également un pipeline de conversion de points de contrôle pour transformer les poids des modèles en structures optimisées pour des moteurs d'inférence haute performance. Le framework supporte l'entraînement à grande échelle sur des clusters GPU via l'utilisation du parallélisme tensoriel et du parallélisme de données fragmenté. Des capacités supplémentaires couvrent la préparation de datasets NLP, le chargement de poids pré-entraînés pour le transfert learning, et le suivi des métriques d'entraînement pour la visualisation de la progression.
Supports interactive command-line sessions for loading models and generating text with configurable sampling parameters.
The TensorFlow Cookbook is a collection of code examples and recipes for building, training, and deploying machine learning models using TensorFlow. It covers the full model lifecycle, from constructing neural networks and training them with configurable parameters to packaging trained models for production deployment with unit tests and multi-device support. The project also integrates TensorBoard for logging and visualizing computational graphs, scalar summaries, and histograms during training. The cookbook demonstrates a wide range of machine learning techniques, including convolutional ne
Manages model training and inference through explicit session creation, variable initialization, and cleanup.
Ce projet est une bibliothèque et une interface en ligne de commande pour l'inférence locale de grands modèles de langage. Il permet la génération de complétions de texte et de réponses de chat à partir de diverses architectures de modèles. Le projet fournit des outils pour la quantification des poids afin de réduire l'empreinte mémoire et intègre l'accélération matérielle via le déchargement GPU pour augmenter la vitesse de calcul. Il inclut également des utilitaires pour l'évaluation des modèles en mesurant la perplexité sur des jeux de données spécifiques. Les capacités couvrent l'ensemble du cycle de vie de l'inférence, incluant le chargement de modèles binaires, la structuration de prompts basée sur des templates et la persistance des sessions pour maintenir le contexte conversationnel. Il prend également en charge l'orchestration des tâches, permettant à plusieurs appels de modèles d'être séquencés dans des pipelines pour des opérations en plusieurs étapes.
Allows saving and loading the state of an interaction to maintain context across sessions.
SAHI est un framework d'inférence par découpage (sliced inference) et un pipeline de vision par ordinateur conçu pour détecter de petits objets dans des images haute résolution. Il fournit un système pour diviser les grandes images en patchs chevauchants afin d'éviter la perte de détails qui se produit généralement lors de la réduction d'échelle standard des modèles, aux côtés d'un utilitaire de tuilage d'image et d'une boîte à outils de jeu de données COCO. Le projet se distingue en offrant un wrapper de prédiction agnostique au modèle qui standardise différents frameworks d'apprentissage automatique dans une interface unifiée. Cela lui permet d'implémenter l'inférence par découpage et la détection d'objets à travers divers backends de modèles tout en maintenant un format de sortie cohérent. Au-delà de l'inférence, le framework couvre la gestion de jeux de données pour les formats COCO et YOLO, incluant des outils pour le découpage d'images annotées, le remapping de catégories et la fusion de jeux de données. Il inclut également une suite pour l'évaluation et la surveillance des performances des modèles, présentant le calcul de métriques pour la précision et le rappel, l'analyse des erreurs de détection et la visualisation des résultats. La boîte à outils est accessible via une interface en ligne de commande pour automatiser les workflows d'inférence à travers les répertoires d'images et les flux vidéo.
Provides a command-line interface for executing object detection predictions and dataset operations.
Ce projet est une stack de développement conteneurisée et un framework d'application pour construire des systèmes de génération augmentée par récupération (RAG). Il fournit un sandbox IA dockerisé qui intègre des runtimes de modèles locaux, des graphes de connaissances et des vector stores pour permettre la création de chatbots contextuels. La stack se distingue par son vector store basé sur les graphes, qui combine des graphes de connaissances structurés avec des index vectoriels pour la récupération de données sémantiques et structurelles. Elle permet l'hébergement de modèles locaux avec accélération CPU ou GPU, permettant des tâches génératives sans dépendance aux API cloud externes. Le framework couvre un large éventail de capacités, notamment le traitement et l'indexation de documents PDF, l'orchestration de services IA basés sur des conteneurs et l'implémentation d'une génération de réponses ancrées (grounded). Il inclut une interface de chat web avec streaming de réponse incrémentiel et une interface standardisée pour basculer entre différents fournisseurs de modèles de langage. L'environnement est amorcé via l'orchestration de conteneurs pour déployer rapidement une stack préconfigurée de modèles et de bases de données.
Automates the download and installation of local language model runtimes and system services.
xtuner est un moteur d'entraînement complet pour les grands modèles de langage, offrant une boîte à outils pour le pré-entraînement, le fine-tuning supervisé et l'optimisation de modèles multimodaux vision-langage. Il sert d'accélérateur d'entraînement distribué et de framework spécialisé pour mettre à l'échelle des modèles Mixture-of-Experts et aligner le comportement du modèle via l'apprentissage par renforcement à partir de feedback humain (RLHF). Le projet se distingue par des optimisations avancées de mémoire et de calcul, telles que le parallélisme de séquence pour des fenêtres de contexte ultra-longues et le parallélisme de pipeline entrelacé pour réduire le temps d'inactivité du GPU. Il fournit une suite dédiée pour l'optimisation des préférences, implémentant des techniques comme Group Relative Policy Optimization et Direct Preference Optimization pour affiner les politiques du modèle et les systèmes de récompense. Les domaines de capacités étendus couvrent l'entraînement de modèle distribué sur plusieurs nœuds, la préparation de jeux de données multimodaux et la gestion du fine-tuning basé sur des adaptateurs. Le moteur inclut également des outils pour l'évaluation de modèle, la fusion de poids et l'exportation des paramètres entraînés vers des moteurs d'inférence. L'entraînement est géré via des fichiers de configuration standardisés et des lanceurs distribués pour assurer des résultats cohérents à travers les clusters de calcul.
Executes interactive chat sessions using specific prompt templates and optional adapter weights.
MedicalGPT is an open-source framework for fine-tuning large language models, with a dedicated focus on adapting general models to the medical domain. It provides a complete pipeline that covers continued pretraining on domain-specific corpora, supervised instruction tuning, tokenizer vocabulary extension with medical terminology, and alignment to clinician preferences through direct preference optimization, reinforcement learning, or knowledge distillation. The framework also supports training models to invoke external tools and functions in multi-turn clinical conversations. The platform di
Loads the fine‑tuned model weights and supports interactive chat or batch text generation from a command‑line session within the framework
Aigcpanel is a visual workflow automation tool and model lifecycle manager designed for generative AI media pipelines. It provides a unified interface to install, launch, and configure both local and remote AI model endpoints, acting as an orchestration platform for large language models and AI tools. The system features a drag-and-drop node editor for chaining AI models and scripts into automated processing pipelines. It distinguishes itself with a breakpoint-aware execution model that allows users to pause and resume long media tasks from specific points in the workflow. Additionally, it in
Provides a command line interface for executing model functions and querying available models for script integration.
ExecuTorch is a lightweight C++ runtime for deploying PyTorch models on mobile, embedded, and edge hardware. It provides an ahead-of-time compilation pipeline that exports, quantizes, and lowers model graphs into compact serialized programs, then executes them through a minimal runtime with hardware acceleration and on-device large language model inference capabilities. The project distinguishes itself through a hardware accelerator delegate system that partitions model subgraphs and offloads computation to specialized backends including NPUs, GPUs, and DSPs from Apple, Arm, Intel, MediaTek,
ExecuTorch loads and executes a language model on-device, wrapping the runtime for text generation tasks.
Distributed-llama is a distributed inference engine and command line tool for running large language models across multiple networked machines. It functions as a compute cluster manager that coordinates worker nodes to share the computational load of a single model. The system utilizes tensor parallelism to shard model weights across different hosts, allowing the execution of models that exceed the memory capacity of a single piece of hardware. It includes a dedicated format converter to transform standard model files into a compatible binary layout optimized for distributed loading. The eng
Provides a command-line interface and server for interactive chat sessions and batch text generation.
Foundry-Local is a machine learning development tool designed to facilitate private, on-device inference and model management. It provides a local server environment that hosts machine learning models directly on the user's hardware, ensuring that all data processing, including prompt handling and audio transcription, remains within the local environment without requiring external cloud connectivity. The project distinguishes itself by automating the entire model lifecycle, including the discovery, downloading, and versioning of assets to maintain compatibility with host hardware. It features
Provides an interactive command-line interface for developers to test inference performance and verify model outputs directly.