LLM as a Chatbot Service
Kubernetes operator for local LLM inference with llama.cpp, vLLM, TGI, and mlx-server — multi-GPU NVIDIA Apple Silicon Metal, autoscaling, air-gapped, production-ready
Flowise is a low-code platform designed for building and deploying complex language model workflows through a visual, node-based interface. It functions as an orchestrator for autonomous multi-agent systems, allowing users to construct conversational pipelines by connecting language models, memory stores, and external tools on a drag-and-drop canvas. The platform distinguishes itself through its support for sophisticated agentic patterns, including supervisor-worker delegation and iterative reasoning strategies. Users can design directed acyclic graphs to manage conditional branching, state p
The Swiss Army Knife of Offline AI. Chat, Speak, and Generate Images - Privacy First, Zero Internet. Download an LLM and use it on your mobile device. No data ever leaves your phone. Supports text-to-text, vision, text-to-image
OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)
The main features of tensorchord/modelz-llm are: Model Serving Engines.
Open-source alternatives to tensorchord/modelz-llm include: deep-diver/alpaca-lora-serve — LLM as a Chatbot Service. defilantech/llmkube — Kubernetes operator for local LLM inference with llama.cpp, vLLM, TGI, and mlx-server — multi-GPU NVIDIA + Apple… flowiseai/flowise — Flowise is a low-code platform designed for building and deploying complex language model workflows through a visual,… fminference/flexgen — FlexGen is an inference engine for large language models designed for high-throughput execution on single or multiple… fujitsuresearch/onecompression — Python package for LLM compression. alichherawalla/off-grid-mobile-ai — The Swiss Army Knife of Offline AI. Chat, Speak, and Generate Images - Privacy First, Zero Internet. Download an LLM…