awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectAboutHow we rankPressMCP server
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
lightseekorg avatar

lightseekorg/tokenspeed

0
View on GitHub↗
1,511 stars·172 forks·Python·MIT·1 viewlightseek.org/blog/lightseek-tokenspeed.html↗

Tokenspeed

TokenSpeed is a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.

Features

  • Inference and Serving - Inference engine optimized for agentic workloads.
  • Inference Engines - Inference engine optimized for agentic workloads.
  • General Productivity Tools - High-performance inference engine for LLMs.

Star history

Star history chart for lightseekorg/tokenspeedStar history chart for lightseekorg/tokenspeed

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Tokenspeed

Similar open-source projects, ranked by how many features they share with Tokenspeed.
  • augustdev/enchantedAugustDev avatar

    AugustDev/enchanted

    5,967View on GitHub↗

    Enchanted is a privacy-focused, cross-platform chat frontend for interacting with self-hosted large language models on iOS and macOS. It serves as a native client for communicating with private model servers, specifically providing integration for the Ollama API. The application supports multimodal interactions, allowing users to combine text, image attachments, and voice prompts. It provides tools for local AI model management, including the ability to define persistent system prompts and switch between different models for specific tasks. The interface includes capabilities for rendering m

    Swift
    View on GitHub↗5,967
  • bentoml/openllmbentoml avatar

    bentoml/OpenLLM

    12,115View on GitHub↗

    OpenLLM is a framework for deploying, managing, and scaling open-source large language models

    Pythonbentomlfine-tuningllama
    View on GitHub↗12,115
  • alexrozanski/llamachatalexrozanski avatar

    alexrozanski/LlamaChat

    1,510View on GitHub↗

    Chat with your favourite LLaMA models in a native macOS app

    Swiftaillamallamacpp
    View on GitHub↗1,510
  • berriai/litellmBerriAI avatar

    BerriAI/litellm

    50,579View on GitHub↗

    LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model providers. It provides a standardized API interface that abstracts vendor-specific schemas, allowing developers to interact with diverse models through a single, consistent format. By acting as a central traffic management layer, it enables organizations to route, secure, and govern model interactions across multiple deployments. The platform distinguishes itself through its policy-driven architecture, which uses configuration-based routing to manage traffic distribution, load balanc

    Pythonai-gatewayanthropicazure-openai
    View on GitHub↗50,579
See all 30 alternatives to Tokenspeed→

Frequently asked questions

What does lightseekorg/tokenspeed do?

TokenSpeed is a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.

What are the main features of lightseekorg/tokenspeed?

The main features of lightseekorg/tokenspeed are: Inference and Serving, Inference Engines, General Productivity Tools.

What are some open-source alternatives to lightseekorg/tokenspeed?

Open-source alternatives to lightseekorg/tokenspeed include: bentoml/openllm — OpenLLM is a framework for deploying, managing, and scaling open-source large language models. bigai-nlco/tokenswift. augustdev/enchanted — Enchanted is a privacy-focused, cross-platform chat frontend for interacting with self-hosted large language models on… alexrozanski/llamachat — Chat with your favourite LLaMA models in a native macOS app. berriai/litellm — LiteLLM is a unified gateway and proxy server designed to centralize access to over one hundred language model… bklieger-groq/g1.