1 repository
Deployment and execution of models integrating text, audio, image, and video across diverse hardware accelerators.
Distinct from Model Deployments: Candidates are too narrow (SBCs or Computer Vision) or too broad (General Cloud), missing the omni-modal aspect
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Omni-Modal Model Deployment. Refine with filters or upvote what's useful.
vllm-omni is a high-throughput serving engine and distributed inference framework designed for omni-modal models. It serves as a multi-modal model API server capable of generating text, image, video, and audio data, providing a standardized interface for remote client access. The system features a non-autoregressive generation engine for parallel media production and a robot policy inference server that acts as a real-time communication bridge to robotic hardware using specialized protocols. It supports hybrid execution models that combine sequential token generation with parallelized media g
Runs models that integrate text, audio, image, and video across various hardware accelerators.