awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
NVIDIA avatar

NVIDIA/cosmos

0
View on GitHub↗
10,494 نجوم·702 تفرعات·Jupyter Notebook·5 مشاهداتwww.nvidia.com/en-us/ai/cosmos↗

Cosmos

Cosmos is an open platform of world models, datasets, and tools for building physical AI systems such as robots and autonomous vehicles. It provides video generation and video understanding models that can generate synthetic videos and world simulations from text, image, video, or action inputs, and analyze videos to produce captions, event timestamps, spatial bounding boxes, and next-action predictions.

The platform includes a world simulation generator that produces images, videos, synchronized audio, and action-conditioned rollouts for synthetic data, alongside a visual content analyzer that extracts structured text outputs for robotics and autonomous systems. Cosmos offers a model fine-tuning framework with checkpoint-based recipes to adapt pre-trained models on custom video, action, or reasoning datasets, and exposes its reasoner and generator models through a standard OpenAI-compatible chat-completions API endpoint for production inference.

The platform's architecture combines multi-modal encoder fusion, a video tokenizer with diffusion backbone, causal transformer reasoning, and synchronized audio-visual generation. It supports application domains including autonomous vehicle development, robotics training data generation, physical AI world simulation, visual content understanding, and production model serving.

Features

  • Physical AI World Generators - An open platform of world models, datasets, and tools for building physical AI systems like robots and autonomous vehicles.
  • Simulation Data Generators - Creating simulated driving scenarios and analyzing visual data to train perception and planning models for self-driving cars.
  • Cross-Attention Fusion Layers - Combines text, image, video, and action inputs into a unified latent space using cross-attention layers for flexible conditioning.
  • Video Tokenizers - Converts raw video frames into discrete latent tokens and reconstructs them using a diffusion-based decoder for high-fidelity generation.
  • Video Content Analyzers - Analyzes images and videos to produce text outputs such as captions, event timestamps, spatial bounding boxes, and next-action predictions for robotics and autonomous systems.
  • OpenAI-Compatible Model Servers - Exposes reasoner and generator models behind an OpenAI-compatible API endpoint for production inference.
  • Synthetic Data Generators - Producing diverse synthetic video and action sequences to train robot manipulation and navigation policies without real-world data.
  • World Simulation Generators - Generates synthetic videos and world simulations from text, image, video, or action inputs for training and testing.
  • World Simulation Generators - Generates images, videos, synchronized sound, and action-conditioned rollouts from text, image, video, or action inputs for world simulation and synthetic data.
  • Video Understanding Models - Analyzes videos to produce captions, event timestamps, spatial bounding boxes, and next-action predictions.
  • Visual Content Analyzers - Analyzing images and video to extract captions, event timestamps, spatial bounding boxes, and next-action predictions for automation.
  • Video Token Reasoners - Processes video tokens autoregressively to output structured text predictions like captions, bounding boxes, and next actions.
  • World Model Rollout Pipelines - Feeds action sequences as conditioning signals into the world model to generate temporally consistent future frames.
  • Action-Conditioned Fine-Tuning - Adapting pre-trained world models on proprietary video or action datasets using supervised fine-tuning recipes for specialized behavior.
  • Checkpoint-Based Fine-Tuning Recipes - Provides supervised fine-tuning recipes to adapt pre-trained checkpoints on custom video, action, or reasoning datasets.
  • World Model Fine-Tuning - Adapts pre-trained checkpoints on custom video, action, or reasoning datasets using supervised fine-tuning recipes for task-specific behavior.
  • Inference API Servers - Exposes reasoner and generator models behind a standard chat-completions endpoint for production inference.
  • Video Checkpoint Fine-Tuning - Provides pre-trained model weights and supervised fine-tuning scripts that adapt the backbone to custom video or action datasets.
  • Model Serving Endpoints - Exposing reasoner and generator models behind a standard chat-completions API endpoint for scalable inference deployment.
  • Generative - Generates temporally aligned audio tracks alongside video frames using a shared latent representation and joint decoder.
  • Foundation Models - NVIDIA's foundation model for world simulation and video generation.
  • World Models and Spatial AI - World foundation model platform for physical AI and robotics.

سجل النجوم

مخطط تاريخ النجوم لـ nvidia/cosmosمخطط تاريخ النجوم لـ nvidia/cosmos

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة nvidia/cosmos؟

Cosmos is an open platform of world models, datasets, and tools for building physical AI systems such as robots and autonomous vehicles. It provides video generation and video understanding models that can generate synthetic videos and world simulations from text, image, video, or action inputs, and analyze videos to produce captions, event timestamps, spatial bounding boxes, and next-action predictions.

ما هي الميزات الرئيسية لـ nvidia/cosmos؟

الميزات الرئيسية لـ nvidia/cosmos هي: Physical AI World Generators, Simulation Data Generators, Cross-Attention Fusion Layers, Video Tokenizers, Video Content Analyzers, OpenAI-Compatible Model Servers, Synthetic Data Generators, World Simulation Generators.

ما هي البدائل مفتوحة المصدر لـ nvidia/cosmos؟

تشمل البدائل مفتوحة المصدر لـ nvidia/cosmos: llava-vl/llava-next — LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images… paddlepaddle/fastdeploy — FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models.… hao-ai-lab/fastvideo — FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine,… thudm/chatglm2-6b — ChatGLM2-6B is an open-weight large language model designed for natural language conversations and text generation in… nvidia/isaac-gr00t. xusenlinzy/api-for-open-llm — This project provides a unified server environment and gateway for hosting and executing open-source large language…

بدائل مفتوحة المصدر لـ Cosmos

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Cosmos.
  • llava-vl/llava-nextالصورة الرمزية لـ LLaVA-VL

    LLaVA-VL/LLaVA-NeXT

    4,695عرض على GitHub↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    عرض على GitHub↗4,695
  • paddlepaddle/fastdeployالصورة الرمزية لـ PaddlePaddle

    PaddlePaddle/FastDeploy

    3,700عرض على GitHub↗

    FastDeploy is a high-performance deployment framework for large language models, vision models, and multimodal models. It provides the infrastructure to launch model services that process combined image, video, and text inputs, exposing these capabilities through a standardized, OpenAI-compatible API for chat and text completions. The project distinguishes itself through advanced inference pipeline engineering and GPU optimization. It employs speculative decoding, tensor parallelism, and a disaggregated execution model that separates prefill and decode phases across different hardware resourc

    Pythonernieernie-45ernie-45-vl
    عرض على GitHub↗3,700
  • hao-ai-lab/fastvideoالصورة الرمزية لـ hao-ai-lab

    hao-ai-lab/FastVideo

    3,743عرض على GitHub↗

    FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a

    Pythondiffusersdiffusion-modelsdistillation
    عرض على GitHub↗3,743
  • thudm/chatglm2-6bالصورة الرمزية لـ THUDM

    THUDM/ChatGLM2-6B

    15,565عرض على GitHub↗

    ChatGLM2-6B is an open-weight large language model designed for natural language conversations and text generation in both English and Chinese. It functions as a bilingual chat model capable of processing and maintaining coherence across text sequences up to 32K tokens. The model is optimized for local deployment through precision quantization, which reduces memory requirements to allow execution on consumer-grade hardware. It supports distributing model weights across multiple graphics cards to handle parameters that exceed the memory of a single device. The project covers capabilities for

    Python
    عرض على GitHub↗15,565
  • عرض جميع البدائل الـ 30 لـ Cosmos→