awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
cliport avatar

cliport/cliport

0
View on GitHub↗
545 نجوم·95 تفرعات·Jupyter Notebook·Apache-2.0·3 مشاهداتcliport.github.io↗

Cliport

CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

Features

  • Manipulation and Control - What and where pathways for robotic manipulation.
  • Multimodal Agents - Vision-language pathways for robotic manipulation tasks.

سجل النجوم

مخطط تاريخ النجوم لـ cliport/cliportمخطط تاريخ النجوم لـ cliport/cliport

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة cliport/cliport؟

CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021

ما هي الميزات الرئيسية لـ cliport/cliport؟

الميزات الرئيسية لـ cliport/cliport هي: Manipulation and Control, Multimodal Agents.

ما هي البدائل مفتوحة المصدر لـ cliport/cliport؟

تشمل البدائل مفتوحة المصدر لـ cliport/cliport: peract/peract — Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2022. simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… qwenlm/qwen3-omni — Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video… allenai/unified-io-2 — This repo contains code for Unified-IO 2, including code to run a demo, do training, and do inference. This codebase…

بدائل مفتوحة المصدر لـ Cliport

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Cliport.
  • peract/peractP

    peract/peract

    0عرض على GitHub↗

    Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2022

    عرض على GitHub↗0
  • othersideai/self-operating-computerالصورة الرمزية لـ OthersideAI

    OthersideAI/self-operating-computer

    10,153عرض على GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    عرض على GitHub↗10,153
  • 11cafe/jaazالصورة الرمزية لـ 11cafe

    11cafe/jaaz

    6,384عرض على GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    عرض على GitHub↗6,384
  • qwenlm/qwen3-omniالصورة الرمزية لـ QwenLM

    QwenLM/Qwen3-Omni

    3,843عرض على GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    عرض على GitHub↗3,843
عرض جميع البدائل الـ 30 لـ Cliport→