awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
cbh123 avatar

cbh123/narrator

0
View on GitHub↗
4,423 نجوم·540 تفرعات·Python·6 مشاهدات

Narrator

Narrator هو نظام ذكاء اصطناعي يحول تدفقات الفيديو في الوقت الفعلي إلى أوصاف صوتية باللغة الطبيعية. يعمل كراوٍ بصري متعدد الوسائط وواصف للمشاهد، مستخدماً الرؤية الحاسوبية لتحويل البيانات البيئية من الكاميرا إلى كلام اصطناعي.

تعمل الأداة كخط معالجة يلتقط صوراً دورية من تدفق الفيديو ويستخدم نموذجاً لغوياً كبيراً متعدد الوسائط لتحليل الأحداث البصرية. يتم بعد ذلك تحويل هذه التحليلات عبر تقنية تحويل النص إلى كلام إلى تعليق صوتي يصف الأنشطة والمحيط في العالم الحقيقي.

يدعم النظام مراقبة البيئة التلقائية والمساعدة البصرية عن طريق أخذ عينات من إطارات الكاميرا وتوليد أوصاف منطوقة لبيئة المستخدم الحالية.

Features

  • Synthetic Narrations - Generates synthetic spoken audio descriptions of live visual events based on AI analysis of camera feeds.
  • Visual Assistance Tools - Provides a comprehensive visual assistance system that transforms live camera feeds into real-time auditory narration using AI.
  • Real-Time Environmental Narration - Transforms live visual data into a natural sounding voiceover that describes real-world activities.
  • AI Scene Descriptors - Functions as an AI-powered system that converts real-time video into natural language descriptions.
  • Multimodal AI Toolkits - Integrates computer vision and text-to-speech to create a multimodal live audio voiceover.
  • Multimodal Analysis Tools - Employs multimodal large language models to interpret visual scenes and generate natural language descriptions.
  • Multimodal Vision Interfaces - Utilizes a multimodal LLM interface to process camera frames and generate auditory descriptions.
  • Real-Time Scene Description - Turns live camera footage into spoken audio descriptions of the user's current physical environment.
  • Text-to-Speech Synthesis - Converts the AI-generated textual descriptions of the environment into spoken audio narration.
  • Computer Vision and Audio - Combines computer vision for scene analysis with audio synthesis for real-time narration.
  • Automated Visual Monitoring - Provides automated monitoring of a physical location by capturing images and generating activity descriptions.
  • Proactive Visual Assistance - Acts as a visual aid by converting environmental visual events into spoken narration for users.
  • PyTorch Computer Vision Pipelines - Implements a complete pipeline that captures images and uses AI to describe physical activities.

سجل النجوم

مخطط تاريخ النجوم لـ cbh123/narratorمخطط تاريخ النجوم لـ cbh123/narrator

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة cbh123/narrator؟

Narrator هو نظام ذكاء اصطناعي يحول تدفقات الفيديو في الوقت الفعلي إلى أوصاف صوتية باللغة الطبيعية. يعمل كراوٍ بصري متعدد الوسائط وواصف للمشاهد، مستخدماً الرؤية الحاسوبية لتحويل البيانات البيئية من الكاميرا إلى كلام اصطناعي.

ما هي الميزات الرئيسية لـ cbh123/narrator؟

الميزات الرئيسية لـ cbh123/narrator هي: Synthetic Narrations, Visual Assistance Tools, Real-Time Environmental Narration, AI Scene Descriptors, Multimodal AI Toolkits, Multimodal Analysis Tools, Multimodal Vision Interfaces, Real-Time Scene Description.

ما هي البدائل مفتوحة المصدر لـ cbh123/narrator؟

تشمل البدائل مفتوحة المصدر لـ cbh123/narrator: dsdanielpark/bard-api — Bard-API is an asynchronous Python wrapper and client for interacting with Google Gemini. It functions as a stateful… openbmb/minicpm-v — MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic… ngxson/smolvlm-realtime-webcam — This is a webcam-based client for a local llama.cpp server that enables real-time object detection and vision-language… jianchang512/chattts-ui — ChatTTS-ui is a web-based interface and API wrapper for the ChatTTS model, designed to convert written text and mixed… idea-research/grounded-segment-anything — Grounded-Segment-Anything is a suite of specialized tools for multimodal visual analysis, text-based segmentation, and… bytedance/ui-tars — UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent…

بدائل مفتوحة المصدر لـ Narrator

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Narrator.
  • dsdanielpark/bard-apiالصورة الرمزية لـ dsdanielpark

    dsdanielpark/Bard-API

    5,196عرض على GitHub↗

    Bard-API is an asynchronous Python wrapper and client for interacting with Google Gemini. It functions as a stateful conversation manager and multimodal interface, allowing users to send text and image prompts to a language model and retrieve responses. The library utilizes a cookie-based authentication system that extracts session tokens from local browser storage to authorize requests. To manage access and connectivity, it includes proxy-based request routing to bypass regional restrictions and avoid IP blocks. The project covers capabilities for multimodal AI analysis and the maintenance

    Pythonai-apiapibard
    عرض على GitHub↗5,196
  • openbmb/minicpm-vالصورة الرمزية لـ OpenBMB

    OpenBMB/MiniCPM-V

    25,653عرض على GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Python
    عرض على GitHub↗25,653
  • ngxson/smolvlm-realtime-webcamالصورة الرمزية لـ ngxson

    ngxson/smolvlm-realtime-webcam

    5,560عرض على GitHub↗

    This is a webcam-based client for a local llama.cpp server that enables real-time object detection and vision-language model inference directly from a browser. It captures frames from the user's webcam at configurable intervals and sends them to a locally running inference server for analysis, displaying both detection results and textual scene descriptions as they are produced. The application distinguishes itself by combining object detection with vision-language scene description in a single real-time interface, all processed through a local llama.cpp server for private, offline operation.

    HTML
    عرض على GitHub↗5,560
  • elevenlabs/elevenlabs-pythonالصورة الرمزية لـ elevenlabs

    elevenlabs/elevenlabs-python

    2,873عرض على GitHub↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    عرض على GitHub↗2,873
  • عرض جميع البدائل الـ 30 لـ Narrator→