awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
cbh123 avatar

cbh123/narrator

0
View on GitHub↗
4,423 Stars·540 Forks·Python·2 Aufrufe

Narrator

Narrator ist ein System der künstlichen Intelligenz, das Echtzeit-Video-Feeds in natürlichsprachliche Audiobeschreibungen umwandelt. Es fungiert als multimodaler Vision-Narrator und Szenenbeschreiber, der Computer Vision nutzt, um Umgebungsdaten von einer Kamera in synthetische Sprache zu transformieren.

Das Tool arbeitet als Pipeline, die periodisch Bilder aus einem Feed erfasst und ein multimodales Large Language Model verwendet, um visuelle Ereignisse zu analysieren. Diese Analysen werden dann mittels Text-to-Speech-Synthese in ein Voiceover umgewandelt, das reale Aktivitäten und die Umgebung beschreibt.

Das System unterstützt die automatisierte Umgebungsüberwachung und visuelle Unterstützung durch das Sampling von Kamerabildern und die Generierung gesprochener Beschreibungen der aktuellen Umgebung des Nutzers.

Features

  • Synthetic Narrations - Generates synthetic spoken audio descriptions of live visual events based on AI analysis of camera feeds.
  • Visual Assistance Tools - Provides a comprehensive visual assistance system that transforms live camera feeds into real-time auditory narration using AI.
  • Real-Time Environmental Narration - Transforms live visual data into a natural sounding voiceover that describes real-world activities.
  • AI Scene Descriptors - Functions as an AI-powered system that converts real-time video into natural language descriptions.
  • Multimodal AI Toolkits - Integrates computer vision and text-to-speech to create a multimodal live audio voiceover.
  • Multimodal Analysis Tools - Employs multimodal large language models to interpret visual scenes and generate natural language descriptions.
  • Multimodal Vision Interfaces - Utilizes a multimodal LLM interface to process camera frames and generate auditory descriptions.
  • Real-Time Scene Description - Turns live camera footage into spoken audio descriptions of the user's current physical environment.
  • Text-to-Speech Synthesis - Converts the AI-generated textual descriptions of the environment into spoken audio narration.
  • Computer Vision and Audio - Combines computer vision for scene analysis with audio synthesis for real-time narration.
  • Automated Visual Monitoring - Provides automated monitoring of a physical location by capturing images and generating activity descriptions.
  • Proactive Visual Assistance - Acts as a visual aid by converting environmental visual events into spoken narration for users.
  • PyTorch Computer Vision Pipelines - Implements a complete pipeline that captures images and uses AI to describe physical activities.

Star-Verlauf

Star-Verlauf für cbh123/narratorStar-Verlauf für cbh123/narrator

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Häufig gestellte Fragen

Was macht cbh123/narrator?

Narrator ist ein System der künstlichen Intelligenz, das Echtzeit-Video-Feeds in natürlichsprachliche Audiobeschreibungen umwandelt. Es fungiert als multimodaler Vision-Narrator und Szenenbeschreiber, der Computer Vision nutzt, um Umgebungsdaten von einer Kamera in synthetische Sprache zu transformieren.

Was sind die Hauptfunktionen von cbh123/narrator?

Die Hauptfunktionen von cbh123/narrator sind: Synthetic Narrations, Visual Assistance Tools, Real-Time Environmental Narration, AI Scene Descriptors, Multimodal AI Toolkits, Multimodal Analysis Tools, Multimodal Vision Interfaces, Real-Time Scene Description.

Welche Open-Source-Alternativen gibt es zu cbh123/narrator?

Open-Source-Alternativen zu cbh123/narrator sind unter anderem: dsdanielpark/bard-api — Bard-API is an asynchronous Python wrapper and client for interacting with Google Gemini. It functions as a stateful… openbmb/minicpm-v — MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic… ngxson/smolvlm-realtime-webcam — This is a webcam-based client for a local llama.cpp server that enables real-time object detection and vision-language… jianchang512/chattts-ui — ChatTTS-ui is a web-based interface and API wrapper for the ChatTTS model, designed to convert written text and mixed… idea-research/grounded-segment-anything — Grounded-Segment-Anything is a suite of specialized tools for multimodal visual analysis, text-based segmentation, and… bytedance/ui-tars — UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent…

Open-Source-Alternativen zu Narrator

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Narrator.
  • dsdanielpark/bard-apiAvatar von dsdanielpark

    dsdanielpark/Bard-API

    5,196Auf GitHub ansehen↗

    Bard-API is an asynchronous Python wrapper and client for interacting with Google Gemini. It functions as a stateful conversation manager and multimodal interface, allowing users to send text and image prompts to a language model and retrieve responses. The library utilizes a cookie-based authentication system that extracts session tokens from local browser storage to authorize requests. To manage access and connectivity, it includes proxy-based request routing to bypass regional restrictions and avoid IP blocks. The project covers capabilities for multimodal AI analysis and the maintenance

    Pythonai-apiapibard
    Auf GitHub ansehen↗5,196
  • openbmb/minicpm-vAvatar von OpenBMB

    OpenBMB/MiniCPM-V

    25,653Auf GitHub ansehen↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Python
    Auf GitHub ansehen↗25,653
  • ngxson/smolvlm-realtime-webcamAvatar von ngxson

    ngxson/smolvlm-realtime-webcam

    5,560Auf GitHub ansehen↗

    This is a webcam-based client for a local llama.cpp server that enables real-time object detection and vision-language model inference directly from a browser. It captures frames from the user's webcam at configurable intervals and sends them to a locally running inference server for analysis, displaying both detection results and textual scene descriptions as they are produced. The application distinguishes itself by combining object detection with vision-language scene description in a single real-time interface, all processed through a local llama.cpp server for private, offline operation.

    HTML
    Auf GitHub ansehen↗5,560
  • elevenlabs/elevenlabs-pythonAvatar von elevenlabs

    elevenlabs/elevenlabs-python

    2,873Auf GitHub ansehen↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    Auf GitHub ansehen↗2,873
  • Alle 30 Alternativen zu Narrator anzeigen→