2 Repos
Tools that process captured screen images to analyze visual content using recognition models.
Distinct from Screen Capture Tools: Focuses on the analysis of the captured image rather than just the utility of taking the screenshot.
Explore 2 awesome GitHub repositories matching development tools & productivity · Visual Analysis Tools. Refine with filters or upvote what's useful.
Douyin-Bot is a Python-based automation tool designed for interacting with Douyin accounts through automated likes, follows, and comments. It functions as a computer vision social bot that uses face recognition and image analysis to filter profiles based on visual criteria. The project distinguishes itself by using aesthetic content filtering to trigger social actions only when a user meets a specified beauty threshold. To reduce the risk of account bans, it incorporates account safety management that mimics human behavior through randomized delay scheduling. The framework covers a broad ran
Processes periodic screenshots using face recognition models to evaluate user appearance against defined thresholds.
AI0x0.com ist ein multimodaler KI-Desktop-Assistent und Cross-Application-Wrapper. Er bietet ein schwebendes Interface-Overlay, das Large Language Models in jede aktive Softwareanwendung integriert, um globale Abfragen und Textautomatisierung zu erleichtern. Das System zeichnet sich durch die Fähigkeit aus, Echtzeit-Bildschirmaufnahmen zur visuellen Analyse zu verarbeiten und eine Voice-Pipeline für freihändige Speech-to-Text- und Text-to-Speech-Interaktion zu nutzen. Es ermöglicht zudem die direkte KI-Content-Injektion durch Simulation von Tastatureingaben, um generierte Antworten in aktive Softwarefelder einzufügen. Das Projekt enthält eine RAG-Wissensdatenbank (Retrieval-Augmented Generation), die lokale Dokumentenbibliotheken und Echtzeit-Webdaten durchsucht. Es unterstützt eine Multi-Modell-Provider-Schnittstelle zum Wechseln zwischen verschiedenen KI-APIs, ein Plugin-System für Drittanbieter-Integrationen und Tools zur Generierung von Multimedia-Artikeln. Benutzer können benutzerdefinierte funktionale Presets verwalten und den Gesprächsverlauf für zukünftige Abrufe mit Lesezeichen versehen.
Captures real-time pixel data from the display to provide visual context for multimodal AI analysis.