1 个仓库
Generating a sequence of interaction steps based on visual analysis of user interfaces.
Distinct from Action Plan Generators: Focuses specifically on the visual analysis of screenshots to plan GUI interactions, rather than general agent action plans.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · GUI Action Planning. Refine with filters or upvote what's useful.
CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system
Analyzes user interface screenshots to generate a logical sequence of interaction steps and coordinates for automation.