awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to pickle-com/glass

Open-source alternatives to Glass

30 open-source projects similar to pickle-com/glass, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Glass alternative.

  • mg-chao/snow-shotmg-chao 的头像

    mg-chao/snow-shot

    4,118在 GitHub 上查看↗

    This project is an AI-powered screenshot manager and visual assistant designed for capturing screen content and processing it through large language models. It functions as an OCR translation application and screen annotation tool, allowing users to extract text from images and perform intelligent analysis of visual data. The software differentiates itself through an AI-driven OCR pipeline and the ability to convert screenshots into structured Markdown or HTML via layout-aware document transformation. It features a visual AI assistant capable of analyzing screen content and a prompt-engineere

    TypeScriptchatbotocrscreen-capture
    在 GitHub 上查看↗4,118
  • aardio/imtipaardio 的头像

    aardio/ImTip

    2,580在 GitHub 上查看↗

    ImTip is a desktop utility that provides input method visualization and large language model integration. It renders real-time indicators at the system text cursor to track language, layout, and punctuation settings. The project connects large language model APIs to a desktop interface to render rich content, including mathematical formulas and syntax-highlighted code. It uses a scriptable logic engine and global hooks to map programmable hotkeys to complex task sequences and external API calls. The software includes tools for desktop UI customization, allowing users to adjust the style, col

    aardioimeinput-method
    在 GitHub 上查看↗2,580
  • soffes/hotkeysoffes 的头像

    soffes/HotKey

    1,077在 GitHub 上查看↗

    HotKey is a developer library for registering system-wide keyboard shortcuts in macOS applications. It binds modifier keys and key codes to system-wide identifiers using native operating system event hooks, allowing applications to respond to key combinations while running in the background. The library executes custom closures asynchronously when native keyboard monitor notifications match registered global shortcut signatures. It handles both press and release events, and automatically unregisters system hot key bindings during object deallocation to prevent stale event listeners and memory

    Swiftappkitcarboncarthage
    在 GitHub 上查看↗1,077

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • damascenorafael/reminders-menubarDamascenoRafael 的头像

    DamascenoRafael/reminders-menubar

    3,790在 GitHub 上查看↗

    This is a native macOS utility that serves as a menu bar client for Apple Reminders. It synchronizes with the system task database to allow users to view, create, and manage reminders directly from the system tray. The application is distinguished by a natural language task processor that extracts due dates, lists, and tags from unstructured text to automate reminder creation. It also features a global hotkey system, enabling the task interface to be triggered instantly via a keyboard shortcut from any active window. The tool provides broader productivity capabilities including task search a

    Swiftapple-remindersmacosmacos-menubar
    在 GitHub 上查看↗3,790
  • suitedaces/computer-agentsuitedaces 的头像

    suitedaces/computer-agent

    583在 GitHub 上查看↗

    This project is an autonomous desktop automation agent that interprets natural language instructions to control applications, browser interfaces, and system terminals. It functions as a cross-platform utility designed to manage complex workflows by integrating visual screen analysis with system-level input simulation. The agent distinguishes itself through its ability to perform tasks asynchronously, ensuring that web and terminal operations run in the background without interrupting the active user session or desktop focus. By combining computer vision to map interface elements with event-dr

    Rustaiai-toolsanthropic
    在 GitHub 上查看↗583
  • steveseguin/vdo.ninjasteveseguin 的头像

    steveseguin/vdo.ninja

    3,910在 GitHub 上查看↗

    VDO.Ninja is a low-latency peer-to-peer media routing service and video streaming platform designed to integrate remote audio and video feeds into professional production workflows. It functions as a WebRTC broadcast integration tool and studio controller, allowing for the direct transmission of high-definition media between publishers and viewers with minimal delay. The platform distinguishes itself through extensive protocol bridging, converting between WebRTC, WHIP, WHEP, SRT, and RTMP to ensure compatibility across diverse network environments and professional studio software. It includes

    JavaScriptlivelow-latencyninja
    在 GitHub 上查看↗3,910
  • shinyflvre/mate-engineshinyflvre 的头像

    shinyflvre/Mate-Engine

    2,809在 GitHub 上查看↗

    Mate-Engine is a 3D desktop avatar engine designed to render interactive characters that float on the desktop and react to system events. It functions as a virtual assistant platform, combining an interactive character renderer with an interface that connects local language models to 3D avatars for desktop conversations. The engine features a custom 3D model loader that imports third-party humanoid models using standard rigging and bone naming conventions. It includes an audio-reactive animation system that monitors live audio output from applications to automatically trigger dance sequences

    ShaderLabanimedesktopdesktop-mate
    在 GitHub 上查看↗2,809
  • winfunc/opcodewinfunc 的头像

    winfunc/opcode

    22,083在 GitHub 上查看↗

    Opcode is a desktop interface designed for managing AI-assisted software development workflows. It provides a centralized workspace to organize interactive programming sessions, configure specialized automated agents, and maintain oversight of development tasks through a visual environment. The platform distinguishes itself by integrating version control for AI conversations, allowing developers to create checkpoints and branches to navigate, compare, and revert between different interaction states. It also functions as a client for standardized context protocols, enabling the connection of e

    TypeScriptanthropicanthropic-claudeclaude
    在 GitHub 上查看↗22,083
  • sdegutis/hydrasdegutis 的头像

    sdegutis/hydra

    5,220在 GitHub 上查看↗

    Hydra is a scriptable productivity platform for macOS that functions as a global hotkey manager, window manager, and automation tool. It allows users to extend system functionality by executing scripts and custom functions to automate repetitive operating system tasks. The system enables the binding of specific keyboard shortcuts to trigger scripts from any active application. It provides capabilities to inspect and modify the state of running application windows to organize the desktop workspace. The platform supports a module-based extension model, allowing for the installation of third-pa

    C
    在 GitHub 上查看↗5,220
  • mushan0x0/ai0x0.commushan0x0 的头像

    mushan0x0/AI0x0.com

    3,945在 GitHub 上查看↗

    AI0x0.com is a multimodal AI desktop assistant and cross-application wrapper. It provides a floating interface overlay that integrates large language models into any active software application to facilitate global querying and text automation. The system distinguishes itself through the ability to process real-time screen captures for visual analysis and utilize a voice pipeline for hands-free speech-to-text and text-to-speech interaction. It further enables direct AI content injection by simulating keyboard input to insert generated responses into active software fields. The project includ

    在 GitHub 上查看↗3,945
  • lihaoyun6/quickrecorderlihaoyun6 的头像

    lihaoyun6/QuickRecorder

    7,969在 GitHub 上查看↗

    QuickRecorder is a screen recording software designed for capturing desktops, application windows, and system audio. It functions as a multi-device video recorder and tutorial capture tool, synchronizing video feeds from a computer and connected mobile devices into a single stream. The system distinguishes itself through an alpha-channel video exporter that produces recordings with transparent backgrounds. It also includes a presenter overlay system that renders a floating camera feed over screen captures and a specialized tutorial toolset that provides mouse movement highlighting and a magni

    Swift
    在 GitHub 上查看↗7,969
  • siddharthvaddem/openscreensiddharthvaddem 的头像

    siddharthvaddem/openscreen

    7,282在 GitHub 上查看↗

    OpenScreen is screen recording and editing software used to capture screen video and audio with an integrated timeline for trimming, cropping, and playback adjustments. It functions as a comprehensive system for recording, annotating, and exporting audio-visual content. The project includes a dynamic zoom editor for applying manual or cursor-following zooms with adjustable depth and easing. It features a local caption generator that uses on-device transcription to create voiceover captions without uploading data to external servers. Additional specialized tools allow for the integration of we

    TypeScriptelectronopen-sourcepixijs
    在 GitHub 上查看↗7,282
  • thudm/cogvlmTHUDM 的头像

    THUDM/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Python
    在 GitHub 上查看↗6,742
  • end-4/dots-hyprlandend-4 的头像

    end-4/dots-hyprland

    12,857在 GitHub 上查看↗

    This project is a configuration suite for the Hyprland Wayland compositor, providing a set of automated scripts and files to deploy a consistent desktop environment across Linux distributions. It functions as an automation tool that synchronizes system settings, software packages, and interface themes to ensure a uniform workspace state. The environment distinguishes itself through deep integration with language models, allowing users to access local or cloud-based AI assistants directly from the desktop interface for tasks such as text translation and content generation. Visual consistency i

    QMLdotfileshyprlandlinux
    在 GitHub 上查看↗12,857
  • deepseek-ai/janusdeepseek-ai 的头像

    deepseek-ai/Janus

    17,746在 GitHub 上查看↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    在 GitHub 上查看↗17,746
  • seadve/koohaSeaDve 的头像

    SeaDve/Kooha

    3,262在 GitHub 上查看↗

    Kooha is a screen recorder for Linux desktops that utilizes the Wayland protocol and XDG Portals for secure recording. It functions as a hardware-accelerated screen capture tool that offloads video compression to the GPU to reduce CPU load and power consumption. The application integrates the PipeWire framework to capture system and microphone audio streams and leverages FFmpeg for muxing video streams and exporting various codecs and containers. Its user interface is a native Linux application built with the GTK toolkit. The software covers screen recording and capture of entire displays, s

    Rustgnomegstreamergtk-rs
    在 GitHub 上查看↗3,262
  • basedhardware/omiBasedHardware 的头像

    BasedHardware/omi

    12,869在 GitHub 上查看↗

    Omi is an open-source wearable AI platform that captures audio and screen data to provide real-time conversational assistance and memory. It integrates a wearable hardware development kit with a vector memory database and large language model capabilities to create a persistent digital record of user interactions. The platform is distinguished by its BLE audio streaming pipeline, which transmits raw audio from wearable hardware for real-time transcription and speaker identification. It utilizes a plugin-based agent tool framework that allows AI assistants to autonomously invoke custom functio

    Dartaiappbci
    在 GitHub 上查看↗12,869
  • paddlepaddle/erniePaddlePaddle 的头像

    PaddlePaddle/ERNIE

    7,717在 GitHub 上查看↗

    ERNIE is a development toolkit for training, fine-tuning, and deploying large language models built on the PaddlePaddle deep learning platform. It provides a comprehensive suite of core components, including an inference server for vision and language models, a training and fine-tuning toolkit, and a framework for building retrieval-augmented generation systems using private knowledge bases. The project features multimodal AI models capable of reasoning across text, images, and video to perform complex visual understanding and information extraction. It distinguishes itself through specialize

    Pythonernieernie-45ernie-45-vl
    在 GitHub 上查看↗7,717
  • rdp/screen-capture-recorder-to-video-windows-freerdp 的头像

    rdp/screen-capture-recorder-to-video-windows-free

    2,288在 GitHub 上查看↗

    This project is an open-source screen and audio capture utility for Windows that functions as a virtual media device. It exposes desktop activity and system sounds as standard input sources, allowing external software to detect and process the data stream as if it were coming from a physical camera or microphone. The tool utilizes a filter architecture to integrate with the Windows multimedia ecosystem, enabling third-party applications to access desktop feeds for recording, broadcasting, or remote transmission. By capturing pixel data directly from the screen buffer and hooking into system a

    C++
    在 GitHub 上查看↗2,288
  • autohotkey/autohotkeyAutoHotkey 的头像

    AutoHotkey/AutoHotkey

    12,601在 GitHub 上查看↗

    AutoHotkey is a Windows automation scripting language and task automator. It serves as a keyboard macro engine and custom hotkey manager designed to map specific key combinations to scripted actions. The project provides a domain-specific language for automating repetitive tasks and operating system functions. It enables the creation of keyboard shortcuts and macros to replace manual input and streamline digital workflows on Windows. The system covers window and process management, virtual input simulation, and the interception of keyboard and mouse input via operating system hooks. It furth

    C++autohotkeyautomationc-plus-plus
    在 GitHub 上查看↗12,601
  • modelcontextprotocol/modelcontextprotocolmodelcontextprotocol 的头像

    modelcontextprotocol/modelcontextprotocol

    8,458在 GitHub 上查看↗

    Model Context Protocol is a standardized framework for connecting large language models to external data sources and executable tools. It enables the creation of a universal interface where servers expose tools, resources, and prompts that can be discovered and utilized by various AI clients. The protocol utilizes a JSON-RPC message system that is transport-agnostic, supporting both standard input/output for local processes and HTTP with server-sent events for remote connections. It emphasizes security and control by delegating model sampling to the client to keep API keys secure from servers

    TypeScript
    在 GitHub 上查看↗8,458
  • mjolnirapp/mjolnirmjolnirapp 的头像

    mjolnirapp/mjolnir

    5,220在 GitHub 上查看↗

    Mjolnir is a macOS automation framework and extensible scripting engine. It provides a system for creating custom productivity workflows, managing application states, and controlling the macOS desktop interface programmatically. The project functions as a global hotkey manager that binds keyboard shortcuts to trigger automated scripts across the operating system. It includes a macOS application controller to inspect active windows and manage system-wide user interface interactions. The environment supports extensibility through a pluggable package management system, allowing for the installa

    C
    在 GitHub 上查看↗5,220
  • chenfei-wu/taskmatrixchenfei-wu 的头像

    chenfei-wu/TaskMatrix

    34,082在 GitHub 上查看↗

    TaskMatrix is a multimodal AI chat interface and visual task orchestrator. It combines language models with visual recognition to enable the exchange, analysis, and modification of images within a conversational environment. The system coordinates multiple foundation models through orchestration pipelines that chain language, detection, and segmentation models. This allows for complex visual operations, such as using text instructions to guide image masking and executing modular inpainting workflows to edit specific image regions. The project includes a computer vision toolset for object det

    Python
    在 GitHub 上查看↗34,082
  • microsoft/visual-chatgptmicrosoft 的头像

    microsoft/visual-chatgpt

    34,079在 GitHub 上查看↗

    Visual-ChatGPT is a visual orchestration framework and multimodal AI pipeline designed to coordinate large language models with visual foundation models. It functions as an integration layer that enables the exchange of text and images between different AI models to automate image analysis and editing tasks without requiring additional model training. The system differentiates itself through model-chain orchestration and prompt-based task dispatching, allowing natural language instructions to trigger specific vision models or tools. It utilizes coordinate-based region mapping and iterative ma

    Python
    在 GitHub 上查看↗34,079
  • open-llm-vtuber/open-llm-vtuberOpen-LLM-VTuber 的头像

    Open-LLM-VTuber/Open-LLM-VTuber

    5,946在 GitHub 上查看↗
    Pythonaiai-companionai-vtuber
    在 GitHub 上查看↗5,946
  • eczarny/spectacleeczarny 的头像

    eczarny/spectacle

    13,631在 GitHub 上查看↗

    Spectacle is a keyboard-driven window manager and organizer that uses system accessibility frameworks to manipulate window coordinates and dimensions. It allows for the arrangement, resizing, and movement of application windows across multiple displays using global keyboard shortcuts. The tool focuses on multi-monitor layout management, enabling users to shift active windows between connected displays and snap windows into predefined screen regions such as halves, thirds, or corners. It also provides the ability to center and maximize windows to optimize screen real estate without using a mou

    Objective-C
    在 GitHub 上查看↗13,631
  • llava-vl/llava-nextLLaVA-VL 的头像

    LLaVA-VL/LLaVA-NeXT

    4,695在 GitHub 上查看↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    在 GitHub 上查看↗4,695
  • dsdanielpark/bard-apidsdanielpark 的头像

    dsdanielpark/Bard-API

    5,196在 GitHub 上查看↗

    Bard-API is an asynchronous Python wrapper and client for interacting with Google Gemini. It functions as a stateful conversation manager and multimodal interface, allowing users to send text and image prompts to a language model and retrieve responses. The library utilizes a cookie-based authentication system that extracts session tokens from local browser storage to authorize requests. To manage access and connectivity, it includes proxy-based request routing to bypass regional restrictions and avoid IP blocks. The project covers capabilities for multimodal AI analysis and the maintenance

    Pythonai-apiapibard
    在 GitHub 上查看↗5,196
  • livekit/agentslivekit 的头像

    livekit/agents

    9,379在 GitHub 上查看↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Pythonagentsaiopenai
    在 GitHub 上查看↗9,379
  • mathewsachin/capturaMathewSachin 的头像

    MathewSachin/Captura

    10,731在 GitHub 上查看↗

    Captura is a desktop screen recording and screenshot utility designed to capture video, webcam feeds, and system audio into multimedia recordings. It functions as a recording suite that can also be operated as a command line video recorder, allowing users to trigger and manage capture workflows via terminal commands. The software distinguishes itself by recording user input, such as mouse movements and keystrokes, as visual overlays on top of the captured video. It further supports automated workflows through the integration of system-level hardware media keys, enabling recording state change

    C#
    在 GitHub 上查看↗10,731