# AI desktop assistant

> AI-ranked search results for `jarvis` on awesome-repositories.com — ordered by an LLM for relevance, best match first. 88 total matches; showing the top 11.

Explore on the web: https://awesome-repositories.com/q/jarvis

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [this search on awesome-repositories.com](https://awesome-repositories.com/q/jarvis).**

## Results

- [home-assistant/core](https://awesome-repositories.com/repository/home-assistant-core.md) (87,753 ⭐) — Home Assistant is a centralized home automation platform designed to orchestrate diverse internet-connected devices and services. It functions as a local-first control system that normalizes heterogeneous hardware protocols into a unified set of entities, attributes, and services. The core architecture relies on an event-driven state bus and a modular integration model, allowing the system to manage state changes and communicate across decoupled components through standardized interfaces.

The platform distinguishes itself through a highly flexible, declarative configuration framework that all
- [idootop/mi-gpt](https://awesome-repositories.com/repository/idootop-mi-gpt.md) (12,458 ⭐) — mi-gpt is a voice assistant bridge and agent orchestrator that connects smart speakers to large language models. It functions as an integration layer that routes audio requests from hardware speakers to AI providers and converts generated text back into speech via a customizable synthesis system.

The project features a retrieval-augmented generation knowledge base that uses embeddings and external documents to provide context-aware responses. It includes a persona definition system for configuring behavioral rules, system prompts, and roleplay characteristics, alongside a plugin architecture
- [dnhkng/glados](https://awesome-repositories.com/repository/dnhkng-glados.md) (5,595 ⭐) — GLaDOS is a multimodal AI agent framework designed to create autonomous systems that process text, speech, and visual data to interact with users and their environment. It centers on an AI personality framework that emulates complex character personas using a multi-agent architecture and configurable behavioral profiles.

The project distinguishes itself through an integrated tool layer that connects language models to external hardware, smart home devices, and system APIs via a standardized protocol. It features a character text-to-speech engine with low-latency playback and interruption hand
- [akshayaggarwal99/jarvis-ai-assistant](https://awesome-repositories.com/repository/akshayaggarwal99-jarvis-ai-assistant.md) (564 ⭐) — Jarvis AI Assistant - Voice-powered AI assistant for Mac
- [sukeesh/jarvis](https://awesome-repositories.com/repository/sukeesh-jarvis.md) (3,528 ⭐) — Personal Assistant for Linux and macOS
- [xiaomi/ha_xiaomi_home](https://awesome-repositories.com/repository/xiaomi-ha-xiaomi-home.md) (21,779 ⭐) — This project is a software integration designed to connect and control local Xiaomi smart home devices within a centralized home automation environment. It functions as a bridge that enables unified monitoring and management of various connected appliances across a local network, providing a standardized interface for IoT device orchestration.

The integration secures communication channels by validating encrypted handshake sequences required to authorize commands between the controller and local hardware. It maintains state consistency by translating proprietary device attributes into standar
- [nvidia/tacotron2](https://awesome-repositories.com/repository/nvidia-tacotron2.md) (5,300 ⭐) — This project is a neural text-to-speech framework and PyTorch model designed to synthesize human speech. It converts written text into synthetic audio by predicting mel spectrograms, which serve as an intermediate representation for voice generation.

The system includes a conditioning model for WaveNet to ensure natural-sounding audio output. It provides a distributed training framework that utilizes multi-GPU processing and automatic mixed precision to optimize training speed and reduce memory usage.

The project covers the full pipeline of neural speech synthesis, from model training using
- [sesameailabs/csm](https://awesome-repositories.com/repository/sesameailabs-csm.md) (14,669 ⭐) — CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.

The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person.

The model covers conversational speech synthesis and text-to-speech generation, transfo
- [microsoft/vibevoice](https://awesome-repositories.com/repository/microsoft-vibevoice.md) (49,394 ⭐) — VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content.

The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow
- [home-assistant/home-assistant](https://awesome-repositories.com/repository/home-assistant-home-assistant.md) (87,771 ⭐) — Home Assistant is a home automation platform and IoT device orchestrator that serves as a central hub for controlling smart devices and executing automated routines. It functions as a local smart home controller, managing device states and automation logic on a local network to provide a private alternative to cloud-based hubs.

The system emphasizes privacy-focused IoT management by prioritizing local control to reduce reliance on external cloud services. It enables multi-vendor device integration, translating diverse third-party hardware signals into a unified interface for consolidated mana
- [netease-youdao/emotivoice](https://awesome-repositories.com/repository/netease-youdao-emotivoice.md) (8,446 ⭐) — EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres.

The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts.

The system covers a broad range of speec
