For an open source alternative to Midjourney, the first results are hlky/stable-diffusion-webui (This project is a feature-rich, self-hostable web interface for Stable Diffusion that provides robust text-to-image generation, local execution, and extensive support for custom extensions like ControlNet and LoRA training), invoke-ai/invokeai (InvokeAI is a self-hostable generative AI platform built on a stable diffusion backend with a rich web user interface, supporting text-to-image workflows, inpainting, and advanced node-based image synthesis) and latentcat/qrbtf (This repository provides a specialized web-based interface and latent diffusion backend tailored for generating artistic AI images and QR codes with ControlNet support, though its scope is narrowly focused on QR synthesis rather than general-purpose image creation). automatic1111/stable-diffusion-webui and voltaml/voltaml-fast-stable-diffusion round out the shortlist. Compare the match explanations and check the project documentation against your requirements.
We curate open-source GitHub repositories matching “open source alternatives to midjourney”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
Stable Diffusion Web UI is a browser-based interface for generating, editing, and upscaling images and videos using latent diffusion models. It functions as a text-to-image generator, an AI image editor, and a tool for increasing image resolution and clarity. The system includes capabilities for custom model training, specifically allowing the creation of textual inversion embeddings to teach a model new concepts and visual styles from user photos. It also provides tools for AI video production, generating short clips from text prompts. The software covers image-to-image transformation, imag
This project is a feature-rich, self-hostable web interface for Stable Diffusion that provides robust text-to-image generation, local execution, and extensive support for custom extensions like ControlNet and LoRA training.
InvokeAI is a self-hosted, professional-grade platform designed for managing generative models and performing complex image synthesis. It provides a local application environment that allows users to execute diffusion models directly on their own hardware, ensuring data privacy and complete ownership of all generated assets. The platform distinguishes itself through a node-based workflow system that enables the construction of reproducible and automated image generation pipelines. By chaining modular functional units into directed acyclic graphs, users can automate intricate production tasks
InvokeAI is a self-hostable generative AI platform built on a stable diffusion backend with a rich web user interface, supporting text-to-image workflows, inpainting, and advanced node-based image synthesis.
qrbtf is an AI QR code generator and image synthesis system that blends machine-readable data with artistic imagery. It uses a latent diffusion model and spatial control networks to produce functional QR codes that incorporate visual art generated from descriptive text prompts. The system provides a dedicated interface and programmatic API for tuning visual output, allowing for the adjustment of control strength, padding ratios, and error correction levels. It supports deterministic sampling via random seeds and the use of negative prompts to refine the final aesthetic of the generated assets
This repository provides a specialized web-based interface and latent diffusion backend tailored for generating artistic AI images and QR codes with ControlNet support, though its scope is narrowly focused on QR synthesis rather than general-purpose image creation.
Stable Diffusion Web UI is a browser-based interface designed for managing text-to-image generation tasks. It provides a centralized dashboard for controlling generative processes, including native support for multi-stage model architectures to facilitate high-quality image refinement. The platform distinguishes itself through granular control over the generation process, offering tools for precise parameter management and advanced prompt engineering. Users can customize generation styles and capabilities by integrating external model-extension formats, such as textual inversions, low-rank ad
This repository provides a self-hosted web user interface for Stable Diffusion with text-to-image generation, ControlNet support, and LoRA training capabilities, perfectly matching your search for a local generative AI art tool.
VoltaML-fast-stable-diffusion is a generative system designed for high-performance image synthesis from text prompts. It provides a comprehensive environment for executing inference tasks, managing pre-trained machine learning models, and integrating visual asset creation into external applications and workflows. The project distinguishes itself through multi-modal interaction capabilities, including a browser-based web interface for direct generation and a messaging platform integration that allows users to trigger and monitor tasks via chat commands. It supports automated creative workflows
This repository provides a self-hosted text-to-image generative system with a web user interface and Stable Diffusion backend, though it lacks explicit native support for ControlNet and LoRA training out of the box.
Fooocus is a generative image interface designed to simplify the creation of high-quality visual content from text descriptions. It functions as a latent diffusion pipeline and model orchestrator, managing the complex interactions between neural network layers, mathematical samplers, and hardware resource allocation to produce professional-grade imagery. The project distinguishes itself through a sophisticated prompt engineering engine and modular style management. Users can dynamically modify output characteristics by injecting style adapters directly into prompts or by utilizing wildcards a
Fooocus provides a local web interface for text-to-image generation backed by Stable Diffusion, making it a well-suited tool for running generative AI models for artistic creation on your own hardware.
This project is a cloud-based AI deployment system and latent diffusion model trainer. It provides a framework for launching image generation interfaces and training pipelines on remote GPU infrastructure, specifically serving as a text-to-image model fine-tuner. The system features a specialized training interface for fine-tuning Stable Diffusion models on custom image datasets. It allows for the creation of personalized visual outputs by training models on specific subjects or artistic styles using a small set of reference images. The software covers generative AI deployment, custom style
This project provides a cloud-based framework and training interface for stable diffusion models, though it is primarily structured as a notebook deployment rather than a traditional self-hosted local application.
ComfyUI is a node-based generative AI orchestration engine designed for constructing, testing, and executing complex image and video synthesis pipelines. By utilizing a directed acyclic graph execution model, the platform allows users to build reproducible workflows through modular, interconnected processing blocks without requiring manual code implementation. It serves as both a local environment for high-performance model inference and a production-ready server for deploying generative capabilities. The platform distinguishes itself through its focus on workflow portability and extensibilit
ComfyUI is a powerful node-based generative AI orchestration engine that runs locally with a web interface, providing deep Stable Diffusion support, ControlNet capabilities, and advanced workflow customization for text-to-image synthesis.
Stable Diffusion is a generative machine learning pipeline that synthesizes high-resolution visual content by performing iterative denoising within a compressed latent space. By mapping natural language embeddings into pixel outputs through conditioned probabilistic processes, the framework enables the generation of images from text prompts and the transformation of existing visual inputs based on semantic instructions. The architecture utilizes a modular execution environment that decouples model loading, scheduler logic, and inference components to support diverse hardware configurations. I
This repository provides the core Stable Diffusion model and latent-space generative pipeline required for text-to-image synthesis, though running it locally with a web user interface and advanced training features typically requires companion frontends or extensions rather than being an out-of-the-box UI.
Stable Diffusion WebUI Forge is a web-based interface and inference engine designed for the generation of AI media. It functions as a platform for executing diffusion-based models, providing a centralized environment to manage image preprocessors, custom generation logic, and hardware-accelerated sampling. The project distinguishes itself through a neural network patching framework that allows for the modification of model layers and the application of spatial conditioning during inference. By injecting custom logic and adapters directly into the network, users can influence output behaviors
This project is a popular self-hostable web interface and inference engine for running Stable Diffusion models locally, though it omits built-in LoRA training out of the box.
StableSwarmUI is a web interface and backend orchestrator for Stable Diffusion image generation. It functions as a distributed GPU image generator and a modular AI image pipeline, providing a centralized controller to manage image generation requests. The system distinguishes itself through the ability to split generation tasks across multiple graphics processors to increase batch throughput. It utilizes a backend-agnostic interface to connect to local servers, remote servers, and cloud APIs, and includes a graph-based visual workflow designer for defining complex image processing operations.
StableSwarmUI is a self-hostable web interface and backend orchestrator for Stable Diffusion that delivers text-to-image generation with advanced node-based workflows, though it lacks direct mentions of native LoRA training in its core feature set.
ComfyUI is a modular generative AI workflow orchestrator and node-based GUI for designing and executing complex diffusion model pipelines. It functions as both a visual interface for building generative logic graphs and a programmable backend API that exposes diffusion model operations for external integration. The system distinguishes itself through a graph-based execution model that supports differential workflow execution, re-running only modified nodes to reduce computation. It features dynamic model offloading to manage memory between system RAM and GPU VRAM and utilizes metadata-embedde
ComfyUI is a powerful node-based web interface and workflow orchestrator for stable diffusion models that lets you run text-to-image generation locally while supporting advanced features like ControlNet pipelines and LoRA training.
HunyuanDiT is a bilingual text-to-image generative model and diffusion transformer image generator. It uses a latent diffusion system to synthesize high-resolution images from text prompts, with a specific focus on understanding and generating content from both Chinese and English language descriptions. The project features a multi-resolution transformer architecture and a bilingual embedding space to map different scripts into a shared semantic area. It supports iterative multi-turn image refinement, which translates conversational dialogue into updated prompts to progressively modify visual
HunyuanDiT is a bilingual text-to-image latent diffusion model capable of generating artistic images from text prompts and running locally, though it functions as a backend model repository rather than a complete self-hosted web interface with ControlNet and LoRA training out of the box.
StableCascade is a generative AI system and latent diffusion framework designed for text-to-image synthesis and image-to-image transformations. It utilizes a multi-stage cascade architecture that encodes and decodes images via a latent space to produce high-fidelity visual imagery. The system includes a cascade diffusion pipeline for controlling image structure through inpainting, outpainting, and super-resolution. It also provides a toolkit for image-to-image generation and the creation of image variations using embeddings. The framework supports model optimization through low-rank adaptati
StableCascade is a generative AI model and latent diffusion framework for text-to-image synthesis that supports LoRA training and stable diffusion workflows, though it provides the underlying model architecture rather than a ready-to-run self-hosted web user interface.
Sygil-webui is a web interface for Stable Diffusion latent diffusion models, providing a creative suite for text-to-image and text-to-video synthesis. It functions as an image generation tool and a latent diffusion image editor, allowing users to create visuals and video sequences from textual descriptions. The project includes a dedicated model training interface for creating custom textual inversion embeddings, which introduces specific new concepts or styles into the diffusion models. It also features specialized tools for generative image editing, including mask-based inpainting, image-to
Sygil-webui is a self-hostable web interface for Stable Diffusion that provides text-to-image generation and model training features, though it lacks explicit built-in support for ControlNet.
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec
Janus is a multimodal framework that integrates text-to-image generation within a unified architecture, though it focuses more on unified foundational research than offering an out-of-the-box local user interface with LoRA training and ControlNet.
OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks through a single system. It functions as a multimodal diffusion framework that treats diverse vision operations as unified image synthesis problems using shared model weights, removing the need for external adapter modules. The system supports subject-driven image generation to preserve the identity of objects from reference photos and allows for multi-reference image synthesis. It also operates as an instruction-based image editor, modifying visual content through natural languag
OmniGen is a unified multimodal diffusion model and text-to-image generator that can be self-hosted, though as a model repository it lacks a dedicated full-featured web user interface, ControlNet integration, and LoRA training out of the box.
This project is a framework for running Stable Diffusion image generation models on Apple Silicon using Core ML hardware acceleration. It provides a local generative AI pipeline for producing images from text prompts using Swift and Python without relying on external cloud APIs. The system includes a model converter to transform deep learning checkpoints into Core ML formats and a model optimizer to quantize weights and activations. It features a ControlNet integration layer to guide image generation using external signals such as edge and depth maps. Capabilities cover text-to-image generat
This repository provides a local text-to-image Stable Diffusion pipeline optimized for Apple Silicon hardware, though it functions as a lower-level framework and converter rather than a full turnkey web user interface.
IF is a text-to-image diffusion system that translates natural language descriptions into visual imagery. The project provides a generative pipeline for creating images, an inpainting tool for modifying specific image sections, and a super-resolution upscaler to increase pixel density and clarity. The system includes a concept fine-tuning framework that allows for the teaching of new visual concepts by updating a small set of parameters. It also supports image style transfer to apply the aesthetic characteristics of a reference image to a new output.
DeepFloyd IF is an open-source text-to-image diffusion system equipped with fine-tuning capabilities and image manipulation tools, though it lacks a built-in web UI and ControlNet support out of the box.
imaginAIry is a system for generating and refining images and videos using diffusion models. It operates as a web-based server that triggers generation requests through standard API calls, allowing for the creation of visuals and video sequences from text prompts or existing files. The project provides a suite for AI image editing and upscaling, enabling the modification of visuals through natural language instructions and super-resolution tools to increase detail and image size. The system includes capabilities for structural image control using depth maps, edge maps, and body poses to main
This project is a diffusion-based generation system with a web server and text-to-image capabilities, though it lacks some specific advanced features like LoRA training and ControlNet support.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| hlky/stable-diffusion-webui | 7.9K | Python | AGPL-3.0 | |
| invoke-ai/invokeai | 27.5K | TypeScript | Apache-2.0 | |
| 7K |
| TypeScript |
| GPL-3.0 |
| automatic1111/stable-diffusion-webui | 163.7K | Python | AGPL-3.0 |
| voltaml/voltaml-fast-stable-diffusion | 998 | Python | GPL-3.0 |
| lllyasviel/fooocus | 50.3K | Python | GPL-3.0 |
| thelastben/fast-stable-diffusion | 7.9K | Python | mit |
| comfy-org/comfyui | 117.2K | Python | GPL-3.0 |
| compvis/stable-diffusion | 73.1K | Jupyter Notebook | NOASSERTION |
| lllyasviel/stable-diffusion-webui-forge | 12.7K | Python | AGPL-3.0 |