awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to modelscope/diffsynth-studio

Open-source alternatives to DiffSynth Studio

30 open-source projects similar to modelscope/diffsynth-studio, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best DiffSynth Studio alternative.

  • huggingface/diffusersAvatar von huggingface

    huggingface/diffusers

    33,872Auf GitHub ansehen↗

    Diffusers is a PyTorch-based library and generative AI framework used to build, train, and deploy diffusion pipelines for producing multi-modal media. It provides a suite of tools for generating images, video, and audio from natural language descriptions, as well as specialized systems for text-to-image generation. The project differentiates itself through a modular architecture that separates noise schedulers, pretrained model blocks, and pipeline compositions. This structure allows for the construction of custom generation workflows and the ability to swap individual components of the diffu

    Pythondeep-learningdiffusionflux
    Auf GitHub ansehen↗33,872
  • zhaochenyang20/awesome-ml-sys-tutorialAvatar von zhaochenyang20

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371Auf GitHub ansehen↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Python
    Auf GitHub ansehen↗5,371
  • microsoft/unilmAvatar von microsoft

    microsoft/unilm

    22,030Auf GitHub ansehen↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Pythonbeitbeit-3bitnet
    Auf GitHub ansehen↗22,030

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
  • hao-ai-lab/fastvideoAvatar von hao-ai-lab

    hao-ai-lab/FastVideo

    3,743Auf GitHub ansehen↗

    FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a

    Pythondiffusersdiffusion-modelsdistillation
    Auf GitHub ansehen↗3,743
  • videoverses/videotunaV

    VideoVerses/VideoTuna

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • thelastben/fast-stable-diffusionAvatar von TheLastBen

    TheLastBen/fast-stable-diffusion

    7,889Auf GitHub ansehen↗

    This project is a cloud-based AI deployment system and latent diffusion model trainer. It provides a framework for launching image generation interfaces and training pipelines on remote GPU infrastructure, specifically serving as a text-to-image model fine-tuner. The system features a specialized training interface for fine-tuning Stable Diffusion models on custom image datasets. It allows for the creation of personalized visual outputs by training models on specific subjects or artistic styles using a small set of reference images. The software covers generative AI deployment, custom style

    Pythona1111aicolab
    Auf GitHub ansehen↗7,889
  • facebookresearch/multimodalAvatar von facebookresearch

    facebookresearch/multimodal

    1,723Auf GitHub ansehen↗

    Multimodal is a machine learning library built on PyTorch for training large-scale models that combine text, image, audio, and video data streams. It functions as a deep learning framework dedicated to generative diffusion models, multi-task training, and vision-language tasks. The library supplies modular building blocks, discrete latent codebook quantization, shared-space embeddings, and stackable adapter layers to handle diverse conditional inputs during training and inference. The framework supports specific architectures for diffusion models, text-to-video generation, image-text retrieva

    Python
    Auf GitHub ansehen↗1,723
  • stability-ai/generative-modelsAvatar von Stability-AI

    Stability-AI/generative-models

    27,189Auf GitHub ansehen↗

    This is a framework for training and sampling diffusion models to generate high-fidelity images, video, and 4D assets. It provides a modular environment for managing generative AI training pipelines, including the handling of datasets, noise sampling, and loss weighting to stabilize the creation of synthetic content. The project features a modular model configuration system that uses YAML-based assembly to define network submodules and conditioners. It also includes a dedicated toolset for AI image watermarking, allowing for the embedding and detection of invisible markers to verify the origi

    Python
    Auf GitHub ansehen↗27,189
  • thudm/cogvideoAvatar von THUDM

    THUDM/CogVideo

    12,792Auf GitHub ansehen↗

    CogVideo is a generative video framework that uses diffusion models and transformer-based architectures to synthesize high-resolution video clips. It functions as both a text-to-video and image-to-video generator, converting textual descriptions or static images into temporal visual sequences. The system integrates large language model capabilities to expand short user prompts into detailed descriptions for better visual alignment. It supports the animation of static images through latent seeding and provides the ability to extend the length of existing video sequences. The project includes

    Python
    Auf GitHub ansehen↗12,792
  • alembics/disco-diffusionAvatar von alembics

    alembics/disco-diffusion

    7,407Auf GitHub ansehen↗

    This project is a diffusion-based AI art generator and animation framework used to create digital images and motion graphics from text prompts. It functions as a system for producing stylized videos and AI art through iterative diffusion sampling and neural network models. The framework distinguishes itself through specialized tools for 3D depth animation, using depth-map transformations to create spatial movement. It also includes neural style transfer capabilities to apply specific artistic looks, such as watercolor or pixel art, and utilizes optical flow frame blending to reduce flickering

    Jupyter Notebook
    Auf GitHub ansehen↗7,407
  • open-mmlab/mmagicAvatar von open-mmlab

    open-mmlab/mmagic

    7,434Auf GitHub ansehen↗

    mmagic is a multimodal training pipeline and framework for generative AI, focusing on visual synthesis and restoration. It provides the infrastructure to build and train models for tasks such as text-to-image and text-to-video generation, 3D-aware content synthesis, and high-fidelity image translation using diffusion models and generative adversarial networks. The project distinguishes itself through specialized capabilities for generative model personalization, including techniques for fine-tuning subjects and styles. It also supports advanced visual manipulations such as latent space interp

    Jupyter Notebookaigccomputer-visiondeep-learning
    Auf GitHub ansehen↗7,434
  • arize-ai/phoenixAvatar von Arize-ai

    Arize-ai/phoenix

    8,605Auf GitHub ansehen↗

    Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

    Jupyter Notebookagentsai-monitoringai-observability
    Auf GitHub ansehen↗8,605
  • lianjiatech/belleAvatar von LianjiaTech

    LianjiaTech/BELLE

    8,273Auf GitHub ansehen↗

    BELLE is a specialized implementation of Chinese conversational large language models, encompassing a full instruction tuning framework. It provides a pipeline for training, evaluating, and deploying models optimized for natural language understanding and dialogue tasks in the Chinese language. The project is distinguished by its integrated approach to model refinement, combining the curation of multi-million entry instruction datasets with a distributed training pipeline. This pipeline supports both full fine-tuning and low-rank adaptation to optimize conversational performance. The system

    HTMLbloomchinese-nlpgpt-evaluation
    Auf GitHub ansehen↗8,273
  • vibrantlabsai/ragasAvatar von vibrantlabsai

    vibrantlabsai/ragas

    12,659Auf GitHub ansehen↗

    Ragas is an evaluation framework designed to measure the performance of retrieval-augmented generation pipelines and autonomous agent workflows. It provides a comprehensive suite of tools for benchmarking system outputs, utilizing language models as automated judges to score performance against defined rubrics and reference data. By standardizing inputs, retrieved contexts, and generated responses into a unified schema, the project enables consistent analysis across complex AI applications. The framework distinguishes itself through its ability to generate synthetic test datasets from existin

    Pythonevaluationllmllmops
    Auf GitHub ansehen↗12,659
  • tele-ai/teletronT

    Tele-AI/TeleTron

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • stepfun-ai/step-video-t2vAvatar von stepfun-ai

    stepfun-ai/Step-Video-T2V

    3,186Auf GitHub ansehen↗

       

    Python
    Auf GitHub ansehen↗3,186
  • tencent/hunyuanvideoT

    Tencent/HunyuanVideo

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • bghira/simpletunerAvatar von bghira

    bghira/SimpleTuner

    2,862Auf GitHub ansehen↗

    A general fine-tuning kit geared toward image/video/audio diffusion models.

    Pythondiffusersdiffusion-modelsfine-tuning
    Auf GitHub ansehen↗2,862
  • wan-video/wan2.1Avatar von Wan-Video

    Wan-Video/Wan2.1

    15,350Auf GitHub ansehen↗

    Wan2.1 is a generative video synthesis framework that provides foundation models for creating high-fidelity video sequences and static images from descriptive text prompts. The system utilizes a unified architecture trained on both static and dynamic datasets, allowing it to function as a comprehensive tool for visual media creation. The framework distinguishes itself through a transformer-based temporal modeling approach that ensures structural coherence and consistent motion across video frames. It supports multi-resolution latent scaling, enabling the generation of content in various aspec

    Pythonaigcvideogeneration
    Auf GitHub ansehen↗15,350
  • x-gengroup/flow-factoryX

    X-GenGroup/Flow-Factory

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • kohya-ss/musubi-tunerAvatar von kohya-ss

    kohya-ss/musubi-tuner

    1,701Auf GitHub ansehen↗
    Python
    Auf GitHub ansehen↗1,701
  • shengshu-ai/minwmS

    shengshu-ai/minWM

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • pku-yuangroup/open-sora-planAvatar von PKU-YuanGroup

    PKU-YuanGroup/Open-Sora-Plan

    12,163Auf GitHub ansehen↗

    Open-Sora-Plan is a text-to-video framework and distributed video training system. It utilizes a diffusion transformer architecture and large language model components to transform written descriptions or image prompts into high-quality video sequences. The system features a distributed infrastructure designed for large-scale video training and inference. It employs sequence parallelism to split high-resolution or long-duration video samples across multiple GPUs and uses a sparse attention mechanism to increase processing speed. The project includes capabilities for both text-to-video and im

    Python
    Auf GitHub ansehen↗12,163
  • lightricks/ltx-videoAvatar von Lightricks

    Lightricks/LTX-Video

    9,324Auf GitHub ansehen↗
    Pythondiffusion-modelsditimage-to-video
    Auf GitHub ansehen↗9,324
  • skyworkai/skyreels-v2Avatar von SkyworkAI

    SkyworkAI/SkyReels-V2

    6,356Auf GitHub ansehen↗

    SkyReels-V2 is a video generation system that creates, extends, and refines video clips from text descriptions, images, or both. It operates as a diffusion-based video generation model that can produce videos of any duration by denoising frames sequentially, with each new frame conditioned on the ones that came before it. The system supports generating videos from scratch using text prompts, starting from a single image and producing subsequent frames, or constraining both the first and last frames to match user-provided images. What distinguishes SkyReels-V2 is its combination of infinite-le

    Python
    Auf GitHub ansehen↗6,356
  • spacepxl/hunyuanvideo-trainingS

    spacepxl/HunyuanVideo-Training

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • modeltc/lightx2vM

    ModelTC/LightX2V

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • hpcaitech/open-soraAvatar von hpcaitech

    hpcaitech/Open-Sora

    29,101Auf GitHub ansehen↗

    Open-Sora is a video generation framework designed to produce cinematic sequences from text prompts and images. It functions as a generative system that transforms written descriptions or reference images into video content featuring realistic textures and lighting. The project includes a dedicated prompt engineering tool that uses large language models to expand simple user inputs into detailed descriptions. It also features a motion controller for adjusting movement intensity in generated sequences and evaluating motion levels in existing video files. The framework incorporates text-to-vid

    Python
    Auf GitHub ansehen↗29,101
  • tdrussell/diffusion-pipeAvatar von tdrussell

    tdrussell/diffusion-pipe

    1,976Auf GitHub ansehen↗

    A pipeline parallel training script for diffusion models.

    Python
    Auf GitHub ansehen↗1,976
  • yaofang-liu/mochi-full-finetunerY

    Yaofang-Liu/Mochi-Full-Finetuner

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0