awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
ยฉ 2026 Bringes Technology SRLยทVAT RO45896025ยทhello@awesome-repositories.com
ByteFlow-AI avatar

ByteFlow-AI/TokenFlow

0
View on GitHubโ†—
465 starsยท10 forksยทPythonยทApache-2.0ยท15 viewsarxiv.org/abs/2412.03069โ†—

TokenFlow

[CVPR 2025] ๐Ÿ”ฅ Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".

Features

  • Computer Vision Research - Unified image tokenizer for multimodal understanding and generation.
  • Generative AI - Unified tokenizer for multimodal understanding and generation.
  • Unified Models - Unified framework for multimodal token processing.
  • Unified Multimodal Models - Framework for multimodal token-based generation.

Star history

Star history chart for byteflow-ai/tokenflowStar history chart for byteflow-ai/tokenflow

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English โ€” the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does byteflow-ai/tokenflow do?

[CVPR 2025] ๐Ÿ”ฅ Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".

What are the main features of byteflow-ai/tokenflow?

The main features of byteflow-ai/tokenflow are: Computer Vision Research, Generative AI, Unified Models, Unified Multimodal Models.

What are some open-source alternatives to byteflow-ai/tokenflow?

Open-source alternatives to byteflow-ai/tokenflow include: bytedance/x-dyna โ€” [CVPR 2025 Highlight] X-Dyna: Expressive Dynamic Human Image Animation. facebookresearch/tuna-2 โ€” Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation. alpha-vllm/lumina-dimoo โ€” Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding. bytedance/lance โ€” A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing. deepseek-ai/janus โ€” Janus is a multimodal large language model and unified framework that integrates visual understanding and imageโ€ฆ hustvl/lightningdit โ€” [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models.

Open-source alternatives to TokenFlow

Similar open-source projects, ranked by how many features they share with TokenFlow.
  • bytedance/lancebytedance avatar

    bytedance/Lance

    1,250View on GitHubโ†—

    A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

    Python
    View on GitHubโ†—1,250
  • bytedance/x-dynabytedance avatar

    bytedance/X-Dyna

    269View on GitHubโ†—

    CVPR 2025 Highlight X-Dyna: Expressive Dynamic Human Image Animation

    Python
    View on GitHubโ†—269
  • alpha-vllm/lumina-dimooAlpha-VLLM avatar

    Alpha-VLLM/Lumina-DiMOO

    1,001View on GitHubโ†—

    Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Python
    View on GitHubโ†—1,001
  • deepseek-ai/janusdeepseek-ai avatar

    deepseek-ai/Janus

    17,746View on GitHubโ†—

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    View on GitHubโ†—17,746
See all 30 alternatives to TokenFlowโ†’