awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
airsplay avatar

airsplay/vokenization

0
View on GitHub↗
191 stars·21 forks·Python·MIT·5 views

Vokenization

PyTorch code for EMNLP 2020 Paper "Vokenization: Improving Language Understanding with Visual Supervision"

Features

  • Multimodal Pretraining - Improving language understanding with visually-grounded supervision.

Star history

Star history chart for airsplay/vokenizationStar history chart for airsplay/vokenization

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Vokenization

These projects share indexed features with Vokenization. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • google-research/big_visiongoogle-research avatar

    google-research/big_vision

    3,363View on GitHub↗

    This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces. The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide

    Jupyter Notebook
    View on GitHub↗3,363
  • evolvinglmms-lab/otterEvolvingLMMs-Lab avatar

    EvolvingLMMs-Lab/Otter

    3,331View on GitHub↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Pythonartificial-inteligencechatgptdeep-learning
    View on GitHub↗3,331
  • jayleicn/clipbertjayleicn avatar

    jayleicn/ClipBERT

    730View on GitHub↗

    Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

    Python
    View on GitHub↗730
  • salesforce/albefsalesforce avatar

    salesforce/ALBEF

    1,758View on GitHub↗

    This is the official PyTorch implementation of the ALBEF paper Blog . This repository supports pre-training on custom datasets, as well as finetuning on VQA, SNLI-VE, NLVR2, Image-Text Retrieval on MSCOCO and Flickr30k, and visual grounding on RefCOCO+. Pre-trained and finetuned checkpoints…

    Python
    View on GitHub↗1,758
Compare all 8 related projects→

Frequently asked questions

What does airsplay/vokenization do?

PyTorch code for EMNLP 2020 Paper "Vokenization: Improving Language Understanding with Visual Supervision"

What are the main features of airsplay/vokenization?

The main features of airsplay/vokenization are: Multimodal Pretraining.

Which projects share features with airsplay/vokenization?

Projects with overlapping indexed features include: google-research/big_vision — This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal… evolvinglmms-lab/otter — Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It… jayleicn/clipbert — Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling. salesforce/albef — This is the official PyTorch implementation of the ALBEF paper [Blog] . This repository supports pre-training on… uclanlp/visualbert. airsplay/lxmert — Our servers break again :(. I have updated the links so that they should work fine now. Sorry for the inconvenience.…