27 open-source projects similar to mobvoi/seq-monkey-data, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Seq Monkey Data alternative.
This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l
MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS
ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
NeurIPS 2025 The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning"
COYO-700M: Large-scale Image-Text Pair Dataset
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration
ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
Official implementation of "Visually Dehallucinative Instruction Generation" (ICASSP 2024)
Official implementation of "Visually Dehallucinative Instruction Generation: Know What You Don't Know"
ECCV2024 Video Foundation Models & Data for Multimodal Understanding
NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents
ECCV 2024 M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.
NeurIPS 2023 MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
2023-06-13 Added tuned linear weights for Vicuna-7b. 2023-05-25 Our paper is available at this link. 2023-05-09 We have launched our project website. 2023-05-08 The first version of DetGPT is available now! Try our demo.
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)
The official GitHub page for ''What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning''
GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.
RedPajama-Data is a toolset for preprocessing large-scale text datasets used to train large language models. It provides a preprocessing pipeline focused on cleaning, deduplicating, and scoring massive collections of text to ensure data quality and diversity. The project utilizes a document quality scoring framework that employs machine learning and statistical heuristics to evaluate if documents are suitable for training. It includes a dataset filtering pipeline that uses classifiers and blocklists to remove undesirable words or URLs. The system features a text deduplication toolset that el
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料