awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to mobvoi/seq-monkey-data

Open-source alternatives to Seq Monkey Data

27 open-source projects similar to mobvoi/seq-monkey-data, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Seq Monkey Data alternative.

  • plexpt/chatgpt-corpusAvatar von PlexPt

    PlexPt/chatgpt-corpus

    964Auf GitHub ansehen↗

    This project provides a comprehensive Chinese language corpus designed to support the training and fine-tuning of large language models. It serves as a structured natural language processing resource, offering a collection of text data that includes dialogue, customer service interactions, and creative writing. The dataset is organized into distinct thematic categories, allowing for targeted model development across specific conversational and narrative contexts. By providing information in standardized, schema-agnostic text formats, the collection ensures portability across various machine l

    awesomecorpuscorpus-data
    Auf GitHub ansehen↗964
  • esbatmop/mnbvcAvatar von esbatmop

    esbatmop/MNBVC

    4,123Auf GitHub ansehen↗

    MNBVC is a dataset pipeline and toolkit designed for the collection, cleaning, and normalization of massive text and code corpora used to train large language models. It provides specialized tools for harvesting source code, commit histories, and repository metadata from version control platforms, alongside a multilingual text corpus collector for gathering parallel text and academic papers. The project distinguishes itself through comprehensive capabilities for processing diverse document types, including a PDF-to-text converter that transforms complex layouts and formulas into structured JS

    chinesechinese-languagechinese-nlp
    Auf GitHub ansehen↗4,123
  • fuxiaoliu/lrv-instructionAvatar von FuxiaoLiu

    FuxiaoLiu/LRV-Instruction

    297Auf GitHub ansehen↗

    ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

    Pythonchatgptevaluationevaluation-metrics
    Auf GitHub ansehen↗297

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
  • guoyang9/unk-vqaAvatar von guoyang9

    guoyang9/UNK-VQA

    7Auf GitHub ansehen↗

    A VQA dataset that includes unanswerable questions TPAMI 2024.

    Auf GitHub ansehen↗7
  • hypjudy/sparklesAvatar von HYPJUDY

    HYPJUDY/Sparkles

    45Auf GitHub ansehen↗

    Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models

    Python
    Auf GitHub ansehen↗45
  • icoz69/stablellavaAvatar von icoz69

    icoz69/StableLLAVA

    95Auf GitHub ansehen↗

    Official repo for StableLLAVA

    Python
    Auf GitHub ansehen↗95
  • inst-it/inst-itAvatar von inst-it

    inst-it/inst-it

    40Auf GitHub ansehen↗

    NeurIPS 2025 The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning"

    Pythoninstruction-tuninglarge-multimodal-modelsmultimodal
    Auf GitHub ansehen↗40
  • kakaobrain/coyo-datasetAvatar von kakaobrain

    kakaobrain/coyo-dataset

    1,256Auf GitHub ansehen↗

    COYO-700M: Large-scale Image-Text Pair Dataset

    Python
    Auf GitHub ansehen↗1,256
  • luodian/otterAvatar von Luodian

    Luodian/Otter

    3,410Auf GitHub ansehen↗

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    Python
    Auf GitHub ansehen↗3,410
  • lyuchenyang/macaw-llmAvatar von lyuchenyang

    lyuchenyang/Macaw-LLM

    1,590Auf GitHub ansehen↗

    Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration

    Pythondeep-learninglanguage-modelmachine-learning
    Auf GitHub ansehen↗1,590
  • mbzuai-oryx/video-chatgptAvatar von mbzuai-oryx

    mbzuai-oryx/Video-ChatGPT

    1,504Auf GitHub ansehen↗

    ACL 2024 🔥 Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    Pythonchatbotclipgpt-4
    Auf GitHub ansehen↗1,504
  • microsoft/llava-medAvatar von microsoft

    microsoft/LLaVA-Med

    2,214Auf GitHub ansehen↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    Python
    Auf GitHub ansehen↗2,214
  • ncsoft/cap2qaAvatar von ncsoft

    ncsoft/cap2qa

    5Auf GitHub ansehen↗

    Official implementation of "Visually Dehallucinative Instruction Generation" (ICASSP 2024)

    Auf GitHub ansehen↗5
  • ncsoft/idkAvatar von ncsoft

    ncsoft/idk

    6Auf GitHub ansehen↗

    Official implementation of "Visually Dehallucinative Instruction Generation: Know What You Don't Know"

    Auf GitHub ansehen↗6
  • opengvlab/internvideoAvatar von OpenGVLab

    OpenGVLab/InternVideo

    2,292Auf GitHub ansehen↗

    ECCV2024 Video Foundation Models & Data for Multimodal Understanding

    Pythonaction-recognitionbenchmarkcontrastive-learning
    Auf GitHub ansehen↗2,292
  • openlamm/lammAvatar von OpenLAMM

    OpenLAMM/LAMM

    317Auf GitHub ansehen↗

    NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents

    Python
    Auf GitHub ansehen↗317
  • openm3d/m3dbenchAvatar von OpenM3D

    OpenM3D/M3DBench

    61Auf GitHub ansehen↗

    ECCV 2024 M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.

    Python3ddatasetinstruction-tuning
    Auf GitHub ansehen↗61
  • openmotionlab/motiongptAvatar von OpenMotionLab

    OpenMotionLab/MotionGPT

    1,938Auf GitHub ansehen↗

    NeurIPS 2023 MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs

    Python3d-generationchatgptgpt
    Auf GitHub ansehen↗1,938
  • optimalscale/detgptAvatar von OptimalScale

    OptimalScale/DetGPT

    788Auf GitHub ansehen↗

    2023-06-13 Added tuned linear weights for Vicuna-7b. 2023-05-25 Our paper is available at this link. 2023-05-09 We have launched our project website. 2023-05-08 The first version of DetGPT is available now! Try our demo.

    Jupyter Notebook
    Auf GitHub ansehen↗788
  • phellonchen/x-llmAvatar von phellonchen

    phellonchen/X-LLM

    318Auf GitHub ansehen↗

    X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

    Python
    Auf GitHub ansehen↗318
  • polyu-chenlab/etbenchAvatar von PolyU-ChenLab

    PolyU-ChenLab/ETBench

    74Auf GitHub ansehen↗

    👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)

    Python
    Auf GitHub ansehen↗74
  • rucaibox/comvintAvatar von RUCAIBox

    RUCAIBox/ComVint

    19Auf GitHub ansehen↗

    The official GitHub page for ''What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning''

    Python
    Auf GitHub ansehen↗19
  • stevengrove/gpt4toolsAvatar von StevenGrove

    StevenGrove/GPT4Tools

    771Auf GitHub ansehen↗

    GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

    Python
    Auf GitHub ansehen↗771
  • togethercomputer/redpajama-dataAvatar von togethercomputer

    togethercomputer/RedPajama-Data

    4,947Auf GitHub ansehen↗

    RedPajama-Data is a toolset for preprocessing large-scale text datasets used to train large language models. It provides a preprocessing pipeline focused on cleaning, deduplicating, and scoring massive collections of text to ensure data quality and diversity. The project utilizes a document quality scoring framework that employs machine learning and statistical heuristics to evaluate if documents are suitable for training. It includes a dataset filtering pipeline that uses classifiers and blocklists to remove undesirable words or URLs. The system features a text deduplication toolset that el

    Python
    Auf GitHub ansehen↗4,947
  • vt-nlp/multiinstructAvatar von VT-NLP

    VT-NLP/MultiInstruct

    135Auf GitHub ansehen↗

    MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

    Python
    Auf GitHub ansehen↗135
  • cluebenchmark/cluecorpus2020Avatar von CLUEbenchmark

    CLUEbenchmark/CLUECorpus2020

    1,012Auf GitHub ansehen↗

    Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

    albertbertchinese
    Auf GitHub ansehen↗1,012
  • zhourax/vegaAvatar von zhourax

    zhourax/VEGA

    38Auf GitHub ansehen↗

    Project Page | Paper | Dataset

    Python
    Auf GitHub ansehen↗38