MOSS-VL's training data consists of (1) open-source datasets collected from the community and (2) in-house synthetic data. This page lists the 205 open-source datasets used in our training.
Note: Before being used in training, these open-source datasets went through our internal curation pipeline, including format unification, quality filtering, deduplication, and additional cleaning and re-annotation. As a result, the actual training samples may differ from the raw public releases.
- AI-MO/NuminaMath-CoT
- AI2D-Caption
- ALLaVA-4V
- allenai/big-reasoning-traces
- allenai/CoSyn-400K
- allenai/Molmo2-AskModelAnything
- allenai/objaverse
- allenai/pixmo-cap
- allenai/tulu-3-sft-mixture
- amphora/QwQ-LongCoT-130K
- AnyWord-3M
- APRIL-AIGC/UltraVideo
- ArXivQA
- AscendKernelGen/Ascend-CoT
- ashraq/tmdb-celeb-10k
- Azu/Handwritten-Mathematical-Expression-Convert-LaTeX
- BAAI/DenseFusion-1M
- BAAI/Infinity-Instruct
- BAAI/Infinity-MM
- BBox-DocVQA
- BenthamQA
- Bluel0la/Creative_Stories_Logical_Reasoning
- CaptionEmporium/TextOCR-GPT4o
- Chart-to-Text
- Chart2Text
- ChartQA
- Chinese Text in the Wild
- CinePile
- CoMM: A Coherent Interleaved Image-Text Dataset
- Daemontatox/LongCOT-Reason
- DAMO-NLP-SG/multimodal_textbook
- DAMO-NLP-SG/VL3-Syn7M
- di-zhang-fdu/R1-Vision-Reasoning-Instructions
- dim/competition_math
- Diving48
- Duplex-UltraChat
- DVQA
- Ego-Exo4D
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- EgoTaskQA
- EPIC-KITCHENS-100
- EST-VQA
- facebook/PLM-Video-Human
- Fancy-MLLM/R1-Onevision
- FastJobs/Visual_Emotional_Analysis
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- FineVideo
- FreedomIntelligence/medical-o1-reasoning-SFT
- FSC-147
- fudan-generative-vision/OpenHumanVid
- FunQA
- GAIR-NLP/MegaScience
- galaxyMindAiLabs/stem-reasoning-complex
- GameQA-140K
- Geo170K
- Geometry3K
- GeomVerse
- GeoQA+
- gghfez/long-cot-4k
- Google Scanned Objects (GSO),经 NVlabs/FoundationPose 打包
- Harshkmr/orca-math-word-reflection
- HiTab
- HoloAssist
- HuggingFaceM4/Docmatix
- HuggingFaceM4/FineVision
- HuggingFaceM4/OBELICS
- HuggingFaceM4/WebSight
- HumanPCR
- HW-SQuAD
- ICDAR 2019-LSVT
- ImageNet-Think-250K
- InfographicVQA
- Interplay-LM-Reasoning
- Jackrong/Competitive-Programming-python-blend
- Jackrong/DeepSeek-V3.2-Exp-reasoning-example
- Jackrong/GLM-5.1-Reasoning-1M-Cleaned
- jeggers/competition_math
- JHU-CROWD++
- jimmycarter/textocr-gpt4v
- Kamizuru00/diagram_image_to_text
- Kinetics
- LLaVA-Video-178K
- LLaVAR-Instruct-16K
- LLM360/MegaMath
- lmms-lab/llava-critic-113k
- lmms-lab/LLaVA-NeXT-Data
- lmms-lab/LLaVA-OneVision-Data
- lmms-lab/LLaVA-OneVision-Mid-Data
- lmms-lab/LLaVA-ReCap-118K
- lmms-lab/LLaVA-ReCap-558K
- lmms-lab/LLaVA-ReCap-CC3M
- lmms-lab/multimodal-open-r1-8k-verified
- LongViTU
- LRV-Instruction
- m-bain/webvid
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B
- MapQA: A Dataset for Question Answering on Choropleth Maps
- MAVIS
- Memotion Analysis
- MER2023 Multimodal Emotion Recognition Challenge
- MET-Meme: A Multi-modal Meme Dataset Rich in Metaphors
- meta-math/MetaMathQA
- microsoft/orca-math-word-problems-200k
- MIMIC-IT
- MMC-Instruction
- MMDU-45k
- MMEvol
- MMInstruction/Clevr_CoGenT_TrainA_R1
- MMK12
- MMTab
- mPLUG/DocReason25K
- MT-Video-Bench
- MTWI 2018
- Mulberry-260k
- multi-domain-reasoning/commonsense_qa
- MultiHiertt
- mutonix/Vript
- mutonix/Vript_Chinese
- mvp-lab/LLaVA-OneVision-1.5-Instruct-Data
- mychen76/invoices-and-receipts_ocr_v1
- Na0s/sft-ready-hendrycks-competition_math
- newfacade/LeetCodeDataset
- nvidia/Nemotron-SFT-Competitive-Programming-v2
- nvidia/OpenMathReasoning
- nyu-visionx/VSI-590K
- OCR-VQA
- OleehyO/latex-formulas
- OmniAlign-V
- OODCV-VQA
- Open-Orca/SlimOrca
- open-r1/OpenR1-Math-220k
- open-thoughts/OpenThoughts3-1.2M
- openbmb/RLAIF-V-Dataset
- OpenGVLab/ShareGPT-4o
- OpenMMReasoner (EvolvingLMMs-Lab) SFT 数据集
- ORAND-CAR
- PathGen-1.6M
- PDFVQA
- Places365
- POIE
- priyank-m/chinese_text_recognition
- ProLongVid
- qingy2024/QwQ-LongCoT-Verified-130K
- qwedsacf/competition_math
- ReCTS
- reilxlx/chinese-meme-description-dataset
- RICO ScreenQA
- RoboMIND
- RobuT
- Roman1111111/gpt-5.4-step-by-step-reasoning
- SceneWalk
- ScienceQA
- SciTSR
- ServiceNow-AI/R1-Distill-SFT
- ShanghaiTech Crowd Counting Dataset
- Share14/ShareGemini
- ShareGPT4Video
- ShareGPTVideo/train_video_and_instruction
- Sherlock
- shreyanshu09/Block_Diagram
- SimChart9K
- SIMS-VSI
- skvarre/movie_posters-100k
- Skyhigh-2203/MiMo-2.5-Pro-Reasoning-Traces-Hard
- SlideVQA
- Something-Something V2
- SpatialVID-HQ
- SROIE
- ST-VQA
- STVQA-7K
- SuperSecureHuman/competition_math_hf_dataset
- TabMWP
- Tarsier2-Recap-585K
- TAT-DQA
- TeichAI
- TencentARC/SEED-Bench-R1
- TextCaps
- thuzhizhi/DAPO-MATH-17k-oss-reasoning
- TIGER-Lab/Mantis-Instruct
- TIGER-Lab/VisualWebInstruct
- TinyChart
- Tongyi-DataEngine/SA1B-Dense-Caption
- UniChart pretrain data
- UReader
- URSA-Alignment-860K
- Vary-600k
- Video-R1-260k
- VideoChat2-IT
- VideoInstruct100K
- videollm-online-chat-ego4d-134k
- VidGen-1M
- vidore/colpali_train_set
- Visual Genome
- VizWiz-VQA
- VQAonBD 2023
- WeatherQA
- WikiArt
- wikimedia/wikipedia
- WildVision/vision-chat-0506_00_00_001
- WordArt
- YesBut
- yuecao0119/MMInstruct-GPT4V
- zake7749/Qwen3-Coder-Next-Open-Code-SFT
- 书生·万卷