Skip to content

Latest commit

 

History

History
211 lines (208 loc) · 15.8 KB

File metadata and controls

211 lines (208 loc) · 15.8 KB

Open-Source Training Datasets

MOSS-VL's training data consists of (1) open-source datasets collected from the community and (2) in-house synthetic data. This page lists the 205 open-source datasets used in our training.

Note: Before being used in training, these open-source datasets went through our internal curation pipeline, including format unification, quality filtering, deduplication, and additional cleaning and re-annotation. As a result, the actual training samples may differ from the raw public releases.