Skip to content
AI HOT (Curated Pool)

OpenBMB Releases Two UltraData Open Datasets, Tops HuggingFace Trending

OpenBMB发布UltraData两大开源数据集,登顶HuggingFace趋势榜

OpenBMB, Tsinghua NLP, and Modelbest released two UltraData open datasets: Ultra-FineWeb-L3 contains 600B+ tokens, including 400B+ English and 200B+ Chinese tokens, while UltraData-SFT-2605 contains 15M+ SFT samples with thinking and non-thinking labels.

Why it matters: HKR-H/K/R pass: two open datasets, 600B+ tokens, and 15M+ SFT samples are concrete practitioner signal. Single-source release with no evals or license detail keeps it at the lower featured band.

Read the original ↗Export Markdown