Skip to content
r/LocalLLaMA

I built a 103B-token Usenet corpus from 1980–2013

I built a 103B-token Usenet corpus (1980–2013) — pre-web, human-only, zero AI contamination. Got strong traction on r/ML, thought this community would find it useful.

OwnerByDane released a 103.1B-token Usenet corpus covering 1980–2013, 408M posts, and 18,347 newsgroups, with free 5K-post-per-hierarchy samples and full-corpus licensing available.

Why it matters: HKR-H/K/R all pass: the zero-contamination corpus has a clear hook, concrete scale, and relevance to training-data scarcity. Score is capped by Reddit-only sourcing, licensed full access, and no third-party validation or benchmark results.

Read the original ↗Export Markdown