US DOJ says training LLMs on copyrighted text is generally fair use
美国司法部在 OpenAI 与纽约时报版权诉讼中支持训练属于合理使用
The US Department of Justice filed its first statement on AI training and copyright, arguing that training LLMs on copyrighted works is generally fair use. It separates the process into data acquisition, training, and output, noting that training does not substitute for the original work. The DOJ also warns that blanket licensing would raise barriers for smaller companies. The filing is advisory and not binding, but if courts adopt this framework, legal pressure will shift toward how data is obtained and what models output.
Why it matters: The DOJ backs fair use in a landmark copyright case, directly touching the legal foundation of model training. The brief offers a three-step analytical framework and flags the anti-competitive effect of mandatory licensing — high information density. Deduction because it's non...