MIT's Gaussian probing flags models fine-tuned on CSAM with 100% accuracy, no images generated
MIT's New Method Flags AI Models Trained on CASM Without Generating It
MIT and child safety nonprofit Thorn built Gaussian probing, an audit method that identifies models fine-tuned to generate CSAM with 100% accuracy. It feeds random data into a model and analyzes internal shifts caused by LoRA adaptors, never producing an image. This sidesteps the legal paradox of generating CSAM to test for CSAM. NCMEC received over 1.5 million AI-generated CSAM reports in 2025, up from 67,000 in 2024. The technique can be integrated into platforms like Hugging Face or Civitai to screen uploads automatically. Evasion would require altering the base model architecture, a much higher bar than prompt tweaking. The paper was presented at ICML.
Why it matters: MIT and Thorn's Gaussian probing method breaks a long-standing deadlock in safety detection: previously, checking whether a model was trained on illegal material required generating illegal content. Now it works by analyzing internal artifacts left by LoRA fine-tuning, with th...