A Reddit user accused Open-OSS/privacy-filter on Hugging Face of delivering a Windows infostealer through loader.py, PowerShell, an EXE, and Task Scheduler.
I'll be blunt: this is not a weird Hugging Face corner case. This is the normal software-supply-chain problem entering the AI model distribution layer. Model repos now bundle weights, Python helpers, demo apps, Gradio code, evaluation scripts, config files, and install instructions. Plenty of users clone a repo and run `pip install -r requirements.txt` or `python app.py` without reading the code. npm and PyPI have been punished by that behavior for years. Hugging Face inherits the same attack surface, with a user base that is often less security-trained than package maintainers.
The source here is thin. The body is only a Reddit 403 page. The original post, screenshots, code snippets, repo commits, hashes, IOC list, Microsoft response, and Hugging Face response are not visible. The title names Open-OSS/privacy-filter. The summary says it mimics OpenAI’s privacy filter, uses loader.py to fetch PowerShell, downloads an EXE, and runs it through Windows Task Scheduler. The summary also says Linux is unaffected. I have not verified the repo, the binary signature, the C2 endpoint, the scheduled task name, or the current takedown state.
Still, the described chain is plausible. Task Scheduler is a standard Windows persistence mechanism. PowerShell as a remote payload launcher is old tradecraft. The AI-specific trick is the label: “privacy filter.” That name borrows trust from OpenAI, privacy anxiety, safety tooling, and open-source culture at once. Local-model users worry about sending data to hosted APIs. That makes them especially receptive to a tool claiming to protect privacy locally. The bait is well chosen.
There is prior context here. Hugging Face has warned for years about pickle and `torch.load` risks, and safetensors exists largely to avoid arbitrary code execution through model loading. The community has also seen malicious model repos, malicious Spaces, and poisoned dependencies. On PyPI, typosquatting and post-install credential stealers are already boringly common. Hugging Face has an extra problem: it is not used only by package-aware developers. Researchers, notebook users, product teams, and automation scripts all click “Use this model.” The UI lowers friction, and attackers love lowered friction.
I do have doubts about the claim as presented. A malware allegation without a VirusTotal link, SHA256, repo snapshot, or platform confirmation is not enough for a final forensic call. Security threads with all-caps “WARNING MALWARE” sometimes spread faster than the evidence. Some Windows helper scripts also look ugly to static scanners. But loader.py fetching PowerShell, dropping an EXE, and registering a scheduled task is far beyond what a normal ML privacy-filter demo should need. Unless the maintainer gives a precise installer rationale, I would treat the repo as hostile.
For practitioners, the operational lesson is concrete. Do not run Hugging Face repo code directly on a primary Windows machine. Inspect loader.py, setup.py, postinstall hooks, PowerShell, batch files, EXEs, DLLs, and any dynamic download path. Prefer safetensors for weights. Run demos in containers or disposable VMs. Restrict outbound network access during first execution. If you touched this repo on Windows, inspect Task Scheduler and startup locations. In a company environment, Hugging Face downloads need software supply-chain scanning. Treating the platform as a “model website” is too naive now.
The uncomfortable part is that open-source AI culture trains people to run first and audit later. Rankings, README snippets, one-click demos, and notebook cells all reward speed. Attackers are simply adapting. Hugging Face can add malware scanning, repo warnings, pickle warnings, and takedown workflows, but arbitrary Python inside model repos means the platform cannot absorb all user risk. A model repo is not automatically trustworthy because the weights are open. Weights are one asset class; executable code is another. Once a repo asks you to run scripts, you are in supply-chain territory, not model-download territory.