Perplexity and OpenAI's PII detectors are not LLMs but bidirectional encoders with classification heads
上云前先脱敏:拆开 Perplexity 和 OpenAI 的 PII 检测器,它其实不是 LLM
Perplexity's open-source pplx-pii-masking is a 0.6B-parameter bidirectional encoder built on Qwen3 with causal masking disabled, topped with a token classification head and a document sensitivity head. It uses Viterbi decoding to output start/end offsets and confidence scores for 9 PII categories. OpenAI's Privacy Filter is a 1.5B sparse MoE model with ~50M active parameters and a nominal 128K context window, but its banded attention limits each token's effective view to 257 tokens. In tests, both models missed bare API keys and produced slice offsets; pplx silently truncates inputs beyond 4096 tokens, while OpenAI mislabeled an account number 550 tokens away from its context label as a phone number. The takeaway: on-device PII protection needs small classifiers for natural-language entities plus regex and entropy checks for fixed-format secrets.
Why it matters: The author ran hands-on tests against Perplexity's open-source detector, documenting misclassification, slice offset, and missed keys, then explained why autoregressive LLMs can't natively output per-span confidence. The second half defines requirements but doesn't unpack Open...