Perplexity open-sources Unigram tokenizer to reduce CPU usage
Perplexity开源Unigram分词器降低CPU占用
Perplexity open-sourced a rebuilt Unigram tokenizer that reduces CPU usage by 5-6x, targeting tokenization latency when small rerankers and embedding models run on GPUs in single-digit milliseconds.
Why it matters: HKR-H/K/R all pass: the 5-6x CPU claim and tokenizer bottleneck are concrete for production RAG/search teams. It stays in the featured-threshold band because the post lacks independent benchmarks, repo details, and deployment scale.