Skip to content
AI HOT (Curated Pool)

Pulpie: Pareto-optimal models for web cleaning match SOTA quality at 1/20 the cost

Pulpie:用于清理网络的Pareto最优模型

Feyn Labs open-sourced Pulpie, a family of models for extracting main content from HTML. The smallest variant, pulpie-orange-small (210M params), scores 0.862 ROUGE-5 F1 on WebMainBench, nearly matching Dripper's 0.864 at 600M params. Speed is the headline: 13.7 pages/sec on an L4 GPU vs. Dripper's 0.68 pages/sec. At $0.39/hr per L4 instance, cleaning 1B pages costs $7,900 with Pulpie and $159,000 with Dripper. Pulpie uses an encoder that labels every HTML block in a single forward pass; Dripper's decoder emits labels token-by-token, bottlenecked by memory bandwidth. Training labels came from DeepSeek V3.2 and Dripper 0.6B cross-labeling 15,880 Common Crawl pages, with 93.3% block-level agreement. The post also cites an AICC experiment: training on model-extracted corpora improved average accuracy by 1.08 pp across 13 benchmarks, beating FineWeb and RefinedWeb. Caveat: only English pages were tested; multilingual performance isn't disclosed.

Why it matters: Feyn Labs open-sourced a family of HTML content extraction models. The smallest variant (210M params) nearly matches the current SOTA Dripper (600M params) at 1/20th the inference cost. Concrete benchmarks and throughput numbers make it actionable for teams doing RAG, crawling...

Read the original ↗Export Markdown