Adding "Do not guess" cut made-up fields from 71% to 20% in web extraction
What happened
Earn an Honest Dollar 用 16 个模型和 3 个付费 API 做了个网页信息提取测试,专门看它们遇到页面上没有的字段时会不会老实说“没有”。不加“别瞎猜”指令时,573 个缺失字段里模型编造了 405 个,比例高达 70.7%;加上这句指令后,574 个缺失字段里只编造了 116 个,降到 20.2%。表现最好的是 Gemini ...
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pagePickAdding "Do not guess" cut made-up fields from 71% to 20% in web extraction
Earn an Honest Dollar tested 16 models and 3 paid APIs on web extraction with missing fields. Without "Do not guess," models invented 405 of 573 missing fields (70.7%). Adding the instruction dropped that to 116 of 574 (20.2%). Gemini 3.8 Flash made up only 1 of 36 missing fields at $0.16. Firecrawl, a paid API, made up 24/36—worse than 13 models. A cheap checker using GPT-6 Luna caught 38 of 49 made-up values with zero false rejections, costing $0.0049. The post notes these are synthetic pages; real-site results may differ.
Heat over time
Not enough continuous observations to draw a trend yet.