ACL 2026 Oral: LLMs still stumble on phrase-level semantic reasoning
ACL 2026 Oral|语义推理如鲠在喉:大模型被「短语」难住了
SemanticQA, an ACL 2026 Oral paper, stress-tests frontier models on phrase semantics. GPT-5 nails idiom classification at 85.4% but drops to 78.7% on extraction and 22.5% on interpretation. DeepSeek-R1's accuracy collapses from 81.7% to 35.4% when moving from 4-way to 16-way classification. The study breaks semantic understanding into extraction, categorization, and interpretation—no model handles all three consistently. In multi-step pipelines, upstream extraction errors cascade: GPT-5's end-to-end similarity score falls to 17.3%. Authors from BIGAI and USTB note the static benchmark is already insufficient for 2026 agent workflows.
Why it matters: ACL 2026 Oral paper with counterintuitive findings on phrase-level semantic understanding in GPT-5 and DeepSeek-R1. Concrete numbers across three tasks. Held back from higher bands because it's a single paper without cross-source pickup, and pure academic benchmarking has limi...