Trending storyWatching
Nonobench v1.2 tests 43 LLMs on nonogram puzzles: open-weight DeepSeek V4 Pro ties for 4th, no model solves 20×20 Hard
1 report1 sourceupdated 3 days ago
What happened
Summary
一个叫Nonobench的基准测试用数织(Nonogram,类似像素填图谜题)考了43个大模型。开源模型DeepSeek V4 Pro跟几个闭源模型并列第四,但新出的20×20困难模式没有一个模型能解出来。正文只给了排名和难度分级,没披露具体得分和测试样本量,所以这个第四名含金量有多高、困难模式到底多难,目前还不好判断。
Coverage
Follow the reports to see the story from different sides.
Sep 27
- r/LocalLLaMANonobench v1.2 tests 43 LLMs on nonogram puzzles: open-weight DeepSeek V4 Pro ties for 4th, no model solves 20×20 Hard
Nonobench v1.2 benchmarks 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th with some closed models. No model solves the new 20×20 Hard mode. The post doesn't disclose scores or sample sizes, only rankings and difficulty tiers.
Heat over time
Not enough continuous observations to draw a trend yet.