Skip to content
Trending storyWatching

Nonobench v1.2 tests 43 LLMs on nonogram puzzles: open-weight DeepSeek V4 Pro ties for 4th, no model solves 20×20 Hard

1 report1 sourceupdated 3 days ago

What happened

Summary

一个叫Nonobench的基准测试用数织(Nonogram,类似像素填图谜题)考了43个大模型。开源模型DeepSeek V4 Pro跟几个闭源模型并列第四,但新出的20×20困难模式没有一个模型能解出来。正文只给了排名和难度分级,没披露具体得分和测试样本量,所以这个第四名含金量有多高、困难模式到底多难,目前还不好判断。

Coverage

Follow the reports to see the story from different sides.

Sep 27
  1. r/LocalLLaMA
    Nonobench v1.2 tests 43 LLMs on nonogram puzzles: open-weight DeepSeek V4 Pro ties for 4th, no model solves 20×20 Hard

    Nonobench v1.2 benchmarks 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th with some closed models. No model solves the new 20×20 Hard mode. The post doesn't disclose scores or sample sizes, only rankings and difficulty tiers.

Heat over time

Not enough continuous observations to draw a trend yet.