Nonobench v1.2 tests 43 LLMs on nonogram puzzles: open-weight DeepSeek V4 Pro ties for 4th, no model solves 20×20 Hard
Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode
Nonobench v1.2 benchmarks 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th with some closed models. No model solves the new 20×20 Hard mode. The post doesn't disclose scores or sample sizes, only rankings and difficulty tiers.