Skip to content
r/LocalLLaMA

DeepSeek released Thinking with Visual Primitives framework

DeepSeek released 'Thinking-with-Visual-Primitives' framework

DeepSeek, Peking University, and Tsinghua released the Thinking with Visual Primitives paper and repository. The framework inserts coordinate points and bounding boxes into chain-of-thought; the post does not disclose benchmark scores.

Why it matters: HKR-H/K/R all pass: the hook is visual primitives inside reasoning, the new fact is point/box CoT plus an open repo, and the audience cares about grounded VLMs. No benchmark scores are disclosed, so it stays at 80, not P1.

Read the original ↗Export Markdown