Skip to content
Hacker News front page

Show HN: A New Benchmark for Testing LLMs for Deterministic Outputs

Show HN: A new benchmark for testing LLMs for deterministic outputs

Interfaze released Structured Output Benchmark, scoring schema pass rate, types, and value accuracy across text, image, and audio. Each record has a JSON Schema and human plus LLM-checked ground truth; GLM-4.7 ranks No. 2 overall. The key bug is field-level value error: GPT-5.4 ranks 3rd on text and 9th on images.

Why it matters: HKR-H/K/R all pass: the ranking has a hook, the methodology is concrete, and structured-output reliability matters to builders. Single-source Show HN launch with no adoption signal keeps it in the 72–77 band.

Read the original ↗Export Markdown