Prompts Aren't Real: Build Evaluation Pipelines Instead
Prompts Aren't Real
Dan McKinley argues that prompt engineering is a distraction. Building consumer-facing agents taught him that even structured output fails on a fraction of requests—models will flood a field with nonsense. His fix was renaming a field from 'title' to 'heading,' which he calls deranged. The talk pushes for pass^k testing and evaluation pipelines to constrain behavior, since prompts alone can't tame the beast. The post is a slide deck; it names no specific eval frameworks or metrics.
Why it matters: Dan McKinley's first-hand production experience with concrete cases and numbers, sharp opinion. But it's a personal talk, not a formal publication, and the post doesn't disclose pass^k test pass rates or scale — slight deduction.