DoorDash uses LLM juries and multimodal AI to tag food items, beating human accuracy by 20%
Building Food Metadata with LLM Juries
DoorDash's catalog has millions of items with wildly inconsistent names, making manual tagging slow and expensive. They built an AI metadata platform that uses multimodal models—text, images, and web search—to infer attributes like spiciness or cuisine type. Instead of human review, an 'LLM jury' of multiple strong models votes independently and aggregates a consensus, lifting annotation accuracy roughly 20% above typical human reviewers. Context-optimization agents iterate prompts in minutes, adding another 20%+ precision gain and speeding up prompt development 10x. They auto-generate training data to fine-tune small models that match frontier LLM quality at 10% of the inference cost. Distributed inference cut backfill time for millions of items from over a month to a few days. The post doesn't disclose which models, latency, or per-item cost.
Why it matters: DoorDash engineering blog shares a practical LLM-jury + multimodal labeling pipeline with concrete accuracy numbers and auto-prompt-iteration. Useful for AI data pipeline builders, but the food-delivery domain limits audience breadth — R axis missed, so it lands at the feature...