The counterintuitive bit here is that OpenAI—the company with the most compute and talent—chose the laziest possible entry into legal AI. No pretraining, no fine-tuning. Just GPT-6 Astra with a 230M-URL legal index and tuned system instructions. On Vals AI's 200-question private set, it hit 54% all-pass, 15.3 points above the base model. But these are self-reported numbers with no third-party verification yet, so I'd discount them a bit.
The real value of this piece is the autopsy of four years of legal AI spending. BloombergGPT burned massive cash pretraining a 50B-parameter model from scratch on 363B tokens of financial data—never made it into any commercial product. Harvey's early approach of full fine-tuning on case law got steamrolled every time a new frontier model dropped. Two waves of failure, two clear lessons: don't bet against general-purpose foundation models, and don't bake fast-changing legal facts into static weights.
Three surviving bets emerged. Thomson Reuters and LexisNexis bet on content moats—1,500+ human editors, citation graphs like KeyCite and Shepard's. Harvey pivoted to post-training with RL, shaping behavioral patterns rather than injecting knowledge, using 150 B300 GPUs over two months. And all four tech giants—OpenAI, Microsoft, Anthropic, Google—chose peripheral config only, touching zero weights. Astra for Law kills simple API wrappers, but Harvey at $400M ARR with deep workflow integration is defensible. The blunt takeaway: most hard problems in legal AI sit outside model weights, in data freshness, citation verification, and workflow fit.