GitHub Copilot cuts AI coding costs to one-third with preference-trained small models
GitHub Copilot 如何在不牺牲任务质量的前提下降低 AI 编码成本
GitHub published an engineering blog detailing how they cut Copilot's AI coding costs to roughly one-third without hurting task quality. The key move: training a 1.8B-parameter model on 1,040 preference pairs to act as a router that decides when to use a cheap model and when to call a stronger one. After rollout, strong-model calls dropped 70% and overall latency stayed under 11 seconds. The post also mentions a training method called DV-DPO that uses preference data to teach a small model a specific response style. One caveat: these numbers come from GitHub's own setup, so your mileage may vary.
Why it matters: GitHub shared a concrete cost-optimization engineering post with real numbers and methods, directly useful for teams shipping AI products. Score capped because it's an engineering optimization, not a new model release.