Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,229 picksRelated topicsAgentsCursorTutorials

Latest picks

1221–1229 of 1,229

Sep 12, 2024Thursday

OpenAI News

Introducing OpenAI o1

OpenAI released o1-preview and o1-mini on Sept. 12, 2024, with access for ChatGPT Plus, Team, and tier-5 API developers. The post cites 83% vs 13% on an IMO qualifier, 84 vs 22 on a jailbreak test, and says o1-mini is 80% cheaper than o1-preview. The tradeoff is clear: the API lacks function calling, streaming, and system messages, and the models do not yet support browsing or file and image uploads.

Why it matters: OpenAI's o1 models trade speed for deliberation, lifting math olympiad accuracy from GPT-4o's 13% to 83% and reaching Codeforces top 11%.

OpenAI News

Learning to reason with LLMs

OpenAI released o1-preview and reported 74% single-sample accuracy on AIME 2024, versus 12% for GPT-4o. The post says o1 reached the 89th percentile on Codeforces and exceeded human PhD experts on GPQA Diamond; it attributes this to large-scale RL and gains from both train-time and test-time compute. The key signal is scaling reasoning with compute, not just pretraining a larger base model.

Why it matters: OpenAI's o1-preview scores 74% on AIME 2024 math problems against GPT-4o's 12%, rising to 83% with repeated sampling.

OpenAI News

OpenAI o1-mini

OpenAI released o1-mini on Sept. 12, 2024 for Tier 5 API users at 80% lower cost than o1-preview. The post reports 70.0% on AIME and 1650 Codeforces Elo, close to o1 at 74.4% and 1673, with about 3-5x faster answers than o1-preview in one word-reasoning example. The key tradeoff is explicit: it targets STEM reasoning, while non-STEM factual knowledge is only comparable to small models like GPT-4o mini.

Why it matters: OpenAI's o1-mini costs a fifth of o1-preview per API call, opening first to paying ChatGPT and Tier 5 API users, so the price-capability tradeoff is the key detail.

Sep 5, 2024Thursday

DeepSeek · API updates

deepseek-coder and deepseek-chat merge into DeepSeek V2.5

DeepSeek V2 Chat and DeepSeek Coder V2 have been merged and upgraded into DeepSeek V2.5. API users reach the new model through either deepseek-coder or deepseek-chat.

Why it matters: DeepSeek merged its chat and code models into V2.5 while keeping both existing API names so current call patterns still work.

Aug 20, 2024Tuesday

OpenAI News

Fine-tuning now available for GPT-4o

OpenAI has opened GPT-4o fine-tuning to developers on all paid tiers, with 1M free training tokens per org per day through September 23. Training costs $25 per 1M tokens, and inference costs $3.75 per 1M input tokens and $15 per 1M output tokens on gpt-4o-2024-08-06. The signal for practitioners: partners reported 43.8% on SWE-bench Verified and 71.83% on BIRD-SQL with fine-tuned GPT-4o.

Why it matters: OpenAI opens GPT-4o fine-tuning to all paying developers at $25 per million training tokens, with 1 million free tokens per organization per day until September 23.

Aug 13, 2024Tuesday

OpenAI News

Introducing SWE-bench Verified

OpenAI released SWE-bench Verified, a human-validated subset built with the benchmark’s authors to assess real software issue resolution more reliably. The post names 3 failure modes in SWE-bench: overly narrow tests, underspecified issue statements, and unreliable environment setup; as of Aug. 5, 2024, top agents scored about 20% on SWE-bench and 43% on SWE-bench Lite. The key point is that the original benchmark can systematically underestimate coding-agent ability.

Why it matters: OpenAI and the original SWE-bench authors built a human-verified subset, arguing the original systematically underestimates real coding ability due to rigid unit tests, vague descriptions and environment issues.

Jul 18, 2024Thursday

OpenAI News

GPT-4o mini: advancing cost-efficient intelligence

OpenAI released GPT-4o mini on July 18, 2024 at $0.15 per 1M input tokens and $0.60 per 1M output tokens, replacing GPT-3.5 in ChatGPT. It supports text and vision, offers a 128K context window and 16K max output, scores 82.0% on MMLU and 87.2% on HumanEval. The key detail for builders is that its API version is the first to use instruction hierarchy against jailbreaks and prompt injection.

Why it matters: GPT-4o mini's price per million input tokens and its replacement of GPT-3.5 in ChatGPT set a new floor for cheap inference, and the instruction hierarchy shows how OpenAI handles jailbreaks.

Jun 28, 2024Friday

DeepSeek · API updates

deepseek-chat upgraded to DeepSeek-V2-0628, with better reasoning and role-play

The deepseek-chat model has been upgraded to DeepSeek-V2-0628. Reasoning improved, and role-play got a clear boost. HumanEval Pass@1 rose from 79.88% to 84.76%, MATH ACC@1 from 55.02% to 71.02%, and BBH from 78.56% to 83.40%.

Why it matters: The math benchmark gain is larger than the code and reasoning gains, and the win rate against GPT-4-0314 on Arena-Hard also improved.

Jun 14, 2024Friday

DeepSeek · API updates

deepseek-coder upgraded to DeepSeek-Coder-V2-0614

The deepseek-coder model has been upgraded to DeepSeek-Coder-V2-0614, with a clear gain in coding ability. The company says its code generation, code understanding, code repair and code completion reach the level of GPT-4-Turbo-0409, with strong math and reasoning, while general ability matches DeepSeek-V2-0517.

Why it matters: DeepSeek upgrades its code model and benchmarks code generation, understanding, repair and completion against GPT-4-Turbo-0409.