Skip to content

AI coding

Everything about AI writing code: coding assistants, vibe coding, code model evals and new developer workflows.

1,196 picksRelated topicsAgentsCursorTutorials

Latest picks

221–240 of 1,196

Aug 23Sunday

Hacker News front page

I gave Qwen 3.8 27B a reverse-engineering job I assumed needed a frontier model, and it finished in 30 minutes

The author ran Qwen 3.8 27B on a single Lenovo ThinkStation PGX and tasked it with reverse-engineering a commercial app's license check. The model initially refused, but after the author posed as the developer, it built a working bypass in 30 minutes and fixed its own mistakes along the way. Inference reached ~50 tokens/s with SGLang, NVFP4, and DFlash2. The post doesn't name the app or detail the license mechanism.

Why it matters: A first-person experiment with concrete numbers, not a marketing piece. Qwen 3.8 27B ran a reverse-engineering job locally in 30 minutes at ~50 tok/s — enough substance. But XDA is a consumer tech outlet, not a primary AI source, and the reverse-engineering angle is niche, so ...

Bloomberg Technology

Mystery model Ox Alpha draws developers with free access

An unknown model called Ox Alpha appeared on the LMSYS leaderboard, beating GPT-5.1 and Gemini 3.0 Pro on math and coding benchmarks, and it's completely free. No one knows who built it—the website is just a cow photo and an email. Developers are speculating it could be an anonymous test release from a major lab, but the article doesn't disclose model size, training data, or who actually runs it.

Why it matters: Anonymous model Ox Alpha beats GPT-5.1 and Gemini 3.0 Pro on LMSYS math/coding benchmarks with free access — the mystery factor is high. Bloomberg coverage adds credibility, but the post doesn't disclose model size, training data, or who runs it, capping the score at the featu...

Hacker News front page

Simon Willison launches Agentic Engineering Patterns to document best practices for coding agents

Simon Willison started a project to document coding patterns for agentic tools like Claude Code and OpenAI Codex. He draws a line between vibe coding—ignoring the code entirely—and agentic engineering, where professional devs use agents to amplify their expertise. The first two chapters cover how near-zero code generation cost reshapes team intuition, and how test-driven development helps agents produce tighter, more reliable code. The content lives as a 'guide' with chapters designed to be updated over time. Willison says all prose is his own; LLMs only assist with proofreading and code examples. The post doesn't specify a timeline for future chapters.

Why it matters: Simon Willison launches a structured guide to agentic engineering patterns, drawing a clear line from vibe coding. Directly useful for devs using Claude Code. Score capped here because it's a project launch with concepts defined but patterns not yet fleshed out—worth revisitin...

Hacker News front page

Prime Intellect ran 153 autonomous NanoGPT optimizer runs across 18 frontier models

Prime Intellect ran a speedrun where 18 frontier models, each paired with a coding agent harness, autonomously optimized a NanoGPT training script to minimize validation loss. Fable 5 with Claude Code reached a loss of 2,726, closing 81.7% of the gap from the human baseline to the optimal record. Opus 5 and Kimi K3 followed at 53.6% and 52.2%. DeepSeek V4 Pro closed only 12.3%, ranking 13th. The experiment consumed hundreds of millions of tokens, with the longest run lasting 8.7 days. Caveat: this benchmark tests a narrow skill—autonomous optimization of known code—not general capability, but it reveals how different models handle stability and exploration over long-horizon tasks. The post does not disclose the specific optimization strategies each model used or why certain runs failed.

Why it matters: Prime Intellect's speedrun pits 18 models against each other on autonomous code optimization, with Fable 5 + Claude Code posting a dominant lead. The concrete numbers, leaderboard, and methodology make this genuinely useful for agent builders and benchmark watchers. Not scorin...

Hacker News front page

LLMs make language choice less consequential, pushing devs toward Rust, Zig and harder tech

Armin Ronacher notes that LLMs erase the friction of learning a language, so devs increasingly pick based on speed marketing. Rust is gaining, and Zig appears in Cloudflare Artifacts (a ~100 KB Wasm Git engine) and Vercel fx, both LLM-assisted. Harder tech like DWARF, eBPF and custom crypto is now accessible to more people. His take: more slop, but also more devs who want things fast and small.

Why it matters: Armin Ronacher's observation is backed by named projects, not just vibes. The core insight—LLMs lower language-switching cost—isn't new, but he traces a downstream effect: devs now pick languages based on speed marketing, and Rust/Zig benefit. Missing piece: how many actually ...

Aug 22Saturday

Hacker News front page

Software has no excuse to be slow anymore — Dan Luu shows AI makes deep perf work cheap

Dan Luu took FRE, a regex engine built by an agent loop over a month, and had AI wire AOT compilation into ripgrep in minutes — 2–4× faster on long queries, ~7% on real mixed queries. He also built the world’s strongest Azul AI with zero game-AI background using GPT-5.1/5.2-era models, winning mostly on cheap-to-him-now optimizations like multithreading and native compilation. Marc Brooker and Michael Malis agree: JIT compilers, custom indexes, and other formerly rare-skill work can now be done by anyone typing a few sentences. The post doesn’t give full holdout benchmark numbers for FRE or Elo/match counts for the Azul AI.

Why it matters: Dan Luu ran a concrete experiment: AI dropped AOT compilation into ripgrep in minutes, yielding 2-4x on long queries but only ~7% on real mixed workloads. Title is provocative but the data is honest. Good for perf-minded readers to debate AI tuning boundaries. Not p1 because i...

Aug 21Friday

AI HOT (Curated Pool)

Anthropic publishes the AI-Native SDLC playbook, showing how it builds software with Claude

Anthropic open-sourced its internal playbook for building software with Claude, covering every phase from requirements and design through coding, testing, and ops. The post lays out concrete practices and team structure shifts. No quantitative benchmarks are disclosed—treat this as a methodology guide, not an independent evaluation.

Why it matters: Anthropic open-sourced their internal SDLC playbook with full-lifecycle practices—directly useful for teams using Claude Code. But zero metrics disclosed, making it a methodology guide rather than an independent evaluation, so it lands right at the featured threshold.

Hacker News front page

Stop Making TUIs: AI-Generated Native GUIs Are the Real Deal

The author built 7 native macOS apps with AI, from a Markdown viewer to an Apple TV remote, without writing a single line of UI code. He argues the TUI era should end: just screenshot a design and give it to Claude. The post doesn't provide performance or compatibility data, but shows real integrations like SQLite backends, virtual filesystems, and embedded LLM agents.

Why it matters: The screenshot-to-SwiftUI workflow is genuinely reproducible and backed by 7 real apps, which is stronger than a pure opinion piece. Score capped at 72 because no performance or compatibility data is provided, and the title reads more like a manifesto than an evaluation.

Hacker News front page

OpenRouter lists anonymous reasoning model Ox Alpha, currently free

Ox Alpha is an anonymous third-party reasoning model focused on coding, long-horizon agentic work, and production workloads. It's currently free on OpenRouter with a 1M context window, 1s P50 latency, and 69 tok/s throughput. The post doesn't disclose who built it, parameter count, or how long the free tier lasts. Usage data shows Claude Code and Nous Research's Hermes Agent already pushing significant token volume—treat it as a coding agent model worth testing, but anonymity and free pricing make long-term reliability uncertain.

Why it matters: Anonymous reasoning model drops free with 1M context, 1s P50 latency, 69 tok/s, targeting coding and sustained agentic work. Developer, param count, and free-tier duration all undisclosed — big info gaps — but Claude Code and Nous Research are already using it, so it's not vap...

Aug 20Thursday

Hacker News front page

Slack Code turns AI coding into a multiplayer team activity

Salesforce added Code channels to Slack so coding agents like Claude Code, Devin, and ChatGPT can write code, show diffs, and run live previews inside a channel visible to the whole team. Channels are project-based and auto-archive when done. The post does not disclose pricing or launch date.

Why it matters: Slack pulls coding agent workflows into channels, solving the 'agent works in a black box' pain point for teams. Product thinking is clear, but the post gives no pricing or launch date — it's an announcement, not a release, so score stays below 80.

Hacker News front page

Building a custom watch face on a $27 PineTime with Claude

Mike Kasberg used OpenCode with open-weight models—Kimi K3, K2.6, DeepSeek v4 Pro and Flash—to build a Casio-style watch face for the $27 PineTime. He started by getting a build working in the InfiniSim simulator, then fed the model a reference photo to replicate the layout. The first attempt was rough: text sizing and positioning were guessed, making elements overlap and unreadable. He switched to giving isolated, concrete feedback and fixed one text element at a time. Later he turned static parts into a fullscreen 240x240 background image so only dynamic elements needed code. It worked in the simulator, but on real hardware the image took 10 minutes to transfer over Bluetooth and screen refreshes lagged 1–2 seconds; the watch can't hold the whole image in memory and streams it from flash. He calls it a working prototype, pushed the code to GitHub, and had the model summarize lessons learned into an AGENTS.md file.

Why it matters: A first-person experiment with concrete debugging details, not a tutorial roundup or promo. H and K are solid, but the niche audience and lack of cross-source coverage keep R from hitting, so it lands right at the featured threshold.

Latent Space

Z.ai CEO Jie Tang on GLM 5.3: The era of parameter counting is over, post-training is the new scaling law

Jie Tang posted a long thread on X arguing that parameter count alone is meaningless—you need data volume, compute allocation, and deployment conditions. GLM-5.3's gains come entirely from RL on long-horizon environments, some simulating days of engineer work. They built synthetic pipelines that auto-generate executable, verifiable environments and reward signals, pushing the model to own complex tasks end-to-end. Tang identified 5 scaling knobs including MoE sparsity, and noted that finding software vulnerabilities requires holding 20+ inference-step causal chains, not memorization. The post does not disclose GLM-5.3's exact parameter count or release date.

Why it matters: Jie Tang personally explains GLM 5.3's post-training scaling law with concrete experimental cases (simulated cluster diagnosis and optimization), not just rhetoric. But the source is a paid newsletter excerpt, and key numbers (specific speedup ratios, task success rates) aren'...

Hacker News front page

LLMs turn static software into something users can extend by asking

Jeremy Morrell argues that most web apps only serve the top of the demand curve, leaving a long tail of unmet needs. LLMs now let users “speak code into existence,” so a product can keep a stable core while AI-generated extensions cover the long tail. He points to Pi and DeepSeek Harness as early examples: users ask for a feature, the system writes a TypeScript extension and hot-reloads it. The post then focuses on how to bring this to the web—traditional webhooks set the bar too high; sandboxed runtimes like Cloudflare Dynamic Workers could let non-developers safely run custom logic. No timeline or performance numbers are disclosed; the piece is a product-direction argument.

Why it matters: Opinion piece backed by two named product examples, not empty theory. Hits all three HKR axes but is more 'thought-provoking' than 'industry-shaking,' so placed at the lower end of the 72–77 band per policy.

Aug 19Wednesday

Hacker News front page

Ornith-1.5 uses self-generated tasks for RL, with three model sizes beating comparable open-source models on coding and agent benchmarks

Ornith-1.5 extends the self-scaffolding idea from Ornith-1.0 into a full self-improvement loop: the model proposes tasks, builds scaffolds, generates solution rollouts, and improves via RL. The 397B MoE flagship scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, matching Claude Opus 4.8 (85.0, 59.0) and beating GLM-5.2 and DeepSeek-V4-Flash-0731. The 35B MoE activates only 3B parameters per token yet outperforms Gemma 4-31B and Meta Muse Glimmer-30B on agentic coding. The 9B dense model has a quantized mobile version that runs on phones and scores 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-Bench Verified, beating many larger models. Task reward multiplies validity, frontier difficulty (targeting a 20% success rate), and novelty. The post does not disclose training compute, data scale, or a release timeline.

Why it matters: Ornith-1.5 turns self-improvement into a full loop, and the 397B variant essentially matches Claude Opus 4.8 on Terminal-Bench and DeepSWE — a real open-source catch-up moment. Score isn't higher because Ornith isn't a tier-1 lab yet; community reproduction and real-world depl...

OpenAI News

Replit launches Free Mode powered by GPT-5.6 Luna, removing token costs for software creation

Replit introduced Free Mode running on GPT-5.6 Luna, so users can plan, ideate, and explore projects without tracking token spend. CEO Amjad Masad credits recent OpenAI price cuts for making the free tier viable at millions-of-users scale. Complex reasoning tasks get routed to GPT-5.6 Sol, then return to Luna while preserving project context. Sam Altman frames it as a step toward anyone with internet building a product or startup. The post does not disclose Free Mode quotas, concurrency limits, or the exact launch date.

Why it matters: Replit's free tier running GPT-5.6 Luna is a concrete product update with a real mechanism (dual-model handoff) and a direct CEO quote on cost economics — enough signal for featured. But it's an OpenAI customer story, not a model release, so the score stays at 72.

Hacker News front page

Bun 1.4 Rust rewrite is three months late and the community is losing trust

Bun has gone three months without a stable release for the first time since 2022. Founder Jarred Sumner has been promising v1.4 since June, but dates keep slipping and community replies now openly mock the repeated 'tomorrow' promises. The rewrite moves the codebase from Zig to Rust. In the past month, 15.8k commits came from robobun, 1.6k from autofix-ci[bot], and only 790 from Jarred; he himself noted most PRs are now Claude prompting Claude. Zig creator Andrew Kelley called the original Bun code 'hacks on top of hacks' and said Jarred was writing slop before LLMs. The author argues the rewrite looks more like an Anthropic ad than a genuine memory-safety fix, pointing to the number of unsafe blocks in the new Rust code. The project now has over 5,000 open PRs, far exceeding GitHub's recommended 1,000 limit.

Why it matters: Bun's Rust rewrite has left it without a stable release for 3 months — the longest gap since 2022. Founder repeatedly missed ship dates, community is openly mocking, and 15.8k commits came from bots, raising real questions about AI-assisted maintenance on a critical tool. Scor...

Computing Life · Share · Yage

$60B for a Data Flywheel, $10M for a Working Memory

The AI race in 2026 is shifting from compute to real-world behavioral data. SpaceX acquired Cursor for $60B in stock, gaining access to coding interaction traces from over 50,000 enterprise clients—data fed directly into Grok 4.5 and 4.6 training. Google bid $10M in Spirit Airlines' bankruptcy auction for roughly 100M emails, 500M Teams messages, and 30M lines of internal code. Atlassian updated its terms to turn Jira and Confluence usage data into a continuous model-training input. DeepSeek is building its own front-end harness to capture local developer workflows that API-only access misses. The article maps these four paths—acquisition, bankruptcy purchase, SaaS terms, and self-built harness—and notes that the actual model gains from data flywheels, signal retention after de-identification, and real-world compliance friction remain unverified.

Why it matters: SpaceX's $60B all-stock Cursor acquisition frames developer behavioral data as the next piece of the model race, with a concrete pricing benchmark from Google's bid on a bankrupt airline's data. Hits all three HKR axes; the data-flywheel logic is clear, but the piece is synthe...

Hacker News front page

Purely AI-generated code has no author and no copyright under US law

This site walks founders and engineering leads through a hard legal reality: under current US copyright law, code generated entirely by AI has no human author, so it can't be copyrighted or defended as an owned asset. It cites four settled anchors—including the Supreme Court's March 2026 denial of cert in Thaler and the first rejection of a fair-use defense for AI training in Thomson Reuters v. Ross—to show the rule is already locked in. It also breaks down four common blind spots: pure AI output isn't yours, vibe coding where the AI makes creative choices leaves code unprotected, mixed codebases only protect the human-authored parts, and open-source licenses are unenforceable on code no one owns. The post doesn't offer fixes; it's a risk primer with a self-assessment quiz.

Why it matters: This piece connects four settled U.S. copyright rulings into one clear takeaway: purely AI-generated code has no author and therefore no copyright. For teams shipping heavily AI-assisted code daily, this is an overlooked but high-stakes legal reality. Not scored higher because...

Hacker News front page

Vercel open-sourced fx, a 6.39MB minimal coding agent in Zig

fx is a Zig-based CLI coding agent that weighs 6.39MB, cold-starts in 10µs, and uses single-digit MB of memory. It's model-agnostic, runs locally or in the cloud, and compiles to WebAssembly for browser use. The design leans Unix: minimal output, no heavy TUI, built to be embedded into larger systems. Currently at v0.0.3 and marked experimental—the team warns of frequent breaking changes, so hold off on production use.

Why it matters: Vercel Labs open-source agent harness: 6.39MB binary, 10µs cold start, Wasm support. Clean technical choices. Not scored higher because it's v0.0.3 experimental with no usage data and no discussion cluster yet.

AI HOT (Curated Pool)

Mojo language is now fully open source under Apache 2.0

Modular open-sourced the entire Mojo compiler and toolchain under Apache 2.0 with LLVM exceptions. All source code is now in the modular GitHub repo. Mojo hit 1.0 last week with source stability guarantees. The permissive license lets developers freely build and distribute Mojo-compiled binaries. The post does not spell out community governance or external contribution workflows.

Why it matters: Full open-sourcing right after the 1.0 release, under Apache 2.0 with an LLVM exception — that removes the commercial distribution friction and sends a real signal to devs who want one language for CPU and GPU. Not scoring higher because we only have the official announcement ...