Skip to content

All news

8 today

Jul 23Thursday

Latent Space

Poolside co-CEO on how a 70-person team ships a 118B MoE model in 8 weeks

Poolside co-CEO Eiso Kant walked through their 'Model Factory' on the Latent Space podcast. A team of fewer than 70 researchers runs 10,000–20,000 experiments per month, cutting model cycles from six months to five to eight weeks. Their new Laguna S 2.1 is a 118B-total, 8B-active MoE model with a 1M context window and dual thinking/no-thinking modes, beating Thinking Machines' ~1T open-weights model. Eiso argued 95% of model building comes down to better data or compute efficiency, called MCP and traditional tool calls 'stupid,' predicted RL will move earlier into pre-training, and said he'd rather live in a world with 100 foundation model companies than five.

Why it matters: Poolside opens up its model factory internals for the first time, with real numbers on the 118B MoE architecture and 5-8 week iteration cadence — useful for anyone doing model training or code tooling. Not scoring higher because Poolside's audience is still code-niche, and thi...

Jul 21Tuesday

Google DeepMind

Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber

Google DeepMind released three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and the security-focused 3.5 Flash Cyber.

Why it matters: It gives pricing, token efficiency and benchmark comparisons for all three models, so readers can judge cost and model choice for agent workflows.

Jul 20Monday

AI HOT (Curated Pool)

Xiaohongshu and Peking University open-source UltraEP for real-time MoE load balancing

Xiaohongshu and Peking University open-sourced UltraEP, a real-time load balancing method for large MoE models. It dynamically replicates hot experts per microbatch and per layer using exact routing info, hitting 94.6% of ideal training throughput. On Qwen3-235B, training throughput is 42% higher than Megatron-LM and prefill throughput is 1.56× SGLang. The post doesn't disclose the license or deployment requirements.

Why it matters: Joint open-source release from Xiaohongshu and Peking University tackles real MoE idle-GPU pain with strong numbers (94.6% ideal training throughput, 1.56x SGLang inference). Missing license and deployment requirements keep it from scoring higher—those determine real-world ado...

Hacker News front page

Xiaomi drops XR-1, a robot foundation model pre-trained on 100K hours of embodiment-free data

Xiaomi Robotics open-sourced XR-1, a ready-to-use robot foundation model. It pre-trains on 100K hours of embodiment-free manipulation videos across 1,700+ scenarios, then post-trains on 7,200 hours of real-robot data for embodiment and instruction alignment. Pre-training shows clean scaling laws—lower action error with more data and larger models—and those gains transfer directly to real-robot success rates with no sign of saturation yet. After post-training, XR-1 picks up new tasks like phone packing and printer refilling from under 10 hours of demos on average, hitting 75% overall success (nearly 2× π 0.5); with under 40 hours it reaches 85%. It also achieves SOTA on four sim benchmarks. Code, weights, and paper are public.

Why it matters: Xiaomi Robotics open-sourced XR-1, a robot foundation model with code, weights, and paper. The two-stage recipe (100K hrs embodiment-free pretraining + 7,200 hrs real-robot post-training) and the pretraining scaling law are hard signals, directly comparable to π 0.5. Scored as...

Jul 19Sunday

Computing Life · Share · Yage

Grok Build open-sourced its client harness, not the model or cloud

xAI released the Rust client harness that handles local files, commands, and permissions for Grok Build under Apache-2.0. The Grok model, cloud services, and the official binary build chain remain closed. The repo doesn't accept external PRs. The commit from the earlier upload controversy isn't in the public history, so the current code can't close that case. The real win: you can now pin a public commit, build it yourself, and compare its behavior against the official binary.

Why it matters: xAI open-sourcing Grok Build's client harness is substantive—Apache-2.0, headless mode, and ACP support go beyond signaling. But the model and build chain remain closed, and the repo rejects PRs, capping it below 85. All three HKR axes hit, so featured.

Jul 17Friday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash Cyber security model

Google DeepMind released Gemini 3.5 Flash Cyber, fine-tuned from 3.5 Flash to find, verify and patch vulnerabilities quickly. With multiple calls, it approaches larger models on benchmarks such as CyberGym.

Why it matters: It reports how a lightweight security model performs on several benchmarks and inside Google's own codebase, so readers can judge the cost-benefit for vulnerability discovery.

Hacker News front page

Moonshot AI launches 2.8T-parameter Kimi K3, calling it the first open 3T-class model

Moonshot AI released Kimi K3, a 2.8T-parameter model and the most expensive from a Chinese lab so far at $3/$15 per million input/output tokens—matching Claude Sonnet pricing. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 but lose to Claude Fable 5 and GPT-5.6 Sol. On Artificial Analysis's private long-horizon knowledge eval, K3 hit an Elo of 1547, +732 over K2.6, behind only Fable 5. Cost per task is $0.94, close to GPT-5.6 Sol's $1.04 and roughly half of Opus 4.8. Output tokens dropped 21% vs K2.6. The model only offers a 'max' reasoning effort; Simon's pelican-on-a-bike SVG cost 25 cents and burned 13,241 reasoning tokens. Input token count suggests an ~85-token hidden system prompt. Vision works well. Open weights promised by July 27.

Why it matters: Moonshot AI released Kimi K3, a 2.8T-param model priced at $3/$15 — matching Claude Sonnet and making it the most expensive Chinese lab model. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 high, but lose to Claude Fable 5 and GPT-5.6 Sol. Simon Willison's ta...

AI HOT (Curated Pool)

Kimi K3 tops frontend coding leaderboard, open weights coming July 27

Kimi K3 scored 1679 on Frontend Code Arena, taking first in 6 of 7 frontend sub-tasks and beating Claude Fable 5 and GPT-5.6 Sol. It's a 2.8-trillion-parameter MoE model with a 1M context window, and open weights are promised for July 27. API pricing is $15 per million tokens—no low-cost play here, it's priced against top closed-source models and aimed at long-context coding and agent workflows.

Why it matters: Moonshot AI's Kimi K3 tops Frontend Code Arena at 1679, winning 6 of 7 subtasks against Claude Fable 5 and GPT-5.6 Sol. 2.8T MoE params, 1M context window, weights opening July 27. A domestic flagship model directly challenging the closed-source duopoly on a concrete coding be...

Jul 16Thursday

Hacker News front page

The LLM Critics Are Right. I Use LLMs Anyway

At Local-First Conf in Berlin, the author noticed a shared dissonance: speakers criticized LLMs while the audience applauded with Claude Code open. He concedes every critique—slop, trust erosion in OSS, broken junior-senior teaching loops, geopolitical supply risks—yet still uses LLMs heavily. The post doesn't resolve the tension; it lays out the contradiction and asks others to share their usage patterns so the community can better understand this collective unease.

Why it matters: An honest personal observation that lays out the collective dissonance devs feel about LLMs, with a concrete on-stage anecdote (Armin Ronacher's reply). Strong resonance, but lacks hard data or actionable takeaways, so the score sits right at the featured threshold.

Hugging Face Blog

Ai2 shares the engineering lessons behind Shippy, a maritime AI agent

Ai2's Skylight team built Shippy, an AI assistant that helps maritime analysts query fishing activity, EEZ boundaries, and vessel tracks. The post breaks its architecture into three parts: a soul (system prompt), skills (markdown files that teach it to call APIs and interpret track data), and config (runtime settings; currently Claude Opus 4.6 with the OpenClaw framework). The core idea is wrapping a non-deterministic model in deterministic tools—every answer includes source, data cutoff, and a deep link to the Skylight map so an analyst can verify it. The post doesn't disclose error rates or latency numbers, but it stresses sandboxed hosting and evaluating the agent as a system, not just the model.

Why it matters: Ai2's three-layer agent architecture (soul/skills/config) and the Markdown-as-skill-sheet pattern are concrete engineering takeaways. But the maritime domain is too niche for broad resonance, landing right at the featured threshold.

Hacker News front page

Unsolved Problems in MLOps

This ACM Queue piece lays out why classical ops practices break down for ML: non-deterministic outputs and data as a system driver make canary deploys, health checks, and alerting nearly useless. Azure validates new models by having LLMs judge LLM output—the SRECon audience was audibly surprised. The authors argue the field must either find a better paradigm or fix the ones we have.

Why it matters: This ACM Queue piece lays out MLOps' core tension: traditional ops relies on deterministic responses for health checks and canary releases, but ML systems are non-deterministic and data-driven. The Microsoft Azure example—using LLMs as judges with employee thumbs-up as fallbac...

Jul 15Wednesday

Computing Life · Share · Yage

AI Trains AI: What a Public Self-Improvement Experiment Actually Closed the Loop On

Dan Austin open-sourced a full AI-trains-AI loop. An outer Qwen3.6 agent designs post-training recipes; an inner Qwen3-0.6B or 1.7B model runs real GPU training, and hidden eval scores feed back as reward to update the outer policy. Over 54 steps, the agent first learned to reduce invalid submissions, then shifted 1.7B model usage from 42% to 95% and began tuning temperature, optimizer, and other hyperparameters. Trained small-model scores rose from noise level into the 0.22–0.48 range, with limited transfer to a held-out triage task. A postmortem also revealed an evaluator bug: the old tool-use detector looked for `.function.name` instead of `.name`, so the 0.4-weighted score never fired—yet the reward curve still climbed. The fix required a full restart. The experiment shows outer RL can reshape a training agent's behavior, but tasks, rewards, and budgets are still human-designed.

Why it matters: Dan Austin open-sourced a full AI-trains-AI experiment: an outer Qwen3.6 agent designs post-training recipes, inner models actually train on GPU, and after 54 steps the agent learned to pick models and tune hyperparams, pushing 1.7B usage from 42% to 95%. All three HKR axes hi...

Jul 14Tuesday

AI HOT (Curated Pool)

Tencent Hunyuan releases 1-bit and 4-bit quantized Hy3, a 295B MoE that runs on a single GPU

Tencent Hunyuan quantized its flagship Hy3 (295B MoE) into 1-bit and 4-bit versions that run on a single GPU. Hy3 is claimed to be best-in-class at this scale and competitive with trillion-parameter models for most agent scenarios. The quantized versions work via llama.cpp with MTP support, drastically lowering hardware requirements. Apache 2.0 license, commercial use allowed, plus two weeks of free API through OpenRouter. The post doesn't disclose quantization accuracy loss or the specific GPU memory needed.

Why it matters: Tencent Hunyuan's quantized Hy3 puts a 295B MoE model on a single GPU — immediately actionable for local deployment and agent builders. Apache 2.0 license plus a two-week free API window lowers the barrier to test. Held below 85 because the post doesn't disclose quantization a...

Jul 13Monday

AI HOT (Curated Pool)

Tencent Hunyuan open-sources HyOCR-1.5: a 1B end-to-end OCR model with 6.37× faster inference

Tencent Hunyuan fully open-sourced HyOCR-1.5—training, inference, and model weights—a first for end-to-end OCR large models. The 1B-parameter model handles 8+ text-centric tasks and scores 94.74 on OmniDocBench v1.6, ranking first end-to-end. DFlash speculative decoding speeds up inference 6.37× under Transformers and 2.14× under vLLM, hitting 1.408s per page. It supports 4K resolution and a 128K context window, and uses Agentic Data Flow to extend low-resource OCR to 331 languages, ancient script recognition, and multi-image QA.

Why it matters: Tencent Hunyuan fully open-sourced an end-to-end OCR model — training code, inference code, weights. 1B params, 94.74 on OmniDocBench v1.6 (#1 among end-to-end models), 6.37x inference speedup via DFlash. A genuine open-source move from a major Chinese lab, not weights-only. S...

Jul 10Friday

Computing Life · Share · Yage

RLM treats context as external data, not a prompt dump

Alex Zhang's Recursive Language Model (RLM) keeps long text outside the model window as external data; a root model queries it via code. With GPT-5-mini, RLM lifted OOLONG-Pairs F1 from 0.04% to 58.0% and BrowseComp-Plus accuracy from 0% to 91.3%. But BrowseComp-Plus has known data contamination, OOLONG-Pairs is author-designed, and baselines were tuned by the authors—discount those numbers. RLM only works at depth=1; depth=2 brings 28x latency and 100x token cost. It performs worse on math and science tasks, and Q95 cost can spike 10x above median. The repo has 5,230 stars; an independent reproduction pushed DeepSeek v3.2 on OOLONG from 0% to 42.1%.

Why it matters: Alex Zhang's RLM flips long-context from 'cram into window' to 'query as external data,' hitting 58.0% and 91.3% on two hard benchmarks at depth=1 with GPT-5-mini. The author's honesty about multi-layer recursion failing is a plus. Cap at 78 because it's still a model-specific...

Jul 9Thursday

AI HOT (Curated Pool)

Ant Lingbo open-sources LingBot-Video, a MoE video base model for embodied AI

Ant Lingbo open-sourced LingBot-Video, the first MoE-based video generation model built for embodied AI. It has 30B total parameters but activates only ~3B during inference, roughly 3× faster than a dense model of similar size. Training used 70,000 hours of robot-related video—dexterous manipulation, navigation, egocentric interaction. On the RBench benchmark for robot manipulation videos it scored 0.620, ahead of Wan2.6 (0.607) and Seedance 1.5 Pro (0.584). Internal tests also place it above NVIDIA Cosmos 3 and Hunyuan Video 1.5 on physical plausibility and motion consistency. The model targets robot action prediction, simulation data generation, and world-model research. Code is public.

Why it matters: Ant Lingbo open-sourced the first MoE video foundation model for embodied AI — 30B total params, ~3B activated during inference, 3x faster than dense models of similar scale, trained on 70k hours of real robot video. HKR all hit, but it's a fresh release with no external repro...

Jul 8Wednesday

OpenAI News

OpenAI audits SWE-Bench Pro, finds ~30% of tasks are broken

OpenAI audited SWE-Bench Pro and estimates ~30% of its tasks are broken. An automated pipeline flagged 286 suspicious tasks; Codex-based investigator agents and five experienced engineers then reviewed them. Engineers identified 249 (34.1%) flawed tasks, mostly due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. OpenAI advises model developers to scrutinize results rather than trust leaderboard scores. The post does not disclose a fix timeline or a revised dataset release.

Why it matters: OpenAI audited SWE-Bench Pro and found ~34% of tasks defective — a ratio that forces the industry to re-examine coding benchmark reliability. The post provides concrete defect categories and a human review pipeline. Not scored higher because this is a benchmark quality report,...

AI HOT (Curated Pool)

Liquid AI open-sources Antidoom, a final-token preference optimization method that fixes reasoning model doom loops

Reasoning models can get stuck in doom loops, repeating useless tokens until the context window fills up. Liquid AI open-sourced Antidoom, which uses Final Token Preference Optimization (FTPO) to fix this. The method trains the model on 1,040 preference pairs to learn when to stop at the end of reasoning. On DeepSeek V4 Pro, the doom-loop rate dropped from 3.2% to 0.3% without hurting math or coding scores. The post doesn't disclose training cost or how well it transfers to non-DeepSeek models.

Why it matters: Liquid AI open-sourced a practical fix for reasoning model doom loops, dropping the rate from 3.2% to 0.3% on DeepSeek V4 Pro — solid numbers. Not scoring higher because it's a single blog post with no paper or third-party validation yet; 78 for a strong single-source piece.

Jul 5Sunday

AI HOT (Curated Pool)

Meituan LongCat-2.0 fully open-sourced under MIT license, releasing 1.6T MoE weights and inference code

Meituan fully open-sourced LongCat-2.0 under MIT license, releasing both weights and inference code. It's a 1.6T-parameter MoE model activating ~48B per token, with 1M-token context. LongCat Sparse Attention handles long sequences, Zero-Compute Experts dynamically activate 33B–56B to avoid wasted compute, and MOPD routes tasks across Agent, Reasoning, and Interaction expert groups. On benchmarks: SWE-bench Pro hits 59.5, edging out GPT-5.5's 58.6; Terminal-Bench 2.1 scores 70.8; multilingual SWE-bench reaches 77.3. It natively integrates with Claude Code, OpenClaw, and Hermes Agent, supports GPU and NPU deployment, and has been validated on large-scale domestic clusters.

Why it matters: Meituan fully open-sources LongCat-2.0, a 1.6T MoE model, under MIT license with weights and inference code — a rare move from a major Chinese tech company. The 1M-token context window and sparse attention design are concrete technical hooks, not just marketing. Score held at ...

Jul 2Thursday

Ben's Bites

Fable 5 is back, and there's a new Claude Sonnet 5

Anthropic re-released Fable 5 for paid users with stronger guardrails, available in subscriptions only through July 7 and capped at 50% of usage limits. Scale's benchmark shows it completes 16% of remote work tasks, double Opus 4.8. Claude Sonnet 5 also launched—benchmarked close to Opus 4.8 on agent tasks, cheaper per token but roughly the same cost per task in practice; the author finds it expensive and slow. Google dropped two new models: Nano Banana 2 Lite for fast, cheap images and Omni Flash for video generation and editing. Bridgewater and Thinking Machines trained a financial triage model hitting 84.7% accuracy at 13.8x lower cost than the best frontier model tested.

Why it matters: Anthropic dropped Fable 5's limited return and Sonnet 5 simultaneously — two signals stacked. Scale benchmark provides hard comparable numbers, not pure marketing. Fable 5's 16% task completion rate doubled but absolute number is still low, so not pushing past 90.

Hacker News front page

Cursor releases CursorBench 3.1, scoring coding agents on real multi-file tasks from user sessions

Cursor built CursorBench 3.1 from real user sessions—ambiguous, multi-file coding tasks—and added codebase understanding, bugfinding, planning, and code review problems. Fable 5 Max leads at 72.9%, averaging $18.02, 63,842 tokens, and 76 steps per task. GPT-5.5 High scores 62.6% at just $3.59 per task, a much better cost-performance ratio. Composer 2.5 hits 63.2% for only $0.55, the cheapest on the board. The post doesn't disclose the total number of tasks or grading details; small score differences may not be statistically meaningful.

Why it matters: Cursor drops CursorBench 3.1, a benchmark built from real user sessions. Fable 5 Max leads at 72.9% but costs $18/task and 76 steps; GPT-5.5 High scores 64.3%. Concrete numbers and head-to-head comparison make this directly useful for Cursor users. Not scoring higher because i...

AI HOT (Curated Pool)

Tencent Hy3 released: matches much larger flagship models with 2-5x fewer parameters

Tencent officially released Hy3 under Apache 2.0. With 1/5 to 1/2 the parameters of competing flagship models, Hy3 matches or beats them on reasoning, agent, and long-context benchmarks. In a 270-person internal blind test, Hy3 scored 2.67/4 vs GLM5.1's 2.51/4. Hallucination rate dropped from 12.5% to 5.4%, multi-turn error rate from 17.4% to 7.9%. WorkBuddy task completion jumped from 72% to 90%, with 34% less time. API pricing: ¥1/M input tokens, ¥4/M output, ¥0.25 cached. The post does not disclose exact parameter count or training details.

Why it matters: Tencent Hunyuan releases Hy3, open-source under Apache 2.0, with parameter counts 1/5 to 1/2 of competitors yet matching or beating them on reasoning, agent, and long-context benchmarks. A 270-person blind test shows it beating GLM5.1. Domestic flagship open-source release is ...

Jul 1Wednesday

AI HOT (Curated Pool)

Meituan releases LongCat-2.0: a 1.6T-parameter model trained on 50,000 domestic GPUs, now open source

Meituan open-sourced LongCat-2.0, a 1.6T total-parameter model with ~48B activated per inference and native 1M context. It was trained and served entirely on a 50,000-card domestic GPU cluster. The architecture combines LSA sparse attention, zero-compute experts, ScMoE, and MOPD multi-expert fusion that blends Agent, Reasoning, and Interaction expert groups. SWE-bench Pro hits 59.5, Multilingual 77.3. A preview is live on OpenRouter and longcat.ai, already ranking top three globally in monthly calls on OpenRouter. The post doesn't disclose training cost, inference latency, or the specific domestic chip model, so I'd hold off on those details.

Why it matters: Meituan's trillion-param model trained end-to-end on domestic GPUs is the headline; code benchmark scores are solid. Not scoring higher because Meituan isn't a tier-1 model lab yet, and real-world usability depends on the open-source release.

Jun 30Tuesday

Hacker News front page

Meituan open-sources LongCat-2.0, a 1.6T MoE model with 48B active params, trained entirely on AI ASIC superpods

Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter MoE model with ~48B active parameters per token. It was pretrained on over 35 trillion tokens using 50K+ in-house AI ASICs with no rollbacks or irrecoverable loss spikes, showing frontier-scale training is viable on non-GPU hardware. The model targets long-context and agentic workloads: it introduces LongCat Sparse Attention to speed up 1M-token processing and was trained on hundreds of billions of 1M-context tokens. Official charts place it alongside Gemini 3.1 Pro, GPT-5.5, and Opus 4.8 on Terminal-Bench 2.1, SWE-bench Pro, and other coding/agent benchmarks, though the post does not provide exact numeric comparisons. An N-gram Embedding module with 135B parameters expands the embedding space roughly 100×, which the team claims outperforms scaling standard MoE experts by the same amount. The model is integrated with Claude Code, OpenClaw, and Hermes; code and weights are available on GitHub and HuggingFace.

Why it matters: Meituan open-sources a 1.6T MoE model trained entirely on in-house AI ASICs across 50k+ cards with zero rollbacks, plus dedicated long-context and agent optimizations. Score held at 82 rather than higher because we only have the official blog post — no third-party evals or rea...

Jun 28Sunday

AI HOT (Curated Pool)

Four top AIs play Civ VI: Claude nukes France and still loses

Liam Wilkinson, a former data scientist at 10 Downing Street, built 76 MCP tools over a weekend and dropped Claude, GPT, Gemini, and another model into 23 games of Civ VI. In the wildest match, Claude's Portugal was two points from a diplomatic victory, panicked at France's cultural surge, spent 50 turns rushing nukes, and leveled Toulouse—only to lose to France's diplomatic win. Two numbers stand out: AIs proactively checked the global state only 1–2% of the time; if they don't query it, it doesn't exist. And 48–66% of their written plans were actually executed within 10 turns—Gemini 3.1 Pro topped out at 65.8%. GPT-5 scored 99.26% on GovBench, but in the game it hit the same sensorium blind spot and knowing-doing gap. The bottleneck isn't intelligence—it's architecture and engineering.

Why it matters: 76 MCP tools, 23 matches, and a nuclear-diplomacy disaster — HKR all hit. Capped at 82 because it's a weekend experiment, not a formal study, but the narrative and insight are strong enough for featured.

Computing Life · Share · Yage

As AI subsidies recede, agents are priced by intelligence per dollar

Hidden token subsidies are fading. GitHub Copilot switched to usage-based billing on June 1, 2026; OpenAI, Anthropic, and others updated prompt caching pricing; the Linux Foundation plans a Tokenomics Foundation for cost standards. The article argues this isn't just tokens getting pricier—it's the old subsidy structure collapsing, shifting agent design goals from adoption to reliable tasks per dollar. Four engineering levers are proposed: prompt caching to avoid paying for repeated prefixes, cleaning up tool-output noise in context, routing simple work to cheaper models, and eval-driven fallback to guard quality. A cost-per-accepted-task formula is provided, factoring in model, tool, retry, and human review costs. The post doesn't include specific benchmark numbers—it's more architectural guidance and industry signal reading.

Why it matters: The piece nails a structural shift—token subsidy retreat—with three concrete signals: Copilot's billing change, caching price tiers, and the Tokenomics Foundation proposal. Not scored higher because it's trend analysis rather than breaking news, and the post doesn't disclose s...

Jun 26Friday

New York Times Chinese

Chinese AI Models Narrow Performance Gap with Anthropic and OpenAI

Zhipu's GLM-5.2 surged in popularity after Anthropic restricted access to Fable and Mythos, entering OpenRouter's top ten. It costs about one-eighth of Claude Opus 4.8 for certain tasks and is fully open-source. Experts estimate China's lag behind US firms has shrunk to six months or less. The post notes Zhipu's compute spending exceeded 7x its revenue in H1 2025, but does not disclose whether GLM-5.2's training involved distillation.

Why it matters: Zhipu's GLM-5.2 quickly filled the gap after Anthropic restricted access, costs one-eighth of Claude Opus 4.8, is fully open source, and the US-China gap estimate has shrunk to six months — three signals stacking up, worth recommending. Not scoring higher because the post does...

Jun 24Wednesday

Hacker News front page

Qwen-AgentWorld: Language World Models That Simulate Environments for General Agents

Qwen team released Qwen-AgentWorld, a language world model that predicts environment dynamics for general agents. It covers 7 domains and uses long chain-of-thought reasoning to forecast next states. Two model sizes are available: 35B-A3B and 397B-A17B, trained on over 10 million real-world interaction trajectories via a three-stage pipeline—CPT injects world modeling from state transitions, SFT activates next-state prediction, and RL sharpens fidelity with hybrid rubric-and-rule rewards. The team also built AgentWorldBench from real interactions of 5 frontier models across 9 benchmarks. Qwen-AgentWorld significantly outperforms existing frontier models. It works in two modes: as a decoupled simulator enabling scalable RL across thousands of environments, surpassing real-environment-only training; and as a unified agent foundation model where world-model training serves as effective warm-up, boosting performance on 7 agentic benchmarks. Code is open-sourced.

Why it matters: Qwen team trains an LLM-based world simulator on 10M interaction traces across 7 environments. Novel approach with concrete scale and benchmarks, relevant for agent builders. Score capped below 85 because it's a paper, not a product release — real-world agent task gains aren't...

Jun 23Tuesday

Hacker News front page

Krea releases Krea 2 technical report: open-source text-to-image models built for aesthetic diversity and creative control

Krea 2 is a series of open-source text-to-image foundation models released under a permissive license. Instead of optimizing for a single polished default look, it aims to cover a broad range of visual styles and give users ways to explore them via text or reference images. The pretraining data deliberately excludes AI-generated images and avoids aesthetic-score filters—only duplicates, samples VLMs can't describe well, harmful biases, and overly complex images are removed. The architecture is a diffusion transformer (DiT) with iREPA, improved VAEs, Qwen3-VL text encoder, and components like GQA and sigmoid-gated attention to speed up convergence. Training runs through pretraining, midtraining, SFT, preference optimization, and RL. To bridge the gap between short user prompts and the model's rich conditioning space, Krea 2 adds a prompt expander (two-stage SFT+RL on open-source LLMs) and a style-reference system that lets users control style and mood from uploaded images, with adjustable strength and weighted mixing. It ranks in the top 10 on the Artificial Analysis text-to-image leaderboard and second among independent labs.

Why it matters: Krea 2 ships open-source with a detailed technical report and a clear data curation stance (no AI-generated images, no aesthetic scoring). Useful for model builders, but the image-gen space is crowded and Krea isn't a tier-1 lab, so it lands at the 78 featured threshold.

Jun 20Saturday

Hacker News front page

GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2, challenging the bigger-model dogma

The author benchmarks GPT-5.5, DeepSeek V4 Pro, and GLM-5.2 on the AA-Omniscience hallucination metric and a Python async coding prompt. GPT-5.5 hits 86% hallucination, DeepSeek V4 Pro 94%, while GLM-5.2 scores 28%. DeepSeek V4 Pro spent nearly 4 minutes and 7.7k reasoning tokens producing a confidently wrong solution; GLM-5.2 needed 12 seconds and ~800 tokens to flag the prompt as technically impossible under single-threaded, no-polling constraints. GLM-5.2 trails GPT-5.5 by only 4 points on the AA Intelligence Index and Claude Fable 5 by 9 points—Fable 5 was restricted by the US government three days post-launch over a single jailbreak. The post argues that scaling parameters and data makes models worse at saying “I don’t know,” and frames an unsolved trilemma: raw capability, hallucination calibration, and compute efficiency. The article does not disclose GLM-5.2’s training data size or exact release date.

Why it matters: First-person benchmark with concrete, counterintuitive numbers; hits all three HKR axes. Score held at 78 rather than 85+ because it's a personal blog, sample size and methodology aren't fully detailed, limiting authority.

Jun 19Friday

AI HOT (Curated Pool)

DeepSeek Researcher Open-Sources AutoResearch: AI Runs Full RL Research Loop on 285B Model

DeepSeek researcher Deli Chen open-sourced AutoResearch, a protocol where an AI agent independently ran a full RL research loop on a 285B model—designing experiments, writing code, submitting GPU jobs, debugging, and summarizing results with zero human intervention. The system used GRPO. The post doesn't disclose the specific task, training duration, or success rate, so I'd hold off on getting too excited until there are reproductions.

Why it matters: A DeepSeek researcher released an experimental protocol where an agent independently ran a full RL research loop on a 285B model—strong premise. But the post doesn't disclose the specific task, training duration, or success rate; all key metrics are missing, so the score stays...

AI HOT (Curated Pool)

Steve Yegge: Fable’s shutdown signals frontier AI will be locked down like nukes

Steve Yegge argues Fable’s brief USG shutdown marks the moment model intelligence became dangerous. He predicts frontier models will be controlled like nuclear weapons within 2–3 generations, with most Fortune 500 companies locked out. Open-source can reach Fable-class but won’t blow past it due to compute walls and supply-chain lockdowns. The capability curve will appear flat to most people—not because progress stops, but because the smartest models will be kept out of public hands.

Why it matters: Steve Yegge's deep analysis of the Fable takedown argues the AI capability curve is about to be flattened by government regulation. Sharp thesis with concrete predictions, but it's commentary, not primary reporting — docked for lacking verifiable new facts.

Jun 18Thursday

AI HOT (Curated Pool)

GPT-5.5 Instant brings frontier health intelligence to free ChatGPT users

OpenAI says GPT-5.5 Instant matches its priciest Thinking models on health benchmarks and is available to free users. In a blind review of 3,500 responses, physicians rated 5.5 Instant higher than human-written answers on accuracy, communication, and completeness, with fewer failure modes like missing red flags or failing to ask for context. Production monitors show a 71% drop in health-response factuality issues over two months. The improvements come from model advances and physician-led evaluations that define what good looks like in real-world health conversations.

Why it matters: OpenAI's official post on GPT-5.5 Instant health QA performance: 3,500 blind-rated responses scored higher than human doctors on accuracy, communication, and completeness, with a 71% drop in factual errors, available to free users. Concrete numbers and physician-led eval keep ...

AI HOT (Curated Pool)

Alibaba open-sources LOGOS, a 1B-param science model that beats Microsoft NatureLM on multiple tasks

Alibaba's ATH-Token Foundry and Renmin University's Gaoling School of AI open-sourced LOGOS, a generative model that uses a unified 'science grammar' to handle seven modalities including proteins, small molecules, and materials. It encodes 3D pocket-ligand contacts as discrete tokens, predicting spatial interactions without explicit 3D coordinates. LOGOS-1B uses only 1/56 the parameters of Microsoft NatureLM (8×7B) and matches or beats domain-specific methods across six science tasks. Pretrained on 44.87B tokens, it shares the same sequence format and next-token prediction objective for both pretraining and downstream tasks, eliminating heavy adaptation. Weights, inference code, and the tech report are fully open on HuggingFace and GitHub.

Why it matters: Alibaba open-sourced LOGOS, a 1B-param scientific model that unifies seven data types into token sequences and beats Microsoft's 56x-larger NatureLM on multiple tasks. Concrete numbers and open code give it strong knowledge value, but the niche domain limits resonance — lands ...

Jun 17Wednesday

OpenAI News

OpenAI releases LifeSciBench: a benchmark built by PhD scientists for real research tasks

OpenAI released LifeSciBench, a 750-task benchmark authored and reviewed by PhD scientists with biotech/pharma experience. It tests real research workflows—interpreting conflicting evidence, designing experiments, assessing translational risk—not fact recall. 53% of tasks require processing attached artifacts like figures or sequence files, averaging four reasoning steps per task. Grading uses 25 rubric criteria per task on average, checking scientific validity and operational usefulness, not just final answers. The post does not disclose model scores.

Why it matters: OpenAI released a PhD-scientist-written benchmark with 750 questions testing experimental design, conflicting-evidence interpretation, and translational risk assessment — closer to real research workflows than existing benchmarks. Score capped here because only a preprint and ...

Jun 16Tuesday

AI HOT (Curated Pool)

Ant Group BaiLing releases Ling & Ring 2.6 tech report, all three models open-sourced

Ant Group BaiLing published full architecture, pretraining, post-training, and agent RL details for Ling-2.6-flash, Ling-2.6-1T, and Ring-2.6-1T. All three use a Hybrid Linear Attention that mixes Lightning Attention and MLA at a 7:1 ratio. Ling-2.6-flash hits 340 tokens/s decoding on 4×H20 hardware. Ling-2.6-1T shows roughly 4× token efficiency gain over its predecessor on the Artificial Analysis Intelligence Index. Ring-2.6-1T high scores 87.60 on PinchBench and 63.82 on ClawEval. Code and weights are open.

Why it matters: Ant Group's BaiLing team open-sourced three models with a Hybrid Linear Attention design blending Lightning Attention and MLA at 7:1, backed by concrete long-context efficiency data. Code and weights are public, making this a verifiable release. Not scoring higher because Ant'...

r/LocalLLaMA

HalBench tests 29 open models on sycophancy and hallucination; Qwen 3.6 and Gemma 4 punch far above their weight

HalBench is an open benchmark that gives models a false premise and measures whether they push back or play along. v2.3 covers 33 models, 29 of them open. Only Sonnet 4.6 (65.1%) and Grok 4.3 (50.9%) clear 50% pushback. The best open model is Qwen 3.6 (~27B dense) at 36.6%, beating GPT-5.4 and Gemini 3.1 Pro. Gemma 4 26B follows at 29.2%. Model size barely predicts performance; phi-4 sits dead last at 2.3%. Dataset, scoring code, and Space are all open.

Why it matters: A community benchmark with concrete numbers and rankings, where Qwen 3.6 outperforms GPT-5.4 on refusal rate, is real signal. Not p1 because it's a self-built eval without peer review yet — treating it as a strong recommendation.

Jun 15Monday

AI HOT (Curated Pool)

MiniMax open-sources M3 model weights (428B total, 23B active) with lower long-context cost

MiniMax open-sourced M3 model weights last Friday—428B total parameters, 23B active—along with the MSA sparse attention paper that cuts long-context inference cost. M3 is the first open-source model trained with interleaved text and image data from the pre-training stage. Two weeks post-release, it ranked #1 among open-source models on the Artificial Analysis Intelligence Index and GDPval-AA, reached Pareto-optimal on Code Arena WebDev, and topped Chinese models on Vals.AI. Output speed improved from ~30 TPS to ~80 TPS, with another 30–40% planned. A usage dashboard was added to the Token Plan backend.

Why it matters: MiniMax open-sourced a 428B MoE model with interleaved image-text pretraining and two #1 open-source rankings in two weeks — enough signal for featured. Held back from p1 because the post is a first-party announcement without third-party benchmarks or concrete MSA cost numbers...

Jun 13Saturday

AI HOT (Curated Pool)

Zhipu launches GLM-5.2 flagship model with 1M context, open-sourcing next week under MIT license

Zhipu's new flagship GLM-5.2 is live for Coding Plan subscribers, emphasizing coding strength and 1M context. API and chatbot access arrive next week, alongside an MIT-licensed open-source release. The post doesn't disclose benchmark scores or pricing details.

Why it matters: Zhipu drops GLM-5.2 with 1M context and MIT open-source next week, coding-focused. No benchmarks or pricing disclosed, so real capability is unverified — hence below 85. But a domestic flagship update plus open-source is strong signal, worth featuring.

AI HOT (Curated Pool)

MiniMax open-sources M3 weights, takes a swipe at Anthropic's export control ban

MiniMax released M3 model weights on HuggingFace. The post says 'M3 would never,' a jab at Anthropic's Fable 5 and Mythos 5 being forcibly disabled under US export controls, blocking all foreign nationals. The post doesn't disclose M3's parameter count, benchmarks, or license.

Why it matters: MiniMax open-sources M3 weights as a direct response to Anthropic's export controls — strong conflict and topicality, but the post lacks parameter count, benchmarks, and license details, capping the score.