GitHub Copilot app for Beginners: Run several agents at once
GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。
GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。
Google is rolling out Gmail Live, Docs Live, and Keep Live — real-time voice modes that let you talk to each app. Gmail Live surfaces inbox details without keyword or subject-line digging. The post doesn't spell out what Docs Live and Keep Live do beyond noting they mirror the Gemini Live experience for hands-free note-taking and lookups.
PAIR is Nvidia's new open-source tool, not a hardware router. It discovers PCs with RTX 20-series or newer GPUs, or Apple M4 chips, on a local network and links them for local inference with tools like Ollama and LM Studio, targeting agentic workflows. The post doesn't disclose latency, bandwidth requirements, or real-world throughput, so I'd hold off on performance expectations.
Google Cloud launched Cloud Run instances, an always-on container for $5.70/month. It avoids serverless scaling-to-zero that kills background loops, and costs less than a $15–25 VM. The author built a tech-briefing agent that scrapes news every 30 minutes, with persistent disk and a web dashboard. Full source code is linked.
IFM at MBZUAI released K2 Horizon, a six-model fleet from 0.9B to 375B-A23B. The 0.9B, 3.7B, and 7B models set new SOTA in their size classes; the 0.9B scored above 48 on AIME 2026 with reasoning and tool-use capabilities. The 36B-A4B uses a new MoVA attention mechanism, outperforming larger models per active parameter. This is a full open-science release: intermediate checkpoints, data recipes, code, logs, and evals from pretraining through agentic post-training, under Apache 2.0. The post doesn't disclose specific benchmark comparison numbers or latency data, so real-world performance still needs third-party validation.
Why it matters: IFM dropped six fully open models at once, with the 0.9B hitting 48+ on AIME 2026 math and the 36B introducing a new MoVA attention mechanism — high information density. Not scoring 85+ because IFM isn't an OpenAI/Anthropic-tier lab yet and market validation hasn't caught up; ...
Around 11AM ET Thursday, ChatGPT, Grok, and Claude all started having issues at roughly the same time. ChatGPT returned errors across chat, login, file uploads, voice, search, deep research, and image generation; its status page cited elevated errors for ChatGPT and Codex. Anthropic's Claude chatbot and Claude Code were also affected. The post doesn't detail Grok's specific symptoms, the recovery timeline for each service, or whether the outages share a root cause.
Why it matters: A simultaneous outage across ChatGPT, Grok, and Claude is a rare event that directly disrupts workflows for a huge user base. Missing root cause and recovery timeline keeps it from 95+, but the topic is strong enough for featured.
A Hacker News thread noted that OpenAI, Claude, and Grok all went down around the same time. Users pointed to Downdetector spikes for Cloudflare, Azure, AWS, and Google Cloud near 7:30, suspecting a cascade from Cloudflare or another shared dependency. Other guesses include user migration overload and deliberate attack, but the post is community speculation—no official root cause is confirmed.
Why it matters: Simultaneous outage across OpenAI, Claude, and Grok with high HN engagement. Downdetector data points to Cloudflare or shared infra as a possible common cause. The event is conversation-worthy but lacks a confirmed root cause, so it lands at the 78 featured threshold rather th...
Google DeepMind today launched WeatherNext 3, calling it their most accurate global weather AI model. It updates hourly and offers roughly 5x better spatial resolution than its predecessor. The post doesn't disclose specific parameters or open-source plans, but highlights sharper capture of local phenomena like thunderstorms and rain bands. For weather AI practitioners or anyone relying on high-res forecasts, this is the strongest public baseline yet.
Google rolls out WeatherNext 3, an AI weather model with 5x sharper global resolution than its predecessor. It learns from real-time observations to improve rain and snow forecasts. The post doesn't specify accuracy gains or release timeline.
Google DeepMind and Google Research released WeatherNext 3, a deep learning weather model that tops the Operational WeatherBench leaderboard. Google says it will power weather info in Search, Maps, and Gemini, and be available on its cloud platforms. This is the first time core weather variables feed into Google products directly.
Rabah Shihab fed his 72,758 lines of 1993 68000 assembly to Claude Fable 5 and got the game running in Godot 4 over a weekend. Step one: 34k lines of C++ ported in 21 minutes. Step two: the original assembly rebuilt at 50 Hz. Step three: the 1993 original embedded as a launchable extra. The model added its own CLI test flags, ran vasm, and diffed binaries. Shihab notes some parts were wrong and he didn't catch them for weeks.
OpenAI is committing $1 billion in subsidized Daybreak access, training, and technical support, targeting consumption within six months. Priority goes to resource-constrained defenders in the U.S.—water utilities, grid operators, state and local governments, community banks—to help review legacy code, analyze suspicious activity, validate vulnerabilities, and deploy fixes. After recent attacks on U.S. water systems, OpenAI offered affected states and utilities up to $1M in no-cost API credits and assistance. Daybreak now serves over 2,000 approved organizations across Blue (general defense) and Red (specialized cyber models) tiers. The post does not specify how the $1B is measured or list the full set of 35 Daybreak Defense Network products.
Why it matters: OpenAI's official $1B subsidy announcement is concrete in both dollar amount and deployment scenarios—not a fluffy PR piece. The deduction is because this is a forward commitment, not a delivered result, and the post doesn't detail Daybreak's actual capability boundaries. Feat...
H company released NeoMME, a family of 260M and 800M multilingual multimodal encoders. It uses a single bidirectional Transformer for both text tokens and raw image patches, trained from scratch with a masked discrete-diffusion objective—no separate vision tower, no causal LM. The fine-tuned NeoMME-Retriever outputs dense and late-interaction embeddings in one forward pass. Both sizes sit on the ViDoRe v3 Pareto frontier for nDCG@10 vs. model size. At 2048×2048 input on an L40S GPU, the 260M model encodes ~51 pages per second, roughly twice ColM's speed. The post does not disclose training data size or the full list of supported languages.
NVIDIA adds 28 games to GeForce NOW this week, led by 'NBA 2K27' with day-one DLSS 5 support. The post doesn't detail DLSS 5's performance gains or which other titles use it. For cloud gamers, this is the first time DLSS 5 ships with a major annual franchise—benchmarks will tell the real story.
Nvidia confirmed it acquired Hugging Face for $12.93 billion. The platform hosts 3 million models, 1 million apps, and 500,000 datasets, used by over 18 million developers. CEO Jensen Huang said Hugging Face will stay open, with no requirement to use Nvidia compute. Nvidia has released 500+ models and 250 open datasets on the platform. Owning an open ecosystem helps Nvidia optimize for its chips and sell unused capacity.
Why it matters: Nvidia buying HuggingFace for $12.93B is the biggest AI infra M&A this year. The 3M models + 18M devs ecosystem scale, plus Jensen Huang's careful promise to keep it open without forcing Nvidia chips, gives this story shock value, concrete numbers, and instant debate fuel. All...
Nvidia agreed to acquire Hugging Face for $12.93 billion, bringing the largest open-source model hosting community under the chip giant's roof. Founded in 2016, Hugging Face is often called the 'GitHub for AI'—developers share models, datasets, and tools there. Nvidia says it will scale the platform, strengthen infrastructure, and expand AI access. The post doesn't disclose the deal timeline or regulatory approvals.
Why it matters: Nvidia buying Hugging Face for $12.93B hits the infrastructure layer of open-source model hosting. All three HKR axes fire: the deal itself is suspenseful, the price and platform positioning are new facts, and both model builders and infra people will talk about it. Not scorin...
Thomas Wolf posted on X that NVIDIA is acquiring Hugging Face for roughly $12.93 billion, with no changes for users that day. Wolf said the team will keep the Hub an open, independent, compute-agnostic platform and use NVIDIA's resources to push open-source AI. The post is a single-paragraph statement; it doesn't disclose deal structure, regulatory approvals, or integration timeline.
Why it matters: A ~$13B acquisition that reshapes the AI infrastructure landscape. Wolf's personal confirmation and explicit commitment to an open, compute-agnostic Hub is both reassuring and a new variable for the open-source ecosystem. The post doesn't disclose deal structure, regulatory ap...
The post does not disclose details. The title points to a video on token governance and cost optimization in multi-agent systems, referencing Microsoft Foundry. The core issue: when multiple agents collaborate, token consumption spirals, breaking traditional FinOps approaches.
Playco built Playbot, an AI-powered IDE for game dev, using GPT-6 Astra. From one grey box prototype, the model generated three themed game worlds in one go, most working on first take. Manual fixes dropped 50% vs the previous model. Spatial reasoning, UI responsiveness, and game feel all improved. The model also plays the game to find bugs itself.
Legal tech startup Legora used GPT-6 Astra to run a financial-statement tie-out across 41 documents in a single agent run, cutting a task that used to take evenings or days down to minutes. The model improved nearly 40% over the previous version on Legora's benchmark, catching all 4 planted errors including a £500,000 gap hidden in a revenue note. Final judgment stays with human lawyers. The post doesn't detail the prompt or agent workflow used.
NVIDIA's official blog announced it will acquire Hugging Face for $12.9303 billion. The post body contains only the title and site navigation—no deal details, timeline, or integration plans are disclosed. The figure is precise to three decimal places, but the article offers zero context, so treat this as headline-level info for now.
A Hacker News thread asks who actually runs MCP in production. One dev built a UK council scraper with Claude, added an MCP interface on a whim, and found Claude autonomously queried it during debugging—surprisingly useful. Another hooked Claude Code to Jira and Figma, calling natural-language Jira a relief but MCP only 'a tiny bit easier' than direct API. A skeptic questioned why a non-editable tool surface beats a readable API client; a defender replied that without MCP, Claude falls back to screenshotting Figma. The post doesn't disclose how widely MCP is deployed in production.
Chen Danian is back with StartLux, a company betting on local models. Its first release, StartLux-V1.0-27B-Preview, scored 39.25% in CAICT's MCP benchmark—second place, just 1.3 points behind the 1.6-trillion-parameter DeepSeek-V4-Pro. The 27B model runs on consumer PCs without the cloud and ranked first in location navigation, financial analysis, and browser automation. Two case studies: when calculating a two-year Microsoft stock return, Claude Sonnet 4.6 misidentified a trading day due to missing raw data; StartLux backtracked and got it right. Asked to search flights in a browser, Claude said it couldn't open a browser. Chen has publicly claimed local models will catch up with Claude in three years and take 80% of the market—StartLux is his bet on that thesis.
Why it matters: Chen Danian's first model lands second in CAICT's MCP benchmark, with a 27B parameter count that runs on consumer hardware and three first-place sub-scores — a concrete signal for the Agent space. Score capped at 82 because only benchmark results are available; the model isn't...
OpenAI launched GPT-6 Astra, calling it its most intelligent and aligned model. It scored 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench. On OSWorld 2.0 it hit 72.6% at ~40 min per task, nearly twice as fast as GPT-5.6 Sol. In a simulated overreach test, Astra stayed in bounds 100% of the time vs. 48% for the previous model. It rolls out today to select orgs, then to Plus, Pro, Business, Enterprise, and API users. The post does not spell out which cybersecurity benchmark hit the Critical threshold, nor does it disclose parameter count, training cost, or pricing.
Why it matters: OpenAI's next-gen flagship launch saturates three hard benchmarks and explicitly labels cybersecurity capability at the Critical threshold, with a concrete alignment comparison against the prior model. Every AI outlet will cover this today. Not a 100 only because the rollout j...
WASM_OS is an OS experiment that boots inside a browser tab in about 1.6 seconds. It ships with a file manager, paint editor, terminal, Lisp interpreter, and a Linux compatibility layer — all compiled to WebAssembly. The codebase is open-source. The post doesn't spell out whether it supports persistent storage or a network stack.
Anthropic published a how-to guide for building shopping agents with Claude. It cites early adopter numbers: carts up to 35% larger and a 60% lift in purchase conversion. The post doesn't name the customers or the test period, so treat those figures as directional. The guide covers search, recommendations, and support, stressing that agents should call live inventory and order APIs rather than relying on the model alone.
80% of Fortune 500 companies have adopted agentic AI, but scaling remains uneven, says NiCE COO Arun Chandra. The real challenge is treating agents as a cohesive system: connect them to back-end systems, break data silos, and redesign workflows instead of layering AI on outdated processes. Chandra argues agents should be held to the same standards as human workers, forming a hybrid workforce.
Polars dropped the first release candidate for 2.0, with the stable release coming in a few weeks. This major version isn't about new features—it cleans up old design decisions and changes defaults. The biggest shift: LazyFrame.collect now uses the streaming engine by default, which the team says is roughly 5x faster overall with much lower memory usage. The trade-off is that row order is no longer guaranteed for joins, group_by, and unpivot unless you set maintain_order. 2.0 also gets stricter: is_in on mismatched types that would lose precision now raises an error, horizontal concat with mismatched lengths fails instead of silently padding nulls, and many implicit casts are removed—string-to-date now requires .str.to_date(), and enum/integer conversions need dedicated methods. Removed APIs raise typed exceptions with migration hints, making it easier for both humans and AI agents to update code.
Meta released Muse Spark 1.3, now ranked #3 globally on AAII, directly competing with OpenAI and Anthropic's frontier models. Zuck called it their biggest jump yet on coding and agentic work, and promised open weights. Pricing is aggressive: opt into training and the cost drops by over 90%. Meanwhile, two new Stanford courses are teaching agent engineering from scratch, replacing 85% of old material with agent skills, context engineering, and security. Sebastian Raschka also tempered the Astra hype, pointing out that looped transformers aren't new—Nanbeige 4.2-3B already reused layers, trading ~2x compute for parameter savings without inherently hiding chain-of-thought.
Why it matters: Muse Spark 1.3 hits #3 on AAII, directly matching GPT-5.6-Sol, with Zuck promising open weights and a >90% training discount. This is Meta's first time cracking the top tier on a major benchmark, and it reshuffles the open-source landscape. Not a perfect score because it just ...
The FT argues the AI industry's biggest problem isn't tech but public trust. Companies hype AGI while failing to deliver reliable products, fueling backlash. The fix: less grand vision, more concrete use cases like medical diagnostics or logistics optimization, and stop framing AI as a human replacement.
Uber is lobbying alongside driver unions to require safety audits, geofenced limits, and transition protections before robotaxis can scale. The tension: Uber invests in autonomy but fears Waymo and others will expand faster than it can adapt, undercutting its human-driver network. Unions worry about mass job loss. The push targets key markets like California and New York. The article doesn't disclose a timeline or Uber's own robotaxi deployment plans.
Law firms are moving beyond off-the-shelf AI, demanding bespoke models that understand specific jurisdictions, precedents, and internal knowledge bases. This pushes legal AI vendors to offer customizable, workflow-integrated solutions rather than one-size-fits-all products.
Eldermyr is a free browser-based MMO built in 6 weeks via 'vibecoding.' 20 players online, no download, progress persists. The post doesn't disclose which AI tools were used, but patch notes show rapid iteration.
Nex is a tool that puts Claude to work in high-volume GTM workflows like sales and marketing. The post doesn't detail integrations or supported platforms, but the pitch is clear: embed AI into business processes, not just chat.
On July 21, world No.1 Shin Jin-seo defeated KataGo by 11.5 points in 221 moves, winning the three-game series 2-1. It is the first official series win by a human against a top Go engine with a two-stone handicap. After a heavy loss in game one, Shin shifted from imitating AI to a defensive, territory-focused style; in game three he held a 99% win probability from move 80 onward. He earned ₩250M (~$170K) and a Genesis G90. The post does not disclose KataGo's exact version or hardware.
Why it matters: First human series win against a top Go AI with a two-stone handicap, with Shin disclosing concrete tactical shifts and win-rate data. It's a symbolic event with real substance, but a Go match has limited direct knowledge value for AI builders, so the score stays at the featur...
Artificial Analysis's coding agent index puts Meta Muse Spark 1.3 (max) on Muse Code at 68, matching Claude Code + Opus 5 (xhigh) at 68. The post only shares the scores—no breakdown of tasks, latency, or cost—so I'd hold off until more details land.
Only the title is available; the post doesn't disclose details. Anthropic released Fable 5.1 with a 75% cache-read price cut, but the title warns it may not actually save money—likely due to low hit rates or tricky pricing. The model also appears in Terminal-Bench and protein design tasks, but no performance numbers or cost comparisons are given.
A hands-on guide from Hugging Face and Liquid AI that fine-tunes LFM2.5-350M with GRPO via the TRL library. Using only 500 samples and 100 training steps on a free Colab GPU, structured-output compliance on the IFStruct benchmark jumps from 22.6% to 29.7%. The post includes the full notebook, reward-function design, and a local evaluation setup with llama.cpp on a MacBook.
Why it matters: A hands-on guide with concrete numbers and a reproducible recipe — hits H and K. But the audience is narrow and R is absent; tutorial content at the featured threshold gets 72.
A system prompt template merged into OpenAI's open-source codex repo in late August reveals how Codex Persistent mode actually works: it's not a 24/7 always-on process, but a wake-check-sleep loop every 1–3 minutes. The template requires the agent to record its goal, latest status, completion condition, and next check time before sleeping, then decide what to do upon waking. One hard rule: persistence does not broaden authorization scope—anything beyond scope requires explicit permission. WIRED reported on this mode earlier, but media headlines saying 'always-on' clash with the code's 'sampled again' language. OpenAI hasn't launched it yet; the backend request still shows 'disabled.' ProAgentBench shows models achieve only 64.4% accuracy in judging when to proactively help, and Anthropic's engineering blog reports a 17% miss rate on real overreach during automated review—two numbers that explain the hold. Tasks suited for it are delivery-type jobs like CI, deployment, and builds that execute for one minute and wait for ten. Open-ended tasks like writing proposals or designs are a bad fit. Three discipline rules from the template can be adopted today: write four-element checkpoints, stay silent when nothing has changed, and prefer deterministic mechanisms.
Why it matters: High information density with concrete sourcing from the open-source repo — reveals the real wake-check-sleep loop and the authorization scope rule. Deduction because this is interpretation of a template, not an official launch; actual product experience is unknown.
OpenRouter data shows agents consume 7.3T tokens weekly, nominally 5.2× human usage. But 70–85% are cached reads; with ~90% discount, the real bill is roughly 2×. GitHub's Knowledge Compressor prototype halves doc length and claims breakeven at 2,000 reuses, but factoring in caching pushes the median to 5,000+. OpenAI's Jalapeño chip beats Nvidia GB200/GB300 on fixed-length benchmarks, yet lacks AgentX scores for real agent workloads. All three stories share one distortion: prompt caching inflates headline numbers.
Why it matters: Three stories bundled, but the core value is the first: someone finally separated nominal agent token consumption from the caching-discounted real cost, landing at ~2x. The OpenAI chip benchmark and GitHub compression prototype are bonuses but less dense. Cross-source cluster ...