Skip to content

#Agent

36 today

Apr 24Friday

X · @claudeai

Claude can now connect to more apps outside work, including Tripadvisor, Booking.com, and Resy

Claude added at least 10 consumer app connections, including Tripadvisor, Booking.com, Resy, Instacart, Spotify, Audible, AllTrails, Thumbtack, and TurboTax. The RSS snippet confirms only a product update; the post does not disclose integration method, supported actions, regions, permission scope, or rollout timing. The key question is whether Claude can act in these apps directly, not just list them.

Why it matters: Official Anthropic product update with clear HKR-H/K/R: consumer app connectors expand Claude beyond workplace tools and widen its assistant surface. The score stays at 75 because the post lists apps only; actions, permissions, regions, and rollout details are not disclosed.

Hacker News front page

GPT-5.5: Mythos-Like Hacking, Open to All

XBOW says GPT-5.5 cut miss rate to 10% on its real-vulnerability benchmark, versus 40% for GPT-5 and 18% for Opus 4.6. It scored 97.5% on visual acuity and used about half the login iterations of the next-best model. The key point is black-box testing: GPT-5.5 without source beat GPT-5 with source.

Why it matters: HKR-H/K/R all pass: a major OpenAI model claim, concrete security benchmark numbers, and a clear practitioner safety nerve. The source is XBOW rather than an OpenAI launch post, so it stays below 95.

X · @OpenAI

Introducing GPT-5.5

OpenAI introduced GPT-5.5, and it is now available in ChatGPT and Codex. The RSS snippet says it targets real work and agents, can understand complex goals, use tools, check its work, and carry more tasks to completion; the post does not disclose parameters, pricing, context window, or benchmark results. What matters is the execution loop, not the headline's “new class of intelligence.”

Why it matters: OpenAI launching GPT-5.5 in ChatGPT and Codex is same-day mandatory coverage. HKR-H/K/R all pass: new model release, concrete agent-workflow claims, and direct impact on daily AI work. Price, context window, params, and benchmarks are undisclosed, so it stays below 95.

The Verge · AI

OpenAI says its new GPT-5.5 model is more efficient and better at coding

OpenAI announced GPT-5.5 and says it is more efficient and stronger at coding than GPT-5.4, which shipped last month. The RSS snippet says it handles coding, debugging, online research, and cross-tool work on spreadsheets and documents; the post does not disclose pricing, context window, or benchmark scores.

Why it matters: An OpenAI model release is same-day coverage, and the angle ties efficiency, coding, and tool use into one clear upgrade, so HKR-H/K/R all pass. The post does not disclose price, context window, or benchmark scores, which keeps it in the high 80s instead of 90+.

Hacker News front page

An update on recent Claude Code quality reports

Anthropic said three product-layer changes degraded Claude Code quality for Sonnet 4.6, Opus 4.6, and Opus 4.7, while the API was unaffected; all were fixed on April 20 in v2.1.116. The changes were lowering default reasoning effort on March 4, a March 26 bug that cleared prior thinking every turn after sessions sat idle for over an hour, and an April 16 prompt tweak to reduce verbosity that hurt coding quality. The signal for practitioners is sharp: product and prompt changes can degrade code performance even when model and inference evals do not reproduce it early.

Apr 23Thursday

The Verge · AI

You’re about to feel the AI money squeeze

Anthropic sharply restricted OpenClaw’s access to Claude this month and pushed heavy third-party agent users toward pricier paid plans. The RSS snippet says system strain and profit pressure drove the move, and Boris Cherny said existing subscriptions do not fit this usage pattern; the post does not disclose pricing, limits, or rollout scope. Watch the monetization shift: agent-style usage is being carved out of flat subscriptions.

Why it matters: Anthropic is turning heavy Claude agent usage into a pricing and access story, which directly affects tool builders and power users. HKR-H/K/R all land, but missing price, quota, and rollout details keep it at the low end of featured.

QbitAI · WeChat

Qwen3.6-27B open-weights, beats its 397B flagship predecessor on agentic coding

Qwen released Qwen3.6-27B and says it beats Qwen3.5-397B on 4 agentic coding benchmarks with about 1/15 the parameters. The post cites SkillsBench rising from 30.0 to 48.2, GPQA Diamond at 87.8, and AIME26 at 94.1; it uses a dense architecture, Thinking Preservation, and Gated DeltaNet, with weights on Hugging Face and ModelScope.

Why it matters: This is a substantive Qwen open-source model release with concrete agent-coding and reasoning scores, so HKR-H/K/R all pass. I keep it at 84, not higher, because the post gives strong benchmarks but no pricing, context window, or independent reproduction yet.

The Verge · AI

Microsoft launches 'vibe working' in Word, Excel, and PowerPoint

Microsoft is rolling out Agent Mode in Word, Excel, and PowerPoint this week, extending Copilot from a Q&A assistant to an agent that can act directly on the document canvas. Sumit Chauhan said earlier foundation models were not strong enough for app control; the post does not disclose rollout scope, pricing, or exact actions.

Why it matters: Microsoft moving Agent Mode into Word, Excel, and PowerPoint clears HKR-H/K/R: the hook is strong, the mechanism is new, and the Office install base makes it resonate. But rollout scope, pricing, and the exact action list are undisclosed, so it stays below the 85+ band.

Xinzhiyuan · WeChat

Historic moment: Anthropic nears $1 trillion on private secondary markets, surpassing OpenAI for the first time

Anthropic was quoted at $1.05T-$1.15T on private secondary markets, above OpenAI’s roughly $880B quotes on similar platforms. The post attributes the rerating to scarce float, a sharp rise from a $380B funding valuation three months earlier, and momentum around Claude Code and revenue growth; it does not disclose trade volume, revenue figures, or company confirmation. Do not confuse this with a new funding valuation: these are secondary-market quotes on platforms such as Forge Global.

Why it matters: The signal is a private-secondary quote of $1.05T-$1.15T for Anthropic, above OpenAI's quoted ~$880B, not a new financing round. HKR-H/K/R all pass, but missing volume, revenue detail, and company confirmation keep it in the good-quality band, not must-write.

Xinzhiyuan · WeChat

Zhejiang University open-sources multi-agent evolution system OpenStory: Sun Wukong turns the Grand View Garden into an empty city

Zhejiang University open-sourced OpenStory, a multi-agent narrative system, and inserted a Sun Wukong agent into a 1:1 Dream of the Red Chamber sandbox; within minutes, agents fled the scene. The memory module broadcast “Sun Wukong killed innocents,” fear overrode daily logic, and Wang Xifeng’s physical removal cascaded into an empty Grand View Garden. What matters is the fragility of memory and consensus links; the post does not disclose the base models, metrics, or reproducible setup.

Why it matters: HKR-H/K/R all pass: the stress test is vivid, and the story includes a specific memory-broadcast failure mode with clear agent-safety relevance. Missing model details, metrics, and reproducible setup keep it in the good-featured band, not 85+.

Bloomberg Technology

Alibaba Adds China Eastern Flight Booking to Flagship Qwen App

Alibaba added China Eastern flight booking to the Qwen app, letting users book flights directly; the snippet says this is the first time its agentic AI tech has opened to a major commercial partner. The RSS snippet does not disclose launch regions, fare classes, payment flow, or revenue terms. The real signal is Qwen moving from chat entry to transaction flow, not just another assistant feature.

Why it matters: Featured on HKR-H/K/R: Qwen moves from answers to booking, a concrete agent-commerce step. Kept at 76 because the brief does not disclose rollout scope, payment flow, rev-share, or fulfillment details.

X · @dotey

Microsoft makes Copilot Agent Mode the default in Word, Excel, and PowerPoint

Microsoft made Copilot Agent Mode the default experience in Word, Excel, and PowerPoint, available now to Microsoft 365 Copilot and Premium subscribers, including Personal and Family plans. Microsoft reports internal test gains: Excel engagement up 67% and approval up 65%, Word engagement up 52%, and PowerPoint new-user retention up 36%. What matters is that multi-step in-document execution is now the default, with preview and rollback controls.

Why it matters: This clears HKR-H/K/R: the default flip is the hook, the post includes rollout scope, concrete pilot metrics, and preview/keep/revert controls, and the Office distribution angle will get discussed. I keep it at 84 because this is a large product update, not a new model release or

Hugging Face Blog

How to Use Transformers.js in a Chrome Extension

Hugging Face published a guide for a Transformers.js Chrome extension using Gemma 4 E2B. It defines three MV3 entry points: background service worker, side panel, and content script. The key design keeps local inference in the background and uses messaging plus a tool loop.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face implementation tutorial, not a model or platform release. Score sits at the featured threshold for a concrete MV3 architecture walkthrough.

Financial Times · Technology

Tesla boosts spending plans to $25bn as Musk doubles down on AI bet

Tesla raised its spending plan to $25bn, with Musk directing more capital toward AI-linked projects. The RSS snippet names self-driving taxis, trucks, robots, and chip factories, and says the increase will be “very significant”; the post does not disclose the time frame, line items, or model details. The key signal is that Tesla is funding a full stack, not just model training.

Why it matters: FT reports a concrete capex jump to $25bn tied to robotaxis, trucks, robots and chip factories. HKR-H/K/R all pass on scale and strategic relevance, but missing timing, line-item spend and model specifics keep it in mid-featured, not must-write.

The Verge · AI

OpenAI now lets teams make custom bots that can do work on their own

OpenAI opened ChatGPT cloud “workspace” agents to Business, Enterprise, Edu, and Teachers plans, letting teams build custom bots for business tasks. Examples include finding product feedback on the web and sending a report to Slack, plus drafting follow-up sales emails in Gmail; the post does not disclose pricing, rollout details, or capability limits. The real shift is workflow execution inside ChatGPT, not a one-off chat reply.

Why it matters: OpenAI moves ChatGPT from chat into team workflow agents, so HKR-H/K/R all pass: the hook is strong, the post adds plan coverage and task examples, and it hits enterprise automation demand. It stops short of P1 because price, rollout detail, and capability boundaries are not disl

X · @dotey

OpenAI launches ChatGPT Workspace Agents for cross-tool enterprise workflow automation

OpenAI released ChatGPT Workspace Agents as a research preview for ChatGPT Business, Enterprise, Edu, and Teachers paid plans. The snippet says agents can connect Slack, Gmail, Google Drive, Salesforce, Notion, Linear, and Atlassian, and take actions like updating tickets, creating docs, and replying in Slack. The key point is admin controls for permissions, approvals, and monitoring; the post does not disclose pricing, quotas, or rollout regions.

Why it matters: This is a substantive OpenAI product update: ChatGPT now offers cross-tool workspace agents with named enterprise integrations, so HKR-H/K/R all pass. I keep it at 86 because it is still a research preview and pricing, quotas, and rollout regions are not disclosed.

Latent Space

Shopify’s AI Phase Transition: 2026 Usage Explosion, Unlimited Opus-4.6 Budget, Tangle

Shopify CTO Mikhail Parakhin detailed its AI stack across 3 projects: Tangle, Tangent, and SimGym. The post says Shopify is a 20-year, $200B software company, but does not disclose exact 2026 usage figures. The key shift is from code generation to review, CI/CD, and deployment stability.

Why it matters: HKR-H/K/R all pass: the CTO interview has a clear hook, names internal tools, and maps the coding-agent bottleneck to review and CI/CD. Missing usage numbers keep it in 78–84, not P1.

Hacker News front page

OpenAI: Workspace agents for business

OpenAI is offering Workspace agents in research preview for ChatGPT Business, Enterprise, Edu, and Teachers plans. The page says agents can run on schedules, use tools like Slack, Google Drive, and Microsoft apps, and support approval gates, audit logs, and role-based access control; pricing, model details, and rollout timing are not disclosed.

Hacker News front page

Website streamed live directly from a model

Flipbook generates an entire clickable website in real time with an image model, where each page is a pixel image and every click spawns a deeper image. The post says all on-screen text is drawn by the image model with no HTML or text overlays, and content comes from agentic web search plus model knowledge. The key point is the interaction model, not a standard generative UI; the live video stream remains an experimental, resource-heavy toggle.

Why it matters: HKR-H/K/R all pass: a live, clickable site rendered entirely as model-generated pixels is a strong hook, and the post explains the mechanism (no HTML/text overlay, agentic web search). Kept at 76 because latency, cost, model stack, and usage are not disclosed.

X · @OpenAI

Introducing workspace agents in ChatGPT—shared agents for complex tasks and long-running workflows

OpenAI announced workspace agents in ChatGPT, described as shared agents that work across tools and teams for complex and long-running workflows. Only the title and RSS snippet are disclosed; the post does not disclose supported tools, pricing, access tier, permission model, or rollout timing. The key issue to watch is the collaboration boundary of shared agents, not the headline claim alone.

Hacker News front page

Introducing Parallel Agents in Zed

Zed released Parallel Agents on April 22, 2026, letting multiple agents run in parallel in one window. The new Threads Sidebar sets per-thread folder and repo access, and supports stop, archive, and new-thread actions; the new default layout is opt-in for existing users. The key detail is permission scoping and thread orchestration, not just “multiple agents.”

Why it matters: First-party product update with clear HKR-H/K/R: parallel agents in one window plus thread-level repo and folder access control. It stays in the mid-70s because the post gives no performance delta, pricing impact, adoption data, or external validation; this is still a single-tool

Hacker News front page

Startups Brag They Spend More Money on AI Than Human Employees

Swan AI CEO Amos Bar-Joseph said his 4-person startup spent $113,000 on Claude in one month and treated that bill as headcount budget spent on AI instead of hires. The post says Swan targets $10M ARR with fewer than 10 people and cites Fundable AI claiming AI can replace a 15-person document team; the real signal is that token spend is being used as a growth metric, not proven ROI.

Why it matters: HKR-H lands on the payroll-vs-AI-bill inversion; HKR-K lands on the $113k/month Claude spend from a 4-person team. HKR-R is strong because it speaks to hiring, burn, and replacement anxiety, but this is still a trend piece with a thin sample, not a market-moving event.

Apr 22Wednesday

The Verge · AI

Meta will track employees’ computer activity to train its AI agents

Meta is installing its MCI tool on US employees’ computers and using mouse movements, clicks, keystrokes, and occasional screenshots from work apps and sites to train AI agents. Reuters says the data is meant to teach models to operate computers more like humans and automate tasks employees already do; Meta says it will not be used for performance reviews. The key gap is scope: the post discloses US staff and work contexts, but not retention, opt-out, or full rollout details.

Why it matters: HKR-H lands on the surveillance-for-agents hook. HKR-K lands on concrete collection details: US staff, mouse/keyboard events, occasional screenshots. HKR-R lands on privacy plus job-automation nerves. Strong reporting, but not a shipped product or model release, and key scope/ret

Hacker News front page

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen released the open-weight 27B dense model Qwen3.6-27B and made it available in Qwen Studio. It scores 77.2 on SWE-bench Verified vs. 76.2 for Qwen3.5-397B-A17B, and 59.3 on Terminal-Bench 2.0 under a 256K context and 3-hour timeout. The real takeaway is deployment: this is not a larger MoE, but a denser 27B model with stronger coding results.

Why it matters: Qwen3.6-27B is a substantive flagship-model release with open weights, concrete coding benchmarks, and a practical dense-deployment angle. HKR-H/K/R all pass, and per policy a major Chinese model launch should score on par with an equivalent US-lab release.

OpenAI News

Introducing workspace agents in ChatGPT

OpenAI introduced workspace agents in ChatGPT, describing them as Codex-powered agents that automate complex workflows in the cloud. The RSS snippet confirms secure work across tools for teams, but the post does not disclose pricing, availability, supported tools, or performance metrics.

Why it matters: This is a substantive OpenAI product update inside ChatGPT. HKR-H lands on the jump from chat to workspace agents, HKR-K on Codex-powered cloud execution across tools, and HKR-R on team workflow automation; the score stops at 86 because pricing, rollout, tool support, and metrics

OpenAI News

Speeding up agentic workflows with WebSockets in the Responses API

OpenAI says WebSockets in the Responses API speed up the Codex agent loop, using connection-scoped caching to cut API overhead and improve latency. The RSS snippet confirms the mechanism, but the post does not disclose latency deltas, throughput numbers, or workload conditions. The key point is transport-layer optimization, not a new model.

Why it matters: This is a developer-facing OpenAI product update at the systems layer: WebSockets plus connection-scoped caching target agent-loop round-trip cost. HKR-H/K/R all pass, but the post does not disclose latency gains, throughput, or workload bounds, so it stays mid-featured rather än

QbitAI · WeChat

SenseAuto's Sage with 3B active params claims to beat GPT-5.4 and Opus 4.6 in cars

SenseAuto released Sage, an in-car multimodal edge model with 32B total params and 3B active params, and says it scored 94% on PinchBench, above Claude Opus 4.6 at 93.3% and GPT-5.4 at 90.5%. The post says Sage runs on Nvidia OrinX with about 0.5s TTFT, 0.03s TPOT, and 80 tok/s throughput; its SCOUT training method cuts GPU hours by about 60%, and ERL raises complex-task completion by 20%. The key point is not the headline race but whether a 3B-active model can sustain multi-step tool use on device.

Why it matters: HKR-H/K/R all pass: the 3B-active-vs-GPT hook is strong, and the post gives concrete OrinX latency, throughput, and benchmark numbers. I keep it at 79 because the evidence is self-reported and the impact is narrower than a general model launch.

Synced · WeChat

Honor preinstalls YOYO Claw on MagicBook, calling it the world's first "agent laptop"

Honor said it preinstalls its YOYO Claw on MagicBook and claims 50% lower total token use than an OpenClaw setup. The post says it ships with 5 primary agents and 23 sub-agents, plus local processing, second-step confirmation, and kernel-level encryption. The practical angle is packaging agents as a device default, but the post does not disclose model names, hardware specs, pricing, or launch timing.

Why it matters: This clears HKR-H/K/R: the factory-installed agent angle is novel, and the post includes concrete details on 5/23 agents, 50% token reduction, local handling, confirmation gates, and kernel-level encryption. It stops at 76 because the model, hardware, price, and ship date are not

The Verge · AI

SpaceX cuts a deal to maybe buy Cursor for $60 billion

SpaceX announced an either-or deal: buy AI coding platform Cursor for $60 billion or pay a $10 billion fee. The RSS snippet says this could help xAI's coding tools chase Anthropic; the post does not disclose the structure, timing, or IPO linkage. Watch the $10 billion breakup fee, not just the tentative acquisition headline.

Why it matters: All three HKR axes land: the headline has a strong unexpected hook, and the report gives two hard facts — a $60B price and a $10B breakup fee. I keep it at featured, not P1, because only top-line terms are disclosed; structure, timing, and the exact xAI linkage are still undiscol

X · @dotey

Google splits Gemini Deep Research into Deep Research and Deep Research Max

Google split Gemini Deep Research into Deep Research and Deep Research Max, with public preview starting today in paid Gemini API tiers. Both run on Gemini 3.1 Pro; one targets speed and cost, while Max runs longer with more compute and repeated search and reasoning. The update adds MCP support for sources such as FactSet, S&P, and PitchBook, plus files, code execution, and File Search; the post does not disclose pricing.

Why it matters: This is a substantive Google product update: Deep Research enters paid Gemini API preview with a standard/Max split for cost-speed vs longer-running compute. HKR-H/K/R all pass, but pricing, rate limits, and performance deltas are not disclosed, so it stays in the 78-84 band.

Apr 21Tuesday

QbitAI · WeChat

Mystery model Elephant: 100B parameters reaches same-scale SOTA with high token efficiency

Ant Group's Inclusion AI team is identified as the maker of Elephant, a 100B-parameter model with 256K context and 32K output shown on OpenRouter. The post reports tests on bug fixing, summarizing a 3,000-word meeting note, and a light agent loop, plus AI BENCHY figures of about 2,500 output tokens, about 1 second average latency, and 9.6/10 consistency; the post does not disclose training details, pricing, or an official model card.

Why it matters: HKR-H/K/R all pass: a 100B model posting same-scale SOTA with token efficiency is a strong hook, and the piece includes 256K/32K, ~1s latency, 9.6/10 consistency, plus failure cases. It stays below p1 because training details, pricing, and an official model card are not disclosed

Hacker News front page

CrabTrap: An LLM-as-a-judge HTTP proxy to secure agents in production

Brex open-sourced CrabTrap, an HTTP proxy that intercepts every agent request and allows or blocks it against a policy in real time. The page shows a dual path of static rules plus an LLM judge, and logs whether each decision came from rule matching or model judgment; the post does not disclose the model, latency overhead, or error rates.

Why it matters: This lands on HKR-K and HKR-R, with HKR-H from the 'LLM-as-a-judge HTTP proxy' hook. The open-source artifact and execution-layer mechanism are concrete, but the post does not disclose the judge model, latency overhead, or false-positive rate, so it stays in the high 70s.

Google DeepMind

Partnering with industry leaders to accelerate AI transformation

Google DeepMind 宣布与 Accenture、Bain & Company、BCG、Deloitte、McKinsey 合作,帮助全球企业规模化落地前沿 AI。合作方将获得包括 Gemini 系列在内的前沿模型早期访问权,并直接对接 Google DeepMind 技术团队,聚焦金融、制造、零售、媒体娱乐等行业的智能体转型。目前仅 25% 的组织成功将 AI 规模化投入生产。

Synced · WeChat

Sergey Brin revives founder mode? Google forms a strike team to focus on AI coding

Google has formed an AI coding strike team led by Sebastian Borgeaud, with Sergey Brin and Koray Kavukcuoglu directly involved, to improve long-context coding and internal code automation. The pressure signal cited is that Google said about 50% of its code is written by coding agents and reviewed by engineers, while Anthropic staff claimed 100% code use by Claude Code and Opus 4.5; the post does not disclose team size, launch timing, or the exact Google model version. The key issue is whether Google can turn private codebase training into stronger public models.

Why it matters: HKR-H/K/R all pass: the founder-return angle is clickable, and the piece includes Google's ~50% agent-written-code claim. It stays below p1 because no public launch is disclosed, and team size, timing, and model version are missing.

Xinzhiyuan · WeChat

More agents don't help: a new survey gives three dimensions for scaling agent teams

Researchers from Emory University, the University of Oxford, and Griffith University propose a 3D framework for large-scale agent networks, classifying 8 system types by topology, memory scope, and update behavior. The survey says the core scaling bottleneck is not only communication protocols but inconsistent world models across agents; it also says current benchmarks stay small while real deployments may involve thousands to millions of agents.

Why it matters: Scores on all HKR axes: a contrarian hook, a concrete 3-axis/8-class framework, and strong resonance with agent-team builders. Kept at 78 because this is a review paper, not a model release or production deployment with fresh measured results.

Xinzhiyuan · WeChat

OpenAI launches Chronicle research preview for Codex with screen context

OpenAI launched Chronicle research preview for Codex on April 21. It is limited to ChatGPT Pro users on Mac and reads recent screen context to reduce repeated background prompts. OpenAI says data is “primarily processed locally,” but the post says some cases use cloud help; The Next Web reports screenshots are uploaded and local memories are unencrypted, while upload share and retention time are not disclosed.

Why it matters: HKR-H lands because Codex can read recent screen state, not just pasted prompts. HKR-K lands on concrete constraints—ChatGPT Pro only, Mac only, local-first with some cloud assist—and HKR-R lands on the workflow/privacy nerve for coding agents. Research-preview scope keeps it at

Xinzhiyuan · WeChat

Huawei launches Pura X Max with debut Xiaoyi companion AI

Huawei launched Pura X Max on April 20 and debuted Xiaoyi companion AI on HarmonyOS 6.1. The post says it can be invoked by double-tapping the nav bar or voice, read screen content with consent, collect tasks across apps into Calendar, and connect with Amap and Didi. The key point is system-level cross-app access and persistent side-panel UX; the post does not disclose price, model specs, or coverage.

Why it matters: It clears all three HKR axes: the OS-side companion AI is a strong hook, and the post gives concrete mechanisms like consent-gated screen reading and cross-app task collection. I kept it in featured, not higher, because price, model details, and rollout coverage are not disclosed

Latent Space

Moonshot Kimi K2.6 open-weight model refresh aims to catch Opus 4.6

Moonshot released Kimi K2.6, a 1T-parameter MoE with 32B active and 256K context. The post cites 58.6 on SWE-Bench Pro, 4,000+ tool calls, 12+ hour runs, and 300 parallel sub-agents. The key signal is long-horizon agent execution, not only open-model scores.

Why it matters: HKR-H/K/R all pass: Kimi K2.6 has a strong race narrative, concrete model and agent metrics, and direct relevance to open-model builders. The domestic flagship release signal lifts it into P1.

X · @dotey

OpenAI adds Chronicle to Codex, letting it read screen context

OpenAI added Chronicle to Codex and is rolling it out to ChatGPT Pro users on macOS; it uses periodic screenshots, OCR, and tool detection to turn recent screen activity into memory. The memory is stored as plain Markdown in ~/.codex/memories_extensions/chronicle, and the EU, UK, and Switzerland are excluded; OpenAI says screenshots are uploaded for processing, deleted afterward, and not used for training. The part to watch is risk: the background agent can burn rate limits, local plain-text files widen exposure, and OpenAI warns it amplifies prompt-injection from malicious webpages.

Why it matters: HKR-H/K/R all pass: the screen-watching memory angle is novel, and the post includes testable details like OCR, plaintext local storage, region limits, and deletion claims. The limited macOS ChatGPT Pro rollout keeps it in the 78–84 band rather than p1.

The Verge · AI

Fortnite developers can make AI characters now — just don’t try to date them

Epic Games is rolling out a “conversations” tool for Fortnite creators, turning island NPCs into AI characters that can talk with players in unscripted ways. The snippet says creators define persona, knowledge, behavior, and voice with prompts; the title says don’t try to date them, but the post does not disclose the exact guardrails or moderation system.

Why it matters: This is a mid-weight product update that gives Fortnite creators AI NPC conversation tooling. It clears all three HKR axes, but moderation rules, pricing, and base model details are not disclosed, so it stays at the low end of featured.