Skip to content

#GitHub

0 today

Sep 25Friday

GitHub Blog · AI & ML

When chat is the wrong UI

GitHub Copilot 应用推出 canvas,一种运行在应用内、无浏览器外壳的全栈小应用,可与 Copilot 智能体双向通信,并能在本地执行代码、调用第三方 API。作者认为聊天只是 AI 的通用兜底界面,用户明确任务时更该让智能体生成可复用工具,而非把智能体本身当工具、白白消耗 token。示例包括 Connect 4 游戏、Winget 包管理、SQLite 操作和开发工作流自动化。

Sep 24Thursday

GitHub Blog · AI & ML

Rendering huge pull requests in the GitHub Copilot app

GitHub Copilot 应用重建了 pull request 视图,以流畅渲染含 2,200 个文件、超百万行改动和 400 多条行内评论的超大 PR。其做法是把文档高度拆成确定性的代码几何与动态评论块两套几何:代码行高提前精确算好,评论高度按块懒测量并锚定到文件、行与侧,避免滚动跳动。

Sep 11Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Using the diff, terminal, and browser

GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让用户无需离开应用即可审查、运行和预览 AI 智能体生成的代码变更。diff 面板以绿色和红色高亮显示代码的增删改,终端面板支持直接运行项目命令并可通过 Run 按钮配置脚本,浏览器面板则提供 Pick & Polish 工具来选取页面元素并让智能体调整。

Sep 4Friday

GitHub Blog · AI & ML

GitHub Copilot app for Beginners: Run several agents at once

GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。

Sep 25, 2025Thursday

OpenAI News

More ways to work with your team and tools in ChatGPT

OpenAI rolled out shared projects for ChatGPT Business on September 25, 2025, and made them available for Enterprise and Edu plans. Shared projects support email or link invites, two access levels, and private project memory; Enterprise and Edu have them off by default under admin control. OpenAI also added Gmail, Google Calendar, Outlook, Teams, SharePoint, GitHub, Dropbox, and Box connectors, and said ChatGPT can now choose connectors automatically per prompt.

Why it matters: HKR-H/K/R all pass: shared projects, 8 connectors, and prompt-routed connector selection are concrete workflow changes with clear admin controls. I keep it below 85 because this is a collaboration-layer product update, not a model release or a broad capability jump.

Sep 15, 2025Monday

OpenAI News

Introducing upgrades to Codex

OpenAI released GPT-5-Codex and made it the default model for Codex cloud tasks and code review; in testing, it worked independently for more than 7 hours on complex tasks. OpenAI says it used 93.7% fewer tokens than GPT-5 on the lowest 10% of employee turns, while spending 2x longer reasoning, editing, and testing on the highest 10%. The key point is one model now spans interactive coding and long-running agentic execution; pricing and full availability details are not fully disclosed in the provided body.

Why it matters: This is a substantive OpenAI developer-tool update: GPT-5-Codex becomes the default for Codex cloud tasks and code review, with concrete numbers on 7-hour autonomy and token use. HKR-H/K/R all pass; pricing and full availability are not fully disclosed in the excerpt, so it stays

OpenAI News

Addendum to GPT-5 system card: GPT-5-Codex

OpenAI published a GPT-5-Codex system card addendum on September 15, 2025, stating the model is optimized for agentic coding in Codex and is available in terminal, IDE, web, GitHub, and the ChatGPT mobile app. The post says it uses reinforcement learning on real-world coding tasks, plus safety training for harmful tasks and prompt injection, with sandboxing and configurable network access. Benchmark scores, pricing, and context window are not disclosed.

Why it matters: HKR-H/K/R all pass: this is an OpenAI coding-agent model spanning terminal, IDE, GitHub, web, and mobile, with concrete training and safety details. I kept it below 85 because benchmarks, pricing, and context window are not disclosed in the body.

May 16, 2025Friday

OpenAI News

Addendum to OpenAI o3 and o4-mini system card: Codex

OpenAI published a May 16, 2025 addendum to the o3 and o4-mini system card, stating that Codex is a cloud coding agent powered by codex-1, an o3 variant tuned for software engineering. Each agent runs in an isolated cloud container preloaded with the user's code and environment, then loses internet access while it reads or edits files and runs tests, linters, and type checkers. The practical detail is the audit trail: Codex cites terminal logs and files, and its output can be exported as a GitHub PR or local diff.

Why it matters: This clears HKR-H/K/R because the addendum adds concrete execution details: isolated cloud containers, user-defined dev envs, internet disabled after setup, and test-running behavior. Strong featured score, but not p1: it is supporting safety documentation, not the primary launch

OpenAI News

Introducing Codex

OpenAI released the Codex research preview on May 16, 2025, a cloud software engineering agent powered by codex-1 that can handle multiple coding tasks in parallel. It runs each task in an isolated sandbox, can read and edit repos, execute tests and commands, and usually finishes in 1 to 30 minutes with terminal logs and test outputs as evidence. It launched for ChatGPT Pro, Business, and Enterprise users, then expanded to Plus on June 3; the post excerpt does not fully disclose pricing or complete limitations.

Why it matters: This is a same-day write: OpenAI moved from code assistance to a cloud software-engineering agent, with launch access for ChatGPT Pro, Business, and Enterprise. HKR-H/K/R all pass, with concrete mechanics and verifiable outputs; incomplete pricing and limits keep it at 88.

Aug 13, 2024Tuesday

OpenAI News

Introducing SWE-bench Verified

OpenAI released SWE-bench Verified, a human-validated subset built with the benchmark’s authors to assess real software issue resolution more reliably. The post names 3 failure modes in SWE-bench: overly narrow tests, underspecified issue statements, and unreliable environment setup; as of Aug. 5, 2024, top agents scored about 20% on SWE-bench and 43% on SWE-bench Lite. The key point is that the original benchmark can systematically underestimate coding-agent ability.

Why it matters: This is a strong benchmark release, not a routine post: OpenAI re-audited SWE-bench with the original authors, named 3 defect classes, and reported new score ceilings of 20% and 43%. HKR-H/K/R all pass because it changes how builders read code-agent leaderboards.