Show HN: TurboGPT: train 22KiB transformer in 13s
TurboGPT 是一个用 CUDA C++ 实现的字节级微型 GPT 训练项目,采用 MIT 协议,可在 13 秒内训练 22KiB 的 Transformer。
TurboGPT 是一个用 CUDA C++ 实现的字节级微型 GPT 训练项目,采用 MIT 协议,可在 13 秒内训练 22KiB 的 Transformer。
PostHog 在 GitHub 开源 Jeeves 项目,通过推理能力改进 Jev 类决策模型。仓库包含 drafter、inference、loader、model、prep、sdk 等模块,并附有 calibrate.py、checkpoint.py 等脚本,采用 master 单分支,已获 49 星、7 次 fork。
Jeff is a set of 0.8B parameter models fine-tuned from Qwen3.5 and Gemma 4 for zero-shot classification. Trained on consumer hardware at home, it runs inference in ~30 ms and is Jev-compatible. The post doesn't disclose dataset size or benchmarks, but the GitHub repo includes code and weights. For teams needing lightweight decision pipelines, the latency and size are practical.
GitHub's security team ran its open-source AI security agent on the Android Open Source Project, automatically found 24 vulnerabilities, and submitted patches. The post doesn't disclose the vulnerability types, false positive rate, or which underlying model was used. The key takeaway: code auditing is shifting from manual review to autonomous agent workflows, and the tool is already open source.
LightCloud is a new cloud console that organizes projects, databases, and containers into a file-system tree. Static sites go to a global CDN; containers scale to zero when idle; Postgres is provisioned from the same project. It connects to GitHub, GitLab, or Bitbucket repos, auto-builds on push, and gives every branch and PR a preview URL. It also integrates with Claude Code via an MCP server: one command to install, then ask “Deploy this project to Light Cloud” and it signs you up, asks two questions, and returns a live URL. The post doesn't spell out full pricing details beyond a free Hobby plan with $5 usage credit.
A GitHub project goes viral: run a 700B-parameter GLM on a laptop without a GPU. The trick is using SSD as VRAM, trading storage for speed. The post doesn't disclose exact latency or precision loss, but the idea is straightforward: swap memory for disk. For developers without a GPU, this is a low-cost way to test large models.
GitHub published a beginner tutorial for the Copilot app, focusing on building custom workflows with canvases. The post is a hands-on guide, not a new feature announcement. No technical specs or model updates are disclosed. Useful for developers new to Copilot workflows.
GitHub's engineering team shares how they cut server-side rendering time by 55% by shipping more CSS upfront instead of lazy-loading it. The key insight: inline critical styles in the initial HTML so the browser doesn't wait for separate CSS downloads. They detail how they analyzed the critical rendering path, extracted above-the-fold CSS, and avoided duplication after migrating to CSS Modules. First Contentful Paint also improved. A practical case study for teams working on front-end performance or SSR optimization.
Jev is an open-source code review tool that prioritizes understanding developer intent before explaining code changes with AI. It offers a local CLI, agent skill, and GitHub extension. The post doesn't disclose which model it uses, pricing, or performance benchmarks.
This GitHub repo offers a full workflow from idea to production for teams using AI coding agents like Copilot. It covers specs, architecture, testing, security, code review, and CI/CD, claiming to be battle-tested. The post doesn't include benchmarks or user stories, so you'll have to try it yourself.
GitHub Copilot 应用推出 canvas,一种运行在应用内、无浏览器外壳的全栈小应用,可与 Copilot 智能体双向通信,并能在本地执行代码、调用第三方 API。作者认为聊天只是 AI 的通用兜底界面,用户明确任务时更该让智能体生成可复用工具,而非把智能体本身当工具、白白消耗 token。示例包括 Connect 4 游戏、Winget 包管理、SQLite 操作和开发工作流自动化。
Parcha.ai open-sourced AgentRun, a declarative DSL for orchestrating multi-agent workflows. It chains agent steps into reusable pipelines with branching, loops, and parallel execution. The post doesn't spell out how it differs from LangChain or CrewAI, but the idea is 'define workflows in code, not prompt hacks.' Worth a look if you're putting agents into production.
GitHub Copilot 应用重建了 pull request 视图,以流畅渲染含 2,200 个文件、超百万行改动和 400 多条行内评论的超大 PR。其做法是把文档高度拆成确定性的代码几何与动态评论块两套几何:代码行高提前精确算好,评论高度按块懒测量并锚定到文件、行与侧,避免滚动跳动。
Jason Page open-sourced Seal, a tool that lets you write digital letters and passwords that only open for your family after you die. You set an inactivity period—say 30 days without logging in—and Seal sends the encrypted content to your designated contacts. The post doesn't spell out how it reliably detects death vs. a forgotten login, so take that with a grain of salt.
Orbital is an open-source alternative to Claude Project that claims 'context is yours, agents are replaceable.' It turns conversation context into reusable assets instead of locking it inside a platform. The project just launched with 277 stars on GitHub. The post doesn't specify which models it supports, whether it's compatible with Claude API, or deployment requirements. If you're frustrated that Claude Project won't let you export context or swap agents, this is worth a look.
GitHub engineers used Copilot's agent mode to migrate the core Copilot runtime from TypeScript to Rust, producing roughly 830K lines of code. The migration ran in three phases: file-by-file translation by agents, test-driven bug fixing, and performance/security review. After migration, service startup dropped from 30s to 3s, memory usage fell to one-third, and time-to-first-token went from 11s to 1.2s. The team stresses that humans stayed in the loop—agents did the heavy lifting, engineers owned architecture, code review, and test coverage. The post doesn't spell out exact cost savings but states the migration was done 'with a smaller team in less time.'
Why it matters: First-party case study from GitHub: a Copilot agent drove an 830K-line Rust migration with hard performance gains. HKR all hit, but it's a product capability showcase rather than an independent breakthrough, so capped at 82 in the featured tier.
Jim Nielsen argues that uptime percentages like 99.9% are meaningless to non-infrastructure users. He cites Jason Gorman's point that each 'nine' is as hard to achieve as the last, but the numbers look nearly identical. His fix: replace '98.31% uptime' with '12 hours affected in the last 30 days (98.31% uptime).' The post uses screenshots from GitHub and Claude status pages but doesn't disclose their raw data sources.
ImpactGate is a GitHub merge gate that scores how much structural decay AI-generated code introduces. It blocks PRs that pass tests but degrade maintainability. The post doesn't disclose the scoring algorithm or thresholds, but the idea is clear: extend quality gates from correctness to structural health.
Panel is a research workspace where the agent dynamically builds and arranges its own panes. It's open-source on GitHub with 594 commits. The post doesn't specify which models it supports or whether it integrates external tools. Worth a look for devs exploring agent-driven UI generation.
All three are early-stage. Runway Solaris generates interactive UIs frame-by-frame with no frontend code—only curated demos and a waitlist so far, no public testing, pricing, or API. Google WikiSkill distills agent failure logs into reusable skill manuals, lifting Gemini-3.5-Flash accuracy from 49.5% to 68.1%, but skills from a small model can hurt a larger one; no official code repo. GitHub GHES 3.22 lets enterprises self-host Copilot CLI inside air-gapped networks with admin-managed model endpoints, though many features are disabled and it's labeled a technical preview.
Why it matters: Three items bundled, with Solaris as the main hook. Runway's frame-by-frame interface rendering is genuinely novel, but there's only a curated demo and waitlist — no public access, no third-party testing, and the cost comparison dodges standard web rendering. That keeps it bel...
Kinesis is an open-source project that lets you control your Mac using Meta's Neural Band. It translates brain-computer interface signals into native macOS actions like gesture scrolling and cursor movement. The project is fresh on GitHub with 23 stars and runs locally without cloud dependency. The post doesn't disclose latency, accuracy, or supported macOS versions, but the code is public for testing.
Liniora is an AI workspace for engineering teams that pulls tickets, branches, PRs, Slack threads, and meeting notes into one place. Its AI builds a semantic graph of your codebase and conversations, so you can ask natural-language questions like “what was the decision on the payment gateway?” and get an answer. It also auto-summarizes pull requests, extracts action items from calendar syncs, and lets you create branches from tickets. Free tier: 3 users, 2 projects, 50 AI actions/month. Pro: $9/user/month, unlimited everything. The post doesn't specify which AI model powers the semantic search or how it's trained.
Someone fine-tuned a 2B LLM on WhatsApp group chat data and open-sourced the full pipeline as a GitHub cookbook. The post body is blocked by Reddit, so no details on base model, training cost, or results. Title confirms the data source (group chat), model size (2B), and goal (mimic chat style). Good starting point if you want to train a small model on your own chat logs.
GitHub's Japan/Korea marketing lead shows how to turn event planning, execution, and follow-up into code using Copilot. The post details generating event pages, automating follow-up emails, and analyzing attendee data. The core idea: treat marketing ops as software engineering, with AI cutting repetitive work.
Herdr Studio is an open-source browser client that gives you a visual workspace for all your AI agent terminals, files, diffs, and worktrees. It relies on the Herdr daemon to keep agent sessions alive even when your browser or laptop disconnects. Supports local and SSH connections, and can be installed as a PWA on mobile. The post doesn't spell out platform support beyond macOS and Linux install scripts.
GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让用户无需离开应用即可审查、运行和预览 AI 智能体生成的代码变更。diff 面板以绿色和红色高亮显示代码的增删改,终端面板支持直接运行项目命令并可通过 Run 按钮配置脚本,浏览器面板则提供 Pick & Polish 工具来选取页面元素并让智能体调整。
The biggest issue with AI-generated code is trust. Engineers must be accountable for what they ship, even if written by AI. The author recommends deterministic tooling (typed languages, linters), hand-written test cases, and enforcing small PRs to fight quality degradation. Code generation is cheap now, but the outcome matters more than the code itself.
GitHub shared a research preview of Project HydraFusion, a runtime model router that sends each request to a different model. Simple tasks hit cheap small models; hard ones go to frontier models like Claude Sonnet 4.5. GitHub claims this keeps Copilot's response quality while cutting inference cost to one-fifth of using frontier models alone. No launch date yet—it's a research preview.
Why it matters: Official GitHub blog research preview with concrete cost figures and named models—not pure marketing. The lack of a launch timeline keeps it at the 78 featured threshold.
GitHub Copilot 应用支持同时运行多个智能体会话,每个会话运行在独立的 Git worktree 上,互不干扰且各自保留上下文,可随时切换并从中断处继续。用户可在会话视图中查看各任务标题与进度,例如在同一项目上并行执行 funded sort 开发、无障碍审查和测试运行。
OpenRouter data shows agents consume 7.3T tokens weekly, nominally 5.2× human usage. But 70–85% are cached reads; with ~90% discount, the real bill is roughly 2×. GitHub's Knowledge Compressor prototype halves doc length and claims breakeven at 2,000 reuses, but factoring in caching pushes the median to 5,000+. OpenAI's Jalapeño chip beats Nvidia GB200/GB300 on fixed-length benchmarks, yet lacks AgentX scores for real agent workloads. All three stories share one distortion: prompt caching inflates headline numbers.
Why it matters: Three stories bundled, but the core value is the first: someone finally separated nominal agent token consumption from the caching-discounted real cost, landing at ~2x. The OpenAI chip benchmark and GitHub compression prototype are bonuses but less dense. Cross-source cluster ...
GitHub published an engineering blog detailing how they cut Copilot's AI coding costs to roughly one-third without hurting task quality. The key move: training a 1.8B-parameter model on 1,040 preference pairs to act as a router that decides when to use a cheap model and when to call a stronger one. After rollout, strong-model calls dropped 70% and overall latency stayed under 11 seconds. The post also mentions a training method called DV-DPO that uses preference data to teach a small model a specific response style. One caveat: these numbers come from GitHub's own setup, so your mileage may vary.
Why it matters: GitHub shared a concrete cost-optimization engineering post with real numbers and methods, directly useful for teams shipping AI products. Score capped because it's an engineering optimization, not a new model release.
Shopify open-sourced Tangle, a drag-and-drop ML pipeline editor. It lets teams visually build workflows, collaborate, use any language/framework, and cache intermediate results. Code is on GitHub; try the online Playground.
Neil Alexander calls out the rise of AI-generated drive-by PRs and vulnerability reports aimed at inflating GitHub profiles. He cites a contributor with near-zero activity since 2018 who suddenly submitted three spelling-fix PRs—all written and signed off by Claude. He closed them without comment. Security reports are also clearly AI-produced, and his team now declines CVE notices for low-severity items. The bottom line: contribute because you care, not to farm green squares.
Why it matters: First-person maintainer rant with concrete examples and pattern analysis, hits all three HKR axes. Capped at 78 because it's a personal blog post, not an industry event, and the problem itself isn't a new discovery.
Plicara scanned 1.87 million agent skill files and found the non-English share jumped from 13.0% in Q1 2026 to 16.3% in Q2—much faster than GitHub docs ever diversified. Chinese skills sit at 6.2%, nearly double the Chinese share of GitHub documentation. European languages more than doubled in the same window, while Japanese and Korean slipped. Published numbers disagree because each study sampled a different population: curated marketplaces, domain slices, or English-seeded crawls. The post does not address whether non-English instructions degrade agent performance, so hold that question open.
Why it matters: Plicara scanned 1.87M agent skill files and found non-English share jumped from 13% to 16.3% in one quarter—far faster than GitHub doc diversification. Chinese skills at 6.2% (2x the GitHub baseline) is a concrete stat. Solid data, fresh angle, but Plicara isn't a household na...
Only the title is available; the body is empty. Cursor has launched Origin, a code hosting platform that competes directly with GitHub. The title mentions Stacked PR, AI Agent, and Copilot, suggesting deep AI integration, but no details on features, pricing, or release date are disclosed.
Cursor is rolling out Origin, its own code hosting service, in early beta for paid users. You can create repos, open pull requests, browse code, and sync existing GitHub repos with real-time two-way PR comments. Agents live inside every repo—ask questions, make changes, or push branches. First app integrations include Vercel for preview deploys, plus Depot and Buildkite for CI. The post doesn't say when free-tier access will arrive.
Why it matters: Cursor's key move from editor to platform: built-in AI assistant per repo, GitHub sync, and Vercel/Depot integrations add real product substance. Capped below 85 because it's early beta with no pricing or GA date disclosed — real-world reliability is still unknown.
Over 60% of AutoGPT pull requests now come from AI tools. Maintainer Reinier van der Leer uses an AGENTS.md file to set rules for AI contributors and adds skill gating so only agents that pass linting and unit tests can submit code. Spam PRs dropped sharply, though the post doesn't say how many human contributors were wrongly blocked.
Why it matters: First-hand maintainer account from AutoGPT with hard numbers (60% AI PRs) and two reproducible mechanisms. Downside: the post doesn't disclose how many human contributors got blocked, and it's a single-project case study — generalizability is unproven.
GitHub engineers share a workflow for taming AI-generated mega-PRs: after letting AI produce an entire feature in one shot, they use stacked PRs to automatically split thousands of lines into logical, independent chunks of 200–400 lines each. The core idea is to generate the full change first, then slice it into a stack based on file dependencies and semantics, so reviewers can focus on one concern per layer. The post includes concrete commands and branch-naming conventions, but doesn't disclose internal adoption rates or review-time comparisons.
Why it matters: GitHub's official engineering blog shares a hands-on workflow for handling large AI-generated code blocks, with concrete commands and splitting logic that teams using AI for coding can directly reference. But the lack of internal usage data and quantified review-time improveme...
Starling is a Wayland compositor that drives the GPU directly and runs Chrome, Slack, and Zoom—not a browser mock-up. One person directed AI to write it over six months, producing roughly 335K lines of Swift, C, and C++, with the desktop and its Wayland/X11 servers at about 62K lines. It supports one-click tiling/floating switching, hot-plug multi-monitor, virtual desktops, and a dock with per-pixel glass computed via fragment shaders. It's an early preview (v0.2.1) on Ubuntu, with all code public on GitHub. The post does not disclose which AI models were used, how coding tasks were divided, or any performance benchmarks.
Why it matters: One person + AI shipped a GPU-driven Wayland desktop in six months that runs Chrome and Slack natively, with all code public and verifiable. It's an extreme case study in AI-assisted development with concrete numbers and a reproducible artifact — not marketing fluff. Not scori...
This paper studies tens of thousands of Microsoft engineers during the early-2026 rollout of Anthropic’s Claude Code and GitHub Copilot CLI. Three findings stand out. First, initial adoption spread mainly through peer social networks, not top-down mandates. Second, retention correlated more with an engineer’s coding activity level than with demographics. Third, adopters merged roughly 24% more pull requests than they otherwise would have, and the lift held across the four-month window. The authors use merged PRs as a proxy for output while noting a merged PR is not the same as delivered value. They also flag that token spend at organizational scale can reach millions of dollars annually, so misjudging adoption or retention makes the rollout expensive without changing engineering velocity.
Why it matters: Large-scale empirical study from inside Microsoft with concrete numbers and counterintuitive findings (peer-driven adoption, retention unrelated to demographics). HKR all hit. Slight ding for being a paper rather than a product launch, but information density clears the featur...