Show HN: TurboGPT: train 22KiB transformer in 13s
TurboGPT 是一个用 CUDA C++ 实现的字节级微型 GPT 训练项目,采用 MIT 协议,可在 13 秒内训练 22KiB 的 Transformer。
TurboGPT 是一个用 CUDA C++ 实现的字节级微型 GPT 训练项目,采用 MIT 协议,可在 13 秒内训练 22KiB 的 Transformer。
美光、三星、SK Hynix 等存储厂商将 50%-70% 产能分配给 5-16 家大客户并签订 3-5 年长期协议,以打破内存价格的周期性波动。自去年 9 月以来,2TB NVMe SSD 均价上涨 137%,32GB DDR5 套装上涨 363%,DDR5-6000 64GB 套装从 $240 涨至 $1,300-$1,400。
AMD announced an all-stock deal to acquire spatial-intelligence startup World Labs for about $8.2 billion. The deal is expected to close by the end of 2026 and still needs regulatory approval. World Labs founder Fei-Fei Li will join AMD's leadership as executive vice president and chief scientist, reporting directly to CEO Lisa Su and leading frontier research.
Why it matters: Beyond the price and where the team goes, the original lays out the world-model route and the hardware logic behind AMD chasing Nvidia.
EFF 援引《纽约时报》报道称,DraftKings 用客户投注记录训练机器学习模型,找出可能下注亏损的赌客,再向其投放促销广告吸引回站下注。EFF 指出 DraftKings 仅使用自己收集的第一方数据,说明只限制第三方数据买卖的政策不足以阻止这类广告,并主张全面禁止在线行为广告。
PostHog 在 GitHub 开源 Jeeves 项目,通过推理能力改进 Jev 类决策模型。仓库包含 drafter、inference、loader、model、prep、sdk 等模块,并附有 calibrate.py、checkpoint.py 等脚本,采用 master 单分支,已获 49 星、7 次 fork。
HPE 提出,当 AI 从试验走向客服、IT、研究等常驻生产负载,按 token 消费的模式会让支出变成难以预测的月度变动项,企业需按工作负载判断是否转向自有算力。
Jeff is a set of 0.8B parameter models fine-tuned from Qwen3.5 and Gemma 4 for zero-shot classification. Trained on consumer hardware at home, it runs inference in ~30 ms and is Jev-compatible. The post doesn't disclose dataset size or benchmarks, but the GitHub repo includes code and weights. For teams needing lightweight decision pipelines, the latency and size are practical.
Vespper (YC F24) launches DOCX MCP, the first model fine-tuned specifically for editing Word documents, shipped as an MCP server. On internal benchmarks, it claims 3x faster, 2x cheaper, and more accurate agentic editing than alternatives. The approach round-trips .docx through Markdown to avoid low-level SDKs and complex DSLs, but uses a fine-tuned model to fix the lossy conversion problem. The post does not disclose benchmark scores, pricing, or the base model.
The author spent a summer at Recurse Center, a programming retreat in Brooklyn. He joined several study groups: in Agentic Adventures, he and a partner built a remote sandbox for a vibecoding agent with dangerously-skip-permissions; they trained a tiny Shakespeare-style text predictor with minGPT and did LoRA fine-tuning for more conversational output. The Practical Deep Learning group worked through the first half of fastbook, from classical ML to building neural nets. Math Monday was a discussion group he co-started—they solved Project Euler problems, drew fractals, built Voronoi-based games, and tried the Rocq proof assistant. His most hardcore project: implementing a DEFLATE decompressor in Rust straight from RFC 1951, using LZ77 + Huffman coding, debugging by writing out bit sequences by hand. He also finished his mini-language dodo, writing the recursive match statement himself while using LLMs only for specs and tests. The post doesn't say whether he finished the emacs magit plugin.
特斯拉工厂工人被要求穿戴动作捕捉服为 Optimus 采集训练数据,部分员工因认为机器人最终将取代自己而抵触。Optimus 目前仍需在受控环境中编程执行特定任务,手部触觉传感器不可靠,特斯拉已改用可替换的传感器手套。特斯拉人形机器人还依赖中国供应商提供零部件,并面临丰田、现代等车企的竞争。
慕尼黑 CESifo 的新工作论文认为,没有证据显示应届大学毕业生的招聘出现显著、广泛的替代或减少。研究者 Robert Fairlie 和 Jane Wu 聚焦应届生,因为劳动力需求变化可能先体现在招聘减少上;这与上月一项斯坦福研究称"AI 影响"职业入门级就业落后的结论相反。
新泽西州对 DataOne 位于 Vineland 的数据中心处以 110 万美元罚款,原因是其未经许可秘密安装并运行 62 台燃气发电机,其中 45 台被无人机热成像拍到正在运转。
Anthropic's Beneficial Deployments and Applied AI teams worked with CEPI, the WHO African Regional Office and INRB to use Claude in the response to the Bundibugyo ebolavirus (BDBV) outbreak in the Democratic Republic of the Congo.
Why it matters: The post discloses how Claude was used in the DRC Ebola outbreak and how timelines changed, a view of AI's limits in public-health emergencies.
TypeSafe AI 发布 Jev,一种被称为 System One 或决策模型的新形态模型,输入文本后不输出文本,而是返回类别、是/否、评分对应的浮点数和置信度。
MIT Technology Review 与 Times of San Diego 用 15 个月调查,绘制出首张美墨边境监控塔附近死亡情况的综合地图与分析。团队分析可追溯至 2015 年的案例,向得州 17 个县警长办公室申请记录,收到超 4000 页文件,并用 Anthropic 的 Claude 通过 API 提取遗骸发现地点坐标后人工核验。
The author paid Gemini 3.1 Pro $9 to label 4,290 Reddit comments with knife brands, models, and steels, then fine-tuned GLiNER large v2.5 on those labels. The resulting model runs locally and hits 0.83 F1 against Gemini's labels after 24 minutes on a Tesla T4. Zero-shot GLiNER scored roughly 0.65 F1. The hardest bug was words_mask: the docs suggest a binary mask, but it's actually a word index; filling it with ones kept loss flat at 70. Five of ten runs produced no usable model—three config failures, two from words_mask. The post doesn't report human accuracy on Gemini's labels, so 0.83 is measured against Gemini, not ground truth.
Why it matters: A hands-on fine-tuning case with real numbers: $9 to distill Gemini labels into a local GLiNER model, lifting F1 from 0.65 to 0.83. Score capped because the domain (knife NER) is narrow, but the method transfers.
Maksim Ilin pulls together LinkedIn, WEF, Stanford, PwC, and Bain data for a September 2026 snapshot of the AI labor market. AI Engineer is the most-hired role; Research Scientist at a frontier lab is the most prestigious, with median pay around $746K at Anthropic and $1.15M at OpenAI L5. The fastest-growing niches are agentic systems and Forward Deployed Engineering—postings for the latter jumped over 1,000% YoY. Prompt engineer has faded as a job title; the skill remains but dissolved into other roles. AI skills now command a 62% wage premium in the US. Bain projects more than 1.3M AI jobs in the US by 2027 against roughly 645K available workers. In Europe, over half of AI postings sit outside tech departments, and Germany shows seven AI-user roles for every AI-developer role. Entry-level hiring got harder: employment for 22–25-year-olds in AI-exposed occupations now trails the rest by 19%.
Why it matters: A multi-source synthesis of the 2026 AI job market with concrete salary figures and role trends — high reference value for practitioners. Capped at 72 because it's a personal blog aggregating secondary data, not an original institutional report with primary research.
Fyxer breaks email into 30-50 specialized models—reply decision, intent analysis, memory retrieval, draft generation. Trained on 500K+ hours of human EA workflows, fine-tuned via LoRA, and improved through a DPO loop from user edits. Current draft acceptance rate: 53%; 90-day retention: 90%. The post doesn't specify which OpenAI model versions are used.
Someone fine-tuned a 2B LLM on WhatsApp group chat data and open-sourced the full pipeline as a GitHub cookbook. The post body is blocked by Reddit, so no details on base model, training cost, or results. Title confirms the data source (group chat), model size (2B), and goal (mimic chat style). Good starting point if you want to train a small model on your own chat logs.
Together AI updated its fine-tuning service, adding models like DeepSeek V4 Pro, MiniMax M3, and Gemma 4 31B. Users can now see live training loss and accuracy curves without waiting for the job to finish. Finer controls include learning rate schedulers, optimizer parameters, and early stopping. The post doesn't disclose pricing or region availability, but the model list and feature descriptions are detailed.
Mistral 与 Cloudera 宣布合作,将 Mistral 模型集成进 Cloudera 混合数据平台,企业可在私有云、公有云、本地及完全气隙环境中部署推理并保持完全控制。企业还能在受控环境中用专有数据训练定制模型,数据与模型所有权均归企业,模型基于开放权重。Cloudera 平台上客户管理的数据规模达 30 exabytes。
Google DeepMind released AlphaGenome Atlas, a platform holding effect predictions for 9 billion single-nucleotide variants across the human genome. It spans 1PB, more than 30 times the size of the AlphaFold Database.
Why it matters: The post gives the 9 billion-variant prediction dataset and its AVI scoring, showing what a new tool for interpreting genomic variants looks like.
Lambda engineer Zach Mueller admits his home GPU rack doesn't save money—the return is skill investment. The article uses H1 2026 data to show open models are viable, but self-hosting math is counterintuitive. Three ledgers: cost (cloud API wins for most, two H100s need ~2B tokens/month to break even), data (commercial agreements often suffice), and capability (fine-tuning and hands-on skills are the real payoff). Three tiers from renting tokens to owning hardware, with a two-to-three-week rental test recommended before buying.
Why it matters: Zach Mueller, a Lambda engineer, debunks the self-hosting cost-saving assumption with a concrete framework — HKR all hit. Deduction because this is a commentary roundup, not a primary release, and the body stops at summary level without full cost breakdown details.
A hands-on guide from Hugging Face and Liquid AI that fine-tunes LFM2.5-350M with GRPO via the TRL library. Using only 500 samples and 100 training steps on a free Colab GPU, structured-output compliance on the IFStruct benchmark jumps from 22.6% to 29.7%. The post includes the full notebook, reward-function design, and a local evaluation setup with llama.cpp on a MacBook.
Why it matters: A hands-on guide with concrete numbers and a reproducible recipe — hits H and K. But the audience is narrow and R is absent; tutorial content at the featured threshold gets 72.
calmrocks published a set of Colab notebooks on GitHub for AI engineers and forward-deployed engineers. They cover model APIs, structured output, tool calling, RAG, evals-as-the-spine, agent loops from scratch, tool design, guardrails, MCP, Skills, fine-tuning vs LoRA, prompt injection, LLMOps, and customer craft. Everything runs on the free Groq API with no frameworks. The post doesn't specify the number of notebooks or an update schedule.
Why it matters: A free Colab notebook suite for frontline engineers covering RAG, agents, fine-tuning, and security — framework-free and evals-first, with high practical value. Score held at 72 because it's a solo open-source project without community validation or cross-source discussion yet.
Anthropic announced a new Claude team plan for scientists, opening 10,000 seats to researchers worldwide. Standard seats are free; a higher-tier seat with 5x usage limits costs $15 per month for one year. Anthropic says it plans to grow the program beyond 10,000 seats in the coming months.
Why it matters: Anthropic disclosed the free and discounted seat count, application bar and usage caps, so research teams can judge their actual path in.
Sentence Transformers v6.0 introduces MultiVectorEncoder, a new model type for ColBERT-style late interaction retrieval. This blog walks through finetuning a multi-vector model that beats general-purpose retrievers on your own data. The author trained mLateOn-medical on a single RTX 3090 in 14.5 hours, and it outperformed every general-purpose retrieval model (dense, sparse, lexical) on a medical retrieval benchmark. The post covers model initialization, dataset format, loss functions, training arguments, evaluators, and the Trainer class, including multi-dataset training.
Mistral 与 HUMAIN 宣布战略合作,覆盖 AI 基础设施、先进模型开发与 AI 解决方案部署,初期聚焦网络安全和语音,并计划开发阿拉伯语表现强劲的前沿模型。合作规模达数亿欧元,Mistral 将探索使用 HUMAIN 数据中心基础设施,双方还将在沙特面向受监管行业制定联合市场策略。
Engineering teams in 2026 are fine-tuning again—not to make models smarter, but to slash inference costs on high-volume narrow tasks. FermiSense fine-tuned Qwen3.5-9B for e-commerce review, cutting cost from tens of dollars to $0.50 per 1k calls. Intercom's customer-support small model hit 73.1% resolution rate at one-fifth the cost of GPT-5.4. On the vision side, a DINOv3 classifier workflow trained a zero-API-cost local classifier with only 839 reviewed samples, reaching AP 0.9731. The article provides a decision matrix: fine-tuning pays off above ~50k daily requests with automatically verifiable outputs; below that, Prompt Caching plus RAG is the better bet. Most vendor-reported high scores lack third-party reproducible test sets, so hybrid routing remains the pragmatic middle ground.
Why it matters: A well-argued engineering trend piece with concrete numbers from FermiSense, Intercom, and a DINOv3 classifier workflow. It earns featured by making a clear, counterintuitive case backed by data. Held at 78 rather than higher because it's a synthesis/observation piece, not a f...
Soup wraps LLM fine-tuning into a single-command workflow that runs on as little as 4 GB of VRAM. It uses QLoRA with Unsloth acceleration and supports mainstream 8B models like Llama, Mistral, and Qwen. You prepare a JSONL dataset, write a YAML config, and run `soup run`. A built-in Streamlit UI lets you test the model in the browser. The README doesn't disclose exact training speed or memory numbers, but it emphasizes that everything works on a regular laptop GPU.
Why it matters: A practical tool that lowers QLoRA fine-tuning to 4 GB VRAM and one command. Hits H and K, but it's a solo dev's fresh project with no community validation, so R is absent — lands right at the featured threshold of 72. If real-world benchmarks or community feedback appear late...
Thinking Machines open-sourced Inkling-Small: 276B total parameters, 12B active, one-quarter the size of the original Inkling yet matching its performance. Full weights are available. You can fine-tune it on Tinker or chat with it via text, image, and audio in Tinker Playground. The post doesn't disclose specific benchmark scores or comparison details, so I'd hold off on the 'matching performance' claim until independent evals appear.
Why it matters: Thinking Machines released Inkling-Small: 276B total params, 12B active at inference, full weights open for fine-tuning. The compression ratio and open release are strong, but the post doesn't disclose specific benchmark numbers, so the score stays below 80.
Google 在 DOE Genesis Mission Summit 2026 上宣布投入 4000 万美元 AI tokens 和云 credits,支持 Genesis Mission 的科研人员。
Thinking Machines launched Inkling on Product Hunt, a 975B MoE open-weights model with 41B active parameters and 1M context window. It handles text, images, and audio natively, with controllable reasoning effort, under Apache 2.0. The companion Tinker API handles LoRA fine-tuning without infrastructure overhead, aimed at researchers who want full control over data and algorithms. The post does not disclose benchmark scores or pricing.
Why it matters: Thinking Machines dropped a 975B MoE open model with only 41B active params — inference cost should be low. 1M context + native multimodal is a strong spec sheet. No benchmarks or real-world latency numbers yet, so holding below 85.
Thinking Machines open-sourced Inkling, a 975B-total-param, 41B-active MoE model with a 1M-token context window, pretrained on 45T multimodal tokens. It handles text, images, and audio natively, with controllable thinking effort, and is positioned as a customizable base rather than a benchmark leader. A 12B-active Inkling-Small preview was also announced. The team demoed self-fine-tuning on their Tinker platform: the model wrote a training job, ran it, and switched to the new weights in about 27 minutes. Weights are on Hugging Face; fine-tuning is available on Tinker.
Why it matters: Thinking Machines' first open-weights model: 975B MoE, native multimodal, 1M context, pitched as a hackable base rather than a benchmark leader—clear differentiation. Not p1 because it's a debut from a new team with no ecosystem or real-world validation yet; featured is the ri...
Inkling is an open ~1T-param model that natively accepts image, audio, and text inputs with a 1M context window. Trained on 45T multimodal tokens, it uses a MoE architecture with 975B total and 41B active parameters. It includes MTP speculative decoding layers for faster inference and ships in BF16 and NVFP4 variants. Hugging Face provides day-0 support in transformers, SGLang, vLLM, and llama.cpp, covering agentic coding, multimodal vision, and audio tasks.
Why it matters: A new player, Thinking Machines, open-sources a trillion-parameter multimodal MoE model with solid specs (975B/41B activated, 1M context, 45T tokens trained) and MTP speculative decoding. H and K both hit, but R is weak — the team has no name recognition, no emotional anchor. ...
Thinking Machines lays out its technical philosophy: AI should extend human will and judgment, not replace them. The post draws on Hayek's distributed knowledge argument and Toyota's return to craftsmanship, arguing models must be as diverse and distributed as the people they serve. No model specs or release timelines are disclosed.
Why it matters: Thinking Machines published a long-form manifesto arguing AI should extend human will, not replace it, backing 'decentralized alignment' with Hayek and a Toyota case study. The argument is substantive and the stance is sharp, but it's ultimately a vision statement with no prod...
DoorDash's catalog has millions of items with wildly inconsistent names, making manual tagging slow and expensive. They built an AI metadata platform that uses multimodal models—text, images, and web search—to infer attributes like spiciness or cuisine type. Instead of human review, an 'LLM jury' of multiple strong models votes independently and aggregates a consensus, lifting annotation accuracy roughly 20% above typical human reviewers. Context-optimization agents iterate prompts in minutes, adding another 20%+ precision gain and speeding up prompt development 10x. They auto-generate training data to fine-tune small models that match frontier LLM quality at 10% of the inference cost. Distributed inference cut backfill time for millions of items from over a month to a few days. The post doesn't disclose which models, latency, or per-item cost.
Why it matters: DoorDash engineering blog shares a practical LLM-jury + multimodal labeling pipeline with concrete accuracy numbers and auto-prompt-iteration. Useful for AI data pipeline builders, but the food-delivery domain limits audience breadth — R axis missed, so it lands at the feature...
A Reddit user fine-tuned Qwen2.5-7B with 1,040 DV-DPO preference pairs, costing about $3 at Claude Haiku rates, and reported 96% composite performance versus Claude Haiku on a domain task, with 11-second latency on a 4-bit T4 setup.
Why it matters: HKR-H/K/R all pass: the hook is cheap fine-tuning, with concrete DV-DPO and latency numbers. Kept at 74 because it is a single Reddit post; task scope and eval details are thin.
Tencent Hunyuan released UniRL, using one post-training loop to cover diffusion and flow-matching models, LLM/VLM systems, and unified multimodal models, while open-sourcing two algorithms, DRPO and Flow-DPPO.
Why it matters: HKR-H/K/R all pass: Tencent Hunyuan names a unified multimodal RL loop and two open-source algorithms. This fits a strong research/open-source infrastructure release, not a flagship model launch, so it stays in the 78–84 band.
xAI brought in an executive from SpaceX’s Starlink service to run the Grok training team, replacing college-aged engineer Diego Pasini; the RSS snippet does not disclose the executive’s name, tenure, or training process.
Why it matters: HKR-H/K/R pass, but this is a training-team leadership change, not a model release or executive-level departure. Bloomberg sourcing and the Diego Pasini detail put it at the featured floor.