Skip to content

#开源/仓库

3 today

May 6Wednesday

Synced · WeChat

Two Chinese open-source projects turn Mac into a private AI workstation

Mininglamp open-sourced Cider and Mano-P 1.0 for Apple Silicon local inference and GUI agents. Cider speeds Qwen3-VL-2B prefill by 57%–61% on M5 Pro; Mano-P 1.0-72B scores 58.2% on OSWorld. The key constraint is W8A8 memory: on 16GB devices accuracy falls from 58.0% to 54.0%, so 32GB+ is recommended.

Why it matters: HKR-H/K/R all pass: the Mac-local workstation angle is clickable, and Cider/Mano-P include testable numbers. Score stays at 80 because the source entity is not a top-tier model lab.

May 5Tuesday

r/LocalLLaMA

SenseNova-U1-8B-MoT open-source multimodal architecture draws LocalLLaMA discussion

SenseNova open-sourced SenseNova-U1-8B-MoT, an 8B native multimodal understanding and image-generation model. Its Hugging Face text says NEO-Unify removes VE and VAE, supports interleaved image-text generation, and high-density rendering; the post does not disclose test scores. The key question is whether the monolithic design yields reproducible gains.

Why it matters: HKR-H/K/R all pass: the open 8B unified multimodal model has a concrete architecture hook. No benchmark scores, license detail, or deployment cost are disclosed, so it stays in the 72–77 band.

r/LocalLLaMA

vibevoice.cpp: Microsoft VibeVoice ported to ggml/C++ with no Python at inference

LocalAI released vibevoice.cpp, a ggml/C++ port of Microsoft VibeVoice for CPU, CUDA, Metal, and Vulkan inference. TTS uses a 30s reference clip for 24kHz cloned speech; ASR uses a 7B model with diarized JSON and was tested on 17min audio. The key constraint is memory: 17min CPU Q8_0 peaks near 26GB, with no streaming output yet.

Why it matters: HKR-H/K/R all pass: a practical open-source VibeVoice C++ port with concrete runtime numbers. Reddit-source scope and niche audio deployment keep it in the 72–77 featured band, not same-day must-write.

r/LocalLLaMA

FastDMS: 6.4X KV-cache compression running faster than vLLM BF16/FP8

FastDMS released an MIT implementation that cuts KV memory to 1/5–1/8 of vLLM BF16 at 8K context. A Llama-3.2-1B replication reports PPL 9.200 with 6.4x compression; Qwen3-8B c=1 drops KV from 1.406 GiB to 0.184 GiB. The key detail is physical reclamation of evicted slots, not just nominal KV-byte reduction.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, with compression, PPL, KV GiB deltas, and physical slot reclamation. Reddit/open-source sourcing keeps it in 78–84, below P1.

May 4Monday

r/LocalLLaMA

Deep research report with Hermes Agent and qwen3.6-35b-a3b Q6_K

A Reddit user used Hermes Agent and qwen3.6-35b-a3b Q6_K to produce a 21-page research report. The run took 6 loops and over 5 hours on an RTX 4060, at about 28 tokens/s. The repo includes prompts, scripts, intermediate artifacts, and the final report.

Why it matters: HKR-H/K/R all pass: this is a local-agent experiment with hardware, runtime, speed, and artifacts. Reddit source limits reach, so it stays in the 72–77 featured-threshold band.

QbitAI · WeChat

DeepSeek-TUI, a “DeepSeek Claude Code,” reaches 2.3k GitHub stars

DeepSeek-TUI reached 2.3k GitHub stars; the Rust project is MIT-licensed. It targets DeepSeek V4 with a 1M-token context, RLM up to 16 V4 Flash subtasks, MCP, Shell, Git, and three control modes. Watch cache misses: uncached tokens cost 10x cached tokens.

Why it matters: HKR-H/K/R all pass: the hook is a DeepSeek-flavored Claude Code, with 2.3k stars, 1M tokens, 16 subtasks, and a 10x cache-miss cost gap. Impact is developer-specific, so it sits in the 72–77 band.

r/LocalLLaMA

450M On-Board VLM Wildfire Detection Pipeline with Sentinel-2 and LFM2.5-VL

PauLabartaBajo shared a wildfire detection PoC using 450M LFM2.5-VL on Sentinel-2 imagery. It pairs RGB and SWIR tiles, simulates orbit with SimSat, and covers 22 fire-prone sites. The key constraint is bandwidth: on-board inference downlinks only a JSON risk profile.

Why it matters: HKR-H/K/R all pass: the story has a counterintuitive edge-VLM hook and concrete numbers. Single-source Reddit PoC and a narrow wildfire-use case keep it below the 78+ band.

May 3Sunday

r/LocalLLaMA

LLM proxy that lets Claude Code talk to any model

DataNebula released open-source rosetta-llm, letting Claude Code call multiple providers through one gateway. It translates Anthropic Messages, OpenAI Chat, and OpenAI Responses, and round-trips encrypted reasoning via the signature field. The key detail is thinking-block fidelity for multi-turn agent prompt-cache hits.

Why it matters: HKR-H/K/R all pass, but this is a Reddit open-source tool post with no adoption, stars, or benchmark data disclosed. Score stays in the mid-weight tooling band, not 78+.

r/LocalLLaMA

Upskill: skill registry your agent consults before it starts, with 10k+ indexed skills

Autoloops released Upskill, an open-source skill registry with 10k+ indexed skills for agents. Search combines Postgres full-text search, 1024-dim embeddings, and reranking by stars, installs, and feedback. LLM adversarial review blocked hundreds of skills at index time.

Why it matters: HKR-H/K/R pass: a useful open-source agent registry with concrete retrieval and safety mechanics. Source authority is low and adoption is unproven, so it stays in the 72–77 featured band.

QbitAI · WeChat

GS-Playground Embodied AI Simulation Framework Open-Sourced with High-Throughput 3DGS Rendering

Tsinghua AIR DISCOVER Lab and partners open-sourced GS-Playground, accepted by RSS 2026. On an RTX 4090, it reports 10,000 FPS at 640×480 and 2,048 parallel scenes; a 50-humanoid benchmark reaches 1,015 FPS. The key point is coupling batch 3DGS rendering with parallel physics.

Why it matters: HKR-H/K/R pass: the open-source RSS 2026 work reports concrete RTX 4090 throughput and parallel-scene numbers. The robotics-simulation scope is narrower than a model launch, so it fits the 78–84 band.

r/LocalLLaMA

Built a C++17 transformer from scratch with 0.83M params and CPU training

Reddit user Suspicious_Gap1121 released Quadtrix.cpp, a C++17 GPT-style model with 0.83M parameters. It uses 4 layers, 4 heads, 200d width, and a 128-character context; one CPU core trained on 31.4M characters for 76.2 minutes to 1.6371 nats val loss. The key detail is handwritten backprop for LayerNorm, attention, Q/K/V, dropout, and AdamW without PyTorch, BLAS, or autograd.

Why it matters: HKR-H/K/R all pass: the no-framework C++17 build is clickable, the training setup is specific, and local-LLM builders care about dependency-free control. It stays in the 72–77 band because it is a small personal project.

May 2Saturday

r/LocalLLaMA

I built Semvec: A constant-cost semantic memory for LLMs, looking for testers

A developer released Semvec, replacing unbounded chat history with fixed-size semantic state. Its 48-turn benchmark claims about 76% token reduction, with identical input footprint at turn 10 and 10,000. It supports OpenAI-compatible LLMs, MCP, Claude Code, Cursor, and multi-agent shared state.

Why it matters: HKR-H/K/R all pass, but this is a Reddit self-release with author benchmarks only. Treat it as an interesting indie memory tool, not a same-day industry story.

r/LocalLLaMA

Qwen3.6-27B hits 72 tok/s on RTX 3090 with native vLLM on Windows

Reddit user One_Slip1455 released a native Windows vLLM launcher for Qwen3.6-27B, reaching 72 tok/s on an RTX 3090. It reports 64.5 tok/s at ~25k tokens, 53.4 tok/s at 127k ctx on one GPU, and 160k ctx with PP=2 on 2×3090. The key detail is no WSL or Docker, an OpenAI-compatible endpoint, and an INT4 quant path.

Why it matters: HKR-H/K/R all pass: native Windows on an RTX 3090 is the hook, the post gives tok/s and ctx figures, and it hits local-inference cost concerns. Reddit single-source limits it to the lower featured band.

QbitAI · WeChat

Tencent Hunyuan open-sources 440MB offline translation model, claims Google Translate quality lead

Tencent Hunyuan open-sourced Hy-MT1.5-1.8B-1.25bit, compressing a 1.8B translation model to 440MB. It supports 33 languages and 1,056 directions, with an Android demo running offline on Snapdragon 888 and 8GB RAM. The key detail is Sherry 1.25-bit quantization: 3 of every 4 weights use 1 bit and 1 is zeroed.

Why it matters: HKR-H/K/R all pass: the story has a strong offline-phone hook, concrete quantization details, and practitioner relevance around edge inference. It stays below P1 because this is a vertical translation model, not a major foundation-model release.

May 1Friday

r/LocalLLaMA

PFlash: 10x prefill speedup over llama.cpp at 128K on an RTX 3090

PFlash cuts Qwen3.6-27B Q4_K_M 128K TTFT to 24.8s on an RTX 3090, versus 248.4s cold for llama.cpp. It uses a Qwen3-0.6B drafter to score token importance, keeps 5% of spans, and runs C++/CUDA without Python, Triton, or PyTorch. The quality caveat is clear: only NIAH single-needle passes from 32K to 128K; RULER and multi-needle results are not disclosed.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit claim with quality evidence limited to single-needle NIAH 32K–128K. RULER and multi-needle results are not disclosed, so it stays at featured threshold.

Xinzhiyuan · WeChat

Developer Builds WorldX, an AI World Generator, During a 10-Day Wedding Leave

An independent developer built WorldX in 10 days, generating a full AI world from one sentence in about 5 minutes. The system uses a 6-step map pipeline, about 30k–180k tokens per world, Tick loops, layered memory, and two-axis emotion. The key mechanism is overlay labeling plus color-difference localization for deterministic coordinates.

Why it matters: HKR-H/K/R all pass, but this is an indie project rather than a platform release, so it stays in the 72–77 band. The concrete pipeline, token range, and agent memory details justify featured.

QbitAI · WeChat

Peking University Open-Sources Unified World Model Framework for Synthesis and Reasoning Tasks

Peking University DCAI and Kuaishou Kling open-sourced OpenWorldLib for four task types: video generation, 3D modeling, VLA control, and multimodal reasoning. Its Pipeline coordinates Operator, Reasoning, Synthesis, Representation, and Memory modules, supporting forward and stream execution. The key test is whether unified interfaces cut cross-task reproduction cost.

Why it matters: HKR-H/K/R all pass: the post gives a concrete open-source framework, task scope, modules, and inference modes. It lacks benchmark results, adoption data, or major ecosystem integration, so it stays at 78.

Hacker News front page

Show HN: Pu.sh – a full coding-agent harness in 400 lines of shell

Pu.sh ships a coding-agent harness in about 400 lines of shell, using only sh, curl, and awk. It supports Anthropic and OpenAI, 7 tools, REPL, auto-compaction, checkpoint/resume, pipe mode, and 90 no-API tests. It excludes TUI, streaming, images, OAuth, and Windows.

Why it matters: HKR-H/K/R all pass, but this is a small Show HN open-source tool, not a model or platform release. HN frontpage plus a reproducible 400-line implementation clears the featured bar.

Apr 30Thursday

r/LocalLLaMA

inclusionAI/Ling-2.6-1T · Hugging Face

inclusionAI open-sourced Ling-2.6-1T on Hugging Face, with 1 trillion parameters. It uses MLA plus Linear Attention and Contextual Process Redundancy Suppression to reduce CoT overhead. The post cites AIME26 and SWE-bench Verified but does not disclose scores.

Why it matters: HKR-H/K/R all pass, but benchmark scores for AIME26 and SWE-bench Verified are not disclosed. A 1T open model with a named architecture mechanism fits featured, not P1.

X · @dotey

Inside Hermes Agent's Memory System and How It Avoids OpenClaw's Pitfalls

Hermes Agent splits memory into 4 layers: prompt files, SQLite session search, skills, and optional Honcho. MEMORY.md is capped at 2,200 chars, USER.md at 1,375; writes apply after a new session or compression. The key design is cache-first: keep system prompts stable and retrieve long-tail history via tools.

Why it matters: HKR-H/K/R all pass: the OpenClaw contrast is clickable, and the memory limits/mechanisms are concrete. Single X-source tutorial, not a product release, keeps it at the featured threshold.

Apr 29Wednesday

QbitAI · WeChat

Avenir-Web Open-Sources Web Agent Harness With 53.7% on ONLINE-MIND2WEB

UCL, Princeton, and Edinburgh open-sourced Avenir-Web, reaching 53.7% success on ONLINE-MIND2WEB. The training-free harness uses EIP, MoGE, checklists, and adaptive memory across 136 sites and 300 live tasks. The key signal: with Gemini 3 Pro, it beats Claude Computer Use 3.7 at 47.3%.

Why it matters: HKR-H/K/R all pass: the story has a sharp SOTA web-agent hook, concrete benchmark numbers, and practitioner resonance around agent reliability. This is a strong open-source research release, not a major lab model launch, so 82 fits the 78–84 band.

r/LocalLLaMA

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-Unify Architecture

SenseNova released SenseNova-U1 with 4 MoT multimodal models. The post lists 8B and A3B variants with GitHub and HuggingFace weight links. The key claim is a monolithic architecture instead of adapters; benchmarks are not disclosed.

Why it matters: HKR-H/K/R pass: open weights, 4 MoT multimodal models, and a single architecture are concrete. Benchmarks are not disclosed, and the source is Reddit, so this stays in the 72–77 featured-threshold band.

X · @dotey

Microsoft VibeVoice-ASR tested on Mac for a one-hour podcast

Simon Willison ran 4-bit VibeVoice-ASR on an M5 Max MacBook Pro and transcribed a one-hour podcast in 8m45s. The 9B MIT-licensed model supports 60-minute audio, 50+ languages, and structured speaker output. Memory is the constraint: prefill peaked at 61.5GB, making 32GB laptops impractical.

Why it matters: HKR-H/K/R all pass: Simon Willison’s local test gives speed, parameter size, and memory peak that practitioners can act on. It is a single benchmark, not a fresh model launch, so it stays at the featured threshold.

r/LocalLLaMA

XiaomiMiMo MiMo-V2.5: Sparse MoE with 310B total and 15B activated parameters

XiaomiMiMo shared MiMo-V2.5 with 310B total parameters and 15B activated parameters. The post only links Hugging Face and says it runs on more “human” configs than its larger sibling. It does not disclose VRAM needs, quantization, or benchmarks.

Why it matters: HKR passes: the 310B/15B Sparse MoE hook is concrete and relevant to local deployment. Detail is thin: the post links Hugging Face but gives no VRAM, quantization, or benchmarks, so it stays near the featured threshold.

X · @dotey

AI terminal tool Warp open-sources client code with OpenAI as founding sponsor

Warp open-sourced its client code under AGPL; only the client is open, while server code stays closed. The Rust terminal has 700,000+ developers, and its Oz cloud AI handles coding, planning, and tests. The key signal is its AI-first contribution workflow.

Why it matters: HKR-H/K/R all pass: the OpenAI sponsorship hook, AGPL/client-only detail, and 700K-developer signal are concrete. This is a strong dev-tool open-source update, not a major model or capability release.

Apr 28Tuesday

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

QbitAI · WeChat

Open-source SenseNova-U1 unifies image understanding and generation

SenseTime open-sourced two SenseNova-U1 models: an 8B version and a 38B-total MoE version using NEO-unify. The architecture removes VE and VAE, processes pixels directly, and generates 2048×2048 images in about 9 seconds on one H100/H200 node. The key item is interleaved text-image reasoning; 32K context, long-text rendering, and beta interleaved creation remain limits.

Why it matters: HKR-H/K/R all pass: the architecture hook is concrete, the post gives model sizes and latency, and open multimodal work matters to builders. It stays in 78–84 because it is not a top-tier general-model launch.

Hacker News front page

Xiaomi releases MiMo-v2.5 weights with strong coding and agent benchmarks

Xiaomi released MiMo-v2.5 family weights; the title cites strong coding and agent benchmarks. The RSS body only lists URLs, 13 HN points and 2 comments; the post does not disclose size, license, or scores.

Why it matters: HKR-H/K/R pass because a Xiaomi coding/agent weights release is concrete and practitioner-relevant. Sparse sourcing holds it near the featured floor: no parameters, license, or benchmark numbers are disclosed.

X · @op7418

Xiaomi open-sources the MiMo-V2.5 model series

Xiaomi open-sourced the MiMo-V2.5 model series under the MIT license for commercial use, retraining, and fine-tuning. It also launched Orbit 100T Token, offering approved AI builders up to 1.6B credits worth 659 yuan. Agent framework teams can apply for free MiMo token access; the post does not disclose model size or benchmark results.

Why it matters: HKR-H/K/R all pass: Xiaomi MiMo-V2.5 open source, MIT terms, and Orbit 100T credits matter to builders. Missing params and benchmarks keep it in the 78–84 band, below P1.

r/LocalLLaMA

Microsoft Presents TRELLIS.2: Open-Source 4B Image-to-3D Model

Microsoft’s title says TRELLIS.2 is an open-source 4B image-to-3D model. The title lists 1536³ PBR assets, native 3D VAEs, and 16× spatial compression; the Reddit body is blocked by 403 and discloses no license or benchmarks.

Why it matters: HKR-H/K/R pass: 4B, 1536³, 16× compression, and open source are concrete. Reddit 403 leaves no paper, license, benchmark, or official link, so the score sits at the featured floor.

Apr 27Monday

Hacker News front page

Show HN: Utilyze — an open-source GPU monitoring tool claiming higher accuracy than nvtop

Systalyze open-sourced Utilyze to measure real GPU compute efficiency in production, with negligible overhead claimed. The post says nvidia-smi and nvtop only check whether any kernel runs during the sampling window; an H100 has 132 SMs and 17,424 cores. The key issue is real throughput headroom, not binary utilization dashboards.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the post explains the sampling flaw, and GPU waste is a real practitioner nerve. Unknown vendor and single-tool scope keep it in the 72–77 band.

Hacker News front page

Show HN: OSS Agent Dirac topped TerminalBench on Gemini-3-flash-preview

Dirac-run released Dirac and says it topped TerminalBench using Gemini-3-flash-preview. The repo claims 50-80% lower API costs via Hash Anchored edits, parallel operations, and AST manipulation; the post does not disclose full scores.

Why it matters: HKR-H/K/R all pass: an OSS coding agent claims a TerminalBench lead with cost and mechanism details. Held to 78 because the post relies on repo claims and lacks full leaderboard scores or reproduction logs.

OpenAI News

An Open-Source Spec for Orchestration: Symphony

OpenAI released Symphony, an open-source spec for Codex orchestration. The RSS snippet says it turns issue trackers into always-on agent systems; the post does not disclose spec details, license, APIs, or benchmarks.

Why it matters: HKR-H and HKR-R pass: an OpenAI open-source Codex orchestration spec is relevant to agent workflows. HKR-K is weak because license, interfaces, and reproducible mechanics are not disclosed.

Apr 25Saturday

Hacker News front page

Open-source memory layer Stash lets any AI agent do what Claude.ai and ChatGPT memory can do

Stash released an open-source persistent memory layer for AI agents, exposing 28 MCP tools and a 6-stage pipeline for long-term memory. The page says it uses PostgreSQL plus pgvector and hierarchical namespaces to separate user, project, and self memory. The real point is a portable memory layer, not the headline claim about matching ChatGPT or Claude.ai.

Why it matters: HKR-H/K/R all pass: the hook is portable long-term memory for any agent, and the page gives concrete architecture details. The score stays in the low featured band because this is an indie OSS infrastructure launch, not a major lab or platform release.

Apr 23Thursday

QbitAI · WeChat

Qwen3.6-27B open-weights, beats its 397B flagship predecessor on agentic coding

Qwen released Qwen3.6-27B and says it beats Qwen3.5-397B on 4 agentic coding benchmarks with about 1/15 the parameters. The post cites SkillsBench rising from 30.0 to 48.2, GPQA Diamond at 87.8, and AIME26 at 94.1; it uses a dense architecture, Thinking Preservation, and Gated DeltaNet, with weights on Hugging Face and ModelScope.

Why it matters: This is a substantive Qwen open-source model release with concrete agent-coding and reasoning scores, so HKR-H/K/R all pass. I keep it at 84, not higher, because the post gives strong benchmarks but no pricing, context window, or independent reproduction yet.

Xinzhiyuan · WeChat

Zhejiang University open-sources multi-agent evolution system OpenStory: Sun Wukong turns the Grand View Garden into an empty city

Zhejiang University open-sourced OpenStory, a multi-agent narrative system, and inserted a Sun Wukong agent into a 1:1 Dream of the Red Chamber sandbox; within minutes, agents fled the scene. The memory module broadcast “Sun Wukong killed innocents,” fear overrode daily logic, and Wang Xifeng’s physical removal cascaded into an empty Grand View Garden. What matters is the fragility of memory and consensus links; the post does not disclose the base models, metrics, or reproducible setup.

Why it matters: HKR-H/K/R all pass: the stress test is vivid, and the story includes a specific memory-broadcast failure mode with clear agent-safety relevance. Missing model details, metrics, and reproducible setup keep it in the good-featured band, not 85+.

Apr 22Wednesday

Hacker News front page

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen released the open-weight 27B dense model Qwen3.6-27B and made it available in Qwen Studio. It scores 77.2 on SWE-bench Verified vs. 76.2 for Qwen3.5-397B-A17B, and 59.3 on Terminal-Bench 2.0 under a 256K context and 3-hour timeout. The real takeaway is deployment: this is not a larger MoE, but a denser 27B model with stronger coding results.

Why it matters: Qwen3.6-27B is a substantive flagship-model release with open weights, concrete coding benchmarks, and a practical dense-deployment angle. HKR-H/K/R all pass, and per policy a major Chinese model launch should score on par with an equivalent US-lab release.

Apr 20Monday

X · @Yuchenj_UW

Kimi K2.6 is open-source

Kimi K2.6 is now open source, and the RSS snippet says it scored 58.6 on SWE-Bench Pro. The snippet also says it beat GPT-5.4 xhigh and Claude Opus 4.6 max effort. What matters is reproducibility; the post does not disclose weights, license, or eval setup.

Why it matters: All three HKR axes land: the open-source release is a strong hook, the SWE-Bench Pro 58.6 claim is testable, and the open-vs-closed coding race resonates. I keep it at 81 because the post appears title-level only; weights, license, and eval conditions are not disclosed.

r/LocalLLaMA

Training LoRA adapters for Apple's on-device 3B model on a free Colab T4 and a Mac

The author built a QLoRA pipeline for Apple’s on-device 3B model, cutting training needs from about 24GB to about 1GB RAM and 5GB GPU, enough for a free Colab T4 or a 24GB Mac. The post says A100 LoRA, T4 QLoRA, and Mac QLoRA adapters perform about the same, raising accuracy from about 40% to 75%, or 86% with retrieval; it also reports a confirmed Apple bug that writes a hidden ~160MB cache copy per CLI call, reaching 269GB over ~300 runs.

Why it matters: A named first-person experiment with reproducible memory and accuracy numbers clears HKR-H/K/R and beats routine tutorial posts. The score stays below the 85 band because this is a single Reddit post with limited source authority and a narrow benchmark scope.

r/LocalLLaMA

TRELLIS.2 image-to-3D now runs on Mac (Apple Silicon) with no NVIDIA GPU required

A developer ported Microsoft's TRELLIS.2 to Apple Silicon and reports generating ~400K-vertex meshes from one photo in about 3.5 minutes on an M4 Pro with 24GB. The port replaces five CUDA-only extensions with PyTorch MPS and custom backends; texture baking takes about 18 seconds, removing the NVIDIA and cloud requirement.

Why it matters: This is a community port, not an official release, but HKR-H/K/R all pass: the hook is NVIDIA-free image-to-3D on Apple Silicon, and the post includes testable details (M4 Pro 24GB, ~400k vertices, 3.5 minutes, 5 CUDA-extension rewrites). Reddit-level source authority keeps it in