A developer replaced five CUDA-only dependencies and got Microsoft’s TRELLIS.2 running on an M4 Pro with 24GB, producing roughly 400K-vertex meshes from one image in about 3.5 minutes, with texture baking in about 18 seconds. My take is simple: the point is not “Mac can do it too.” The point is that a CUDA-locked 3D generation stack just got peeled apart, dependency by dependency, and that changes where the barrier really sits.
The interesting part here is architectural, not brand tribalism. TRELLIS.2 originally depended on five CUDA-only compiled extensions: flex_gemm, flash_attn, o_voxel, cumesh, and nvdiffrast. The port swaps those for pure PyTorch sparse 3D convolution, Python mesh extraction with spatial hashing, SDPA attention, and MPS-friendly sampling. That tells you something important about a lot of “GPU-required” open-source AI software: part of the lock-in is not raw compute need, it is the extension stack that assumed CUDA as the operating system for AI. Once someone does the ugly systems work, a surprising amount of capability becomes merely slower, not impossible.
I’ve felt for a while that local multimodal tooling was following the same pattern image generation followed in 2023 and small-code models followed in 2024. First, everything looked data-center-native. Then people started removing hard CUDA assumptions. Then the center of gravity shifted from “best benchmark” to “good enough on hardware people already own.” This port fits that pattern. You can compare it to how Stable Diffusion on Apple Silicon went from novelty to usable after the community stopped waiting for first-party perfection and just rewired kernels, schedulers, and memory paths. Same with llama.cpp on Metal: the strategic value was never top throughput. It was distribution.
That is why I think this matters more than the raw 3.5-minute runtime. For an artist, indie game developer, CAD hobbyist, or robotics tinkerer, 3.5 minutes offline with zero cloud cost is a different product category from seconds on an H100 behind an API. Those are not the same buying decision. One is “do I have budget and infra.” The other is “do I already own a MacBook.” Once image-to-3D crosses into the second category, experimentation rates jump even if throughput stays mediocre.
I do have some pushback, though. The post gives one hardware datapoint and one output datapoint, but not the parts that decide whether this is broadly useful. We do not get side-by-side quality comparisons against the original CUDA path. We do not get memory peak, batch behavior, failure rate across different images, or fidelity metrics for geometry and texture. “~400K vertices” sounds substantial, but vertex count is not quality. A messy mesh can be large and still unusable. The author also says “not as fast as H100,” which is fair, but the actual speed gap is not disclosed. Without that, it is hard to know whether the port is 3x slower, 10x slower, or worse on harder scenes.
There is another reason to be careful: ports like this often work first on the happy path and only later reveal the maintenance tax. Flash attention replacements, sparse ops rewritten in plain PyTorch, and Python-side mesh extraction are all plausible for a strong demo. Keeping them stable across upstream model updates, PyTorch releases, and MPS quirks is a different job. Apple’s Metal stack has improved a lot, but anyone who has shipped serious local inference on MPS knows the long tail: silent fallbacks, memory fragmentation, kernels that regress between versions, and performance cliffs on odd tensor shapes. I buy the accomplishment. I am less ready to declare the stack “portable” in the boring production sense.
There is also a bigger market read here. Microsoft publishing research code that is effectively CUDA-shaped has been normal. The open-source community doing the portability work has also been normal. Nvidia still captures the high-end training and fastest inference layer, but the community keeps eroding exclusivity at the application layer. That happened with diffusion. It happened with LLM inference. Now it is happening in image-to-3D. Apple benefits from that without needing to win the frontier model race, because every successful port turns Apple Silicon into a viable endpoint for creative AI workflows.
So I would not frame this as “TRELLIS.2 on Mac beats Nvidia.” It doesn’t. I’d frame it as the beginning of a more annoying reality for CUDA-centric projects: if the model is valuable enough, someone will eventually reimplement the missing pieces. And once that happens, the moat shifts upward. It stops being “can this run at all” and becomes “how much faster, cleaner, and easier is the Nvidia path.” That is still a strong moat. It is just a narrower one than the ecosystem used to assume.
If the repo holds up across more machines and more scenes, this kind of work pushes image-to-3D toward the same place local image generation already reached: not best-in-class speed, but wide enough accessibility to matter.