Computer Build Using Intel Optane Persistent Memory Runs a 1T-Parameter Model at Over 4 Tokens/s
Computer build using Intel Optane Persistent Memory - Can run 1 trillion parameter model at over 4 tokens/sec
Reddit user APFrisco ran the 1T-parameter Kimi K2.5 Q2_K_XL quant locally at about 4 tokens/s using 768GB Intel Optane PMem, 192GB DDR4 ECC DRAM, and a 12GB RTX 3060 with llama.cpp hybrid GPU/CPU inference.
Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, the post gives concrete hardware and speed numbers, and it hits local-inference cost concerns. Single Reddit anecdote and limited replication detail keep it at the featured floor.