Quad 20GB 3080s beat quad 5060 Tis for Qwen3.6-27B code generation on Vast AI
I benched quad 20GB 3080s on Vast AI for code generation with Qwen3.6-27B so you don't have to (it's even better than quad 5060Tis)
Someone rented four 20GB RTX 3080s on Vast AI and ran Qwen3.6-27B for code generation. With MTP on, decode hit 69 t/s near 256K context; prefill dropped to 893 t/s. The author priced used cards at ~$400 each and an X99 board+CPU+64GB RAM combo at ~$275, totaling just over $2K for a high-accuracy, lightly quantized dense-model rig. It beat a quad 5060 Ti setup on speed and cost. The post doesn't disclose specific code benchmarks or accuracy scores, so I'd discount the speed-only claim a bit.
Why it matters: First-person benchmark with concrete numbers and a cost comparison that's directly useful for the local LLM crowd. Held back by Reddit sourcing, non-rigorous test conditions (power-limited, no prompt caching), and the fact that it's hardware selection advice rather than a mode...