Skip to content
r/LocalLLaMA

Weird to get near-linear scaling by adding another GPU?

Weird to get near linear scaling by adding another GPU?

A Reddit user benchmarked qwen3.6-27b-autoround-int4 on 1x3090 versus 2x3090. Narrative decode rose from 53 TPS to 94 TPS, and code decode rose from 62 TPS to 120 TPS, under no NVLink, 8x/8x PCIe, P2P automatically enabled, tensor parallelism set to 2, and different KV-cache settings.

Why it matters: HKR-H/K/R all pass: the result is counterintuitive, includes concrete TPS and TP conditions, and speaks to local-inference cost. Single Reddit test lacks multi-model replication and full setup details, so it stays near the featured threshold.

Read the original ↗Export Markdown