Skip to content
r/LocalLLaMA

95+ TPS through 100K tokens on Qwen 27B with a single 3090

95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

A Reddit user reports running Qwen3.8 27B on a single RTX 3090, achieving 95+ tokens/sec throughput through 100K generated tokens with a 262K context window. This suggests local long-text generation is nearing practical speeds. However, the post body is blocked by Reddit, so the implementation details—quantization, inference framework, or whether this is a real benchmark—are not disclosed.

Read the original ↗Export Markdown