Skip to content
Trending storyDeveloping

95+ TPS through 100K tokens on Qwen 27B with a single 3090

1 report1 sourceupdated 21 hours ago

What happened

Summary

Reddit用户发帖说用一张RTX 3090跑Qwen 27B模型,生成了10万个token,吞吐量超过95 token/秒,上下文窗口开到262K。这个速度意味着本地跑长文本生成已经接近可用了——10万字级别的输出不用等太久。但帖子正文被Reddit屏蔽了,看不到具体用了什么量化、什么推理框架,也没法确认这是真实跑分还是某种优化后的极限测试。如果是真...

Coverage

Follow the reports to see the story from different sides.

Sep 29
  1. r/LocalLLaMA
    95+ TPS through 100K tokens on Qwen 27B with a single 3090

    A Reddit user reports running Qwen3.8 27B on a single RTX 3090, achieving 95+ tokens/sec throughput through 100K generated tokens with a 262K context window. This suggests local long-text generation is nearing practical speeds. However, the post body is blocked by Reddit, so the implementation details—quantization, inference framework, or whether this is a real benchmark—are not disclosed.

Heat over time

Not enough continuous observations to draw a trend yet.