Skip to content
AI Chat-Group Daily (群聊日报)

Jev caught up by open source in a week, M5 Ultra local agent benchmarks land

2026-09-21 群聊日报

A community-built Jev Bench of several hundred questions shows Jev's confidence calibration fails on hard problems—average confidence differs by just 0.006 between correct and incorrect answers. DeepSeek V4.1 Flash hits 95% accuracy; the open-source reflex-27b reaches 76%, beating Jev's 74% with lower cost and no fine-tuning. The group's takeaway: the best way to train a classifier is to train a conversational LLM first. Jev skips chain-of-thought calibration and gives up accuracy. The same day, M5 Ultra Mac Studio reviews dropped: 256GB unified memory hits 2,887 tok/s prefill on Flash-Next, and a reviewer ran an agent team 24/7 for 99 days at zero cost. Group members flagged that the 5090 comparison didn't use NVFP4 quantization—real-world prefill can reach 8,000 tok/s, making the listed 59 tok/s decode suspiciously low. Grok 4.7 launched with coding and legal bench gains, same pricing as 4.6. RTX 5090 prices in China hit ¥50,000; someone was fined ¥12,000 bringing two cards through Shenzhen customs. Kimi Code Desktop went live, with official confirmation of no auto git backup.

Read the original ↗Export Markdown