Skip to content
Trending storyPast story

A Jev-like wrapper for LLMs, including vision models

1 report1 sourceupdated 4 days ago

What happened

Summary

作者受 Jev 启发写了个 Python 封装,核心思路是让 LLM 只输出一个 token(比如 A/B/C),然后读它的 logprobs 来快速做选择题。这样避免了模型长篇大论,速度很快。作者还加了图片支持,用 Gemma 4 12B 在 RTX 3090 上跑摄像头画面,每秒能处理 1 帧,每帧同时问三个问题(有没有人、室内还是室外、亮度如何)...

Coverage

Follow the reports to see the story from different sides.

Sep 26
  1. Hacker News front page
    A Jev-like wrapper for LLMs, including vision models

    The author built a Python wrapper inspired by Jev that makes LLMs answer multiple-choice questions quickly by forcing single-token output and reading logprobs. It also handles images. On an RTX 3090 with Gemma 4 12B, it processes webcam frames at 1 FPS with three questions per frame (person visible, indoor/outdoor, brightness). OpenAI gpt-6-luna runs at 0.2 FPS due to connection overhead per request. The tradeoff: specialized CV models are faster, but LLMs let you change conditions by editing plain text. Code supports llama.cpp and OpenAI backends.