Skip to content
Hacker News front page

A Jev-like wrapper for LLMs, including vision models

A single function Jev-like wrapper for LLMs, including vision models

The author built a Python wrapper inspired by Jev that makes LLMs answer multiple-choice questions quickly by forcing single-token output and reading logprobs. It also handles images. On an RTX 3090 with Gemma 4 12B, it processes webcam frames at 1 FPS with three questions per frame (person visible, indoor/outdoor, brightness). OpenAI gpt-6-luna runs at 0.2 FPS due to connection overhead per request. The tradeoff: specialized CV models are faster, but LLMs let you change conditions by editing plain text. Code supports llama.cpp and OpenAI backends.

Read the original ↗Export Markdown