Skip to content
r/LocalLLaMA

jinfer: an open-source inference engine that brings LLMs to the JVM, no Python needed

jinfer: An open-source AI inference engine for the JVM. Finally, AI in jar.

mukel90 released jinfer, a pure-Java inference engine covering chat, vision, audio transcription, embeddings, reranking, and TTS. It runs without Python, ONNX, or sidecar processes, and reads gguf/safetensors natively. CPU performance is claimed to be competitive with llama.cpp; GPU support via the jota backend is still in progress. It integrates with Spring AI and LangChain4j and supports GraalVM Native Image. The author previously built llama3.java and gemma4.java. This is an early release—I'd wait to see how GPU pans out before getting too excited.

Read the original ↗Export Markdown