Skip to content
r/LocalLLaMA

I turned an Android phone into a Vulkan-accelerated local LLM node

I turned an Android phone into a Vulkan-accelerated local LLM node (GGUF + LiteLLM + Tailscale)

Reddit user GsxrGuy80s configured a Z Fold 6 as a GGUF inference node using Vulkan, LiteLLM, and Tailscale; the post discloses gpu_layers=89, an OpenAI-compatible endpoint, and fallback routing to larger local nodes.

Why it matters: HKR-H/K/R all pass: a concrete phone-as-node hack with reproducible knobs. Source authority is limited to a Reddit post, so it fits the lower featured band rather than a broader industry update.

Read the original ↗Export Markdown