Skip to content
r/LocalLLaMA

Update on a 12×32GB SXM V100 Cluster for Local Legal Drafting

Update on 12x32gb sxm v100 cluster / local AI for legal drafting

A lawyer runs a local legal-drafting pipeline across 16 GPUs, with Qwen3.5-122B-A10B reaching about 50 tok/s on four V100s, while a verifier blocks ungrounded citations, dates, and Bates numbers before any final document is used.

Why it matters: HKR-H/K/R all pass: this is a first-person local-LLM experiment with concrete numbers, not a vendor post. Reddit source limits authority, so it stays at the low featured band rather than p1.

Read the original ↗Export Markdown