Skip to content
r/LocalLLaMA

We tried vectors, ASTs, and brute-force context stuffing for code retrieval; LLM semantic graphs worked best

We tried vectors, ASTs, and brute-force context stuffing for code retrieval. Graphs with LLM-generated semantics worked best. Here's what we learned.

ByteBell open-sourced a code indexing system that stores per-file LLM-generated purpose, summary, business context, entities, classes, functions, keywords, and imports in a Neo4j graph, then uses full-text search instead of vector similarity, with SHA-256 diffing to reindex only changed files and keep LLM calls proportional to churn.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, and the post gives a concrete Neo4j semantic-graph mechanism with SHA-256 incremental rebuilds. Reddit sourcing and missing metrics keep it at the 72–77 featured threshold.

Read the original ↗Export Markdown