Skip to content

#RAG

0 today

May 21Thursday

AI HOT (Curated Pool)

Context compression improves search efficiency and accuracy

Perplexity has deployed query-aware compression in production, reducing context tokens by up to 70% while improving search answer quality.

Why it matters: HKR-H/K/R all pass: a counterintuitive production search update with a 70% token-cut claim and a direct cost-latency-quality hook. Single-source X post lacks benchmarks and reproducible setup, so it stays in the lower good-quality band.

May 20Wednesday

TechCrunch · AI

You Can Now Talk to Your Gmail Inbox, as Seen at Google I/O 2026

Google expanded Gmail’s AI Inbox with conversational voice search, letting users ask Gemini to find details buried in email. The RSS snippet does not disclose rollout scope, supported languages, pricing, latency, or the retrieval mechanism behind Gmail search.

Why it matters: HKR-H/K pass: a Google-scale Gmail voice inbox feature is concrete and clickable. HKR-R is weak because rollout, language support, pricing, and retrieval mechanics are not disclosed.

May 19Tuesday

Xinzhiyuan · WeChat

CUHK and Zhejiang University Question Whether AI Agent Memory Is Just a Memo

CUHK and Zhejiang University researchers argue that mainstream Agent memory is retrieval-based memo storage, not true memory, citing an Ω(k²) case requirement for compositional tasks and a PoisonedRAG result where 5 adversarial texts reached a 90% attack success rate.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the summary gives Ω(k²) and 90% attack success, and the issue matters to agent-memory and RAG-security builders. Strong research signal, not a same-day model-release event.

May 17Sunday

Hacker News front page

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

MinishLab open-sourced Semble, a code-search tool for agents that combines Model2Vec embeddings, BM25, RRF fusion, and reranking; on a 63-repo benchmark, it used 98% fewer tokens than grep+read, reached 0.854 NDCG@10, and ran CPU queries in about 1.5 ms.

Why it matters: HKR-H/K/R all pass: the 98% token claim is clickworthy, the 63-repo benchmark adds substance, and coding-agent context cost is a real practitioner nerve. Impact is still toolchain-level, so it stays below must-write.

May 16Saturday

QbitAI · WeChat

A new AI for 5 million doctors in China: exclusive journal partnership focuses on evidence sources

Alibaba Health launched the medical AI product Qinglizi for China’s 5 million doctors, with access to ten years of content from 70 BMJ Group journals and an evidence workflow constrained by PICO, GRADE, and review from more than 300 clinical experts.

Why it matters: HKR-H/K/R all pass: Alibaba Health and BMJ add concrete evidence sources and review mechanisms to a medical AI product. It remains a vertical product/partnership update, not a foundation-model or platform release.

May 15Friday

r/LocalLLaMA

Evaluated a RAG Chatbot: The Most Expensive Model Was the Worst Performer

The author evaluated a customer-support RAG bot and raised the quality score from 6.62 to 7.88 while cutting per-session cost from $0.002420 to $0.000509, using retrieval logging, LLM-as-judge scoring, chunk deduplication, stricter grounding, and a five-model sweep.

Why it matters: HKR-H/K/R all pass: counterintuitive model ranking, concrete quality and cost deltas, and direct RAG production relevance. Reddit source authority keeps it near the featured floor despite the first-person experiment signal.

AI HOT (Curated Pool)

Databricks brings GPT-5.5 to enterprise agent workflows

Databricks made GPT-5.5 available through AI Unity Gateway for AgentBricks and Agent Supervisor API workflows; on OfficeQA Pro, it became the first model above 50% accuracy and reduced errors by 46% versus GPT-5.4.

Why it matters: HKR-H/K/R all pass: GPT-5.5 enters Databricks workflows with 50% OfficeQA Pro accuracy and 46% fewer errors than GPT-5.4. It stays below a full model-release score because the page is a sales-led OpenAI customer story using Databricks’ own benchmark.

AI HOT (Curated Pool)

Granite Embedding Multilingual R2: Open Multilingual Embedding Model with 32K Context

IBM Granite released Granite Embedding Multilingual R2 on Hugging Face under Apache 2.0, with fewer than 100 million parameters, a 32K-token context length, and top same-scale retrieval performance on MTEB according to the post.

Why it matters: HKR-H/K/R pass: the 32K-context, sub-100M multilingual embedding model gives RAG builders a concrete open-source option. Impact is narrower than a frontier-model release, so it sits at the featured threshold.

AI HOT (Curated Pool)

OpenEvidence reaches 65% of U.S. doctors, drawing attention to shadow AI use

OpenEvidence reaches 65% of U.S. doctors and recorded 27 million clinical uses in April; doctors registered on mobile with license numbers, while hospitals were initially unaware of the shadow AI adoption pattern.

Why it matters: HKR-H/K/R all pass: 65% doctor reach and 27M April uses are unusually strong adoption data, with the hospital-unaware angle adding shadow-AI tension. Kept below 85 because methodology, revenue, and liability controls are not disclosed.

r/LocalLLaMA

MOOSE-Star (ICML 2026): 7B Model and 108K-Paper Dataset for Scientific Hypothesis Discovery

MiroMind researchers released the MOOSE-Star collection with three 7B models and TOMATO-Star, a dataset of 108,717 NCBI papers. MS-IR-7B reaches 54.37% inspiration-retrieval accuracy, uses DeepSeek-R1-Distill-Qwen-7B as its base, runs at about 14GB fp16, and supports llama.cpp, vLLM, and SGLang.

Why it matters: HKR-H/K/R all pass via the local 7B research-agent hook and concrete dataset metrics. Single Reddit source and limited lab gravity keep it below the must-write band.

May 14Thursday

AI HOT (Curated Pool)

OpenSquilla Open-Source Project Uses Smart Routing and Local Retrieval to Cut LLM Costs

OpenSquilla combines local model routing, vector retrieval, incremental sending, and cache hits to reduce transmitted tokens by more than 90%, while routing simple tasks to cheaper models and complex tasks to stronger models without spending tokens on the routing decision.

Why it matters: HKR-H/K/R all pass, but the source appears to be a single X project post; repo traction, test setup, and limits are not disclosed. Score lands at the featured threshold for practical open-source cost tooling.

May 13Wednesday

AI HOT (Curated Pool)

Open-source psql_bm25s speeds up PostgreSQL retrieval for multi-agent systems by 23x

The team open-sourced psql_bm25s, a native PostgreSQL access method for exact BM25 retrieval, and says it runs about 23x faster than pg_search on standard benchmarks.

Why it matters: HKR-H/K/R pass via the 23x retrieval-speed hook, named Postgres access method, and RAG latency pressure. Single-source release details lack independent reproduction and production constraints, so it stays in the lower featured band.

TechCrunch · AI

The AI legal services industry is heating up — Anthropic is getting in on the action

Anthropic introduced tools for law firms that cover five clerical workflows: document search and review, case law resources, deposition preparation, document drafting, and related tasks; the RSS snippet does not disclose pricing, launch timing, or model details.

Why it matters: HKR-H/K/R pass: Anthropic is moving into a high-value legal workflow with five named use cases. No pricing, customer scale, or new model capability is disclosed, so this stays just above the featured threshold.

May 10Sunday

r/LocalLLaMA

We tried vectors, ASTs, and brute-force context stuffing for code retrieval; LLM semantic graphs worked best

ByteBell open-sourced a code indexing system that stores per-file LLM-generated purpose, summary, business context, entities, classes, functions, keywords, and imports in a Neo4j graph, then uses full-text search instead of vector similarity, with SHA-256 diffing to reindex only changed files and keep LLM calls proportional to churn.

Why it matters: HKR-H/K/R all pass: the hook is counterintuitive, and the post gives a concrete Neo4j semantic-graph mechanism with SHA-256 incremental rebuilds. Reddit sourcing and missing metrics keep it at the 72–77 featured threshold.

May 7Thursday

Computing Life · Share · Yage

Agent Filesystems: From Feeding Models Memory to Letting Models Browse Files

The article frames agent filesystems as a three-stage shift from raw context to memory systems to filesystem-as-context, covering design choices from Turso, Anthropic, Vercel, and Manus, and listing four overlooked blind spots.

Why it matters: HKR-H/K/R all pass, but this is design commentary rather than a product or research release. Named comparisons across Turso, Anthropic, Vercel, and Manus justify featured, not the 78+ band.

May 6Wednesday

The Verge · AI

Google’s AI Search Summaries Will Now Quote Reddit

Google updated AI Search to include firsthand views from Reddit, social media, and forums in summaries. The post says a “perspectives” preview links queries to related online discussions; it does not disclose rollout scope or timing. For search teams, the key issue is how AI summaries cite and rank UGC sources.

Why it matters: HKR-H is strong because Google AI summaries quoting Reddit alters the search surface. HKR-K has the perspectives mechanism, and HKR-R hits SEO/UGC traffic concerns; missing rollout scope keeps it in the 72–77 product-update band.

r/LocalLLaMA

An Open Benchmark for Testing RAG on Realistic Company-Internal Data

EnterpriseRAG-Bench released a 500k-document corpus for testing RAG on company-internal data. It simulates Redwood Inference across 9 sources and includes 500 questions over 10 retrieval failure modes. Baselines show BM25 beats vector search overall, while agentic/bash retrieval has the best completeness at higher cost and latency.

Why it matters: HKR-H/K/R all pass: the benchmark targets a real enterprise RAG pain point, with 500k docs and testable BM25-vs-vector results. Single Reddit-source benchmark release keeps it below same-day must-write.

May 5Tuesday

Hacker News front page

Show HN: Airbyte Agents – context for agents across multiple data sources

Airbyte launched Airbyte Agents, using Context Store to index operational data for agents. Its public benchmark reports up to 80% fewer tokens for Gong and 90% for Zendesk versus vendor MCPs. The key point is pre-indexed context, not another MCP wrapper.

Why it matters: HKR-H/K/R all pass: a concrete pre-indexing angle, reproducible claims, and agent data-access pain. Airbyte is not a frontier lab, so this stays at the lower featured band.

May 4Monday

r/LocalLLaMA

Gemma 4 E2B runs well on an 8GB Android phone, powering a private voice notes app

A Reddit user ran Gemma 4 E2B locally on an 8GB OnePlus CE 5 and built a private voice notes app. Whisper Small 244MB transcribes, Gemma 4 E2B 2.4GB splits and tags, and a 10-15s note takes 12-15s end to end. Search uses query expansion, FTS lanes, RRF, and optional Gemma top-K reranking with a 15s fallback.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person build, not an official Google release. Concrete hardware, latency, model size, and retrieval details place it near the top of the tutorial band.

May 3Sunday

r/LocalLLaMA

Upskill: skill registry your agent consults before it starts, with 10k+ indexed skills

Autoloops released Upskill, an open-source skill registry with 10k+ indexed skills for agents. Search combines Postgres full-text search, 1024-dim embeddings, and reranking by stars, installs, and feedback. LLM adversarial review blocked hundreds of skills at index time.

Why it matters: HKR-H/K/R pass: a useful open-source agent registry with concrete retrieval and safety mechanics. Source authority is low and adoption is unproven, so it stays in the 72–77 featured band.

Apr 27Monday

Hacker News front page

TurboQuant: A First-Principles Walkthrough

TurboQuant walkthrough explains compressing AI vectors to 2–4 bits per coordinate. It uses random rotation to map high-dimensional coordinates to a fixed distribution, then reuses one codebook with no scale overhead, training, or calibration.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the mechanisms are new, and the cost angle is relevant. It stays below 78 because this is a technical walkthrough, not a model or product release.

Apr 23Thursday

TechCrunch · AI

AI Overviews are coming to your work Gmail

Google is bringing AI Overviews to work Gmail to generate instant summaries across multiple emails. The RSS snippet confirms only cross-email summarization; the post does not disclose rollout timing, pricing tier, or model details. The key shift is aggregation beyond a single thread.

Why it matters: HKR-K and HKR-R pass: Google adds cross-email summaries to enterprise Gmail, a core work workflow. HKR-H is weaker, and rollout timing, plan scope, and model details are undisclosed, so this stays low-featured rather than P1.

Apr 22Wednesday

X · @dotey

Google splits Gemini Deep Research into Deep Research and Deep Research Max

Google split Gemini Deep Research into Deep Research and Deep Research Max, with public preview starting today in paid Gemini API tiers. Both run on Gemini 3.1 Pro; one targets speed and cost, while Max runs longer with more compute and repeated search and reasoning. The update adds MCP support for sources such as FactSet, S&P, and PitchBook, plus files, code execution, and File Search; the post does not disclose pricing.

Why it matters: This is a substantive Google product update: Deep Research enters paid Gemini API preview with a standard/Max split for cost-speed vs longer-running compute. HKR-H/K/R all pass, but pricing, rate limits, and performance deltas are not disclosed, so it stays in the 78-84 band.

Apr 18Saturday

QbitAI · WeChat

RAG retrieves the right docs but still answers wrong? Saarland University team diagnoses why | ACL 2026

A Saarland University-led team introduced Disco-RAG, adding a 3-step “reading” layer between retrieval and generation, and says the paper was accepted as an ACL 2026 main-conference long paper. The post says it uses RST-based argument trees, cross-passage relation graphs, and outline generation with zero training; it reports gains on Loong, ASQA, and SciNews, but does not fully disclose the exact scores. The key claim is that many RAG failures come from reading and discourse understanding, not retrieval recall.

Why it matters: This is a solid research release with HKR-H, HKR-K, and HKR-R: a strong practical hook, a concrete mechanism, and a pain point RAG builders know well. I keep it at 80, not higher, because the post does not fully disclose benchmark numbers and external replication is still missing

Apr 17Friday

TechCrunch · AI

Google now lets you explore the web side-by-side with AI Mode

Google said on April 16 that clicking a link in AI Mode on Chrome desktop now opens the web page side-by-side with AI Mode. The feature keeps search context and uses page context plus web information for follow-up answers; the post does not disclose rollout scope, timing details, or regional limits. The practical shift is that Google is merging search chat and site browsing into one workflow.

Why it matters: This is a mid-weight Google search workflow update with HKR-H/K/R all present, but it is still a single-feature change. The story gives the context-retention and page-plus-web follow-up mechanism; rollout scope, regions, and timing are not disclosed, so it lands at the low end of

Apr 7Tuesday

X · @dotey

Milla Jovovich and Ben Sigman release open-source AI memory system MemPalace, claim perfect LongMemEval score

Milla Jovovich and Ben Sigman released the open-source memory system MemPalace and claimed a perfect LongMemEval score. The project runs fully local with no cloud or API key, says AAAK compresses context 30x, and uses 19 MCP tools for retrieval. The key issue is evaluation: Penfield Labs says the “perfect” result measured retrieval only, not end-to-end QA, and AAAK dropped retrieval accuracy from 96.6% to 84.2%.

Why it matters: HKR-H lands on the celebrity/open-source hook and the 'perfect score' dispute. HKR-K/R land on concrete metrics and the familiar nerve of eval gaming vs real memory utility; source authority is still just an X post, so this stays featured, not higher.

Apr 4Saturday

X · @dotey

Mintlify uses ChromaFs to make AI document retrieval look like a file system

Mintlify routes its AI doc assistant’s grep, cat, and ls calls through ChromaFs into database queries, cutting session startup from 46s to 100ms and pushing marginal compute cost per chat near zero. Built on Vercel Labs’ just-bash, it maps pages to files and sections to directories; at 850,000 chats per month, replacing real sandboxes saves over $70,000 a year in compute. The real shift is retrieval design: not faster vector RAG, but model-led exploration of structured docs, and the post says this may not fit messy knowledge bases.

Why it matters: This is a substantive engineering write-up, not a routine product note. HKR-H/K/R all pass: the fake-filesystem angle is novel, the post includes hard numbers (46s→100ms, 850k chats/month, >$70k/yr), and it hits operator concerns around latency, cost, and retrieval design; strong

Apr 3Friday

X · @claudeai

Microsoft 365 connectors are now available on every Claude plan

Anthropic made Microsoft 365 connectors available on every Claude plan, covering Outlook, OneDrive, and SharePoint. The post confirms plan coverage and supported apps; it does not disclose pricing, permission boundaries, regional limits, or admin requirements. The real signal is broad rollout across all plans, not a new standalone connector.

Why it matters: This is a mid-weight Claude product update: Anthropic expanded Microsoft 365 connectors to every Claude plan, which changes real Outlook, OneDrive, and SharePoint access. HKR-H/K/R all pass, but missing price, permission, region, and admin details keeps it at low-end featured.

X · @op7418

Karpathy shared how he builds a local AI knowledge base

Karpathy uses Obsidian and local Markdown to build a personal wiki, stores source material in a RAW folder, then has an LLM generate summaries, indexes, concept pages, links, and visualizations. The setup can answer questions over the wiki and write reports or new files, but the post also says AI-generated content can pollute the corpus and should be separated from trusted sources; the post does not disclose the model, scale, or automation details.

Why it matters: HKR-H and HKR-R land because Karpathy’s local-first wiki workflow is inherently clickable and discussable for AI practitioners. HKR-K lands on the RAW→LLM→summary/index/link mechanism, but missing model, corpus size, and automation details keep it in the mid-70s.

Feb 27Friday

36Kr (direct RSS)

From short video to long-form: Douyin is also handing news to AI

Douyin launched long-form posts in late 2025, raising the cap from 4,000 to 8,000 Chinese characters, and added “AI-selected news” summaries in its Hot topics tab. Long-form publishing is web-only for now, and the post says AI news will enter the main feed, but it does not disclose ranking weight, licensing scope, or fact-checking rules. The real issue is distribution and accountability: AI summaries and original articles will compete in the same traffic pool.

Why it matters: This clears HKR-H/K/R: Douyin putting AI summaries into its hot-news surface is a strong hook, and the piece includes concrete mechanics like the 8,000-word cap, web-only publishing, and follow-up queries. The real industry angle is distribution, copyright, and fact-checking, but

Feb 9Monday

36Kr (direct RSS)

Voice Ask is live: why is Xiaohongshu pushing search-by-question?

Xiaohongshu fully launched Voice Ask on Jan. 27, letting users long-press to speak on the search page and get structured answers distilled from in-app user experience posts. The post says it can handle 3-minute spoken queries, foreign languages, and dialects, but does not disclose the model, ASR stack, latency, or accuracy. The real shift is from 3-4 character keyword search to longer spoken questions, widening search intent capture and scenario coverage.

Jan 19Monday

Import AI (Jack Clark)

Import AI 441: My agents are working. Are yours?

Jack Clark says his research agents processed thousands of papers while he hiked or slept, and Claude finished site scraping, embeddings, local vector search, and a GUI in under one hour. The post confirms multi-agent retrieval, cross-checking, and report generation; it does not disclose model versions, cost, failure rate, or benchmark data. The point to watch is workflow friction dropping enough for AI to shift from single prompts to ongoing delegated work.

Why it matters: HKR-H lands with the challenge in the headline; HKR-K lands because Clark describes a <1 hour workflow with retrieval, cross-checking, and report generation. Missing model version, cost, failure rate, and evaluation keep it in featured, not p1.

Jan 6Tuesday

NVIDIA Blog

NVIDIA DGX Spark and DGX Station power the latest open-source and frontier models from the desktop

NVIDIA showed at CES that DGX Spark and DGX Station can run 100B to 1T-parameter models locally on deskside systems. The post cites a 35% average llama.cpp speedup, up to 70% NVFP4 compression, 775GB coherent memory on DGX Station, and a 250,000 token/sec pretraining demo. The real signal is the local dev loop: fine-tuning, inference, RAG, coding assistants, and robotics demos all target replacing some cloud iteration with deskside compute.

Why it matters: HKR-H/K/R all pass: the story pairs a strong desktop-scale hook with concrete specs and demo numbers, and it speaks directly to the local-vs-cloud workflow debate. Still, this is an NVIDIA product post and most performance evidence comes from vendor-run demos, so it stays at 75,.

May 28, 2025Wednesday

Mistral AI

Codestral Embed

Mistral AI 发布首个代码专用嵌入模型 Codestral Embed,官方称其在真实代码数据检索上显著优于 Voyage Code 3、Cohere Embed v4.0 和 OpenAI 的大型嵌入模型。

May 27, 2025Tuesday

Mistral AI

Mistral releases Agents API with built-in connectors and MCP tools

Mistral released an Agents API that pairs its language models with built-in connectors for code execution, web search, image generation and MCP tools. It also offers persistent memory across conversations and agent orchestration.

Why it matters: Mistral details the connectors, memory and orchestration of its Agents API, letting readers judge how an agent platform would be deployed.

Apr 9, 2025Wednesday

Mistral AI

Evaluating RAG with LLM as a Judge

Mistral 介绍用 LLM as a Judge 评估 RAG 系统,由 judge LLM 按数值、二元或定性量表为 generator LLM 的回答打分,再对评测数据集求加权总分。

Mar 11, 2025Tuesday

OpenAI News

New tools for building agents

OpenAI released the Responses API, three built-in tools, and an Agents SDK on March 11, 2025 for single-agent and multi-agent workflows. The post confirms web search, file search, and computer use, says the API is available to all developers today, and says billing stays at standard token and tool rates. The key platform signal is migration: OpenAI plans an Assistants API sunset in mid-2026 after full feature parity with Responses API.

Why it matters: This is a substantive OpenAI developer-platform launch, not a routine feature add. HKR-H/K/R all pass: new entry point, concrete tools and pricing, plus a sunset timeline that will affect agent frameworks and API choices immediately.

Mar 7, 2025Friday

Mistral AI

Mistral AI releases document-understanding OCR API Mistral OCR

Mistral AI released Mistral OCR, an optical character recognition API that takes images and PDFs and outputs interleaved text and images in order. It handles complex layouts such as tables, formulas and LaTeX.

Why it matters: Mistral gives benchmark comparisons, multilingual performance and pricing for Mistral OCR, showing how usable document parsing is in a RAG pipeline.

Feb 10, 2025Monday

OpenAI News

OpenAI partners with Schibsted Media Group

OpenAI partnered with Schibsted Media Group to bring content from titles including VG, Aftenposten, Aftonbladet, and Svenska Dagbladet into ChatGPT for news summaries across its 300 million users. OpenAI says responses will include clear attribution to Schibsted brands for verification; the post does not disclose term length, licensing scope, or revenue sharing. The key signal is that licensed news is moving into ChatGPT’s main answer flow, not just referral traffic.

Why it matters: Primary-source OpenAI partnership with a concrete product effect: Schibsted titles will feed attributed news summaries in ChatGPT for 300m users. HKR-K and HKR-R pass because it expands licensed news inside ChatGPT's answer flow; HKR-H is weak since terms, scope, and economics go

Jan 15, 2025Wednesday

OpenAI News

Partnering with Axios expands OpenAI’s work with the news industry

OpenAI announced a content partnership with Axios and funding to expand Axios Local into 4 U.S. cities. OpenAI says it now works with nearly 20 media organizations, covering 160+ outlets, hundreds of brands, and 20+ languages. ChatGPT Search shows select summaries, excerpts, citations, and source links from partners; the post does not disclose deal value or Axios-specific technical terms.

Why it matters: This passes HKR-K and HKR-R: OpenAI gives concrete scope numbers and a specific Search distribution mechanism. It stays near the featured floor because the post is still partnership PR, and the grant size plus technical terms are not disclosed.