A Citation Can Be Right Without Being Grounded
Mechanistic evidence from Llama-3.1-8B-Instruct shows why a correct-looking RAG citation should not be treated as proof that the cited source actually drove the answer.
Mechanistic evidence from Llama-3.1-8B-Instruct shows why a correct-looking RAG citation should not be treated as proof that the cited source actually drove the answer.
HALO reframes hallucination control as a layered enterprise assurance problem, with evidence showing that verification signals must be assigned and combined according to workload rather than accumulated indiscriminately.
BAR-RAG suggests that RAG systems may improve more by training on evidence matched to generator competence than by adding another always-on production reranker.
Adaptive context compression suggests a tiered way to control prompt growth in persistent assistants while protecting the history most likely to matter later.
A training-free RAG reranker can suppress keyword-stuffed false positives, but its value depends on identifying the retrieval regime before deployment.
A CTI benchmark shows why retrieval architecture should be chosen by question type, failure profile, and operational cost rather than average answer quality alone.
SelfGraphRAG shows how an unlabeled knowledge graph can generate its own retriever-training data, shifting part of RAG quality control from query time to indexing.
HintMR shows that targeted guidance can unlock much more reasoning performance from compact models, offering a structured alternative to larger models or brute-force sampling.
A medical RAG benchmark shows why retrieval components should earn their latency and compute cost through measured marginal gains, not architectural sophistication.
UrduBench shows why Urdu model procurement should test workload fit, prompting behavior, and language stability instead of relying on parameter count or reasoning labels.