{
  "slug": "domain-specific-rag-and-cross-corpus-question-answering",
  "name": "Domain-specific RAG & cross-corpus question answering",
  "tier": "leading-edge",
  "trend": "steady",
  "blockerType": null,
  "tools": [
    {
      "name": "Azure AI Search",
      "url": "https://learn.microsoft.com/en-us/azure/search/retrieval-augmented-generation-overview"
    },
    {
      "name": "Databricks AI Search",
      "url": "https://docs.databricks.com/en/ai-search/"
    },
    {
      "name": "Qdrant",
      "url": "https://qdrant.tech/"
    }
  ],
  "evidence": [
    {
      "title": "Knowledge-Graph-Based Augmentation versus Retrieval-Augmented Generation on Culturally Grounded Question Answering",
      "url": "https://arxiv.org/abs/2609.18317v1",
      "date": "2026-09-16",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Benchmark research comparing Graph-RAG against standard RAG on culturally specific QA (LatamQA dataset), showing 72-78% error reduction versus base LLM and zero-shot multilingual transfer to Portuguese."
    },
    {
      "title": "Accountable NLP for Evidence-Grounded Decision Briefings: Critical Review and Evaluation Framework",
      "url": "https://www.techscience.com/cmc/v89n2/68847/html",
      "date": "2026-09-15",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed synthesis of 104 sources documenting RAG grounding failures (citation-to-claim entailment gaps, suppressed uncertainty, causal language creep) and proposing accountability evaluation framework for evidence-grounded systems."
    },
    {
      "title": "CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual Retrieval-Augmented Generation",
      "url": "https://lanfrica.com/en/record/coral-adaptive-retrieval-loop-for-culturally-aligned-multilingual-rag",
      "date": "2026-09-11",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed ACL 2026 Findings paper on adaptive multilingual RAG with evidence-quality feedback and corpus reselection, achieving up to 3.58 percentage point accuracy improvement on low-resource languages."
    },
    {
      "title": "Health Education Conversational Agent for Gastric Cancer Patients: RAG Knowledge Hit Rate Progression in Iterative Alpha Testing",
      "url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13561191/",
      "date": "2026-09-10",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed healthcare deployment across 44 patients and clinical staff showing RAG knowledge hit rates improving from 86% to 100% across test rounds and answer accuracy reaching 82% through iterative refinement."
    },
    {
      "title": "Progressive Evidence Acquisition for Regulatory Compliance: Production RAG Evolution at Ontario Power Generation",
      "url": "https://techpulse.ro/en/news/items/from-naive-rag-to-deep-agentic-retrieval-an-evolving-context-engineering-pipeline-for-regulatory-compliance-qxgbjs",
      "date": "2026-09-10",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Self-reported production deployment formalising RAG evolution through four architectural stages into the PEA-CAE pattern (Progressive Evidence Acquisition with Cost-Aware Escalation) for regulatory compliance QA over evolving corpora."
    },
    {
      "title": "Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion and Per-Chunk Grounding: Telecom Document Search in Production",
      "url": "https://techpulse.ro/en/news/items/hybrid-retrieval-augmented-generation-with-knowledge-graph-expansion-rrf-fusion-and-per-chunk-grounded-evaluation-for-en-7dq4a0",
      "date": "2026-09-10",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Self-reported production telecom network operations deployment achieving Precision@10 0.69, Recall@10 0.79, and 89.6% grounding rate through hybrid retrieval (BM25+dense+KG); +15-18 percentage point gains over vector-only baseline."
    },
    {
      "title": "VikingRAG: Accurate and Token-Efficient Retrieval-Augmented Generation over Structured Documents",
      "url": "https://arxiv.org/html/2609.11390v1",
      "date": "2026-09-10",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Research on directory-aware semantic storage for structured documents achieving state-of-the-art RAG accuracy while consuming only 11.6-51.9% of tokens required by competing systems, addressing efficiency on enterprise corpora."
    },
    {
      "title": "Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source Small Language Models",
      "url": "https://techpulse.ro/en/news/items/multimodal-hybrid-retrieval-augmented-generation-for-scientific-document-understanding-using-open-source-slms-qxghhc",
      "date": "2026-09-10",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Research applying Qwen2-VL-2B-Instruct multimodal extraction to generate textual summaries of scientific tables and figures, reporting 157% retrieval-quality improvement over naive RAG baseline with only 50ms additional latency."
    },
    {
      "title": "Enterprise RAG Pipeline Failure and Recovery: Benchmark-to-Production Accuracy Gap at Salesforce",
      "url": "https://engineering.salesforce.com/enterprise-ai-accuracy-building-a-more-trustworthy-rag-application/",
      "date": "2026-09-09",
      "type": "case-study",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Production case study showing RAG systems exceeding 90% on benchmarks dropping to ~46% on real enterprise documents; remediation via intelligent parsing and reranking raised performance to 84.4%, documenting persistent benchmark-production gaps."
    },
    {
      "title": "Guaranteeing Faithful Evidence Extraction in Speculative Retrieval-Augmented Generation via Constrained Decoding",
      "url": "https://arxiv.org/abs/2609.10046",
      "date": "2026-09-09",
      "type": "research-paper",
      "added": "2026-09-23",
      "superseded_by": null,
      "window": null,
      "explanation": "Research demonstrating exact-span extraction accuracy below 40% in technical domains with current semi-extractive RAG methods; proposes constrained hybrid decoding (CHyD) to enforce hard decoding constraints on quoted spans."
    },
    {
      "title": "Long-Context AI vs. RAG for Document Analysis Compared",
      "url": "https://intuitionlabs.ai/articles/long-context-ai-vs-retrieval-rag",
      "date": "2026-09-05",
      "type": "industry-report",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific comparison for evidence-sensitive industries (pharma, biotech, legal): traceability and citation discipline are domain requirements distinct from accuracy; FDA/biopharma deployments mandate human verification and source provenance."
    },
    {
      "title": "An agentic AI framework connecting language models to electronic health records and a biomedical knowledge graph for real-world evidence",
      "url": "https://www.frontiersin.org/articles/10.3389/frai.2026.1883853",
      "date": "2026-09-04",
      "type": "research-paper",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "UCSF production deployment: agentic RAG on 7M+ subject EHR with curated biomedical knowledge graph improved accuracy on 100 clinical tasks for GPT-5.5 and Claude Opus, accelerating research protocol timelines from months to hours."
    },
    {
      "title": "Fine-Tuning Embedding Models on Domain-Specific Data to Improve Retrieval Accuracy in RAG Pipelines",
      "url": "https://www.wickedsmartdata.com/articles/fine-tuning-embedding-models-on-domain-specific-data-to-improve-retrieval-accuracy-in-rag-pipelines",
      "date": "2026-09-04",
      "type": "tutorial",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Technical guide: fine-tuning embeddings on domain data improves retrieval precision 20-40% on specialized corpora; demonstrates generic embeddings fail on domain terminology (clinical, legal, financial) requiring specialized semantic spaces."
    },
    {
      "title": "A Novel Benchmark Dataset of Banks' financial statements",
      "url": "https://arxiv.org/html/2609.03654v1",
      "date": "2026-09-02",
      "type": "research-paper",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "FinRAG-QA benchmark: 999 practitioner-curated financial Q&A on 209 regulatory documents; domain-specific optimization (retrieval-adapted embeddings, reranking) achieves +38.8% NDCG and +34.4pp generation accuracy."
    },
    {
      "title": "Your RAG Cites the Wrong Policy Because Your Docs Disagree",
      "url": "https://usqrd.com/insights/rag-contradictory-outdated-corpus-conflict-detection",
      "date": "2026-08-28",
      "type": "case-study",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Production failure mode: corpus contradictions retrieved cleanly but outdated source ranked higher by semantic similarity, yielding confidently-cited wrong answers; solution requires source-authority governance, demonstrating domain-specific RAG success depends on content governance, not retrieval tuning."
    },
    {
      "title": "Why Context Graphs Are Replacing Traditional RAG in Enterprise AI",
      "url": "https://correctcontext.com/why-context-graphs-are-replacing-traditional-rag-in-enterprise-ai/",
      "date": "2026-08-28",
      "type": "opinion",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Industry analysis of RAG maturity ceiling: 72% of enterprise implementations fail first year; Gartner forecasts 40% agentic AI projects canceled by 2027 due to cost, unclear ROI, governance gaps; signals architectural evolution toward context-layer governance and knowledge-graphs."
    },
    {
      "title": "Towards trustworthy LLMs for project management: Benchmarking a vector hybrid and graph RAG architectures",
      "url": "https://www.itcon.org/paper/2026/42",
      "date": "2026-08-27",
      "type": "research-paper",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed benchmark comparing vector-hybrid and graph RAG on government project management data; graph RAG superior for trustworthiness (faithfulness +3%, context recall +4%, answer relevancy +5%)."
    },
    {
      "title": "Why RAG hits a wall on structured data",
      "url": "https://www.getgrist.com/blog/why-rag-hits-a-wall-on-structured-data/",
      "date": "2026-08-27",
      "type": "case-study",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Epoch8 production deployment: hybrid structured-query + vector retrieval for structured domains achieves 90%+ accuracy vs 40-60% for vector-only RAG; demonstrates architectural brittleness, not fixable via retrieval tuning."
    },
    {
      "title": "6个步骤将RAG幻觉率降至1%",
      "url": "https://www.hubwiz.com/blog/6-steps-to-reduce-rag-hallucination-rate-to-1-percent/",
      "date": "2026-08-27",
      "type": "case-study",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Production financial RAG: systematic 6-stage debugging (context reordering, lost-in-middle fix, claim verification, deterministic math) reduced hallucination from 15% to 1%, demonstrating pipeline discipline achieves reliability for regulated domains."
    },
    {
      "title": "Experiential Knowledge Powered by AI",
      "url": "https://industry-science.com/en/articles/experiential-knowledge-powered-ai/",
      "date": "2026-08-27",
      "type": "research-paper",
      "added": "2026-09-09",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed case studies: four corporate domain-specific RAG deployments (assembly, maintenance, commissioning) identify critical success factors (knowledge base strategy, metadata enrichment, employee participation) and persistent challenges (hallucination, data inconsistency)."
    },
    {
      "title": "Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA",
      "url": "https://www.opentrain.ai/papers/same-agent-different-answers-a-repeat-aware-audit-of-corpus-induced-answer-churn--arxiv-2608.22856/",
      "date": "2026-08-24",
      "type": "research-paper",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Preregistered empirical study documenting production RAG fragility: corpus expansion causes silent answer churn (6.44pp exact-match excess, 10.25pp semantic excess on NQ) despite fixed model/prompt/retrieval. Compatibility failures can mask accuracy metrics—critical signal for deployment discipline."
    },
    {
      "title": "Domain-tailored RAG framework improves industrial LLM question answering in new engineering study",
      "url": "https://www.eurekalert.org/news-releases/1141055",
      "date": "2026-08-21",
      "type": "research-paper",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific RAG system for industrial QA (manufacturing, maintenance, fault diagnosis) with multi-module pipeline: domain classifier, GTE-DPR dense retrieval, BGE reranker; demonstrates domain-adapted components essential for specialized technical corpora."
    },
    {
      "title": "Introducing Agentic Search - Mistral",
      "url": "https://mistral.ai/news/agentic-search/",
      "date": "2026-08-20",
      "type": "product-ga",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Mistral Agentic Search GA replaces one-shot RAG with multi-step retrieval loop (search, open, navigate, read, grep); FinanceBench: +59.3pp (26.7%→86%), OfficeQA: +45.6pp (6.3%→51.9%); 39.6% latency reduction, up to 1/3 token savings."
    },
    {
      "title": "Improving RAG Pipeline with BM25 Retrieval and Hybrid Ranking",
      "url": "https://www.linkedin.com/posts/akshatsaxena1701_github-saxena1701building-rag-from-scratch-activity-7495636615378423808-lli5",
      "date": "2026-08-19",
      "type": "case-study",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "SmartReco domain-specific RAG for personalized recommendations: hybrid retrieval (BM25+dense), cross-encoder reranking, query rewriting; 96.3% Precision@3, 0.9739 NDCG@3 (21.16% vs semantic-only); demonstrates retrieval stratification effectiveness at production scale."
    },
    {
      "title": "Designing a Persistent Knowledge Layer That Refuses to Guess",
      "url": "https://daily.dev/posts/designing-a-persistent-knowledge-layer-that-refuses-to-guess-0ecjc6hct",
      "date": "2026-08-16",
      "type": "opinion",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific RAG architecture for regulated domain (insurance) addressing six failure modes: scoped supersession, contradictions, terminology drift, temporal misalignment, rationale loss, multi-hop brittleness; Azure implementation with Cosmos DB knowledge layer and entity resolution patterns."
    },
    {
      "title": "Where Does Retrieval Fail? Evaluating RAG Architectures for Agricultural Advisory",
      "url": "https://arxiv.org/abs/2608.14886v1",
      "date": "2026-08-14",
      "type": "research-paper",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed evaluation of domain-specific RAG across five architectures on Bengali agricultural advisory (1,000 queries, 2,882 KB items); BM25 R@10=0.506, hybrid RRF R@10=0.539; dense retrieval collapses 10× by query type, establishing domain-specific strategy necessity."
    },
    {
      "title": "新卒2年目のAI活用事例 ~前編~ 保守運用現場にRAGを持ち込んだ話",
      "url": "https://zenn.dev/acntechjp/articles/f0715ee2fd9b3d",
      "date": "2026-08-14",
      "type": "case-study",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Accenture Japan IT operations RAG deployment: ticket triage + procedure synthesis via hybrid search (BM25+vector), LLM metadata attachment, Azure AI Search indexing; outcomes: faster incident response, new engineer autonomy, eliminated document search overhead."
    },
    {
      "title": "How Chemcopilot AI Agents Turn Scientific Literature into Active Lab Intelligence",
      "url": "https://www.chemcopilot.com/blog/ai-chem-agents-scientific-literature-lab-intelligence",
      "date": "2026-08-12",
      "type": "case-study",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific RAG deployment processing 1,000+ scientific documents (journals, patents, lab reports) for materials R&D via vector embeddings and semantic segmentation; zero-data-retention security with measurable time savings in literature review bottleneck."
    },
    {
      "title": "Insurance AI Shifts to Core Infrastructure with Claude Integration",
      "url": "https://www.linkedin.com/posts/imhiren_insuranceai_insurtech_dataengineering_activity-7493212058637824001--CcJ",
      "date": "2026-08-12",
      "type": "case-study",
      "added": "2026-08-26",
      "superseded_by": null,
      "window": null,
      "explanation": "Named domain-specific agentic RAG deployments (AIG: SLA reduced 5x+, accuracy 75%→90%; Allianz: autonomous workflows on motor/health claims; Travelers: 10K engineers; Verisk: domain-specific data access); signals production maturity in regulated insurance sector."
    },
    {
      "title": "August 2026 - Azure Databricks (ai_search function)",
      "url": "https://learn.microsoft.com/en-us/azure/databricks/release-notes/product/2026/august",
      "date": "2026-08-10",
      "type": "product-ga",
      "added": "2026-08-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Azure Databricks ai_search() SQL function in Beta enables batch domain-specific RAG pipelines with multi-source retrieval, deduplication, reranking, and grounded answer synthesis—signals data-platform convergence of RAG infrastructure at enterprise scale."
    },
    {
      "title": "7 RAG Failure Modes Crippling Enterprise Deployments in 2026",
      "url": "https://ragaboutit.com/7-rag-failure-modes-crippling-enterprise-deployments-in-2026/",
      "date": "2026-08-06",
      "type": "opinion",
      "added": "2026-08-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of production RAG failures: 73% of 143 enterprise deployments experienced critical failure within first quarter; 41% undetected by standard evals; documents cross-document entity resolution, temporal drift, and multi-hop gaps as blocking adoption in domain-specific systems."
    },
    {
      "title": "SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG",
      "url": "https://arxiv.org/abs/2608.03860",
      "date": "2026-08-04",
      "type": "research-paper",
      "added": "2026-08-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed study of domain-specific RAG on CORD-19 scientific corpus; hybrid retrieval (BM25 + BGE-M3) achieves Recall@10 1.0, domain-mismatched rerankers reduce precision—direct evidence that domain-adapted components are essential in specialized corpora."
    },
    {
      "title": "Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG",
      "url": "https://arxiv.org/abs/2608.02011",
      "date": "2026-08-03",
      "type": "research-paper",
      "added": "2026-08-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical study of 12,000 multi-hop QA trajectories identifies procedural failure in agentic RAG: agents retrieve evidence but skip reading before answering; Read-Gate invariant improves accuracy 14.9–19.9 points, decomposes pre- vs post-evidence failure modes."
    },
    {
      "title": "A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery",
      "url": "https://arxiv.org/abs/2607.27748",
      "date": "2026-07-30",
      "type": "case-study",
      "added": "2026-08-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Xiaohongshu production deployment: domain-specific RAG for 5,300+ table data warehouse; knowledge-graph-guided retrieval improved Hit@10 from 19.1% to 96.6%, reduces tokens 71.6×, demonstrates enterprise retrieval solving real asset discovery failures."
    },
    {
      "title": "BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms",
      "url": "https://arxiv.org/abs/2607.26497",
      "date": "2026-07-30",
      "type": "research-paper",
      "added": "2026-08-12",
      "superseded_by": null,
      "window": null,
      "explanation": "Controlled study spanning 1.7M to 600M tokens (28 tiers) comparing BM25, dense, graph, and agentic RAG; BM25 overtakes agents around 10M tokens, maintains 20-point margin at 600M; establishes lexical retrieval as scalable default for enterprise domain-specific corpora."
    },
    {
      "title": "PAR2-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering",
      "url": "https://aclanthology.org/2026.acl-industry.118/",
      "date": "2026-07-25",
      "type": "research-paper",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "ACL 2026 Industry Track framework addressing multi-hop QA as practical enterprise bottleneck; achieves 23.5% accuracy improvement and 10.5% NDCG gain over baselines without retraining, applicable to domain-specific cross-corpus synthesis."
    },
    {
      "title": "Embedding Model Upgrades Silently Break Production RAG - Here Is How to Handle Them",
      "url": "https://chiraghasija.cc/posts/embedding-model-versioning-rag-production-2026/",
      "date": "2026-07-24",
      "type": "opinion",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner analysis of silent production failures from embedding model swaps; proposes dual-index migration pattern for safe upgrades without coordinate system mismatch; essential governance practice for domain-specific RAG systems under continuous evolution."
    },
    {
      "title": "5 Multi-Agent RAG Fixes That Cut Enterprise Failures 68%",
      "url": "https://ragaboutit.com/5-multi-agent-rag-fixes-that-cut-enterprise-failures-68/",
      "date": "2026-07-23",
      "type": "opinion",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "MIT/Microsoft research synthesizes CoRe-RAG framework reducing multi-hop query failures from 71% to 23% through specialized agent roles with closed-loop verification; demonstrates architectural solution to domain-specific cross-corpus reasoning bottleneck."
    },
    {
      "title": "How Cortex saves 300+ hours a month with Kapa across support, Slack, and docs",
      "url": "https://www.kapa.ai/customer-examples/cortex",
      "date": "2026-07-21",
      "type": "case-study",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Named company (Cortex) deployed domain-specific RAG on 60+ technical sources (docs, GitHub, APIs, tickets); answered 18,000+ queries with 86% certainty and 48,500+ citations, demonstrating production-scale technical domain deployment."
    },
    {
      "title": "LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes",
      "url": "https://www.microsoft.com/en-us/research/publication/lakequest-a-three-domain-benchmark-for-grounded-question-answering-across-data-lakes/",
      "date": "2026-07-21",
      "type": "research-paper",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Microsoft Research benchmark on 9,846 QA pairs across heterogeneous domain-specific data lakes identifies critical failure modes in cross-corpus reasoning; baseline RAG and agentic systems fail on relation chaining, policy grounding, and tabular QA."
    },
    {
      "title": "FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering",
      "url": "https://arxiv.org/abs/2607.18102v1",
      "date": "2026-07-20",
      "type": "research-paper",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Production multi-agent RAG framework for SEC filings with corpus-aligned retrieval and online A/B testing on 1,000 real users; demonstrates deployment-ready domain-specific multi-agent orchestration for financial document QA."
    },
    {
      "title": "Salience Induction against Multi-Hop RAG Agents: Threat and Defense",
      "url": "https://arxiv.org/abs/2607.17535v1",
      "date": "2026-07-20",
      "type": "research-paper",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "Security research identifies vulnerability in agentic multi-hop RAG; truth-preserving edits redirect reasoning toward wrong answers with 83.3% success rate; tested on 5 frontier models and 3 agent architectures, exposing hidden failure mode in production systems."
    },
    {
      "title": "92% of RAG Systems Fail Multi-hop Queries: 5 Fixes",
      "url": "https://ragaboutit.com/92-of-rag-systems-fail-multi-hop-queries-5-fixes/",
      "date": "2026-07-17",
      "type": "opinion",
      "added": "2026-07-29",
      "superseded_by": null,
      "window": null,
      "explanation": "UC Berkeley SkyLab and Google DeepMind MultiHopRAG benchmark evaluates 23 enterprise configurations on 10,000 multi-hop questions; only 8% achieve correct answers; 63% of failures from incomplete retrieval chains, 41% from hallucinated bridges."
    },
    {
      "title": "Enterprise RAG in production: governance & monitoring",
      "url": "https://domino.ai/blog/enterprise-rag-production",
      "date": "2026-07-09",
      "type": "opinion",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Analysis of enterprise RAG operational degradation: embedding drift, access control failure, hallucination from stale context; identifies silent failures with no error signals until user impact observed."
    },
    {
      "title": "RAG Is the New Legacy: Why 2026 Teams Are Shipping Agentic Search Instead of Chatbots",
      "url": "https://icmd.app/article/rag-is-the-new-legacy-why-2026-teams-are-shipping-agentic-search-instead-of-chat-1783213970500",
      "date": "2026-07-05",
      "type": "opinion",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical analysis identifying shift from vector-DB-first RAG toward agentic search with permission-aware retrieval and tool-driven orchestration; establishes new production pattern for multi-step domain-specific reasoning."
    },
    {
      "title": "What's new in Microsoft Foundry | Build Edition",
      "url": "https://devblogs.microsoft.com/foundry/whats-new-in-microsoft-foundry-build-2026/",
      "date": "2026-07-04",
      "type": "product-ga",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Microsoft Foundry IQ achieved GA as unified knowledge platform for multi-agent access; serverless agentic retrieval delivers 46–54% evidence recall improvement and 34% token cost reduction."
    },
    {
      "title": "RAG Development Services: How 8 Industries Are Using Retrieval-Augmented Generation in 2026",
      "url": "https://www.b2bnn.com/2026/07/rag-development-services-how-8-industries-are-using-retrieval-augmented-generation-in-2026/",
      "date": "2026-07-03",
      "type": "industry-report",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Production survey documenting 8-sector RAG deployments with domain-specific architectures: banking >95% citation precision, insurance 70%+ touchless claims processing, healthcare >95% faithfulness."
    },
    {
      "title": "Agentic RAG: Complete Enterprise Implementation Guide (2026)",
      "url": "https://sumatosoft.com/blog/blog-agentic-rag-enterprise-implementation-guide",
      "date": "2026-07-03",
      "type": "tutorial",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Production deployment guide establishing agentic RAG as standard enterprise pattern for multi-source synthesis in finance, legal, support; quantifies 3–10× token cost increase vs traditional RAG."
    },
    {
      "title": "A Systematic Taxonomy of Failure Modes in Retrieval-Augmented Generation Systems",
      "url": "https://aclanthology.org/2026.trustnlp-main.27/",
      "date": "2026-07-02",
      "type": "research-paper",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed taxonomy of 33 RAG failure modes across 7 pipeline stages; identifies critical gaps—12 modes lack empirical evidence including all 8 agentic orchestration modes—revealing research frontier."
    },
    {
      "title": "CHARLIE: An On-Premise Multi-Agent Retrieval-Augmented Generation System for Evidential Reasoning in Forensic Science",
      "url": "https://arxiv.org/html/2607.05428v1",
      "date": "2026-07-01",
      "type": "research-paper",
      "added": "2026-07-15",
      "superseded_by": null,
      "window": null,
      "explanation": "Multi-agent RAG deployed at Brazilian Federal Police Forensic Institute for cross-corpus evidence reasoning; demonstrates domain-specific maturity in regulated, high-stakes environment with traceability requirements."
    },
    {
      "title": "Why RAG Fails for AI Agents (and What Replaces It)",
      "url": "https://www.sentra.app/articles/why-rag-fails",
      "date": "2026-06-28",
      "type": "opinion",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Documents four structural RAG failure modes (similarity ≠ correctness, synthesis hallucination, meaning re-derived at query time, temporal unawareness) affecting domain-specific systems; provides balanced negative signal on architectural limitations blocking broader adoption."
    },
    {
      "title": "GaRAGe: A benchmark with grounding annotations for RAG evaluation",
      "url": "https://www.amazon.science/publications/garage-a-benchmark-with-grounding-annotations-for-rag-evaluation",
      "date": "2026-06-24",
      "type": "research-paper",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Amazon Science benchmark with 35K+ human-curated grounding annotations across 2366 questions enables fine-grained end-to-end RAG evaluation; signals Tier-1 vendor investment in production-grade domain-specific RAG benchmarking."
    },
    {
      "title": "5 Ways Microsoft's Agentic RAG Cuts Hallucination 89%",
      "url": "https://ragaboutit.com/5-ways-microsofts-agentic-rag-cuts-hallucination-89/",
      "date": "2026-06-21",
      "type": "case-study",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "MLflow RAG Agents framework reduces hallucination from 60% to <7% on enterprise clusters via five patterns: query decomposition, self-reflection, context chaining DAGs, tool-augmented reasoning, and evidence validation; demonstrates production-scale cross-corpus reliability."
    },
    {
      "title": "Retrieval-Augmented Generation for Departmental Test & Evaluation",
      "url": "https://itea.org/journals/volume-47-2/retrieval-augmented-generation-t-and-e/",
      "date": "2026-06-20",
      "type": "case-study",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "U.S. Navy warfare center production RAG deployment on AWS GovCloud demonstrates operationally mature domain-specific RAG with security controls (NIST SP 800-53, DoD RMF, CMMC) for classified data handling; validates regulated-domain viability."
    },
    {
      "title": "Google Research Introduces Agentic RAG for Gemini Enterprise",
      "url": "https://quasa.io/media/google-research-introduces-agentic-rag-for-gemini-enterprise-rag-that-doesn-t-give-up-after-the-first-search",
      "date": "2026-06-20",
      "type": "product-ga",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Gemini Enterprise Agentic RAG with Sufficient Context Agent (93% accuracy determining answer sufficiency) achieves 34% factuality improvement over standard RAG and 90.1% accuracy on cross-corpus scenarios; production GA feature."
    },
    {
      "title": "Dissecting Agentic RAG: A Component Ablation for Multi-Hop QA with a Local 7B Model",
      "url": "https://arxiv.org/abs/2606.21553v1",
      "date": "2026-06-19",
      "type": "research-paper",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Ablation study on agentic RAG multi-hop reasoning shows fixed hybrid retrieval outperforms adaptive routing by +1.8 EM; two iterations capture 95% of gains, validating architectural simplification for cross-corpus reasoning."
    },
    {
      "title": "Evaluate AI Search retrieval quality - Azure Databricks",
      "url": "https://learn.microsoft.com/en-us/azure/databricks/ai-search/retrieval-quality-eval",
      "date": "2026-06-17",
      "type": "product-ga",
      "added": "2026-07-01",
      "superseded_by": null,
      "window": null,
      "explanation": "Databricks AI Search now includes native retrieval quality evaluation generating synthetic queries and scoring via LLM-as-judge on graded relevance scale (DCG@10 primary metric); maturity signal that systematic evaluation moved to production infrastructure."
    },
    {
      "title": "Semantic Search for AI Agents at Scale: Retrieval and Ranking for LinkedIn's Hiring Assistant",
      "url": "https://www.linkedin.com/blog/engineering/ai/semantic-search-for-ai-agents-at-scale-retrieval-and-ranking-for-linkedins-hiring-assistant",
      "date": "2026-06-11",
      "type": "case-study",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Production domain-specific RAG at billion-scale (1.3B+ profiles) using MUSE custom domain embeddings; achieves +4% HRR, −5% false positive rate, +24% Kappa quality improvement—evidence of leading-edge deployment maturity."
    },
    {
      "title": "CQC-RAG: Robust Retrieval-Augmented Generation via Cross-Query Consistency",
      "url": "https://arxiv.org/abs/2606.13438",
      "date": "2026-06-11",
      "type": "research-paper",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Cross-query consistency framework addressing cross-corpus robustness: demonstrates semantically equivalent queries retrieve different results in multi-document corpora; solution achieves +4.76 pp EM on TriviaQA, +9.12 pp on MuSiQue."
    },
    {
      "title": "uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking",
      "url": "https://arxiv.org/abs/2606.11945v1",
      "date": "2026-06-10",
      "type": "research-paper",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Multi-turn RAG evaluated across four distinct domains (finance, cloud docs, government, Wikipedia) with significant cross-domain performance variation; demonstrates domain-specific retrieval strategy selection is essential rather than universal."
    },
    {
      "title": "When More Documents Hurt RAG: Mitigating Vector Search Dilution with Domain-Scoped, Model-Agnostic Retrieval",
      "url": "https://arxiv.org/abs/2606.11350",
      "date": "2026-06-09",
      "type": "research-paper",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed research identifying vector search dilution as critical failure mode in large cross-corpus RAG systems; proposes domain-scoped retrieval solution validated across 5 LLM backbones, 6 corpora, and named Wyoming DOTD deployment."
    },
    {
      "title": "Beyond Probabilistic Similarity: Structural, Temporal, and Causal Limitations of Retrieval-Augmented Generation in the Legal Domain",
      "url": "https://arxiv.org/abs/2606.09724v1",
      "date": "2026-06-08",
      "type": "research-paper",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Landmark analysis identifying three architectural pathologies (mereological, diachronic, causal blindness) showing RAG failures in legal domain are structural mismatches, not confabulation—directly applicable to regulated domains requiring hierarchy and temporal dynamics."
    },
    {
      "title": "Google Research Adds Agentic RAG to Gemini Enterprise Agent Platform with a Sufficient Context Agent for multi-hop queries",
      "url": "https://www.marktechpost.com/2026/06/08/google-research-adds-agentic-rag-to-gemini-enterprise-agent-platform-with-a-sufficient-context-agent-for-multi-hop-queries/",
      "date": "2026-06-08",
      "type": "product-ga",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Google Agentic RAG now in public preview on Gemini Enterprise Agent Platform with iterative multi-agent workflow; 34% factuality accuracy improvement vs standard RAG, FramesQA benchmark 90.1% on multi-corpus scenarios."
    },
    {
      "title": "Why Bigger Models Will Not Fix Enterprise RAG",
      "url": "https://www.santiagocompany.com/insights/why-bigger-models-will-not-fix-enterprise-rag",
      "date": "2026-06-08",
      "type": "opinion",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Critical assessment of enterprise RAG failures: identifies failure occurs in retrieval layer (not generation); documents Enterprise RAG Gold Standard benchmark for realistic failure modes (multi-hop dependencies, conflicting regional policies, structured data in PDFs)."
    },
    {
      "title": "Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval",
      "url": "https://arxiv.org/abs/2606.05658",
      "date": "2026-06-04",
      "type": "research-paper",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical comparison across domain-specific RAG (DevOps KB) and multi-hop reasoning benchmarks showing architecture effectiveness depends on domain characteristics; validates necessity of adaptive retrieval strategies rather than universal approaches."
    },
    {
      "title": "Grounded RAG on Azure AI Foundry: From POC to Auditable Production",
      "url": "https://www.kriv.ai/articles/grounded-rag-on-azure-ai-foundry-from-poc-to-auditable-production",
      "date": "2026-06-04",
      "type": "case-study",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Production domain-specific RAG deployment (health insurance): 40% reduction in document lookups, 91% attribution accuracy, 6% claims accuracy improvement, handling time 14→8 min, 4–6 month payback demonstrates regulated industry viability."
    },
    {
      "title": "【Build 2026 速報】Foundry IQ（Azure AI Search）に新たなプラン Serverless が追加！ナレッジソースに Web IQ、Work IQ、Fabric Ontology が追加 ほか",
      "url": "https://qiita.com/nohanaga/items/7acb452b3e08b5dfdb5f",
      "date": "2026-06-04",
      "type": "product-ga",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Azure Build 2026 announcements for Foundry IQ: agentic retrieval quality improvements (evidence recall +46–54%, token cost −34%, answer quality +8–20%), new knowledge sources (Work IQ, Fabric Ontology) enabling cross-source retrieval."
    },
    {
      "title": "Databricks AI Search Quality Guide",
      "url": "https://docs.databricks.com/aws/pt/ai-search/retrieval-quality",
      "date": "2026-06-04",
      "type": "product-ga",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Vendor GA documentation providing systematic methodology for optimizing domain-specific RAG retrieval quality; emphasizes metadata filtering (>90% search space reduction), hybrid search, and reranking as domain-specific optimization priorities."
    },
    {
      "title": "Гибридный поиск в RAG: как мы подняли Top-1 с 62% до 88% на базе из 50 000 документов",
      "url": "https://habr.com/ru/amp/publications/1042980/",
      "date": "2026-06-03",
      "type": "case-study",
      "added": "2026-06-17",
      "superseded_by": null,
      "window": null,
      "explanation": "Production domain-specific RAG: 50K Russian corporate documents improved Top-1 from 62% to 88% via hybrid search (BM25+vector RRF) and cross-encoder reranking; demonstrates retrieval architecture ladder and domain-specific optimization strategies."
    },
    {
      "title": "Do legal AI tools that use RAG still hallucinate?",
      "url": "https://whatstrending.fenwick.com/post/do-legal-ai-tools-that-use-rag-still-hallucinate",
      "date": "2026-06-02",
      "type": "research-paper",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Stanford research shows RAG reduces legal AI hallucinations from 58-82% to 17-34%, demonstrating RAG's significance but exposing persistent limitations in regulated domains; critical negative signal for high-stakes applications."
    },
    {
      "title": "TechGraphRAG: An Agentic Graph-Augmented RAG Framework for Technical Literature Reasoning",
      "url": "https://arxiv.org/html/2606.01613v1",
      "date": "2026-06-02",
      "type": "case-study",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Goodyear production agentic RAG deployment over 2,100 academic papers (24K chunks) with evidence sufficiency scoring and cross-corpus academic search; demonstrates integration of internal knowledge graphs with external academic databases."
    },
    {
      "title": "Agentic retrieval in Azure AI Search",
      "url": "https://docs.azure.cn/en-us/search/agentic-retrieval-overview",
      "date": "2026-06-01",
      "type": "product-ga",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Official Azure AI Search agentic retrieval documentation (public preview 2026-06-01) designed for cross-corpus question answering via LLM-driven query decomposition and parallel execution; production infrastructure for complex domain-specific QA."
    },
    {
      "title": "5 Failure Modes I Found in My Financial RAG (And the One That Actually Mattered)",
      "url": "https://dev.to/joaopaulotr/5-failure-modes-i-found-in-my-financial-rag-and-the-one-that-actually-mattered-4b1p",
      "date": "2026-05-30",
      "type": "case-study",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Financial RAG production analysis quantifying failure modes: corpus audit yield +11% vs algorithmic improvements +5%; recall improved 0.83→0.94; demonstrates data quality bottleneck and bottleneck shift to generation once retrieval improves."
    },
    {
      "title": "RAG Is Dead? 5 Enterprise Failures Say Otherwise",
      "url": "https://ragaboutit.com/rag-is-dead-5-enterprise-failures-say-otherwise/",
      "date": "2026-05-27",
      "type": "case-study",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Five named Fortune 500 RAG failure post-mortems ($12.2M trading loss, $4.1M legal settlement, $2.3M manufacturing loss) document temporal grounding and retrieval necessity; demonstrates RAG criticality in specialized domains and cost of removal."
    },
    {
      "title": "FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing",
      "url": "https://arxiv.org/abs/2605.26476",
      "date": "2026-05-26",
      "type": "research-paper",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Novel framework for domain-specific RAG evaluation with six diagnostic metrics and 200 curated QA pairs; evaluates context-scaling behavior across models revealing attention dilution as primary degradation mechanism in specialized corpora."
    },
    {
      "title": "Benchmarking Large Language Models and Prompt Engineering Strategies in Microsatellite Instability Cancers: Evaluation Study",
      "url": "https://pubmed.ncbi.nlm.nih.gov/42166792/?fc=20220524054416&ff=20260521203316&v=2.20.0",
      "date": "2026-05-21",
      "type": "research-paper",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed medical benchmark (JMIR) with 511 MSI cancer questions; RAG as most effective intervention; demonstrates critical success factors: retrieval precision and knowledge base quality determine system safety."
    },
    {
      "title": "Structured domain knowledge enables trustworthy materials science question-answering with large language models",
      "url": "https://pubs.rsc.org/en/content/articlehtml/2026/dd/d6dd00028b",
      "date": "2026-05-20",
      "type": "case-study",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed materials science RAG achieving 85.6% accuracy vs 21.3% baseline (4x improvement) via structured knowledge + query reformulation on water-splitting catalysis domain; production-scale deployment with quantified ROI."
    },
    {
      "title": "When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering",
      "url": "https://arxiv.org/abs/2605.21807",
      "date": "2026-05-20",
      "type": "research-paper",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Clinical RAG benchmark (OGCaReBench) on 513 expert-validated questions shows 56% baseline to 82% with retrieval augmentation; demonstrates RAG necessity for medical domain evidence-grounded reasoning beyond guidelines."
    },
    {
      "title": "Fine-grained Claim-level RAG Benchmark for Law",
      "url": "https://arxiv.org/abs/2605.21071v2",
      "date": "2026-05-20",
      "type": "research-paper",
      "added": "2026-06-03",
      "superseded_by": null,
      "window": null,
      "explanation": "Bilingual (French/English) legal RAG benchmark (ClaimRAG-LAW) targeting experts and non-experts with fine-grained evaluation framework and hallucination analysis; addresses critical gap in domain-specific legal AI evaluation."
    },
    {
      "title": "Query-Conditioned Knowledge Alignment for Reliable Cross-System Medical Reasoning",
      "url": "https://arxiv.org/abs/2605.18570",
      "date": "2026-05-18",
      "type": "research-paper",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "Cross-domain medical knowledge alignment via Query-Conditioned Entity Alignment (QCEA) improves evidence retrieval and answer accuracy in heterogeneous domains (Traditional Chinese vs Western Medicine); demonstrates semantic integration requirement for cross-corpus medical RAG."
    },
    {
      "title": "Two Retrieval Methods Are Better Than One: Evidence from 500 Clinical Queries",
      "url": "https://dev.to/nomad-link-id/two-retrieval-methods-are-better-than-one-evidence-from-500-clinical-queries-4g41",
      "date": "2026-05-13",
      "type": "case-study",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical medical domain case study across 6 specialties: hybrid BM25+dense retrieval achieves Recall@5 0.86 vs BM25-only 0.71; complementary strengths confirmed with statistical significance (McNemar p<0.001)."
    },
    {
      "title": "Context Convergence Improves Answering Inferential Questions",
      "url": "https://arxiv.org/abs/2605.12370",
      "date": "2026-05-12",
      "type": "research-paper",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "SIGIR 2026: passage structure and quality directly impact domain-specific QA for inferential reasoning; convergence metric for constructing reasoning-optimal passages outperforms cosine similarity selection across six LLM architectures."
    },
    {
      "title": "MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval",
      "url": "https://arxiv.org/abs/2605.11254",
      "date": "2026-05-11",
      "type": "research-paper",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "SIGIR 2026 benchmark for category-aware ranking across heterogeneous sources (Publications, Data, Variables, Tools); addresses production requirement of unified retrieval interfaces over diverse domain-specific data types."
    },
    {
      "title": "Assessment of RAG and Fine-Tuning for Industrial Question-Answering Applications",
      "url": "https://arxiv.org/abs/2605.09533",
      "date": "2026-05-10",
      "type": "research-paper",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "AAAI 2026 empirical study on automotive domain-specific QA: RAG proves most cost-efficient adaptation for both premium and open-source models, with detailed cost-quality tradeoff analysis for industrial closed-domain datasets."
    },
    {
      "title": "7 Enterprise RAG Audit Failures You Should Know",
      "url": "https://ragaboutit.com/7-enterprise-rag-audit-failures-you-should-know/",
      "date": "2026-05-09",
      "type": "industry-report",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "Independent audit of 50 live deployments documenting 7 production failure modes (context gaslighting 76% finance, citation fabrication 81% legal) across adversarial testing; critical negative signal confirming production readiness demands meticulous deployment discipline."
    },
    {
      "title": "Clinical Context Variables Collectively Rival Model Choice in Embedding-Based Retrieval: Multi-Corpus Benchmark Study",
      "url": "https://medinform.jmir.org/2026/1/e94241",
      "date": "2026-05-07",
      "type": "research-paper",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed clinical benchmark: domain context variables (corpus type, query format) explain 49% variance in retrieval performance vs 47.6% for model choice; model rankings unstable across clinical corpora, confirming domain-specific validation mandatory."
    },
    {
      "title": "Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems",
      "url": "https://huggingface.co/papers/2605.04018",
      "date": "2026-05-07",
      "type": "research-paper",
      "added": "2026-05-20",
      "superseded_by": null,
      "window": null,
      "explanation": "Novel agentic retrieval evaluation framework (BRIGHT-Pro, RTriever-Synth, RTriever-4B) for multi-aspect evidence portfolios across search steps, addressing leading-edge requirement to support reasoning-intensive cross-corpus QA beyond single-shot retrieval."
    },
    {
      "title": "Benchmarking Retrieval Strategies for Biomedical Retrieval-Augmented Generation: A Controlled Empirical Study",
      "url": "https://arxiv.org/abs/2605.02520v1",
      "date": "2026-05-04",
      "type": "research-paper",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed empirical comparison of 5 retrieval strategies (Dense, Hybrid, Cross-Encoder, Multi-Query, MMR) on BioASQ benchmark showing cross-encoder reranking achieves 0.827 composite score (0.852 contextual precision), establishing domain-specific retrieval strategy selection guidance."
    },
    {
      "title": "I Built an AI Knowledge Bot. Here's the Silent Bug That Was Breaking It.",
      "url": "https://www.frankyao.com/blog/rag-chatbot-embedding-drift",
      "date": "2026-05-01",
      "type": "case-study",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Real domain-specific RAG deployment (LLC formation, logistics, AI automation with 4,290 vectors) suffered silent embedding model mismatch (ONNX-quantized indexing vs Hugging Face API queries) below similarity threshold. Demonstrates embedding consistency as critical infrastructure requirement."
    },
    {
      "title": "When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation",
      "url": "https://arxiv.org/abs/2605.00911",
      "date": "2026-04-29",
      "type": "research-paper",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Benchmark revealing structural/semantic errors from high-accuracy OCR (11 document types tested) still cause RAG retrieval failures. Challenges assumption that OCR accuracy alone ensures downstream RAG viability in document-heavy domains."
    },
    {
      "title": "Health System Scale Semantic Search Across Unstructured Clinical Notes",
      "url": "https://arxiv.org/abs/2604.25605",
      "date": "2026-04-28",
      "type": "research-paper",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Full production deployment of domain-specific semantic search at children's hospital serving 1.68M patients across 166M clinical notes. Qwen3 embeddings achieved 94.6% clinical QA accuracy with 237ms latency and $4K/month cost, confirming health-system-scale viability."
    },
    {
      "title": "Prism-Reranker: Beyond Relevance Scoring — Jointly Producing Contributions and Evidence for Agentic Retrieval",
      "url": "https://arxiv.org/html/2604.23734v1",
      "date": "2026-04-26",
      "type": "research-paper",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Production reranker models (0.8B-9B) generating relevance + contribution statement + evidence passage for agentic/cross-corpus RAG, improving BEIR scores. Addresses cross-corpus context precision requirement."
    },
    {
      "title": "How We Built an Advanced RAG System for Documents",
      "url": "https://www.kalviumlabs.ai/blog/how-we-built-advanced-rag-document-intelligence/",
      "date": "2026-04-26",
      "type": "case-study",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Legal contracts RAG production deployment (900+ documents): naive 512-token chunking achieved 0.61 precision; parent-child chunking + hybrid search + reranking reached 0.87, demonstrating domain-specific document structure as primary architectural driver."
    },
    {
      "title": "RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora",
      "url": "https://www.themoonlight.io/en/review/rare-redundancy-aware-retrieval-evaluation-framework-for-high-similarity-corpora",
      "date": "2026-04-23",
      "type": "research-paper",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "RAG framework for high-redundancy enterprise domains (finance, legal, patents). Quantifies robustness degradation on enterprise corpora: dense retrievers drop from 66.4% to 5.0-27.9% recall across domain-specific hops, establishing evaluation methodology for specialized domains."
    },
    {
      "title": "Popularity Bias in Vector Retrieval: Why the Same Five Chunks Dominate Every Query",
      "url": "https://tianpan.co/blog/2026-04-23-popularity-bias-vector-retrieval-rag-diversity",
      "date": "2026-04-23",
      "type": "opinion",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Production analysis of hubness phenomenon in vector retrieval where high-dimensional geometry causes disproportionate chunk retrieval. Provides diagnostic framework (Gini coefficient, k-occurrence metrics) and stacked mitigations (MMR, reranking) for domain-specific RAG systems."
    },
    {
      "title": "Domain-oriented RAG Assessment (DoRA): Synthetic Benchmarking for RAG-based Question Answering on Defense Documents",
      "url": "https://www.themoonlight.io/ko/review/domain-oriented-rag-assessment-dora-synthetic-benchmarking-for-rag-based-question-answering-on-defense-documents",
      "date": "2026-04-22",
      "type": "research-paper",
      "added": "2026-05-06",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific RAG benchmark for sensitive domains with quantified deployment results: domain adaptation fine-tuning achieved 26% task success improvement and 47% hallucination reduction on defense document QA."
    },
    {
      "title": "Grounded in Law: A Multi-Stage Anti-Hallucination Pipeline for Legal RAG Systems in Brazilian Portuguese",
      "url": "https://aclanthology.org/2026.propor-2.9/",
      "date": "2026-04-17",
      "type": "research-paper",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed production RAG deployment for legal domain with real metrics: 184,895 audited answers, 81.7% legislation reference resolution, 47.1% jurisprudence resolution, 6.5% hallucination correction rate."
    },
    {
      "title": "Enhanced Agentic-RAG: What If Chatbots Could Deliver Near-Human Precision?",
      "url": "https://www.uber.com/in/en/blog/enhanced-agentic-rag/",
      "date": "2026-04-16",
      "type": "case-study",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Uber's production case study: Genie on-call copilot achieved 27% relative improvement in acceptable answers and 60% reduction in incorrect advice through enhanced agentic RAG for domain-specific (security/privacy) policies."
    },
    {
      "title": "Corpus Curation at Scale: Why Your RAG Quality Ceiling Is Your Document Quality Floor",
      "url": "https://tianpan.co/blog/2026-04-14-corpus-curation-rag-document-quality-floor",
      "date": "2026-04-14",
      "type": "opinion",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Practitioner analysis identifying document quality as critical RAG failure mode. Reports specific metric: document quality scoring improved search accuracy from 62% to 89% with no embedding/retrieval changes."
    },
    {
      "title": "Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method",
      "url": "https://arxiv.org/abs/2604.11209",
      "date": "2026-04-13",
      "type": "research-paper",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "SIGIR 2026 paper introducing ConflictQA benchmark for cross-source knowledge conflicts (text vs. knowledge graphs) and XoT reasoning framework—directly addresses cross-corpus QA challenge."
    },
    {
      "title": "MedDiscover: A Domain-Specific Retrieval-Augmented Generation System for Medical Question Answering",
      "url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13082540/",
      "date": "2026-04-09",
      "type": "research-paper",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific RAG case study in medical/biomedical domain using specialized embedding models (MedCPT) and domain-specific evaluation. Tests retrieval augmentation on biomedical corpora with measured grounding and answerability metrics, demonstrating domain adaptation of RAG architectures for knowledge-intensive medical QA."
    },
    {
      "title": "Embedding Models in Production: Selection, Versioning, and the Index Drift Problem",
      "url": "https://tianpan.co/blog/2026-04-09-embedding-models-production-versioning-index-drift",
      "date": "2026-04-09",
      "type": "opinion",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Production-focused guidance on embedding selection for domain-specific RAG. Highlights MTEB benchmark limitations and index drift risks from silent model versioning changes."
    },
    {
      "title": "A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering",
      "url": "https://arxiv.org/abs/2604.07274",
      "date": "2026-04-08",
      "type": "research-paper",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed systematic evaluation of domain-specific RAG pipeline components for medical QA, showing empirical optimization of retrieval strategies within domain constraints."
    },
    {
      "title": "BEIR Benchmark Leaderboard 2025 & 2026: NDCG@10 Scores & Rankings",
      "url": "https://app.ailog.fr/en/blog/news/beir-benchmark-update",
      "date": "2026-04-08",
      "type": "industry-report",
      "added": "2026-04-22",
      "superseded_by": null,
      "window": null,
      "explanation": "Empirical benchmark data showing domain-specific fine-tuning improvements: medical +29% (48%→62%), code +34% (44%→59%), legal +24% (46%→57%). Direct signal that general models fail on specialized domains."
    },
    {
      "title": "Octen-Embedding-0.6B: Domain-Specific Embedding Model for Vertical Domains",
      "url": "https://techcommunity.microsoft.com/tag/generative%20ai",
      "date": "2026-04-06",
      "type": "product-ga",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Microsoft domain-adapted embedding model (0.6B params) trained for legal, finance, healthcare, code; outperforms voyage-3.5 (0.7241 vs 0.7139) despite tiny parameter count, signals ecosystem shift toward specialized embeddings."
    },
    {
      "title": "Vector Drift in Azure AI Search: Three Hidden Reasons Your RAG Accuracy Degrades After Deployment",
      "url": "https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/vector-drift-in-azure-ai-search-three-hidden-reasons-your-rag-accuracy-degrades-/4493031",
      "date": "2026-04-04",
      "type": "product-ga",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Production RAG maintenance challenge: vector drift causes silent retrieval degradation from embedding model version mismatch, incremental content updates without re-embedding, inconsistent chunking strategies."
    },
    {
      "title": "Why RAG Pipelines Fail at Production Scale (And What We Fixed)",
      "url": "https://dev.to/ailoitte_sk/why-rag-pipelines-fail-at-production-scale-and-what-we-fixed-18mf",
      "date": "2026-04-01",
      "type": "case-study",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Production RAG failures across 12+ systems (fintech, healthcare, SaaS): naive chunking (54%→81% precision fix), wrong embeddings (domain-adapted fixes), missing re-ranker, context compression; demonstrates specialization necessity."
    },
    {
      "title": "RAG for Product Catalogs 2026: When It Works, When It Doesn't, and What I Learned Building Three",
      "url": "https://www.claudio-novaglio.com/en/papers/rag-product-catalogs-2026",
      "date": "2026-03-31",
      "type": "case-study",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Three production domain-specific RAG systems: multilingual furniture (10K+ products, 85%+ cross-language precision), B2B configurators with domain logic, KG integration; pragmatic analysis of domain-specific tradeoffs vs context stuffing."
    },
    {
      "title": "From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents",
      "url": "https://chatpaper.com/paper/264067",
      "date": "2026-03-27",
      "type": "research-paper",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Domain-specific RAG benchmark on financial documents (23,088 queries): hybrid retrieval + neural reranking achieves 0.816 Recall@5; BM25 outperforms dense retrieval in financial domain, challenging semantic-search universality."
    },
    {
      "title": "Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering",
      "url": "https://arxiv.org/abs/2603.26567",
      "date": "2026-03-27",
      "type": "research-paper",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "EACL 2026 StackRepoQA benchmark: domain-specific code cross-corpus QA on 134 Java projects (1,318 questions); structural signals improve RAG but overall accuracy limited, demonstrating repository-scale comprehension complexity."
    },
    {
      "title": "What's new in Azure AI Search (2026)",
      "url": "https://docs.azure.cn/en-us/search/whats-new",
      "date": "2026-03-25",
      "type": "product-ga",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Azure AI Search 2025-2026 evolution: November 2025 GA for agentic retrieval, answer synthesis, semantic ranker on free tier; signals mainstream vendor recognition of domain-specific RAG as core capability."
    },
    {
      "title": "Retrieval Improvements Do Not Guarantee Better Answers: A Study of RAG for AI Policy QA",
      "url": "https://arxiv.org/abs/2603.24580",
      "date": "2026-03-25",
      "type": "research-paper",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Peer-reviewed study on domain-specific RAG for AI policy using AGORA corpus (947 documents): domain fine-tuning improves retrieval metrics but not end-to-end QA; stronger retrieval paradoxically increases confident hallucinations."
    },
    {
      "title": "Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models",
      "url": "https://www.jmir.org/2026/1/e82997",
      "date": "2026-03-25",
      "type": "research-paper",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Healthcare domain-specific RAG: fine-tuned multilingual embeddings on 400K clinical documents (163K patients) achieved mAP@100 0.27 vs. baselines 0.14-0.11, demonstrating clinical embedding specialization necessity."
    },
    {
      "title": "CVPD at QIAS 2026: RAG-Guided LLM Reasoning for Al-Mawarith (Islamic Inheritance Law)",
      "url": "https://www.catalyzex.com/s/Cross%20Encoder%20Reranking",
      "date": "2026-03-25",
      "type": "case-study",
      "added": "2026-04-08",
      "superseded_by": null,
      "window": null,
      "explanation": "Competition-validated domain-specific legal RAG system achieves 0.935 MIR-E score on QIAS 2026 leaderboard via rule-grounded synthesis, hybrid retrieval, neural reranking, schema-constrained generation on Islamic law."
    },
    {
      "title": "Redefining Retrieval Evaluation in the Era of LLMs",
      "url": "https://aclanthology.org/2026.eacl-long.391/",
      "date": "2026-03-24",
      "type": "research-paper",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Recent EACL conference paper identifying fundamental misalignment in RAG evaluation: traditional IR metrics assume sequential human examination, but LLMs process all documents holistically; proposes UDCG metric improving correlation by 36%."
    },
    {
      "title": "Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Question Answering Task",
      "url": "https://aclanthology.org/2026.findings-eacl.35/",
      "date": "2026-03-24",
      "type": "research-paper",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Direct match: peer-reviewed research on RAG for knowledge-intensive multilingual QA. Proposes solutions (CrossRAG) for cross-corpus consistency challenges, published in top-tier venue (EACL 2026 Findings)."
    },
    {
      "title": "Questions Are All You Need To Train A Dense Passage Retriever",
      "url": "https://fr.scribd.com/document/773363467/Questions-Are-All-You-Need-to-Train-a-Dense-Passage-Retriever",
      "date": "2026-03-21",
      "type": "research-paper",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Peer-reviewed research from Meta AI and Google DeepMind proposing ART (unsupervised dense retriever training) for QA retrieval. Advances unsupervised methods for domain-specific question answering systems."
    },
    {
      "title": "Knowledge-Adaptive Question Answering with Retrieval-Augmented Generation on Amazon Bedrock",
      "url": "https://ijctjournal.org/knowledge-adaptive-question-answering/",
      "date": "2026-03-13",
      "type": "research-paper",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Production implementation of domain-specific RAG across regulated domains (medical, legal, financial) with substantial empirical improvements and enterprise governance features demonstrating leading-edge maturity."
    },
    {
      "title": "SEXTANT Research: RAG Enterprise Adoption, Evolution & Market",
      "url": "https://vargazoltan.ai/en/blog/ragfuture-hat-iranyu-elemzes/",
      "date": "2026-03-11",
      "type": "industry-report",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Comprehensive market research synthesizing 85+ sources on RAG adoption rates, ROI metrics, failure modes, and architectural evolution; signals leading-edge maturity with named case studies and cost-benefit data."
    },
    {
      "title": "SimRAG: Self-improving retrieval-augmented generation for adapting large language models to specialized domains",
      "url": "https://www.amazon.science/publications/simrag-self-improving-retrieval-augmented-generation-for-adapting-large-language-models-to-specialized-domains",
      "date": "2026-03-11",
      "type": "research-paper",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Domain-specific RAG adaptation framework for specialized fields (science, medicine). Proposes self-training approach equipping LLMs with joint question answering and question generation capabilities for domain-specific knowledge systems."
    },
    {
      "title": "How We Built a Competitive Memory Retrieval System using Open-Source Models",
      "url": "https://ensue.dev/blog/beating-memory-benchmarks/",
      "date": "2026-03-03",
      "type": "case-study",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Real deployment of multi-stage retrieval system for cross-corpus agent memory achieving 96-100% single-session accuracy; demonstrates domain-specific retrieval at scale (115K+ tokens) with measured model comparisons."
    },
    {
      "title": "Machine Translation Digest for Feb 26 2026 (Towards Faithful Industrial RAG)",
      "url": "https://buttondown.com/daily-mt-picks/archive/machine-translation-digest-for-feb-26-2026/",
      "date": "2026-03-03",
      "type": "case-study",
      "added": "2026-03-25",
      "superseded_by": null,
      "window": "2026-03",
      "explanation": "Production case study with quantified outcomes. Domain-specific RAG system deployed at scale, reducing hallucination by 72% and serving millions of interactions. GraphRAG architecture for relational knowledge."
    },
    {
      "title": "When Not to Use RAG: Scenarios and Alternatives",
      "url": "https://dasroot.net/posts/2026/02/when-not-to-use-rag-scenarios-alternatives/",
      "date": "2026-02-26",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Critical analysis citing real-world failures: healthcare harmful guidance from outdated knowledge base, finance compliance violation, telecom $2.3M service credits for incorrect information—demonstrates deployment risks in high-stakes domains."
    },
    {
      "title": "announcements",
      "url": "https://learn.microsoft.com/en-us/azure/search/whats-new",
      "date": "2026-02-11",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Azure AI Search February 2026 release adds agentic retrieval portal support with new knowledge sources (OneLake, SharePoint, Web) and retrieval options, signaling vendor infrastructure evolution for domain-specific RAG."
    },
    {
      "title": "My Journey Learning To Build AI Apps on Azure (March 2025 to Feb 2026)",
      "url": "https://roykim.ca/2026/02/10/my-journey-learning-to-build-ai-apps-on-azure-march-2025-to-feb-2026/",
      "date": "2026-02-10",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Practitioner Azure MVP deployed multiple domain-specific RAG systems (product catalog, resume search) with embeddings and Azure AI Search, demonstrating 11-month real-world implementation journey."
    },
    {
      "title": "Retrieval-augmented Generation (RAG) in Azure AI Search",
      "url": "https://docs.azure.cn/en-us/search/retrieval-augmented-generation-overview",
      "date": "2026-02-09",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Official Azure documentation detailing agentic and classic RAG patterns with guidance on solving domain-specific challenges including query understanding, multi-source access, and token constraints."
    },
    {
      "title": "Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention",
      "url": "https://www.arxiv.org/abs/2602.04711",
      "date": "2026-02-04",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Research addressing RAG security vulnerabilities via sparse document attention mechanism; shows defenses substantially outperform standard approaches, indicating maturity in security-critical domain deployments."
    },
    {
      "title": "Inferential Question Answering - arXiv",
      "url": "https://www.arxiv.org/abs/2602.01239",
      "date": "2026-02-01",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-02",
      "explanation": "Research reveals RAG and reranking fail at inference-based reasoning on QUIT dataset; current QA pipelines not ready for indirect evidence reasoning, exposing fundamental capability limitations."
    },
    {
      "title": "CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning",
      "url": "https://arxiv.org/abs/2601.14952",
      "date": "2026-01-21",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Research benchmark showing standard RAG systems collapse on 10M token corpora (1.20% accuracy vs. 55.85% at 128K), while memory-augmented agents demonstrate resilience; indicates architectural limits for large-scale cross-corpus retrieval."
    },
    {
      "title": "Retrieval Augmented Generation for Natural Language Querying of Biologics Immunogenicity Data",
      "url": "https://pubmed.ncbi.nlm.nih.gov/41566090/",
      "date": "2026-01-21",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Domain-specific RAG evaluated on 663 biologics package inserts and 206 monoclonal antibodies; retriever showed inaccuracies despite coherent LLM outputs, highlighting retriever-generation gaps in specialized medical domains."
    },
    {
      "title": "An Empirical Evaluation of RAG Architectures for Policy Question Answering",
      "url": "https://arxiv.org/abs/2601.15457",
      "date": "2026-01-21",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Advanced RAG with cross-encoder re-ranking on CDC policy documents achieves 0.797 faithfulness vs. 0.347 baseline; demonstrates two-stage retrieval effectiveness and document segmentation constraints for multi-step domain-specific reasoning."
    },
    {
      "title": "5 Critical Limitations of RAG Systems Every AI Builder Must Understand",
      "url": "https://www.chatrag.ai/blog/2026-01-21-5-critical-limitations-of-rag-systems-every-ai-builder-must-understand",
      "date": "2026-01-21",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Critical assessment of RAG limitations in high-stakes domains: retrieval quality, evolved hallucinations, context constraints, freshness issues, and evaluation complexity; cites healthcare risks and 'lost in the middle' phenomenon."
    },
    {
      "title": "A Transferable Retrieval-Augmented Generation Framework for Vertical-Domain Question Answering",
      "url": "https://www.clausiuspress.com/article/17231.html",
      "date": "2026-01-14",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Framework for cross-domain RAG with domain routing and query rewriting shows improved accuracy and citation consistency across academic, tourism, and finance domains; addresses cross-domain generalization via transfer-oriented training."
    },
    {
      "title": "The Death of Standard RAG: Cache vs. Hypergraph in 2026",
      "url": "https://www.mmntm.net/articles/rag-bifurcation",
      "date": "2026-01-04",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2026-01",
      "explanation": "Critical analysis of RAG architecture bifurcation: caching for small corpora (<1M tokens) with 15-20% accuracy gains vs. agentic hypergraphs for complex domains; signals emerging consensus that standard RAG is obsolete for modern workloads."
    },
    {
      "title": "RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG",
      "url": "https://arxiv.org/abs/2511.04502",
      "date": "2025-11-06",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Framework for rigorous domain-specific RAG evaluation using synthetic QA generation and LLM-as-a-judge metrics; demonstrates evaluation challenges across military, cybersecurity, and engineering domains with context-dependent performance."
    },
    {
      "title": "Azure AI GPT-4.1 agent not able to search search... - Microsoft Learn",
      "url": "https://learn.microsoft.com/en-us/answers/questions/5599564/azure-ai-gpt-4-1-agent-not-able-to-search-search-i",
      "date": "2025-10-27",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Real deployment failure: GPT-4.1 agent with Azure AI Search fails to consistently invoke search tool, exposing challenges in enforcing retrieval behavior and potential tool schema/configuration issues."
    },
    {
      "title": "The 5 best RAG evaluation tools in 2025",
      "url": "https://www.braintrust.dev/articles/best-rag-evaluation-tools",
      "date": "2025-10-23",
      "type": "industry-report",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Industry analysis reporting that RAG powers 60% of production AI applications in 2025; signals broad enterprise adoption of RAG across customer support, knowledge bases, and internal systems."
    },
    {
      "title": "Phase 2: The Query Workflow - Building a Production-Ready Private RAG System",
      "url": "https://www.ucssolutions.com/blog/building-a-production-ready-private-rag-system/",
      "date": "2025-10-10",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Production RAG deployment for SalesWorx product documentation using LlamaIndex, Qdrant vector DB, and multi-agent orchestration with 262K context; demonstrates real-world implementation architecture for domain-specific knowledge retrieval."
    },
    {
      "title": "That Time an Azure AI Search Indexer Failed Silently",
      "url": "https://www.mehmetseckin.com/posts/that-time-an-azure-ai-search-indexer-failed-silently/",
      "date": "2025-10-08",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q4",
      "explanation": "Configuration pitfall in Azure AI Search: Text Splitting skill silently failed due to incorrect context property; highlights specific implementation barriers and need for precise configuration in production domain-specific RAG systems."
    },
    {
      "title": "Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs",
      "url": "https://arxiv.org/html/2601.15429v1",
      "date": "2025-09-16",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Healthcare RAG study: scope-matched knowledge graphs deliver consistent gains, while indiscriminate graph unions introduce distractors; establishes principle that precision-first, scope-focused KG-RAG outperforms breadth-first approaches."
    },
    {
      "title": "Domain-Specific Data Generation Framework for RAG Adaptation",
      "url": "https://arxiv.org/html/2510.11217v1",
      "date": "2025-09-16",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "RAGen framework for generating domain-specific QA training data without labeled datasets; enables RAG adaptation via semantic chunking and hierarchical concept extraction, addressing key barrier of limited domain-specific training data."
    },
    {
      "title": "QuARK: LLM-Based Domain-Specific Question Answering Using Retrieval Augmented Generation and Knowledge Graphs",
      "url": "https://aclanthology.org/2025.ranlp-1.25/",
      "date": "2025-09-15",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "RANLP 2025: domain-specific QA framework integrating KGs with RAG achieves 13% accuracy improvement on financial datasets; demonstrates automated KG construction and question decomposition for domain-specific retrieval."
    },
    {
      "title": "RAG in the Wild: When More Knowledge Hurts",
      "url": "https://cognaptus.com/blog/2025-07-29-rag-in-the-wild-when-more-knowledge-hurts/",
      "date": "2025-07-29",
      "type": "news-coverage",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Study across diverse corpora (Wikipedia, PubMed, GitHub, StackExchange) shows RAG helps small LLMs (+22.87%) but provides minimal gains for large models; reveals that corpus routing remains unsolved with LLMs poor at dynamic source selection."
    },
    {
      "title": "The Elusive Quest for RAG Accuracy: What to Know Up Front",
      "url": "https://graphrag.info/2025/07/28/the-elusive-quest-for-rag-accuracy-what-to-know-up-front/",
      "date": "2025-07-28",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Critical analysis documenting domain-specific RAG accuracy plateau at 75% (25% error rate) in production; advocates graph-based semantic layers over vector-only RAG, citing >95% success on policy document domain vs. vector-only underperformance."
    },
    {
      "title": "Boost Embedding Model Accuracy for Custom Information Retrieval",
      "url": "https://developer.nvidia.com/blog/boost-embedding-model-accuracy-for-custom-information-retrieval/?ncid=so-link-748771",
      "date": "2025-07-24",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q3",
      "explanation": "Coxwave Align production deployment: domain-specific embedding fine-tuning via NeMo Curator achieved 12% RAG accuracy improvement and 6x training speedup, demonstrating practical performance gains in real-world analytics platform."
    },
    {
      "title": "Knowledge-Aware Diverse Reranking for Cross-Source Question Answering",
      "url": "https://www.arxiv.org/abs/2506.20476",
      "date": "2025-06-25",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "SIGIR 2025 LiveRAG competition winner: knowledge-aware reranking RAG pipeline achieving first place on 15M-document FineWeb corpus evaluation, demonstrating cross-source QA capability and retrieval effectiveness at scale."
    },
    {
      "title": "Evaluating Long-Context Question & Answer Systems",
      "url": "https://eugeneyan.com/writing/qa-evals/",
      "date": "2025-06-22",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Practitioner analysis of multi-document long-context QA evaluation challenges: information overload, positional variance, hallucinations at scale; identifies tension between faithfulness and helpfulness in domain-specific deployments."
    },
    {
      "title": "How RAG is Changing Knowledge Base Search",
      "url": "https://blog.helpdocs.io/rag-knowledge-base/",
      "date": "2025-05-15",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Case narrative illustrating domain-specific RAG deployment friction: SaaS knowledge base search returning 15 irrelevant articles for specific configuration queries; highlights practical effectiveness gaps in production knowledge-base RAG systems."
    },
    {
      "title": "EKRAG: Benchmark RAG for Enterprise Knowledge Question Answering",
      "url": "https://aclanthology.org/2025.knowledgenlp-1.13/",
      "date": "2025-05-06",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Enterprise knowledge QA benchmark spanning diverse corporate document types (product releases, technical blogs, financial reports); validates domain-specific RAG effectiveness for real enterprise knowledge bases with frequently updated content."
    },
    {
      "title": "Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains",
      "url": "https://aclanthology.org/2025.naacl-long.575/",
      "date": "2025-04-07",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "NAACL 2025 research on domain adaptation: SimRAG self-training approach enabling LLMs to perform QA and generate questions for science/medicine domains, addressing distribution shift and limited access to specialized data."
    },
    {
      "title": "Evaluating Knowledge Graph Based Retrieval Augmented Generation Methods under Knowledge Incompleteness",
      "url": "https://www.arxiv.org/abs/2504.05163",
      "date": "2025-04-07",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q2",
      "explanation": "Evaluates KG-RAG failure modes when knowledge graphs have missing information; demonstrates practical limitations of domain-specific KG-based retrieval in real-world scenarios where corpus completeness cannot be guaranteed."
    },
    {
      "title": "Amazon Bedrock now supports RAG Evaluation (generally available)",
      "url": "https://aws.amazon.com/about-aws/whats-new/2025/03/amazon-bedrock-rag-evaluation-generally-available/",
      "date": "2025-03-20",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "AWS GA announcement (March 2025) for RAG Evaluation enabling domain-specific application assessment via context relevance, coverage, faithfulness, and responsible AI metrics; signals major cloud vendor platform maturity."
    },
    {
      "title": "Azure AI Search - Case Study - Facet Technologies",
      "url": "https://facettech.com/ai-search-case-study/",
      "date": "2025-02-26",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "FacetTrak field service management platform deployed Azure AI Search with semantic search for domain-specific ticket retrieval; production rollout with automated migration and daily indexing, improving search experience in operational domain."
    },
    {
      "title": "Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models",
      "url": "http://arxiv.org/abs/2502.15854",
      "date": "2025-02-21",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Framework for domain-specific RAG optimization on finance (SEC 10-K), biomedical (PubMed), and cybersecurity (APT) corpora; smaller chunks (<10 tokens) improve precision by 31-42%, with domain-specific embedding yielding 22% variance in optimal sizing."
    },
    {
      "title": "Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA",
      "url": "https://arxiv.org/abs/2502.10497",
      "date": "2025-02-14",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Large-scale empirical evaluation of RAG on 20,000 FAQ queries with 400,000 KB entries; RAG showed knowledge grounding but DoRA outperformed on accuracy (90.1%) and latency (110ms), revealing limitations in accuracy-critical domains like healthcare and finance."
    },
    {
      "title": "Retrieval-Augmented Generation: Challenges & Solutions",
      "url": "https://www.chitika.com/rag-challenges-and-solution/",
      "date": "2025-01-28",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "Practitioner analysis of real-world RAG implementation barriers: poor retrieval in healthcare, semantic mismatches in finance, outdated knowledge bases; highlights critical adoption obstacles in high-stakes domains requiring extensive fine-tuning and domain-specific indexing."
    },
    {
      "title": "Reference-Aligned Retrieval-Augmented Question Answering over Heterogeneous Proprietary Documents",
      "url": "https://www.emorynlp.org/publications/cikm-2025-choi-et-al",
      "date": "2025-01-01",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2025-Q1",
      "explanation": "CIKM 2025 research on domain-specific RAG for automotive industry (crash test docs); pilot deployment showed +1.79 factual correctness, +1.33 informativeness vs. baseline, confirming real-world adoption in manufacturing domain."
    },
    {
      "title": "Everything Wrong with Retrieval-Augmented Generation",
      "url": "https://www.leximancer.com/blog/everything-wrong-with-retrieval-augmented-generation",
      "date": "2024-12-20",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Critical analysis documents nine RAG limitations: retrieval quality dependency, latency/scalability, context limits, transparency gaps, bias vulnerability, maintenance costs, ethical concerns, and domain-specific brittleness."
    },
    {
      "title": "What's new in Azure AI Search",
      "url": "https://learn.microsoft.com/en-au/azure/search/whats-new",
      "date": "2024-11-19",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Microsoft releases agentic retrieval features, multimodal RAG support, and knowledge base source expansions; demonstrates continued platform investment in domain-specific RAG infrastructure."
    },
    {
      "title": "Seven Failure Points When Engineering a Retrieval Augmented Generation System",
      "url": "https://www.gabormelli.com/RKB/Barnett_et_al.,_2024",
      "date": "2024-11-04",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Experience report identifies seven critical failure points across three case studies (cognitive reviewer, AI tutor, biomedical QA); highlights that RAG validation is only feasible in operation and robustness evolves rather than being designed in."
    },
    {
      "title": "Retrieval-Augmented Generation for Domain-Specific Question Answering about Pittsburgh and Carnegie Mellon University",
      "url": "https://arxiv.org/html/2411.13691v1",
      "date": "2024-11-01",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Case study builds RAG system over 1,800+ web pages with hybrid BM25/FAISS retrieval; F1 improves from 5.45% to 42.21%, demonstrating practical domain-specific application."
    },
    {
      "title": "Leveraging the Domain Adaptation of Retrieval Augmented Generation Models for Question Answering and Reducing Hallucination",
      "url": "https://www.arxiv.org/abs/2410.17783",
      "date": "2024-10-23",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q4",
      "explanation": "Domain adaptation on hotel customer service domain shows fine-tuning improves QA performance and significantly reduces hallucinations across all evaluated RAG architectures."
    },
    {
      "title": "Enhancing Large Language Models with Domain-specific Retrieval Augment Generation: A Case Study on Long-form Consumer Health Question Answering in Ophthalmology",
      "url": "https://arxiv.org/abs/2409.13902",
      "date": "2024-09-20",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Case study of domain-specific RAG for ophthalmology with 70,000 clinical documents; RAG reduced hallucinations to 26.7% (vs. 45.3% without), increased evidence accuracy to 54.5%, demonstrating real-world medical domain deployment."
    },
    {
      "title": "The Challenges of Implementing Retrieval Augmented Generation in Production",
      "url": "https://www.marktechpost.com/2024/08/18/the-challenges-of-implementing-retrieval-augmented-generation-rag-in-production/",
      "date": "2024-08-18",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Analysis of seven production RAG failure points (missing content, incorrect specificity, ranking, context, format, extraction, output); signals deployment barriers requiring extensive testing and fine-tuning beyond basic setup."
    },
    {
      "title": "RAG for Question-Answering for Vocal Training Based on Domain Knowledge",
      "url": "https://scholars.ln.edu.hk/en/publications/rag-for-question-answering-for-vocal-training-based-on-domain-kno/",
      "date": "2024-08-16",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "IEEE BESC 2024 conference paper on domain-specific RAG applied to niche domain (vocal training); demonstrates knowledge segmentation and semantic similarity approaches for highly specialized corpora."
    },
    {
      "title": "Domain-Specific Retrieval-Augmented Generation Using Vector Stores, Knowledge Graphs, and NLP",
      "url": "https://arxiv.org/html/2410.02721v1",
      "date": "2024-07-20",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "SMART-SLIC framework integrates RAG with domain-specific knowledge graphs and vector stores; builds specialized corpora without LLM hallucinations, enabling source attribution and reducing fine-tuning burden."
    },
    {
      "title": "Mindful-RAG: A Study of Points of Failure in Retrieval Augmented Generation",
      "url": "https://arxiv.org/abs/2407.12216",
      "date": "2024-07-16",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "Identifies eight critical failure points in KG-based RAG; proposes Mindful-RAG framework targeting intent-based retrieval and contextual alignment to address fundamental retrieval brittleness in domain-specific systems."
    },
    {
      "title": "Is RAG all you need? A look at the limits of retrieval-augmented generation",
      "url": "https://ep2024.europython.eu/session/is-rag-all-you-need-a-look-at-the-limits-of-retrieval-augmented-generation/",
      "date": "2024-07-10",
      "type": "conference-talk",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q3",
      "explanation": "EuroPython 2024 talk examining RAG limitations and failure modes across domains; argues domain-specific RAG is not one-size-fits-all, with vastly different quality and transparency requirements by use case."
    },
    {
      "title": "Conversational Document Search Using Azure AI Search - ClearPeaks",
      "url": "https://www.clearpeaks.com/conversational-document-search-using-azure-ai-search/?lang=es",
      "date": "2024-06-27",
      "type": "case-study",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Business intelligence deployment of conversational document search using Azure AI Search with hybrid retrieval and OpenAI embeddings; real production system for domain-specific RAG."
    },
    {
      "title": "Azure AI Search vector search scalability limitation (results beyond 1000 items)",
      "url": "https://blog.azure.moe/2024/06/25/azure-ai-search%E3%81%A71000%E4%BB%B6%E7%9B%AE%E4%BB%A5%E9%99%8D%E3%81%AE%E7%B5%90%E6%9E%9C%E3%81%8C%E5%BE%97%E3%82%89%E3%82%8C%E3%81%AA%E3%81%84/",
      "date": "2024-06-25",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Practitioner report of Azure AI Search vector search limitation: returns empty results beyond 1000 items, revealing scalability constraint affecting large-corpus domain-specific RAG deployments."
    },
    {
      "title": "DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation",
      "url": "https://arxiv.org/abs/2406.05654v2",
      "date": "2024-06-09",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "College enrollment domain benchmark identifies six required RAG capabilities; finds existing LLMs struggle with domain-specific questions, confirming need for specialised retrieval-augmented approaches."
    },
    {
      "title": "Retrieval Augmented Generation for Domain-specific Question Answering",
      "url": "https://arxiv.org/abs/2404.14760",
      "date": "2024-04-23",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Adobe products QA system shows fine-tuning retrievers reduces hallucinations and improves generation; demonstrates production-stage domain-specific RAG deployment in corporate environment."
    },
    {
      "title": "Enhancing Q&A with Domain-Specific Fine-Tuning and Iterative Reasoning: A Comparative Study",
      "url": "https://www.arxiv.org/abs/2404.11792",
      "date": "2024-04-17",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Financial filings RAG study shows fine-tuning embedding models and LLMs improves accuracy; iterative reasoning further boosts performance toward human-expert quality on domain-specific corpora."
    },
    {
      "title": "RAG-QA Arena: Evaluating Domain Robustness for Long-Form Question Answering",
      "url": "https://arxiv.org/html/2407.13998v1",
      "date": "2024-04-09",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "LFRQA benchmark across 7 domains finds only 41.3% of top LLM RAG answers preferred over human references, revealing cross-domain robustness as persistent challenge limiting generalization."
    },
    {
      "title": "Announcing updates to Azure AI Search to help organizations build and scale generative AI applications",
      "url": "https://azure.microsoft.com/en-us/blog/announcing-updates-to-azure-ai-search-to-help-organizations-build-and-scale-generative-ai-applications/",
      "date": "2024-04-04",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q2",
      "explanation": "Azure AI Search announces 11x vector index capacity increase, 6x storage increase, 2x throughput improvements at no cost, with Fortune 500 adopters (OpenAI, KPMG, PETRONAS) confirming production deployment at scale."
    },
    {
      "title": "Improving Retrieval for RAG based Question Answering Models on Financial Documents",
      "url": "https://arxiv.org/html/2404.07221v1",
      "date": "2024-03-23",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Domain-specific RAG techniques for financial documents: chunking, query expansion, metadata annotation, re-ranking, and embedding fine-tuning for 10-Ks and earnings transcripts."
    },
    {
      "title": "Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases",
      "url": "https://arxiv.org/abs/2403.10446",
      "date": "2024-03-15",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "End-to-end system design for domain-specific RAG on private knowledge bases (CMU case); demonstrates approach to hallucination mitigation but reveals limitations of fine-tuning with small, skewed datasets."
    },
    {
      "title": "Retrieval-Augmented Generation for AI-Generated Content: A Survey",
      "url": "https://arxiv.org/abs/2402.19473",
      "date": "2024-02-29",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Comprehensive 353-paper survey providing unified taxonomy of RAG foundations, enhancements, and applications; signals research consolidation and maturity in Q1 2024."
    },
    {
      "title": "Issues on RAG application over large data - Microsoft Q&A",
      "url": "https://learn.microsoft.com/en-sg/answers/questions/1603668/issues-on-rag-application-over-large-data",
      "date": "2024-02-29",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Enterprise RAG deployment (SharePoint, 1100 files) shows inconsistent answer retrieval and parameter-tuning difficulties; signals adoption barrier where increasing retrieval volume reduces accuracy."
    },
    {
      "title": "GitHub - hymie122/RAG-Survey: Collecting awesome papers of RAG for AIGC",
      "url": "https://github.com/hymie122/RAG-Survey",
      "date": "2024-02-23",
      "type": "significant-repo",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Community-curated RAG research collection (1.8k stars) supporting the AIGC survey; demonstrates sustained open-source engagement and practitioner interest in RAG taxonomies."
    },
    {
      "title": "Azure AI Search Integrated Vectorization errors - Microsoft Q&A",
      "url": "https://learn.microsoft.com/en-us/answers/questions/1520659/azure-ai-search-integrated-vectorization-errors",
      "date": "2024-01-31",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Deployment barrier: indexing large PDF documents fails with vectorization timeouts; reveals token limits (8,191 tokens/6k words) and need for document chunking in production systems."
    },
    {
      "title": "[Poor results for RAG with Vector/Hybrid/Semantic search] · Issue #40983 · Azure/azure-sdk-for-net",
      "url": "https://github.com/Azure/azure-sdk-for-net/issues/40983",
      "date": "2024-01-03",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2024-Q1",
      "explanation": "Real-world deployment challenge: Azure AI Search RAG system returning inconsistent results across similar queries; highlights fundamental retrieval accuracy barriers in early 2024."
    },
    {
      "title": "Retrieval-Augmented Generation for Large Language Models - A Survey",
      "url": "https://arxiv.org/abs/2312.10997",
      "date": "2023-12-18",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Comprehensive survey consolidating RAG paradigms (Naive, Advanced, Modular), retrieval/generation/augmentation components, and evaluation frameworks; signals academic maturation and recognition of domain-specific information integration benefits."
    },
    {
      "title": "MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation",
      "url": "https://arxiv.org/html/2601.15487v1",
      "date": "2023-10-27",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Framework for generating domain-specific, multimodal, multi-hop QA datasets to evaluate RAG systems; addresses gap in evaluation of specialized technical documents with >2.3 average reasoning hops."
    },
    {
      "title": "Knowledge Corpus Error in Question Answering",
      "url": "https://github.com/xfactlab/emnlp2023-knowledge-corpus-error",
      "date": "2023-10-12",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "EMNLP 2023 study showing corpus incompleteness limits retrieval-based QA; LLM-generated passages outperform retrieved ones by 10-13%, highlighting fundamental limitations of corpus-only RAG without generation augmentation."
    },
    {
      "title": "How do domain-specific chatbots work? A retrieval augmented generation overview",
      "url": "https://news.ycombinator.com/item?id=37261198",
      "date": "2023-08-25",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "Hacker News practitioner discussion highlighting RAG challenges: embedding failures on domain terminology, testing complexity, tool immaturity (LangChain criticism), expensive vector databases; signals implementation barriers."
    },
    {
      "title": "RobustQA: Benchmarking the Robustness of Domain Adaptation for Open-Domain Question Answering",
      "url": "https://virtual2023.aclweb.org/paper_P4722.html",
      "date": "2023-07-12",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "ACL 2023 benchmark across 8 domains (finance, medicine, law) showing ODQA models trained on Wikipedia fail to generalize; demonstrates significant gap in cross-domain robustness for domain-specific RAG."
    },
    {
      "title": "PrimeQA: The Prime Repository for State-of-the-Art Multilingual Question Answering Research and Development",
      "url": "https://research.ibm.com/publications/primeqa-the-prime-repository-for-state-of-the-art-multilingual-question-answering-research-and-development--1",
      "date": "2023-07-09",
      "type": "significant-repo",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H2",
      "explanation": "IBM Research's open-source toolkit supporting domain-specific QA with retrieval and reading comprehension; enables replication of state-of-the-art methods and front-end applications."
    },
    {
      "title": "RAG: Retrieval-Augmented Generation for Knowledge-Intensive Tasks",
      "url": "https://daje0601.tistory.com/359",
      "date": "2023-06-26",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Discussion of the foundational RAG framework combining parametric and non-parametric memory for knowledge-intensive tasks, enabling QA over external corpora."
    },
    {
      "title": "Task-Aware Specialization for Efficient and Robust Dense Retrieval for Open-Domain Question Answering",
      "url": "https://www.microsoft.com/en-us/research/publication/task-aware-specialization-for-efficient-and-robust-dense-retrieval-for-open-domain-question-answering/",
      "date": "2023-06-19",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "ACL 2023 paper demonstrating task-aware specialization in dense retrieval models for efficient open-domain QA, advancing domain-specific retrieval."
    },
    {
      "title": "A Pilot Empirical Study on When and How to Use Knowledge Graphs as Retrieval Augmented Generation",
      "url": "https://arxiv.org/html/2502.20854",
      "date": "2023-06-15",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Empirical study examining KG-RAG effectiveness across diverse domains and scenarios, evaluating knowledge graph quality impact on RAG performance."
    },
    {
      "title": "Retrieval-Augmented Generation - Paper Reading and Discussion",
      "url": "https://arize.com/blog/retrieval-augmented-generation-paper-reading-and-discussion/",
      "date": "2023-06-09",
      "type": "opinion",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Technical analysis of RAG hybrid architecture combining pre-trained parametric models with explicit non-parametric memory for improved knowledge retrieval."
    },
    {
      "title": "Knowledge Base Question Answering (KBQA)",
      "url": "https://docs.deeppavlov.ai/en/1.2.0/features/models/kbqa.html",
      "date": "2023-06-06",
      "type": "product-ga",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "DeepPavlov's KBQA production system enables domain-specific QA over structured knowledge bases like Wikidata, supporting complex multi-step queries."
    },
    {
      "title": "ERNIE-Search: Bridging Cross-Encoder with Dual-Encoder for Open-Domain Question Answering",
      "url": "https://axi.lims.ac.uk/paper/2205.09153",
      "date": "2023-06-05",
      "type": "research-paper",
      "added": "2026-03-20",
      "superseded_by": null,
      "window": "2023-H1",
      "explanation": "Research on neural retrievers using cross-architecture knowledge distillation to improve open-domain QA performance, advancing dense retrieval methods."
    }
  ],
  "tierHistory": [
    {
      "tier": "research",
      "from": "2023-06-01",
      "to": "2024-04-01"
    },
    {
      "tier": "bleeding-edge",
      "from": "2024-04-01",
      "to": "2025-10-01"
    },
    {
      "tier": "leading-edge",
      "from": "2025-10-01",
      "to": null
    }
  ],
  "trendHistory": [
    {
      "trend": "steady",
      "blockerType": null,
      "from": "2026-09-26",
      "to": null
    }
  ],
  "description": "AI that performs retrieval-augmented generation over proprietary domain-specific corpora and answers questions across multiple knowledge bases. Includes specialised embedding and retrieval for technical domains; distinct from enterprise search which targets general internal documentation.",
  "overview": "Domain-specific RAG applies retrieval-augmented generation to specialized knowledge corpora—research databases, technical documentation, proprietary knowledge bases, domain-specific research, and cross-corpus question answering. Unlike general-purpose QA, which retrieves from broad web indexes, domain-specific RAG requires precise embedding and ranking tailored to technical terminology, domain conventions, and structured data formats. The core tension is irreducible engineering complexity: domain-specific systems deliver higher accuracy and relevance for expert queries but demand custom embedding models, corpus curation, meticulous deployment discipline, and continuous production monitoring. By mid-2026, the field has achieved operational maturity: major cloud vendors ship production infrastructure with agentic retrieval for cross-corpus reasoning, practitioners have documented critical failure modes with mitigation patterns, and multiple regulated-industry deployments (legal, medical, financial) confirm viability for carefully scoped, continuously curated applications. However, no generic solution exists. Cross-domain generalization remains structurally brittle—vector search dilution in large corpora, domain-specific architectures that fail on adjacent domains, and retriever-generation misalignment persist as unresolved challenges. Production deployments succeed through discipline: domain-specific embeddings, hybrid retrieval (sparse+dense), cross-encoder reranking, rigorous evaluation frameworks, and acceptance of per-domain tuning burden. Broader adoption remains blocked by this irreducible per-domain engineering investment and the architectural brittleness that affects all approaches equally.",
  "currentLandscape": "By September 2026, domain-specific RAG demonstrates operational viability in narrowly-scoped, carefully curated domains but faces persistent enterprise-scale barriers. Production deployments expand across sectors: healthcare systems validate patient-education conversational agents (reaching 82% accuracy and 100% RAG knowledge hit rates after iterative alpha testing), regulatory compliance (Ontario Power Generation's progressive evidence acquisition pattern formalised as PEA-CAE for energy-board filings), telecommunications support (71% ticket deflection), and scientific document analysis over tables and figures. Architectural innovation accelerates: graph-RAG outperforms vector-only retrieval on multi-document reasoning and cross-document trustworthiness; multimodal retrieval improves structured-document accuracy by extracting and summarising tables and figures; query-conditioned routing adapts indexing and retrieval strategy by query type. Multilingual and culturally grounded RAG emerges as a distinct challenge, with evidence showing embeddings degrade 10-40% on underrepresented languages and that generic retrieval fails on culture-specific facts. However, critical limitations remain unresolved. Benchmark-to-production accuracy gaps persist: systems exceeding 90% on standard benchmarks drop to ~46% on real enterprise documents due to silent parsing failures, superficially similar wrong answers ranking high by semantic similarity, and header/table separation at page breaks. Structured-data handling—queries mixing database records with unstructured text—remains architecturally distinct and underserved. Extraction hallucination in technical domains stays below 40% exact-match accuracy; multilingual boundaries degrade performance 10-23 percentage points. Production success derives from discipline: source-authority governance, corpus curation, domain-specific embeddings, hybrid retrieval (BM25+dense+cross-encoder reranking), per-domain tuning and rigorous evaluation frameworks. Broader adoption remains blocked by irreducible per-domain engineering burden and structural brittleness under corpus evolution.",
  "history": "- **2023-H1:** Foundational RAG research (parametric + non-parametric memory hybrids) gains traction. DeepPavlov releases production-grade KBQA system for structured knowledge bases. ACL 2023 publications showcase dense retrieval specialization and cross-encoder knowledge distillation. Early empirical studies on knowledge-graph RAG effectiveness across domains. No major enterprise deployments yet; practice remains in research and proof-of-concept phase.\n- **2023-H2:** RAG field consolidates with comprehensive surveys and benchmarks. PrimeQA (IBM), RobustQA (multi-domain benchmark), and MiRAGE (evaluation framework for specialized corpora) advance tooling ecosystem. Research reveals persistent gaps: domain adaptation remains challenging across finance, medicine, law; corpus incompleteness limits retrieval-only approaches; embedding models struggle with domain terminology. Practitioner adoption hindered by testing complexity, immature tooling, and high customization costs. Cloud vendors provide RAG templates but target general documentation; specialized domain applications remain high-engineering-overhead endeavours.\n- **2024-Q1:** Cloud vendors (Microsoft Azure) release production RAG tooling and agentic retrieval features, signaling ecosystem maturity. Research community consolidates foundations with comprehensive 353-paper surveys mapping RAG taxonomy. Practitioners immediately encounter real-world deployment barriers: vector search returns inconsistent results; indexing large documents fails at token limits; enterprise systems at scale show degraded accuracy. Domain-specific financial RAG research demonstrates techniques (chunking, query expansion, embedding fine-tuning) but requires significant customization. Signal balance: positive news on tooling and research consolidation offset by concrete evidence of deployment blockers—parameter tuning cannot resolve fundamental retrieval brittleness at enterprise scale.\n- **2024-Q2:** Cloud vendors scale infrastructure aggressively: Azure announces 11x index capacity, 6x storage, 2x throughput improvements with Fortune 500 production adopters (OpenAI, KPMG, PETRONAS). Fine-tuning approaches gain traction—financial RAG and Adobe products demonstrate improved accuracy via embedding fine-tuning, iterative reasoning, and query expansion. However, cross-domain generalization remains poor (41.3% RAG answers beat human refs on multi-domain benchmark). Scalability barriers emerge: Azure vector search fails beyond 1000 items; domain-specific fine-tuning doesn't transfer; retrieval degrades with corpus size. Signal balance: vendor investment and targeted successes offset by evidence of persistent cross-domain brittleness and hidden scalability cliffs.\n- **2024-Q3:** Practitioner deployments mature in niche, well-scoped domains. Ophthalmology case study demonstrates 70K clinical documents with 54.5% accuracy and hallucination reduction via domain-specific RAG; vocal training system shows successful application to ultra-specialized corpora. SMART-SLIC framework advances KG+vector-store approaches to reduce LLM hallucinations without fine-tuning. However, fundamental failure modes persist: eight critical KG-RAG failure points identified (intent understanding, context alignment); EuroPython talk argues RAG success is highly domain-dependent with no universal solution; production deployments reveal seven recurring failure points (missing content, ranking, format, extraction) requiring extensive testing and optimization. By Q3 2024, domain-specific RAG is operationally viable for narrowly-scoped, curated corpora but remains fragile, parameter-sensitive, and failure-prone for broader applications or cross-domain generalization.\n- **2024-Q4:** Cloud vendors release advanced RAG infrastructure—Microsoft announces multimodal RAG, agentic retrieval features, and expanded knowledge base sources (November 2024). Practitioner case studies show strong results in bounded domains: CMU/Pittsburgh domain case study achieves 42.21% F1 (vs 5.45% baseline) with 1,800+ documents; domain-adapted hotel customer service system reduces hallucinations significantly through fine-tuning. Research confirms persistent cross-domain brittleness: EMNLP 2024 benchmark (LFRQA, 26K queries, 7 domains) again finds only 41.3% of RAG answers preferred over human baselines. Critical analysis documents nine RAG limitations (retrieval quality, latency, scalability, transparency, bias, domain brittleness). By year-end, domain-specific RAG achieves reliable results only in narrowly-scoped, carefully curated use cases; broader enterprise deployment remains blocked by parameter sensitivity, scalability cliffs, and fundamental retrieval brittleness. No novel solutions emerged in Q4; vendors invested in infrastructure while practitioners continued workarounds.\n- **2025-Q1:** Ecosystem maturity consolidates with AWS RAG Evaluation GA (March 2025) enabling systematic assessment of domain-specific applications. Real-world deployments expand across domains: automotive industry (CIKM 2025) achieves +1.79 factual correctness improvement from pilot RAG system; financial, biomedical, and cybersecurity corpora demonstrate 31-42% precision gains via token-aware chunking and domain-specific embeddings. Field-service management and business intelligence production deployments confirm practical adoption. However, comparative evaluations reveal limitations: large-scale study (20K queries, 400K KB) shows RAG underperforms DoRA on accuracy and latency in accuracy-critical domains. Practitioner reports document persistent barriers: semantic mismatches, outdated knowledge bases, domain-specific fine-tuning complexity. Optimal chunk sizing varies dramatically by domain (5 tokens for cybersecurity vs. 20+ for finance), requiring extensive per-domain experimentation. Signal balance: platform maturity and demonstrated production deployments offset by evidence that broader enterprise adoption remains constrained by customization intensity and parameter sensitivity. Domain-specific RAG capability proven for carefully scoped, expert-managed domains; generalization barriers unresolved.\n- **2025-Q2:** Enterprise benchmarking validates domain-specific RAG applicability: EKRAG benchmark (May 2025) spans diverse corporate documents; SimRAG (April 2025) demonstrates self-training domain adaptation for science and medicine without labeled data. Competition-validated SIGIR 2025 LiveRAG winner achieves first place on 15M-document cross-corpus evaluation. However, Q2 2025 again surfaces persistent limitations: multi-document evaluation analysis documents hallucinations at scale and faithfulness-helpfulness tensions; KG-RAG failure modes emerge under knowledge incompleteness; production knowledge-base systems show semantic mismatch gaps (knowledge base searches returning 15 irrelevant results for specific queries). Signal balance: enterprise applicability for well-scoped domains confirmed through research benchmarks and vendor feature expansion, but production implementation friction and multi-domain generalization barriers persist. Broader enterprise adoption remains blocked by parameter sensitivity, evaluation complexity, and domain-specific tuning requirements.\n- **2025-Q3:** Domain-specific RAG achieves production maturity with measurable deployment gains: embedding fine-tuning (Coxwave Align, July 2025) delivers 12% accuracy improvement and 6x training speedup; research frameworks (RAGen, QuARK, healthcare KG-RAG) advance automated domain adaptation via semantic chunking and knowledge graph integration (+13% on financial domain). However, critical accuracy plateau emerges: production systems stall at 75% accuracy (25% error unacceptable in finance/telecom/healthcare); vector-only RAG insufficient, graph-based semantic layers required (>95% vs. vector-only on policy domain). Cross-corpus studies reveal model-dependent effectiveness: small LLMs benefit (+22.87%) while large models underperform; no universal corpus selection; intelligent routing remains unsolved. Signal balance: embedding fine-tuning and KG integration enable gains in narrowly-scoped domains, but production accuracy ceiling and corpus selection brittleness block broader adoption in accuracy-critical applications.\n- **2025-Q4:** Platform vendors scale infrastructure investment: Azure AI Search releases agentic retrieval (public preview), AWS GA's RAG Evaluation. Adoption reaches 60% of production AI applications. However, real deployments expose persistent friction: Azure AI Search text splitting silent failures, GPT-4.1 agents inconsistently invoke retrieval tools, configuration precision remains critical. Evaluation becomes tier-1 concern: RAGalyst framework demonstrates context-dependent performance across domains with no universal configuration. Production deployments show architectural variation: local-first RAG (SalesWorx, October 2025) using LlamaIndex and Qdrant achieves stability without vendor lock-in. Hallucination reduction claims (70-90%) coexist with evidence of narrower practical applicability; cross-domain generalization poor; accuracy plateau at 75% persists. Domain-specific RAG achieves mainstream tooling and enterprise scale, but production readiness demands meticulous tuning, careful corpus curation, and acceptance of deployment complexity.\n\n- **2026-Jan:** Critical architectural limitations emerge at scale: CorpusQA benchmark reveals standard RAG collapse on 10M token corpora (1.20% vs. 55.85% at 128K), driving bifurcation into caching for small corpora and agentic hypergraphs for complex reasoning. Domain-specific deployments (immunogenicity, policy QA) expose persistent retriever-generation misalignment despite coherent LLM outputs. Transferable RAG frameworks demonstrate cross-domain gains via domain routing but confirm domain-specific fine-tuning remains essential. Field consensus solidifies: monolithic RAG is obsolete; architecture must match corpus scale and reasoning complexity. Evaluation and risk assessment become mission-critical in high-stakes domains (healthcare, finance) where \"lost in the middle\" and outdated knowledge bases create tangible deployment risks.\n\n- **2026-Feb:** Platform vendors accelerate infrastructure: Azure adds agentic retrieval portal support with expanded knowledge sources (OneLake, SharePoint, Web). Practitioner deployments succeed in narrowly-scoped domains with specialized architectures: product catalogs, resume search, field service management via caching/vector search; security maturity advances with sparse attention defenses against knowledge poisoning attacks. However, critical limitations surface: inferential QA research reveals current pipelines fail at indirect evidence reasoning; documented real-world failures expose tangible risks—healthcare harmful guidance from outdated KBs, finance compliance violations, telecom $2.3M in service credits for incorrect RAG outputs. Retriever-generation misalignment persists as unresolved deployment challenge. Field consensus crystallizes around realistic assessment: domain-specific RAG delivers value only in carefully curated, continuously managed applications; broader adoption remains blocked by architectural brittleness, retriever reliability gaps in specialized domains, and irreducible per-domain tuning burden.\n- **2026-Mar:** EACL 2026 introduced the UDCG evaluation metric achieving 36% improvement in cross-domain QA correlation, while CrossRAG advanced multilingual retrieval across fragmented corpora. Amazon Bedrock Knowledge Adaptive QA and Meta AI/Google DeepMind's ART unsupervised dense retriever training pushed deployment infrastructure forward in regulated domains (medical, legal, financial). A production GraphRAG deployment at scale reduced hallucination by 72% serving millions of interactions; a multi-stage memory retrieval system achieved 96-100% single-session accuracy on cross-corpus QA at 115K+ token scale. These advances address architectural brittleness incrementally but do not resolve the core production accuracy ceiling; per-domain tuning burden and retriever-generation misalignment remain the primary adoption constraints.\n\n- **2026-Apr:** Microsoft released Octen-Embedding-0.6B, a domain-adapted model for legal, finance, healthcare, and code that outperforms larger generic embeddings (0.7241 vs. 0.7139 on domain benchmarks), signalling an ecosystem shift toward specialized embeddings as standard practice. BEIR benchmark data quantified the domain-specialization advantage: fine-tuned embeddings improve retrieval by medical +29% (48%→62%), code +34% (44%→59%), legal +24% (46%→57%)—confirming general models systematically fail on specialized domains. A practitioner postmortem found that document quality scoring alone improved search accuracy from 62% to 89% with no embedding or retrieval changes, establishing corpus curation as a first-order variable. Production failure modes gained sharper documentation: Azure AI Search analysis named three causes of silent vector drift (embedding model version mismatch, incremental corpus updates without re-embedding, inconsistent chunking), and a cross-industry postmortem of 12+ RAG deployments showed naive chunking and wrong embeddings as primary precision failures—fixed from 54% to 81% precision with domain-adapted approaches. Practitioner evidence confirmed the viability threshold: RAG is justified for corpora above 500-1000 items with domain-specific entities but overengineered for smaller homogeneous datasets, reinforcing that domain-specific RAG succeeds only with meticulous specialization and continuous index governance. Peer-reviewed production deployments demonstrate viability in high-stakes domains: Uber's Genie on-call copilot for security/privacy policies achieved 27% relative improvement in acceptable answers and 60% reduction in incorrect advice via agentic RAG; legal RAG system in Brazilian Portuguese (PROPOR 2026) deployed with metrics across 184,895 audited answers (81.7% legislation resolution, 47.1% jurisprudence resolution, 6.5% hallucination correction rate); systematic evaluation of medical RAG pipeline components (retrieval strategies, domain-specific embeddings) published peer-reviewed findings. SIGIR 2026 benchmark (ConflictQA) exposed cross-corpus QA brittleness: knowledge conflicts between textual documents and knowledge graphs reduce accuracy without explicit handling. April 2026 evidence synthesis: domain-specific RAG deployment succeeds across legal, medical, and security domains when architecture, embeddings, and corpus curation are meticulous; production readiness demands continuous monitoring (index drift), systematic evaluation (domain-specific benchmarks), and acceptance that generalization failures are structural, not fixable via retrieval optimization alone. Field consensus: domain-specific RAG is operationally mature for narrowly-scoped, high-stakes domains but remains a disciplined engineering practice, not a generic solution.\n- **2026-May:** Retrieval infrastructure research and documented failure modes reinforced the discipline-over-defaults thesis. Biomedical RAG benchmarking (BioASQ, 5 retrieval strategies) confirmed cross-encoder reranking as the empirically superior retrieval approach (0.827 composite score, 0.852 contextual precision), establishing domain-specific strategy selection as mandatory rather than optional. Health-system-scale semantic search deployment—1.68M patients, 166M clinical notes, Qwen3 embeddings, 94.6% clinical QA accuracy at 237ms latency and $4K/month—demonstrated production viability at genuine enterprise scale. May 2026 research synthesis reinforced critical findings: (1) peer-reviewed clinical embedding benchmark showed domain context variables (corpus type, query format) explain 49% variance in retrieval performance, rivaling model choice (47.6%) and revealing that MTEB leaderboard rankings do not transfer to specialized clinical domains; (2) industrial automotive QA study (AAAI 2026) confirmed RAG as the most cost-efficient and effective adaptation method for domain-specific closed-domain QA across both premium and open-source models; (3) medical domain case study across 6 specialties demonstrated hybrid BM25+dense retrieval achieving 86% recall@5 vs 71% for BM25 alone (p<0.001 statistical significance), confirming complementary strengths; (4) leading-edge agentic retrieval research (BRIGHT-Pro benchmark) introduced aspect-aware evaluation and multi-step evidence portfolio construction for cross-corpus reasoning, shifting paradigm from single-shot relevance matching to iterative evidence gathering. Simultaneously, OCR robustness benchmarking revealed that high-accuracy OCR still causes downstream RAG failures due to structural and semantic errors across 11 document types, while a practitioner case study documented silent embedding model mismatch (ONNX-quantized indexing vs. API queries) degrading retrieval below similarity thresholds for months without detection. Critical production audit of 50 live deployments exposed seven failure modes with high prevalence (context gaslighting 76% finance, citation fabrication 81% legal, low-confidence drift 62% medical) under adversarial testing, reinforcing that sophisticated architectures and evaluation frameworks remain insufficient without rigorous deployment discipline. The consistent signal: domain-specific RAG delivers measurable results in industrial, medical, and financial domains when retrieval strategy, embedding consistency, corpus structure, and agentic reasoning patterns are treated as co-equal engineering concerns rather than defaults; production readiness demands meticulous tuning, continuous evaluation, and acceptance of fundamental cross-corpus brittleness.\n\n- **2026-Jun:** Domain-specific RAG maturation reaches full operational consensus. Multiple production deployments confirm specialized domain viability: materials science RAG achieved 85.6% accuracy (4x baseline improvement); clinical off-guideline QA showed 56%→82% accuracy jump; Goodyear's agentic TechGraphRAG integrated 2,100+ academic papers with internal knowledge graphs; LinkedIn Hiring Assistant (1.3B+ profiles) achieved +24% Kappa quality improvement and −5% false positive rate via custom MUSE domain embeddings at billion scale; a Russian corporate document system improved Top-1 from 62% to 88% via hybrid BM25+vector RRF search and cross-encoder reranking. Vendor infrastructure advanced materially: Azure Build 2026 Foundry IQ Serverless agentic retrieval delivered 46–54% evidence recall improvement and 34% token cost reduction; Google Gemini Enterprise shipped agentic RAG in public preview achieving 34% factuality improvement over standard RAG with FramesQA benchmark at 90.1% on multi-corpus scenarios. Critical benchmarking and failure-mode research sharpened the discipline-over-defaults thesis: CQC-RAG demonstrated semantically equivalent queries retrieve different results in multi-document corpora (15–40% degradation), addressable via cross-query consistency filtering; vector search dilution at corpus scale (75% accuracy at 54 documents degrading to <40% at 1,128 documents) validated domain-scoped retrieval as solution; legal domain analysis identified three structural architectural pathologies (mereological, diachronic, causal blindness) that are not fixable via retrieval tuning alone—applicable to any regulated domain requiring hierarchy and temporal reasoning. Negative signals matched positive: Stanford documented legal hallucinations persisting at 17-34% despite RAG (58-82% without); Fortune 500 post-mortems revealed catastrophic production failures ($12.2M trading loss, $4.1M legal settlement) when RAG removed; enterprise root-cause analysis traced failures to retrieval layer collapse under multi-hop dependencies and conflicting regional policies, not generation model quality. Consensus unchanged: no generic solution exists; success requires domain-adapted embeddings, hybrid retrieval, cross-encoder reranking, and continuous corpus governance as co-equal engineering concerns.\n- **2026-Jul:** Agentic RAG infrastructure matured to GA on two fronts: Google Gemini Enterprise Agentic RAG (Sufficient Context Agent, 93% sufficiency-classification accuracy) achieved 34% factuality improvement over standard RAG and 90.1% accuracy on cross-corpus scenarios, and Databricks AI Search added native retrieval quality evaluation with LLM-as-judge scoring as production infrastructure. Microsoft MLflow RAG Agents demonstrated 89% hallucination reduction from 60% to under 7% via query decomposition and self-reflection patterns. Research reinforced structural failure modes: four architectural failure modes (similarity does not equal correctness, synthesis hallucination, meaning re-derived at query time, temporal unawareness) were documented as blocking broader agentic adoption; Amazon GaRAGe benchmark with 35K+ human-curated grounding annotations across 2,366 questions confirmed Tier-1 vendor investment in rigorous domain-specific RAG evaluation. Agentic multi-hop RAG ablation validated that fixed hybrid retrieval outperforms adaptive routing by +1.8 EM, and two iterations capture 95% of gains—pointing toward simpler, more reliable production architectures. A US Navy warfare center AWS GovCloud deployment confirmed regulated-domain viability with full NIST/DoD security controls. Vendor infrastructure GA accelerated further: Microsoft Foundry IQ reached unified-knowledge-platform GA with serverless agentic retrieval (46-54% evidence-recall improvement, 34% token-cost reduction), and an 8-industry production survey (banking, insurance, healthcare, legal, SaaS, finance, manufacturing, R&D) confirmed domain-specific maturity—banking >95% citation precision, insurance 70%+ touchless claims processing, healthcare >95% faithfulness—at the cost of 3-10x higher token spend versus traditional RAG. A peer-reviewed taxonomy of 33 RAG failure modes found 12 modes, including all 8 agentic-orchestration modes, still lack empirical evidence, marking agentic orchestration as the field's active research frontier. CHARLIE, a multi-agent on-premise RAG system, went into production at Brazil's Federal Police Forensic Institute for cross-corpus evidential reasoning, extending regulated-domain viability beyond the US Navy deployment. Commentary consolidated around a broader architectural shift from vector-DB-first RAG toward permission-aware agentic search, while a governance-focused analysis catalogued silent enterprise failure modes—embedding drift, access-control desynchronization, and hallucination from stale context—that produce no error signal until user-visible impact occurs. Late-July evidence (through July 29) sharpened the multi-hop reliability problem specifically: a UC Berkeley/Google DeepMind benchmark found 92% of enterprise RAG configurations fail multi-hop queries (only 8% fully correct across 10,000 questions, split 63% incomplete-retrieval-chain failures and 41% hallucinated bridges), while Microsoft's LakeQuest benchmark identified relation-chaining and policy-grounding failures across heterogeneous data lakes and a security study found multi-hop RAG agents vulnerable to salience-induction attacks with an 83.3% success rate despite all underlying facts remaining true. Mitigation evidence accumulated alongside: MIT/Microsoft's CoRe-RAG cut multi-hop failure rates from 71% to 23% via specialized agent roles with closed-loop verification, Cortex's production deployment across 60+ technical sources answered 18,000+ queries at 86% certainty, and a governance pattern for safe embedding-model upgrades (dual-index migration) addressed silent coordinate-system breakage in production RAG.\n\n- **2026-Aug:** Evidence through August 12 confirms maturity bifurcation: production deployments thrive in narrow, well-curated domains (Xiaohongshu: 5,300+ data warehouse with knowledge-graph-guided retrieval achieved 96.6% Hit@10 recall; operations/library systems showing adoption) while enterprise-wide deployment barriers persist. Rigorous benchmarking (Wang et al., 28-tier scaling study across 1.7M–600M tokens) established BM25 as scalable default at enterprise scale, overtaking agentic approaches around 10M tokens and maintaining 20-point accuracy margin—contradicting vector-first assumptions. Domain-specific research (SciRet on CORD-19 scientific corpus) validated hybrid retrieval achieving Recall@10 1.0, but generic cross-encoder rerankers trained on web data degrade scientific precision, reinforcing that domain adaptation is architectural requirement, not optional. Cloud infrastructure GA milestones: Databricks ai_search() SQL function (batch RAG pipelines with multi-source deduplication and synthesis), Neo4j GraphRAG context provider (entity-relationship traversal), and Azure Foundry knowledge bases signal ecosystem convergence toward managed domain-specific retrieval as standard enterprise capability. However, agentic orchestration remains operationally fragile: new research (Roh & Han, 12,000 trajectory analysis) identified pre-evidence procedural failures in multi-hop reasoning where agents retrieve evidence but skip reading before finalization—Read-Gate invariant improves accuracy 14.9–19.9 points, suggesting control discipline rather than model capability is limiting factor. Counter-signal: production failure audit of 143 enterprise RAG deployments found 73% experienced critical failure within first quarter (41% undetected by standard evaluations), with cross-document entity resolution, temporal drift, and multi-hop gaps cited as persistent adoption blockers. Field consensus remains unchanged: domain-specific RAG achieves viability only within disciplined, narrowly-scoped domains with meticulous tuning and continuous governance; cross-domain generalization and large-scale enterprise deployment remain structurally constrained by architecture brittleness, retriever reliability in specialized contexts, and operational complexity of maintaining embedding consistency and corpus freshness.\n- **2026-Aug (late):** Production deployment breadth widened further while reliability research sharpened known limits. Named regulated-sector deployments matured: AIG cut SLA 5x+ with accuracy improving 75%→90%, Allianz deployed autonomous motor/health claims workflows, and Travelers scaled to 10K engineers; Mistral's Agentic Search GA replaced one-shot RAG with a multi-step retrieval loop, lifting FinanceBench accuracy from 26.7% to 86%. Domain-specific deployments also landed in materials R&D (Chemcopilot, 1,000+ scientific documents), IT operations (Accenture Japan ticket triage), and agricultural advisory (peer-reviewed five-architecture Bengali evaluation showing dense retrieval collapsing 10x by query type). A preregistered audit documented a new production fragility mode—corpus expansion causing silent answer churn (6-10pp excess) despite fixed models and prompts—reinforcing that compatibility discipline, not just retrieval architecture, gates production RAG reliability.\n- **2026-Sep:** Domain-adaptation evidence consolidated across sectors. UCSF deployed agentic RAG over a 7M+ subject EHR fused with a curated biomedical knowledge graph, improving GPT-5.5 and Claude Opus accuracy on 100 clinical research tasks, while a new FinRAG-QA benchmark (999 practitioner-curated Q&A on 209 regulatory documents) showed domain-adapted embeddings and reranking lifting NDCG by 38.8 points and generation accuracy by 34.4pp over generic pipelines—reinforcing evidence that fine-tuned embeddings improve domain retrieval precision by 20-40% over off-the-shelf models. Architectural limits sharpened: production case studies documented corpus-contradiction failures (outdated documents outranking current ones via semantic similarity) and structured-data brittleness (vector-only RAG at 40-60% accuracy versus 90%+ for hybrid structured-query approaches), while graph RAG outperformed vector-hybrid architectures on trustworthiness metrics in a government project-management benchmark. Industry skepticism hardened in parallel: an analysis citing 72% of enterprise RAG implementations failing within their first year and Gartner's forecast of 40% agentic AI project cancellations by 2027 argued for context-graph architectures over traditional RAG, even as a production financial-RAG case study demonstrated hallucination reduction from 15% to 1% through systematic pipeline debugging (context reordering, claim verification, deterministic math)—underscoring that domain-specific RAG remains viable only with sustained architectural and governance investment. A Salesforce production case study showed RAG accuracy dropping from 90%+ on benchmarks to 46% on real enterprise documents, recovered to 84.4% only after intelligent parsing and reranking—reinforcing the benchmark-production gap. Elsewhere, telecom and utility deployments (hybrid KG-RAG at 89.6% grounding; Ontario Power Generation's progressive-evidence compliance pattern) showed production hardening, while research found exact-span evidence extraction below 40% accuracy in technical domains.",
  "historyEntries": [
    {
      "period": "2023-H1",
      "text": "Foundational RAG research (parametric + non-parametric memory hybrids) gains traction. DeepPavlov releases production-grade KBQA system for structured knowledge bases. ACL 2023 publications showcase dense retrieval specialization and cross-encoder knowledge distillation. Early empirical studies on knowledge-graph RAG effectiveness across domains. No major enterprise deployments yet; practice remains in research and proof-of-concept phase."
    },
    {
      "period": "2023-H2",
      "text": "RAG field consolidates with comprehensive surveys and benchmarks. PrimeQA (IBM), RobustQA (multi-domain benchmark), and MiRAGE (evaluation framework for specialized corpora) advance tooling ecosystem. Research reveals persistent gaps: domain adaptation remains challenging across finance, medicine, law; corpus incompleteness limits retrieval-only approaches; embedding models struggle with domain terminology. Practitioner adoption hindered by testing complexity, immature tooling, and high customization costs. Cloud vendors provide RAG templates but target general documentation; specialized domain applications remain high-engineering-overhead endeavours."
    },
    {
      "period": "2024-Q1",
      "text": "Cloud vendors (Microsoft Azure) release production RAG tooling and agentic retrieval features, signaling ecosystem maturity. Research community consolidates foundations with comprehensive 353-paper surveys mapping RAG taxonomy. Practitioners immediately encounter real-world deployment barriers: vector search returns inconsistent results; indexing large documents fails at token limits; enterprise systems at scale show degraded accuracy. Domain-specific financial RAG research demonstrates techniques (chunking, query expansion, embedding fine-tuning) but requires significant customization. Signal balance: positive news on tooling and research consolidation offset by concrete evidence of deployment blockers—parameter tuning cannot resolve fundamental retrieval brittleness at enterprise scale."
    },
    {
      "period": "2024-Q2",
      "text": "Cloud vendors scale infrastructure aggressively: Azure announces 11x index capacity, 6x storage, 2x throughput improvements with Fortune 500 production adopters (OpenAI, KPMG, PETRONAS). Fine-tuning approaches gain traction—financial RAG and Adobe products demonstrate improved accuracy via embedding fine-tuning, iterative reasoning, and query expansion. However, cross-domain generalization remains poor (41.3% RAG answers beat human refs on multi-domain benchmark). Scalability barriers emerge: Azure vector search fails beyond 1000 items; domain-specific fine-tuning doesn't transfer; retrieval degrades with corpus size. Signal balance: vendor investment and targeted successes offset by evidence of persistent cross-domain brittleness and hidden scalability cliffs."
    },
    {
      "period": "2024-Q3",
      "text": "Practitioner deployments mature in niche, well-scoped domains. Ophthalmology case study demonstrates 70K clinical documents with 54.5% accuracy and hallucination reduction via domain-specific RAG; vocal training system shows successful application to ultra-specialized corpora. SMART-SLIC framework advances KG+vector-store approaches to reduce LLM hallucinations without fine-tuning. However, fundamental failure modes persist: eight critical KG-RAG failure points identified (intent understanding, context alignment); EuroPython talk argues RAG success is highly domain-dependent with no universal solution; production deployments reveal seven recurring failure points (missing content, ranking, format, extraction) requiring extensive testing and optimization. By Q3 2024, domain-specific RAG is operationally viable for narrowly-scoped, curated corpora but remains fragile, parameter-sensitive, and failure-prone for broader applications or cross-domain generalization."
    },
    {
      "period": "2024-Q4",
      "text": "Cloud vendors release advanced RAG infrastructure—Microsoft announces multimodal RAG, agentic retrieval features, and expanded knowledge base sources (November 2024). Practitioner case studies show strong results in bounded domains: CMU/Pittsburgh domain case study achieves 42.21% F1 (vs 5.45% baseline) with 1,800+ documents; domain-adapted hotel customer service system reduces hallucinations significantly through fine-tuning. Research confirms persistent cross-domain brittleness: EMNLP 2024 benchmark (LFRQA, 26K queries, 7 domains) again finds only 41.3% of RAG answers preferred over human baselines. Critical analysis documents nine RAG limitations (retrieval quality, latency, scalability, transparency, bias, domain brittleness). By year-end, domain-specific RAG achieves reliable results only in narrowly-scoped, carefully curated use cases; broader enterprise deployment remains blocked by parameter sensitivity, scalability cliffs, and fundamental retrieval brittleness. No novel solutions emerged in Q4; vendors invested in infrastructure while practitioners continued workarounds."
    },
    {
      "period": "2025-Q1",
      "text": "Ecosystem maturity consolidates with AWS RAG Evaluation GA (March 2025) enabling systematic assessment of domain-specific applications. Real-world deployments expand across domains: automotive industry (CIKM 2025) achieves +1.79 factual correctness improvement from pilot RAG system; financial, biomedical, and cybersecurity corpora demonstrate 31-42% precision gains via token-aware chunking and domain-specific embeddings. Field-service management and business intelligence production deployments confirm practical adoption. However, comparative evaluations reveal limitations: large-scale study (20K queries, 400K KB) shows RAG underperforms DoRA on accuracy and latency in accuracy-critical domains. Practitioner reports document persistent barriers: semantic mismatches, outdated knowledge bases, domain-specific fine-tuning complexity. Optimal chunk sizing varies dramatically by domain (5 tokens for cybersecurity vs. 20+ for finance), requiring extensive per-domain experimentation. Signal balance: platform maturity and demonstrated production deployments offset by evidence that broader enterprise adoption remains constrained by customization intensity and parameter sensitivity. Domain-specific RAG capability proven for carefully scoped, expert-managed domains; generalization barriers unresolved."
    },
    {
      "period": "2025-Q2",
      "text": "Enterprise benchmarking validates domain-specific RAG applicability: EKRAG benchmark (May 2025) spans diverse corporate documents; SimRAG (April 2025) demonstrates self-training domain adaptation for science and medicine without labeled data. Competition-validated SIGIR 2025 LiveRAG winner achieves first place on 15M-document cross-corpus evaluation. However, Q2 2025 again surfaces persistent limitations: multi-document evaluation analysis documents hallucinations at scale and faithfulness-helpfulness tensions; KG-RAG failure modes emerge under knowledge incompleteness; production knowledge-base systems show semantic mismatch gaps (knowledge base searches returning 15 irrelevant results for specific queries). Signal balance: enterprise applicability for well-scoped domains confirmed through research benchmarks and vendor feature expansion, but production implementation friction and multi-domain generalization barriers persist. Broader enterprise adoption remains blocked by parameter sensitivity, evaluation complexity, and domain-specific tuning requirements."
    },
    {
      "period": "2025-Q3",
      "text": "Domain-specific RAG achieves production maturity with measurable deployment gains: embedding fine-tuning (Coxwave Align, July 2025) delivers 12% accuracy improvement and 6x training speedup; research frameworks (RAGen, QuARK, healthcare KG-RAG) advance automated domain adaptation via semantic chunking and knowledge graph integration (+13% on financial domain). However, critical accuracy plateau emerges: production systems stall at 75% accuracy (25% error unacceptable in finance/telecom/healthcare); vector-only RAG insufficient, graph-based semantic layers required (>95% vs. vector-only on policy domain). Cross-corpus studies reveal model-dependent effectiveness: small LLMs benefit (+22.87%) while large models underperform; no universal corpus selection; intelligent routing remains unsolved. Signal balance: embedding fine-tuning and KG integration enable gains in narrowly-scoped domains, but production accuracy ceiling and corpus selection brittleness block broader adoption in accuracy-critical applications."
    },
    {
      "period": "2025-Q4",
      "text": "Platform vendors scale infrastructure investment: Azure AI Search releases agentic retrieval (public preview), AWS GA's RAG Evaluation. Adoption reaches 60% of production AI applications. However, real deployments expose persistent friction: Azure AI Search text splitting silent failures, GPT-4.1 agents inconsistently invoke retrieval tools, configuration precision remains critical. Evaluation becomes tier-1 concern: RAGalyst framework demonstrates context-dependent performance across domains with no universal configuration. Production deployments show architectural variation: local-first RAG (SalesWorx, October 2025) using LlamaIndex and Qdrant achieves stability without vendor lock-in. Hallucination reduction claims (70-90%) coexist with evidence of narrower practical applicability; cross-domain generalization poor; accuracy plateau at 75% persists. Domain-specific RAG achieves mainstream tooling and enterprise scale, but production readiness demands meticulous tuning, careful corpus curation, and acceptance of deployment complexity."
    },
    {
      "period": "2026-Jan",
      "text": "Critical architectural limitations emerge at scale: CorpusQA benchmark reveals standard RAG collapse on 10M token corpora (1.20% vs. 55.85% at 128K), driving bifurcation into caching for small corpora and agentic hypergraphs for complex reasoning. Domain-specific deployments (immunogenicity, policy QA) expose persistent retriever-generation misalignment despite coherent LLM outputs. Transferable RAG frameworks demonstrate cross-domain gains via domain routing but confirm domain-specific fine-tuning remains essential. Field consensus solidifies: monolithic RAG is obsolete; architecture must match corpus scale and reasoning complexity. Evaluation and risk assessment become mission-critical in high-stakes domains (healthcare, finance) where \"lost in the middle\" and outdated knowledge bases create tangible deployment risks."
    },
    {
      "period": "2026-Feb",
      "text": "Platform vendors accelerate infrastructure: Azure adds agentic retrieval portal support with expanded knowledge sources (OneLake, SharePoint, Web). Practitioner deployments succeed in narrowly-scoped domains with specialized architectures: product catalogs, resume search, field service management via caching/vector search; security maturity advances with sparse attention defenses against knowledge poisoning attacks. However, critical limitations surface: inferential QA research reveals current pipelines fail at indirect evidence reasoning; documented real-world failures expose tangible risks—healthcare harmful guidance from outdated KBs, finance compliance violations, telecom $2.3M in service credits for incorrect RAG outputs. Retriever-generation misalignment persists as unresolved deployment challenge. Field consensus crystallizes around realistic assessment: domain-specific RAG delivers value only in carefully curated, continuously managed applications; broader adoption remains blocked by architectural brittleness, retriever reliability gaps in specialized domains, and irreducible per-domain tuning burden."
    },
    {
      "period": "2026-Mar",
      "text": "EACL 2026 introduced the UDCG evaluation metric achieving 36% improvement in cross-domain QA correlation, while CrossRAG advanced multilingual retrieval across fragmented corpora. Amazon Bedrock Knowledge Adaptive QA and Meta AI/Google DeepMind's ART unsupervised dense retriever training pushed deployment infrastructure forward in regulated domains (medical, legal, financial). A production GraphRAG deployment at scale reduced hallucination by 72% serving millions of interactions; a multi-stage memory retrieval system achieved 96-100% single-session accuracy on cross-corpus QA at 115K+ token scale. These advances address architectural brittleness incrementally but do not resolve the core production accuracy ceiling; per-domain tuning burden and retriever-generation misalignment remain the primary adoption constraints."
    },
    {
      "period": "2026-Apr",
      "text": "Microsoft released Octen-Embedding-0.6B, a domain-adapted model for legal, finance, healthcare, and code that outperforms larger generic embeddings (0.7241 vs. 0.7139 on domain benchmarks), signalling an ecosystem shift toward specialized embeddings as standard practice. BEIR benchmark data quantified the domain-specialization advantage: fine-tuned embeddings improve retrieval by medical +29% (48%→62%), code +34% (44%→59%), legal +24% (46%→57%)—confirming general models systematically fail on specialized domains. A practitioner postmortem found that document quality scoring alone improved search accuracy from 62% to 89% with no embedding or retrieval changes, establishing corpus curation as a first-order variable. Production failure modes gained sharper documentation: Azure AI Search analysis named three causes of silent vector drift (embedding model version mismatch, incremental corpus updates without re-embedding, inconsistent chunking), and a cross-industry postmortem of 12+ RAG deployments showed naive chunking and wrong embeddings as primary precision failures—fixed from 54% to 81% precision with domain-adapted approaches. Practitioner evidence confirmed the viability threshold: RAG is justified for corpora above 500-1000 items with domain-specific entities but overengineered for smaller homogeneous datasets, reinforcing that domain-specific RAG succeeds only with meticulous specialization and continuous index governance. Peer-reviewed production deployments demonstrate viability in high-stakes domains: Uber's Genie on-call copilot for security/privacy policies achieved 27% relative improvement in acceptable answers and 60% reduction in incorrect advice via agentic RAG; legal RAG system in Brazilian Portuguese (PROPOR 2026) deployed with metrics across 184,895 audited answers (81.7% legislation resolution, 47.1% jurisprudence resolution, 6.5% hallucination correction rate); systematic evaluation of medical RAG pipeline components (retrieval strategies, domain-specific embeddings) published peer-reviewed findings. SIGIR 2026 benchmark (ConflictQA) exposed cross-corpus QA brittleness: knowledge conflicts between textual documents and knowledge graphs reduce accuracy without explicit handling. April 2026 evidence synthesis: domain-specific RAG deployment succeeds across legal, medical, and security domains when architecture, embeddings, and corpus curation are meticulous; production readiness demands continuous monitoring (index drift), systematic evaluation (domain-specific benchmarks), and acceptance that generalization failures are structural, not fixable via retrieval optimization alone. Field consensus: domain-specific RAG is operationally mature for narrowly-scoped, high-stakes domains but remains a disciplined engineering practice, not a generic solution."
    },
    {
      "period": "2026-May",
      "text": "Retrieval infrastructure research and documented failure modes reinforced the discipline-over-defaults thesis. Biomedical RAG benchmarking (BioASQ, 5 retrieval strategies) confirmed cross-encoder reranking as the empirically superior retrieval approach (0.827 composite score, 0.852 contextual precision), establishing domain-specific strategy selection as mandatory rather than optional. Health-system-scale semantic search deployment—1.68M patients, 166M clinical notes, Qwen3 embeddings, 94.6% clinical QA accuracy at 237ms latency and $4K/month—demonstrated production viability at genuine enterprise scale. May 2026 research synthesis reinforced critical findings: (1) peer-reviewed clinical embedding benchmark showed domain context variables (corpus type, query format) explain 49% variance in retrieval performance, rivaling model choice (47.6%) and revealing that MTEB leaderboard rankings do not transfer to specialized clinical domains; (2) industrial automotive QA study (AAAI 2026) confirmed RAG as the most cost-efficient and effective adaptation method for domain-specific closed-domain QA across both premium and open-source models; (3) medical domain case study across 6 specialties demonstrated hybrid BM25+dense retrieval achieving 86% recall@5 vs 71% for BM25 alone (p<0.001 statistical significance), confirming complementary strengths; (4) leading-edge agentic retrieval research (BRIGHT-Pro benchmark) introduced aspect-aware evaluation and multi-step evidence portfolio construction for cross-corpus reasoning, shifting paradigm from single-shot relevance matching to iterative evidence gathering. Simultaneously, OCR robustness benchmarking revealed that high-accuracy OCR still causes downstream RAG failures due to structural and semantic errors across 11 document types, while a practitioner case study documented silent embedding model mismatch (ONNX-quantized indexing vs. API queries) degrading retrieval below similarity thresholds for months without detection. Critical production audit of 50 live deployments exposed seven failure modes with high prevalence (context gaslighting 76% finance, citation fabrication 81% legal, low-confidence drift 62% medical) under adversarial testing, reinforcing that sophisticated architectures and evaluation frameworks remain insufficient without rigorous deployment discipline. The consistent signal: domain-specific RAG delivers measurable results in industrial, medical, and financial domains when retrieval strategy, embedding consistency, corpus structure, and agentic reasoning patterns are treated as co-equal engineering concerns rather than defaults; production readiness demands meticulous tuning, continuous evaluation, and acceptance of fundamental cross-corpus brittleness."
    },
    {
      "period": "2026-Jun",
      "text": "Domain-specific RAG maturation reaches full operational consensus. Multiple production deployments confirm specialized domain viability: materials science RAG achieved 85.6% accuracy (4x baseline improvement); clinical off-guideline QA showed 56%→82% accuracy jump; Goodyear's agentic TechGraphRAG integrated 2,100+ academic papers with internal knowledge graphs; LinkedIn Hiring Assistant (1.3B+ profiles) achieved +24% Kappa quality improvement and −5% false positive rate via custom MUSE domain embeddings at billion scale; a Russian corporate document system improved Top-1 from 62% to 88% via hybrid BM25+vector RRF search and cross-encoder reranking. Vendor infrastructure advanced materially: Azure Build 2026 Foundry IQ Serverless agentic retrieval delivered 46–54% evidence recall improvement and 34% token cost reduction; Google Gemini Enterprise shipped agentic RAG in public preview achieving 34% factuality improvement over standard RAG with FramesQA benchmark at 90.1% on multi-corpus scenarios. Critical benchmarking and failure-mode research sharpened the discipline-over-defaults thesis: CQC-RAG demonstrated semantically equivalent queries retrieve different results in multi-document corpora (15–40% degradation), addressable via cross-query consistency filtering; vector search dilution at corpus scale (75% accuracy at 54 documents degrading to <40% at 1,128 documents) validated domain-scoped retrieval as solution; legal domain analysis identified three structural architectural pathologies (mereological, diachronic, causal blindness) that are not fixable via retrieval tuning alone—applicable to any regulated domain requiring hierarchy and temporal reasoning. Negative signals matched positive: Stanford documented legal hallucinations persisting at 17-34% despite RAG (58-82% without); Fortune 500 post-mortems revealed catastrophic production failures ($12.2M trading loss, $4.1M legal settlement) when RAG removed; enterprise root-cause analysis traced failures to retrieval layer collapse under multi-hop dependencies and conflicting regional policies, not generation model quality. Consensus unchanged: no generic solution exists; success requires domain-adapted embeddings, hybrid retrieval, cross-encoder reranking, and continuous corpus governance as co-equal engineering concerns."
    },
    {
      "period": "2026-Jul",
      "text": "Agentic RAG infrastructure matured to GA on two fronts: Google Gemini Enterprise Agentic RAG (Sufficient Context Agent, 93% sufficiency-classification accuracy) achieved 34% factuality improvement over standard RAG and 90.1% accuracy on cross-corpus scenarios, and Databricks AI Search added native retrieval quality evaluation with LLM-as-judge scoring as production infrastructure. Microsoft MLflow RAG Agents demonstrated 89% hallucination reduction from 60% to under 7% via query decomposition and self-reflection patterns. Research reinforced structural failure modes: four architectural failure modes (similarity does not equal correctness, synthesis hallucination, meaning re-derived at query time, temporal unawareness) were documented as blocking broader agentic adoption; Amazon GaRAGe benchmark with 35K+ human-curated grounding annotations across 2,366 questions confirmed Tier-1 vendor investment in rigorous domain-specific RAG evaluation. Agentic multi-hop RAG ablation validated that fixed hybrid retrieval outperforms adaptive routing by +1.8 EM, and two iterations capture 95% of gains—pointing toward simpler, more reliable production architectures. A US Navy warfare center AWS GovCloud deployment confirmed regulated-domain viability with full NIST/DoD security controls. Vendor infrastructure GA accelerated further: Microsoft Foundry IQ reached unified-knowledge-platform GA with serverless agentic retrieval (46-54% evidence-recall improvement, 34% token-cost reduction), and an 8-industry production survey (banking, insurance, healthcare, legal, SaaS, finance, manufacturing, R&D) confirmed domain-specific maturity—banking >95% citation precision, insurance 70%+ touchless claims processing, healthcare >95% faithfulness—at the cost of 3-10x higher token spend versus traditional RAG. A peer-reviewed taxonomy of 33 RAG failure modes found 12 modes, including all 8 agentic-orchestration modes, still lack empirical evidence, marking agentic orchestration as the field's active research frontier. CHARLIE, a multi-agent on-premise RAG system, went into production at Brazil's Federal Police Forensic Institute for cross-corpus evidential reasoning, extending regulated-domain viability beyond the US Navy deployment. Commentary consolidated around a broader architectural shift from vector-DB-first RAG toward permission-aware agentic search, while a governance-focused analysis catalogued silent enterprise failure modes—embedding drift, access-control desynchronization, and hallucination from stale context—that produce no error signal until user-visible impact occurs. Late-July evidence (through July 29) sharpened the multi-hop reliability problem specifically: a UC Berkeley/Google DeepMind benchmark found 92% of enterprise RAG configurations fail multi-hop queries (only 8% fully correct across 10,000 questions, split 63% incomplete-retrieval-chain failures and 41% hallucinated bridges), while Microsoft's LakeQuest benchmark identified relation-chaining and policy-grounding failures across heterogeneous data lakes and a security study found multi-hop RAG agents vulnerable to salience-induction attacks with an 83.3% success rate despite all underlying facts remaining true. Mitigation evidence accumulated alongside: MIT/Microsoft's CoRe-RAG cut multi-hop failure rates from 71% to 23% via specialized agent roles with closed-loop verification, Cortex's production deployment across 60+ technical sources answered 18,000+ queries at 86% certainty, and a governance pattern for safe embedding-model upgrades (dual-index migration) addressed silent coordinate-system breakage in production RAG."
    },
    {
      "period": "2026-Aug",
      "text": "Evidence through August 12 confirms maturity bifurcation: production deployments thrive in narrow, well-curated domains (Xiaohongshu: 5,300+ data warehouse with knowledge-graph-guided retrieval achieved 96.6% Hit@10 recall; operations/library systems showing adoption) while enterprise-wide deployment barriers persist. Rigorous benchmarking (Wang et al., 28-tier scaling study across 1.7M–600M tokens) established BM25 as scalable default at enterprise scale, overtaking agentic approaches around 10M tokens and maintaining 20-point accuracy margin—contradicting vector-first assumptions. Domain-specific research (SciRet on CORD-19 scientific corpus) validated hybrid retrieval achieving Recall@10 1.0, but generic cross-encoder rerankers trained on web data degrade scientific precision, reinforcing that domain adaptation is architectural requirement, not optional. Cloud infrastructure GA milestones: Databricks ai_search() SQL function (batch RAG pipelines with multi-source deduplication and synthesis), Neo4j GraphRAG context provider (entity-relationship traversal), and Azure Foundry knowledge bases signal ecosystem convergence toward managed domain-specific retrieval as standard enterprise capability. However, agentic orchestration remains operationally fragile: new research (Roh & Han, 12,000 trajectory analysis) identified pre-evidence procedural failures in multi-hop reasoning where agents retrieve evidence but skip reading before finalization—Read-Gate invariant improves accuracy 14.9–19.9 points, suggesting control discipline rather than model capability is limiting factor. Counter-signal: production failure audit of 143 enterprise RAG deployments found 73% experienced critical failure within first quarter (41% undetected by standard evaluations), with cross-document entity resolution, temporal drift, and multi-hop gaps cited as persistent adoption blockers. Field consensus remains unchanged: domain-specific RAG achieves viability only within disciplined, narrowly-scoped domains with meticulous tuning and continuous governance; cross-domain generalization and large-scale enterprise deployment remain structurally constrained by architecture brittleness, retriever reliability in specialized contexts, and operational complexity of maintaining embedding consistency and corpus freshness."
    },
    {
      "period": "2026-Aug (late)",
      "text": "Production deployment breadth widened further while reliability research sharpened known limits. Named regulated-sector deployments matured: AIG cut SLA 5x+ with accuracy improving 75%→90%, Allianz deployed autonomous motor/health claims workflows, and Travelers scaled to 10K engineers; Mistral's Agentic Search GA replaced one-shot RAG with a multi-step retrieval loop, lifting FinanceBench accuracy from 26.7% to 86%. Domain-specific deployments also landed in materials R&D (Chemcopilot, 1,000+ scientific documents), IT operations (Accenture Japan ticket triage), and agricultural advisory (peer-reviewed five-architecture Bengali evaluation showing dense retrieval collapsing 10x by query type). A preregistered audit documented a new production fragility mode—corpus expansion causing silent answer churn (6-10pp excess) despite fixed models and prompts—reinforcing that compatibility discipline, not just retrieval architecture, gates production RAG reliability."
    },
    {
      "period": "2026-Sep",
      "text": "Domain-adaptation evidence consolidated across sectors. UCSF deployed agentic RAG over a 7M+ subject EHR fused with a curated biomedical knowledge graph, improving GPT-5.5 and Claude Opus accuracy on 100 clinical research tasks, while a new FinRAG-QA benchmark (999 practitioner-curated Q&A on 209 regulatory documents) showed domain-adapted embeddings and reranking lifting NDCG by 38.8 points and generation accuracy by 34.4pp over generic pipelines—reinforcing evidence that fine-tuned embeddings improve domain retrieval precision by 20-40% over off-the-shelf models. Architectural limits sharpened: production case studies documented corpus-contradiction failures (outdated documents outranking current ones via semantic similarity) and structured-data brittleness (vector-only RAG at 40-60% accuracy versus 90%+ for hybrid structured-query approaches), while graph RAG outperformed vector-hybrid architectures on trustworthiness metrics in a government project-management benchmark. Industry skepticism hardened in parallel: an analysis citing 72% of enterprise RAG implementations failing within their first year and Gartner's forecast of 40% agentic AI project cancellations by 2027 argued for context-graph architectures over traditional RAG, even as a production financial-RAG case study demonstrated hallucination reduction from 15% to 1% through systematic pipeline debugging (context reordering, claim verification, deterministic math)—underscoring that domain-specific RAG remains viable only with sustained architectural and governance investment. A Salesforce production case study showed RAG accuracy dropping from 90%+ on benchmarks to 46% on real enterprise documents, recovered to 84.4% only after intelligent parsing and reranking—reinforcing the benchmark-production gap. Elsewhere, telecom and utility deployments (hybrid KG-RAG at 89.6% grounding; Ontario Power Generation's progressive-evidence compliance pattern) showed production hardening, while research found exact-span evidence extraction below 40% accuracy in technical domains."
    }
  ],
  "historyFallback": false,
  "lastUpdated": "2026-09-23",
  "domain": {
    "id": "research-analysis",
    "label": "Research & Knowledge",
    "icon": "🔬"
  },
  "url": "https://www.thestateofplay.ai/practice/domain-specific-rag-and-cross-corpus-question-answering",
  "license": "CC BY 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "generatedAt": "2026-10-01"
}