Javascript must be enabled to continue!
Beyond Sequential Hybrid Retrieval: A Parallel Framework for Accurate and Scalable RAG
View through CrossRef
Abstract
Large Language Models (LLMs) frequently generate hallucinated or outdated responses when operating without external knowledge. Retrieval Augmented Generation (RAG) grounds model outputs in retrieved documents, yet existing hybrid RAG systems execute multiple retrievers sequentially, creating latency bottlenecks that limit scalability. Current research focused on PH-RAG, a Parallel Hybrid RAG framework that runs BM25 sparse retrieval and FAISS based dense retrieval simultaneously using a persistent thread pool, fuses their ranked outputs via Reciprocal Rank Fusion (RRF) or Weighted Score Fusion, and applies a cross-encoder reranker for improved top rank precision. This framework coupled a strong instruction tuned bi-encoder (BGE large en-v1.5) with an exact FAISS inner-product index and a lightweight MiniLM cross-encoder, forming a modular four-stage pipeline in which every component can be independently ablated. Evaluated on TriviaQA with 1,000 validation questions and a corpus of 5,470 Wikipedia passages, PH-RAG achieves Hit@1 of 0.656, Hit@5 of 0.797, Recall@20 of 0.833, and MRR of 0.718, surpassing the HybGRAG baseline (ACL 2025) on all four reported metrics without requiring a knowledge graph or an iterative critic module. A theoretical latency model, validated against per retriever measurements across seven corpus sizes, demonstrates parallel speedups of 1.01× to 1.64× for corpora between 1,000 and 20,000 documents, with the advantage growing as sparse-retrieval cost begins to dominate. Ablation studies quantify the individual contributions of fusion strategy, encoder strength, and reranking, showing that the weighted fusion of complementary sparse and dense signals is the primary driver of retrieval quality, while parallel execution provides a deployment time efficiency benefit that is orthogonal to accuracy. Our results indicated that a carefully engineered but architecturally simple hybrid retriever match or exceed substantially more complex agentic RAG systems on open-domain question answering.
Title: Beyond Sequential Hybrid Retrieval: A Parallel Framework for Accurate and Scalable RAG
Description:
Abstract
Large Language Models (LLMs) frequently generate hallucinated or outdated responses when operating without external knowledge.
Retrieval Augmented Generation (RAG) grounds model outputs in retrieved documents, yet existing hybrid RAG systems execute multiple retrievers sequentially, creating latency bottlenecks that limit scalability.
Current research focused on PH-RAG, a Parallel Hybrid RAG framework that runs BM25 sparse retrieval and FAISS based dense retrieval simultaneously using a persistent thread pool, fuses their ranked outputs via Reciprocal Rank Fusion (RRF) or Weighted Score Fusion, and applies a cross-encoder reranker for improved top rank precision.
This framework coupled a strong instruction tuned bi-encoder (BGE large en-v1.
5) with an exact FAISS inner-product index and a lightweight MiniLM cross-encoder, forming a modular four-stage pipeline in which every component can be independently ablated.
Evaluated on TriviaQA with 1,000 validation questions and a corpus of 5,470 Wikipedia passages, PH-RAG achieves Hit@1 of 0.
656, Hit@5 of 0.
797, Recall@20 of 0.
833, and MRR of 0.
718, surpassing the HybGRAG baseline (ACL 2025) on all four reported metrics without requiring a knowledge graph or an iterative critic module.
A theoretical latency model, validated against per retriever measurements across seven corpus sizes, demonstrates parallel speedups of 1.
01× to 1.
64× for corpora between 1,000 and 20,000 documents, with the advantage growing as sparse-retrieval cost begins to dominate.
Ablation studies quantify the individual contributions of fusion strategy, encoder strength, and reranking, showing that the weighted fusion of complementary sparse and dense signals is the primary driver of retrieval quality, while parallel execution provides a deployment time efficiency benefit that is orthogonal to accuracy.
Our results indicated that a carefully engineered but architecturally simple hybrid retriever match or exceed substantially more complex agentic RAG systems on open-domain question answering.
Related Results
DARE-RAG: Difficulty-Aware Retrieval Expansion for Retrieval-Augmented Generation
DARE-RAG: Difficulty-Aware Retrieval Expansion for Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) systems face a fundamental trade-off: query expansion can improve retrieval effectiveness for ambiguous or underspecified queries, yet indiscri...
JADE: jawbone lesion diagnosis and decision supporting system
JADE: jawbone lesion diagnosis and decision supporting system
Abstract
Objectives
To develop and evaluate JADE, a proof-of-concept retrieval-augmented generation (RAG) diagnostic assi...
Self-updating Retrieval Mechanisms for Scalable and Personalized RAG Systems in Dynamic Environments
Self-updating Retrieval Mechanisms for Scalable and Personalized RAG Systems in Dynamic Environments
Retrieval-augmented generation (RAG) has become a practical mechanism for improving the factuality, traceability, and domain relevance of large language model (LLM) systems. Howeve...
Automating Information Retrieval from Biodiversity Literature Using Large Language Models: A Case Study
Automating Information Retrieval from Biodiversity Literature Using Large Language Models: A Case Study
Recently, Large Language Models (LLMs) have transformed information retrieval, becoming widely adopted across various domains due to their ability to process extensive textual data...
Diagnosing RAG Failures: A Taxonomy of Failure Modes in Retrieval-Augmented Generation
Diagnosing RAG Failures: A Taxonomy of Failure Modes in Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding large language models in external knowledge, yet RAG systems still fail in production in ways that...
A Systematic Literature Review of Retrieval-Augmented Generation Implementation for Enhancing Large Language Models in Education
A Systematic Literature Review of Retrieval-Augmented Generation Implementation for Enhancing Large Language Models in Education
The rapid advancement of Large Language Models (LLM) has led to the creation of increasingly adaptive intelligent learning systems. However, many educational implementations of LLM...
Development and validation of Retrieval Augmented Generation (RAG) and GraphRAG for complex clinical cases
Development and validation of Retrieval Augmented Generation (RAG) and GraphRAG for complex clinical cases
Abstract
Objective
Chronic Kidney Disease (CKD) is a progressive condition requiring evidence-based management, but adherence t...
Unconventional Method of Subsea Umbilical Retrieval Using Anchor Handling Vessel
Unconventional Method of Subsea Umbilical Retrieval Using Anchor Handling Vessel
Abstract
A deepwater field in West Africa was decommissioned and subsea facilities retrieval operation was carried out as part of the Abandonment and Decommissioning...

