Awesome Testing

Production systems · reviewed · reviewed Oct 6, 2026 · 4 min

How do hybrid search and reranking work?

Hybrid search combines candidates from complementary retrieval methods, often keyword and vector search. Fusion orders their results; a reranker then compares the query with that smaller candidate set. Reranking can improve ordering, but it cannot recover a passage absent from its candidates.

A product code and a meaning both matter

A reader asks: How long is the warranty for ZX-42? The exact product identifier matters, but so does the relationship between that model and its coverage duration. A passage merely mentioning ZX-42 is not necessarily an answer; a general discussion of warranties can also miss the model-specific fact.

Keyword retrieval uses lexical matching, often with a scoring method such as BM25. Vector retrieval compares learned representations and can find related wording without an exact string match. These methods can produce different useful candidates and different mistakes.

Hybrid search combines their results. A later reranker examines the query and the smaller candidate set more closely. These stages have separate jobs: find possibilities, combine them, then prioritize the supplied possibilities.

Follow two lists into one candidate set

The experiment has four authored passages: a product index, a warranty overview, a return window, and the actual ZX-42 warranty. Its keyword and semantic orders are fixed teaching fixtures, not outputs of search engines.

Start with hybrid retrieval at depth three. Both lists supply the warranty passage, but their fused order places the overview first. Turn reranking on to prioritize the model-specific passage under the authored relevance scores. Then reduce each retriever’s depth to two and predict whether reranking can still select it.

Explore the mechanism

Retrieve candidates before you rerank them

Fixed question: How long is the warranty for ZX-42?

Keyword candidates

  1. Product index
  2. Warranty overview
  3. ZX-42 warranty

Semantic candidates

  1. Warranty overview
  2. Return window
  3. ZX-42 warranty

Fused candidate union

  1. Warranty overview

    Coverage varies by model and purchase date.

    Computed RRF 0.03252
  2. ZX-42 warranty

    ZX-42 carries a 24-month warranty from purchase.

    Computed RRF 0.03175
  3. Product index

    ZX-42 identifies the device model; this index contains no warranty duration.

    Computed RRF 0.01639
  4. Return window

    Unused purchases can be returned within 30 days.

    Computed RRF 0.01613
Unique candidates
4
Relevant passage in candidates
Yes

First passage: Warranty overview.

The warranty passage is available to this ranking stage.

Inspect fusion scores for the candidate union
PassageKeyword rankSemantic rankRRF
Warranty overview210.03252
ZX-42 warranty330.03175
Product index1Absent0.01639
Return windowAbsent20.01613

Each appearance adds 1 / (60 + rank). Absent items contribute zero. A higher fusion score is not evidence that a passage answers the question.

The ZX-42 corpus, keyword/semantic orders, and reranker scores are authored teaching fixtures. Reciprocal rank fusion, candidate membership, and ordering are computed. No BM25 engine, embedding model, neural reranker, or answer generator runs; the displayed scores are not calibrated probabilities or benchmark results.

At depth two, the warranty passage never enters the candidate union. The reranker only sees the index, overview, and return window. Its best available choice can still be insufficient to answer the question.

Removing the relevant passage from the corpus exposes a different failure. Increasing depth can repair a cutoff miss; it cannot restore content that is absent. These conditions call for different interventions even if the final answer would look equally unsupported.

Fusion combines rankings rather than incompatible scores

A keyword score and a vector similarity can live on different numerical scales. Adding raw values without a justified normalization or learned combination can give one retriever unintended influence.

Reciprocal Rank Fusion, or RRF, instead uses each passage’s position in each list. Every appearance contributes 1 / (constant + rank); contributions are added and sorted. The demo uses constant 60 and one-based ranks. An absent passage contributes nothing, and duplicate appearances produce one candidate with a combined score.

RRF does not read the passage or establish its relevance. Agreement between retrieval lists can elevate a shared mistake. The fusion inspector makes that distinction visible: computed rank arithmetic and the passage’s actual words remain separate evidence.

The constant, retrieval depth, filters, and any retriever weighting are configuration choices. The original RRF experiments support the mechanism; they do not make one configuration optimal for every collection or query population.

Reranking spends more attention on fewer possibilities

A common neural reranker is a cross-encoder: it processes the query and a candidate passage together and scores their relevance. In contrast, a conventional vector retriever can compare representations computed separately and reuse stored passage vectors.

The BERT passage-reranking work illustrates this second-stage design. A more detailed query/passage comparison can improve ordering, but running it across an entire large corpus is usually more expensive than scoring a shortlist. Candidate depth therefore affects both what the reranker can see and how much work it receives.

The demo substitutes authored scores for neural predictions so candidate membership remains the visible lesson. Its score of 0.95 is not a measured probability of truth, and selecting the warranty passage is not a benchmark result.

Keep the missing stage visible

Retrieval evaluation should distinguish whether relevant passages entered the candidates from how well the supplied candidates were ordered. BEIR compares retrieval approaches across varied tasks and domains, showing why a strong result in one setting is insufficient evidence for another.

For the ZX-42 example, ask two concrete questions: was the warranty passage included at the chosen depth, and where did the final ranking place it? A ranking metric cannot explain the problem on its own if the candidate set is hidden.

Selected passages still need current versions, access controls, useful context, and support for the requested claim. Retrieval and reranking do not grant authorization. RAG then owns context construction, generation, citation handling, and the decision to ask or abstain when evidence is insufficient.

The complete teaching fixture

Keyword order: product index, warranty overview, ZX-42 warranty, return window. Semantic order: warranty overview, return window, ZX-42 warranty, product index. Each active list is truncated separately at the selected depth before fusion.

Authored reranker scores are 0.08, 0.35, 0.95, and 0.12 for those respective passages. Reranking sorts only the fused candidate union; it cannot inspect another record behind the cutoff. Removing the policy excludes it before list truncation and rank calculation.

Only the policy passage states a 24-month warranty for ZX-42. The product index deliberately contains the identifier without the duration; the overview provides no model-specific amount; the 30-day return window answers a different question. No model generates an answer in this experiment.

Sources and further reading

  1. 01
    Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning MethodsCormack, Clarke, and Buettcher · research · source checked Oct 6, 2026

    Original rank-fusion formula sums reciprocal rank contributions with a smoothing constant; ranks avoid directly mixing incompatible raw retrieval scores.

  2. 02
    Passage Re-ranking with BERTNogueira and Cho · research · published Jan 13, 2019 · source checked Oct 6, 2026

    Query/passage cross-encoder relevance scoring supplies the second-stage reranking mechanism and explicitly depends on first-stage candidates.

  3. 03
    BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsThakur et al. · research · published Apr 17, 2021 · source checked Oct 6, 2026

    A heterogeneous retrieval benchmark demonstrating why retriever quality and out-of-domain behaviour need explicit evaluation.