Skip to content

Hybrid Search

Last updated:

What hybrid search is

Keyword search, usually based on full-text indexes and ranking functions such as BM25, is strong at exact matches: product codes, names, acronyms and quoted phrases. Semantic search, based on embeddings, is strong at meaning: synonyms, paraphrases and loosely worded questions. Each has blind spots that the other covers.

Hybrid search runs both methods for the same query and combines what they find. The result is a single list of passages that ranks well for exact terms and for intent.

How it works

The query goes to two retrievers. The keyword retriever ranks passages by how well their words match the query. The vector retriever ranks them by embedding similarity. The two ranked lists are then merged, most often with one of these methods:

  • Reciprocal rank fusion (RRF): each passage gets a score based on its position in each list, so items ranked well by both rise to the top.
  • Weighted score blending: normalized scores from both retrievers are added with chosen weights.
  • Re-ranking: a separate model re-orders the merged candidates by relevance to the question.

An example

An electronics retailer's chatbot receives the question “Does the XR-220 work with USB-C chargers?”. The keyword retriever finds the product sheet because it contains “XR-220”. The vector retriever finds a general article on charging compatibility. With hybrid search, the product sheet ranks first and the charging article is available as supporting context.

Semantic search alone might have returned the sheet for a similar model. Keyword search alone would miss the charging article if it never mentions the model number.

Why it matters for business chatbots

Customer questions mix both kinds of language: everyday phrasing and exact identifiers. If retrieval misses the right passage, even a capable language model cannot answer correctly and may fill the gap with a guess. Hybrid search reduces those misses, which makes grounded answers more reliable.

Rank fusion is a common choice for chatbots because keyword scores and similarity scores are on different scales and hard to compare directly. Using positions instead of raw scores avoids that problem and needs little tuning.

How intoCHAT uses it

intoCHAT uses hybrid retrieval for every question. It runs a vector similarity search and a keyword full-text search over your training content and merges the two rankings with rank fusion. Q&A answers you write take priority over other sources, and sources you star rank higher. There is nothing to configure: you improve results by improving your content.

Frequently asked questions

Is hybrid search better than vector search?

For most business content, yes, because real questions mix everyday language with exact names and codes. Vector search alone can miss exact identifiers, and keyword search alone can miss paraphrases. Results still depend on content quality and chunking.

What is reciprocal rank fusion?

Reciprocal rank fusion is a method for merging ranked lists. Each item gets a score from its position in each list, and items that rank high in several lists end up on top. It is widely used in hybrid search because it does not require comparing raw scores from different systems.

Do I need to configure hybrid search in intoCHAT?

No. intoCHAT runs hybrid retrieval automatically for every agent. You can influence results by writing Q&A pairs for important questions, starring key sources and retraining a source when the original content changes.

See it answer from your own website

Paste your website address and chat with an agent built from your pages. It takes about a minute.

Create your agent free

Free plan, no credit card needed.