A Hybrid Search Engine Combining Classical Information Retrieval with Pluggable Semantic Search

Authors

  • Sabnam Ghimire

DOI:

https://doi.org/10.65091/icicset.v3i1.100

Abstract

Most search features in production today sit on top
of a mature platform, typically Elasticsearch, Apache Lucene,
or a managed vector database, and the algorithms doing the
actual retrieval work are rarely touched by the engineer using
them. This paper describes a hybrid search engine whose core
retrieval, ranking, and fusion components, an inverted index,
TF-IDF, Okapi BM25, dense vector search, and score fusion,
are implemented directly in Python rather than delegated to an
existing search or vector-database framework, so that the tradeoffs
those platforms hide can be measured directly. Standard
supporting libraries (a web framework, document parsers, an
optional embedding backend) are used where they are incidental
to that goal, but none of them performs the retrieval or ranking
work itself. The system pairs classical lexical retrieval, an inverted
index supporting Boolean and phrase queries, TF-IDF, and Okapi
BM25, with a pluggable semantic layer built on dense embeddings
and cosine-similarity vector search. The two signals are combined
through a normalized linear fusion into a single hybrid ranking
function, and the retrieval stack is further extended into a small
retrieval-augmented generation (RAG) pipeline. The ranking,
fusion, and generation components are implemented and tested
at the library and benchmark level; the deployed command-line
and HTTP interfaces currently serve BM25 only, a boundary this
paper states explicitly rather than blurs. Development proceeded
across eleven incremental phases, moving from custom data
structures (hash tables, tries, heaps, graphs) through the ranking
algorithms, a custom binary persistent index format, parallel and
asynchronous performance work, a breadth-first web crawler, a
FastAPI backend, and a React frontend. Of 414 automated unit
tests, 408 pass (98.6%); the six that fail trace back to a genuine,
reproducible infinite loop in the Knuth–Morris–Pratt substringmatching
routine, documented here as an engineering finding
rather than fixed quietly and forgotten. A retrieval-effectiveness
benchmark, run over a small 15-document relevance-judged corpus
with only four queries (too small a scale to support a strong
effectiveness claim), shows BM25 (NDCG@5 = 0.964) ahead of
TF-IDF (0.910) and every tested hash-based hybrid configuration.
A follow-up sweep over α ∈ {0.1, 0.3, 0.5, 0.7, 0.9} with
the dependency-free hash embedding provider shows NDCG@5
rising monotonically with α, from 0.753 at α = 0.1 to a
best of 0.917 at α = 0.9, still below BM25’s 0.964; a hashembedding-
only run (NDCG@5 = 0.753) and a Reciprocal Rank
Fusion baseline combining BM25 with hash-based semantic
scores (NDCG@5 = 0.866) likewise fall short of BM25 alone.
On this corpus, then, no tested hybrid or fusion configuration
improves on plain BM25. A separate attempt to evaluate a
trained sentence-embedding model (all-MiniLM-L6-v2) was
blocked by the execution environment, insufficient disk space
for the CUDA/PyTorch dependency stack, missing CUDA shared
libraries in a reduced-dependency install, and blocked network
access to the model hub, so this paper draws no conclusion about
real dense semantic retrieval and reports the blocker itself as a
reproducible engineering finding rather than a result. A separate
parallel-indexing benchmark confirms, on hardware with a single
available CPU core, that the multiprocessing speedups reported
in the project’s internal documentation (3.14× on a multicore
machine) turn into a slowdown (0.77×) once no physical
parallelism is available, an empirical illustration of Amdahl’s
law. Taken together, the results position the system as a reference
implementation for understanding hybrid retrieval and systems
engineering, validated to the explicit scope described in Table X,
and not as a production-ready replacement for existing search
infrastructure.

Downloads

Published

2026-10-02

How to Cite

[1]
S. Ghimire, “A Hybrid Search Engine Combining Classical Information Retrieval with Pluggable Semantic Search”, ICICSET2025, vol. 3, no. 1, Oct. 2026.