A national DIY retailer replaced the keyword index behind its in-store product assistant with a vector index built from the same catalogue, embedding each product as one short record as before. Accuracy on evaluation questions such as which drill suits masonry rose, but accuracy on questions naming a part number, such as a replacement filter for model HX-4410, fell from 94 percent to 61 percent. Retrieval traces for the failing questions show the wrong records, for example HX-4401 and HX-4140, already ranked first before generation begins, and the index was rebuilt from the current catalogue the night before the evaluation. What is the most likely cause of the part-number failures?
- ADense embeddings capture meaning rather than exact character sequences, so codes differing by a character or two sit close together and lose exact matching Correct
- BThe model is altering part numbers while drafting its answer, because its sampling settings let it substitute similar-looking product codes
- CThe vector index is serving vectors from an older catalogue build, so newly listed part numbers are missing and the nearest older products are returned
- DSemantic retrieval needs a larger top-k than keyword search did, so the correct product is retrieved but cut off before it reaches the model's context
Why A is correct: Correct. The only change was swapping lexical for dense retrieval, and dense vectors place near-identical identifiers close together without rewarding an exact token match, which is precisely what the keyword index supplied for part-number queries.
Why B is wrong: Tempting because the visible symptom is a wrong code in the answer, and generation can corrupt identifiers. It is wrong because the traces show the wrong records already ranked first before generation begins, so the fault is upstream of the model.
Why C is wrong: Tempting because a stale index is a classic cause of confident wrong retrieval after a change. It is wrong because the stem states the index was rebuilt from the current catalogue the night before the evaluation, so freshness is ruled out.
Why D is wrong: Tempting because raising k is a common first lever when the right item is missing from context. It is wrong because the traces show the wrong products ranked first; a larger k would add more candidates without fixing a ranking signal that cannot tell near-identical codes apart.