LAB 6

Re-ranking

Stage 1 scored every chunk before it ever saw your question. That is why its order cannot be trusted. A re-ranker reads your question and each chunk together, then re-orders them.

Stage 1 — Fast Retrieval
🧬
Bi-encoder (Vector Search)

Every chunk was embedded on its own, long before your question existed. Comparing two ready-made vectors is very fast, so this scales to millions of chunks. But it is a guess made without the question in hand.

Stage 2 — Precision Re-ranking
🏆
Pair scoring (reads both together)

Reads your question and one chunk together, and scores that pair. Far more accurate, and far too slow to run over millions. So we retrieve 10 fast, then re-rank only those 10.

🏆This lab scores each (question, chunk) pair with gpt-4o-mini at temperature 0. In production you would use a dedicated cross-encoder such as Cohere Rerank or ms-marco. Same principle: read the pair together.