RAG can retrieve better context when an index contains more documents and embeddings. A broad corpus covers queries that a small collection does not contain, but it also makes each query search through many more vectors.

With a small number of vectors, exact search compares a query with every embedding and ranks the results. The cost of that brute-force approach grows with the corpus, and becomes unsustainable with millions or billions of vectors when the system must return results at low latency. Approximate nearest-neighbor (ANN) indexes reduce the space explored before retrieving candidates, which, in return, can miss the exact nearest neighbor.

Two index families often appear in this scenario. HNSW builds a navigable graph of vectors and usually achieves good recall with low latency. Its graph and the vectors needed to traverse it grow with the corpus and need to stay in RAM in many configurations. IVF indexes divide the space into regions and first choose which regions are worth exploring. They can keep the vector lists outside memory and retain the layer that routes the search in RAM.

In our paper, Memory Optimization of IVF Indices for RAG Systems Using FP8 Centroid Quantization, we focus on the IVF routing layer because of that separation. With IVF, centroids can remain in memory while posting lists containing the document vectors can be stored on SSDs.

We study whether centroids can move from float32 to FP8 without changing the document vectors or their posting lists and how this affects the routing quality.

Why centroids are on the critical path

An IVF index splits embeddings into nlist regions. Each region has a centroid. To answer a query, the index compares its embedding with every centroid and selects the nprobe closest regions. It then searches only the corresponding posting lists. This reduces the number of comparisons needed.

The index starts with no assigned clusters.

Animated IVF search exampleAnimated IVF search example

This separation allows the large corpus to live in cheaper storage, but the centroid table is consulted for every query and needs to remain in memory. Its cost is:

M=nlist×d×bM = nlist \times d \times b

where MM is routing memory, dd is the embedding dimensionality, and bb is bytes per coordinate. In float32, bb is 4. In FP8, it is 1.

The point here is not to compress the entire index, but only the representations used to decide which lists to explore. Once the lists have been selected, the document vectors and the ranking continue to use the index’s original representation, as they are loaded from cheaper storage

Precision or range

We tested the E4M3FN and E5M2 FP8 formats. Both use 8 bits, but distribute them differently between exponent and mantissa. E5M2 has more dynamic range. E4M3FN allocates one additional bit to the mantissa, so it represents nearby values more precisely.

FP32 and FP8 quantization in two dimensionsA two-dimensional scatter plot of values sampled from a normal distribution around zero.

For a centroid table, that local precision proved more useful than wider range. Cluster assignment depends on fine comparisons between similar distances. A small shift in a centroid can change which list is explored even when the stored vectors have not changed.

Evaluation

For each dataset, search engine and FP8 format, we trained a separate IVF index. We extracted its float32 centroids, cast them to E4M3FN or E5M2, then converted them back to float32 before putting them back into the index. The search engine therefore follows the same float32 path in every run. Only the rounding error left by the FP8 conversion changes the routing table.

We ran the protocol in both FAISS and ScaNN, with the same corpus and index configuration for each FP32 baseline and its FP8 counterpart.

For geometric retrieval, we used glove-100-angular, which asks the index to recover the nearest vectors under angular distance. We report Recall@10, and compare results to see whether FP8 centroids change which parts of the vector space IVF explores.

For semantic retrieval, we encoded three MTEB Retrieval datasets with the 384-dimensional all-MiniLM-L6-v2 model: FiQA2018, MLQuestions and QuoraRetrieval. They cover financial questions, multilingual question retrieval and duplicate questions, with different corpus sizes. We report nDCG@10, which rewards relevant documents near the top of the ranking, and MRR@10, which measures the position of the first relevant result. Together, the two benchmarks test geometric nearest-neighbour recovery and retrieval quality on semantic embeddings.

Results

E4M3FN has the tighter recall distribution

We measured absolute Recall@10 degradation over the full glove-100-angular configuration sweep. E4M3FN stays closer to the FP32 baseline in both engines, while E5M2 produces a wider spread and the largest outliers. The extra mantissa bit matters more than E5M2’s additional exponent range when IVF decides which lists to probe.

E4M3FN preserves recall more consistentlyAbsolute Recall@10 degradation (%)FAISS0.000.380.751.11.5FP8 E4M3FN · outlier: 0.410%FP8 E4M3FN · outlier: 0.330%FP8 E4M3FN · outlier: 0.469%FP8 E4M3FN · outlier: 0.385%FP8 E4M3FNFP8 E5M2 · outlier: 1.047%FP8 E5M2 · outlier: 1.261%FP8 E5M2 · outlier: 1.164%FP8 E5M2 · outlier: 1.386%FP8 E5M2ScaNNFP8 E4M3FN · outlier: 0.254%FP8 E4M3FN · outlier: 0.353%FP8 E4M3FNFP8 E5M2 · outlier: 1.322%FP8 E5M2

Geometric retrieval

Across the full configuration sweep, mean Recall@10 degradation with E4M3FN was 0.085% in FAISS and 0.095% in ScaNN, while E5M2 averaged 0.304% and 0.405%, respectively. The QPS-recall frontiers remain almost on top of their FP32 equivalents. QPS is included to show the configurations that were evaluated, not as a claim that FP8 made the engines faster. Centroids are restored to FP32 before search, so the experiment preserves the search path and measures the routing error left by the conversion.

FP32FP8 E4M3FNFP8 E5M2
Geometric retrieval keeps the same QPS-recall trade-offQueries per second (log scale)FAISS1001k10k100k0.310.540.771.00FP32 Recall@10: 0.3283 QPS: 253670.4FP32 Recall@10: 0.3639 QPS: 182303.0FP32 Recall@10: 0.5539 QPS: 109727.3FP32 Recall@10: 0.5952 QPS: 75959.4FP32 Recall@10: 0.6494 QPS: 65779.9FP32 Recall@10: 0.6903 QPS: 41314.9FP32 Recall@10: 0.7291 QPS: 22626.7FP32 Recall@10: 0.8238 QPS: 16002.4FP32 Recall@10: 0.8556 QPS: 9137.5FP32 Recall@10: 0.8799 QPS: 8424.8FP32 Recall@10: 0.8879 QPS: 4870.3FP32 Recall@10: 0.9068 QPS: 4707.6FP32 Recall@10: 0.9235 QPS: 4369.8FP32 Recall@10: 0.9329 QPS: 2494.2FP32 Recall@10: 0.9458 QPS: 2409.5FP32 Recall@10: 0.9564 QPS: 1271.0FP32 Recall@10: 0.9660 QPS: 1265.6FP32 Recall@10: 0.9755 QPS: 657.4FP32 Recall@10: 0.9826 QPS: 639.7FP32 Recall@10: 0.9940 QPS: 327.7FP32 Recall@10: 0.9995 QPS: 162.7FP32 Recall@10: 1.0000 QPS: 130.1FP8 E4M3FN Recall@10: 0.3268 QPS: 249076.7FP8 E4M3FN Recall@10: 0.3627 QPS: 181839.3FP8 E4M3FN Recall@10: 0.5526 QPS: 109837.0FP8 E4M3FN Recall@10: 0.5938 QPS: 75983.8FP8 E4M3FN Recall@10: 0.6478 QPS: 65415.2FP8 E4M3FN Recall@10: 0.6891 QPS: 41362.7FP8 E4M3FN Recall@10: 0.7288 QPS: 22621.3FP8 E4M3FN Recall@10: 0.8231 QPS: 16149.2FP8 E4M3FN Recall@10: 0.8553 QPS: 9137.2FP8 E4M3FN Recall@10: 0.8794 QPS: 8438.3FP8 E4M3FN Recall@10: 0.8873 QPS: 4870.3FP8 E4M3FN Recall@10: 0.9062 QPS: 4691.3FP8 E4M3FN Recall@10: 0.9230 QPS: 4376.7FP8 E4M3FN Recall@10: 0.9327 QPS: 2494.2FP8 E4M3FN Recall@10: 0.9455 QPS: 2415.9FP8 E4M3FN Recall@10: 0.9563 QPS: 1270.3FP8 E4M3FN Recall@10: 0.9656 QPS: 1263.8FP8 E4M3FN Recall@10: 0.9755 QPS: 658.4FP8 E4M3FN Recall@10: 0.9825 QPS: 639.7FP8 E4M3FN Recall@10: 0.9940 QPS: 328.2FP8 E4M3FN Recall@10: 0.9995 QPS: 162.7FP8 E4M3FN Recall@10: 1.0000 QPS: 130.1FP8 E5M2 Recall@10: 0.3245 QPS: 244140.2FP8 E5M2 Recall@10: 0.3593 QPS: 182603.6FP8 E5M2 Recall@10: 0.5498 QPS: 110568.1FP8 E5M2 Recall@10: 0.5902 QPS: 76758.2FP8 E5M2 Recall@10: 0.6451 QPS: 65688.4FP8 E5M2 Recall@10: 0.6874 QPS: 41571.0FP8 E5M2 Recall@10: 0.7276 QPS: 22741.6FP8 E5M2 Recall@10: 0.8213 QPS: 16076.3FP8 E5M2 Recall@10: 0.8541 QPS: 9225.5FP8 E5M2 Recall@10: 0.8784 QPS: 8463.4FP8 E5M2 Recall@10: 0.8863 QPS: 4895.1FP8 E5M2 Recall@10: 0.9056 QPS: 4727.1FP8 E5M2 Recall@10: 0.9224 QPS: 4377.5FP8 E5M2 Recall@10: 0.9320 QPS: 2505.6FP8 E5M2 Recall@10: 0.9449 QPS: 2420.7FP8 E5M2 Recall@10: 0.9560 QPS: 1273.0FP8 E5M2 Recall@10: 0.9655 QPS: 1269.4FP8 E5M2 Recall@10: 0.9751 QPS: 659.4FP8 E5M2 Recall@10: 0.9821 QPS: 640.8FP8 E5M2 Recall@10: 0.9939 QPS: 328.7FP8 E5M2 Recall@10: 0.9996 QPS: 162.7FP8 E5M2 Recall@10: 1.0000 QPS: 130.2Recall@10ScaNN0.350.570.781.00FP32 Recall@10: 0.3652 QPS: 64312.3FP32 Recall@10: 0.4876 QPS: 59201.6FP32 Recall@10: 0.5945 QPS: 52323.3FP32 Recall@10: 0.6779 QPS: 42325.5FP32 Recall@10: 0.8512 QPS: 20403.3FP32 Recall@10: 0.8617 QPS: 19324.4FP32 Recall@10: 0.8668 QPS: 17987.8FP32 Recall@10: 0.8748 QPS: 16483.8FP32 Recall@10: 0.8817 QPS: 15997.4FP32 Recall@10: 0.8923 QPS: 14354.7FP32 Recall@10: 0.9005 QPS: 13227.8FP32 Recall@10: 0.9048 QPS: 12581.5FP32 Recall@10: 0.9134 QPS: 11796.6FP32 Recall@10: 0.9230 QPS: 10062.6FP32 Recall@10: 0.9348 QPS: 8657.0FP32 Recall@10: 0.9458 QPS: 7503.4FP32 Recall@10: 0.9547 QPS: 6664.4FP32 Recall@10: 0.9591 QPS: 5838.1FP32 Recall@10: 0.9676 QPS: 5091.2FP32 Recall@10: 0.9714 QPS: 4429.2FP32 Recall@10: 0.9756 QPS: 3893.6FP32 Recall@10: 0.9805 QPS: 3583.4FP32 Recall@10: 0.9863 QPS: 2851.7FP32 Recall@10: 0.9920 QPS: 2201.5FP32 Recall@10: 0.9976 QPS: 1362.3FP8 E4M3FN Recall@10: 0.3648 QPS: 66226.6FP8 E4M3FN Recall@10: 0.4869 QPS: 60636.0FP8 E4M3FN Recall@10: 0.5930 QPS: 52991.4FP8 E4M3FN Recall@10: 0.6755 QPS: 43778.2FP8 E4M3FN Recall@10: 0.8505 QPS: 20788.0FP8 E4M3FN Recall@10: 0.8604 QPS: 19310.5FP8 E4M3FN Recall@10: 0.8661 QPS: 18545.7FP8 E4M3FN Recall@10: 0.8739 QPS: 16785.0FP8 E4M3FN Recall@10: 0.8809 QPS: 16111.2FP8 E4M3FN Recall@10: 0.8912 QPS: 14476.6FP8 E4M3FN Recall@10: 0.8990 QPS: 13444.0FP8 E4M3FN Recall@10: 0.9035 QPS: 12784.1FP8 E4M3FN Recall@10: 0.9129 QPS: 11900.4FP8 E4M3FN Recall@10: 0.9222 QPS: 10250.1FP8 E4M3FN Recall@10: 0.9337 QPS: 8964.8FP8 E4M3FN Recall@10: 0.9451 QPS: 7774.2FP8 E4M3FN Recall@10: 0.9543 QPS: 6712.6FP8 E4M3FN Recall@10: 0.9587 QPS: 6119.2FP8 E4M3FN Recall@10: 0.9672 QPS: 5172.3FP8 E4M3FN Recall@10: 0.9714 QPS: 4453.5FP8 E4M3FN Recall@10: 0.9753 QPS: 3948.7FP8 E4M3FN Recall@10: 0.9802 QPS: 3530.0FP8 E4M3FN Recall@10: 0.9860 QPS: 2915.4FP8 E4M3FN Recall@10: 0.9919 QPS: 2251.1FP8 E4M3FN Recall@10: 0.9975 QPS: 1402.1FP8 E5M2 Recall@10: 0.3615 QPS: 66387.3FP8 E5M2 Recall@10: 0.4828 QPS: 59903.9FP8 E5M2 Recall@10: 0.5867 QPS: 52996.3FP8 E5M2 Recall@10: 0.6710 QPS: 43072.0FP8 E5M2 Recall@10: 0.8476 QPS: 20716.9FP8 E5M2 Recall@10: 0.8578 QPS: 19357.4FP8 E5M2 Recall@10: 0.8619 QPS: 18477.6FP8 E5M2 Recall@10: 0.8702 QPS: 16781.2FP8 E5M2 Recall@10: 0.8770 QPS: 16104.5FP8 E5M2 Recall@10: 0.8881 QPS: 14535.7FP8 E5M2 Recall@10: 0.8970 QPS: 13515.3FP8 E5M2 Recall@10: 0.9016 QPS: 12717.6FP8 E5M2 Recall@10: 0.9096 QPS: 11717.7FP8 E5M2 Recall@10: 0.9192 QPS: 10118.3FP8 E5M2 Recall@10: 0.9321 QPS: 8950.2FP8 E5M2 Recall@10: 0.9436 QPS: 7768.2FP8 E5M2 Recall@10: 0.9529 QPS: 6535.6FP8 E5M2 Recall@10: 0.9576 QPS: 6117.0FP8 E5M2 Recall@10: 0.9667 QPS: 5179.9FP8 E5M2 Recall@10: 0.9711 QPS: 4365.4FP8 E5M2 Recall@10: 0.9752 QPS: 3949.5FP8 E5M2 Recall@10: 0.9796 QPS: 3584.9FP8 E5M2 Recall@10: 0.9854 QPS: 2875.9FP8 E5M2 Recall@10: 0.9916 QPS: 2231.3FP8 E5M2 Recall@10: 0.9976 QPS: 1389.0Recall@10

Semantic retrieval

Across all metric values from the semantic benchmark, E4M3FN had a mean absolute difference of 0.262 percentage points from FP32 and E5M2 reached 0.295 percentage points. The differences are small enough that routing preserves the ranking. Occasional small improvements can appear when approximate routing swaps nearly tied candidates, rather than because quantization improves retrieval. The plots below break this down by engine and task for Recall@10, nDCG@10 and MRR@10.

FP32FP8 E4M3FNFP8 E5M2

When this is useful

FP8 centroids are a localized optimization. They do not automatically reduce the size of the full index, because posting lists, vectors, metadata and auxiliary structures can use much more memory. The approach is useful when the routing table itself is large enough to pressure RAM, especially in hybrid architectures where the corpus lives outside memory.

These experiments make E4M3FN the more consistent FP8 option. It should still be tested with the model, metric and embedding distribution used in production. IVF either searches a posting list or skips it, so centroid quantization depends on the gap between candidate centroids.

The full paper, including the experimental protocol and result tables, is available on Springer.