<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title><![CDATA[Manas Singh - Elasticsearch Labs]]></title>
    <description><![CDATA[Articles and tutorials from the Search team at Elastic]]></description>
    <copyright><![CDATA[© 2026. Elasticsearch B.V. All Rights Reserved]]></copyright>
    <image>
      <title><![CDATA[Manas Singh - Elasticsearch Labs]]></title>
      <url>https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1121c0bf0e8a6e65/6a88da6340a1841030ef456f/search-labs-thumbnail.png</url>
      <link>https://www.elastic.co/search-labs/author/manas-singh</link>
    </image>
    <link>https://www.elastic.co/search-labs/author/manas-singh</link>
    <atom:link href="https://www.elastic.co/search-labs/rss/author/manas-singh.xml" rel="self" type="application/rss+xml"/>
    <language><![CDATA[en]]></language>
    <lastBuildDate>Thu, 08 Oct 2026 00:32:58 GMT</lastBuildDate>
  <item>
    <title><![CDATA[GPU-accelerated vector indexing in Elasticsearch with NVIDIA cuVS: 138M vectors in under 10 minutes]]></title>
    <description><![CDATA[Moving index builds to the GPU leaves the CPU free for queries, which is how vector indexing throughput went up 7x and p90 search latency fell 6x while indexing ran, with no change to recall.]]></description>
    <content:encoded><![CDATA[<p>Modern enterprise applications are ingesting terabyte- to petabyte-scale unstructured data to power semantic search, large language model–based (LLM-based) retrieval augmented generation (RAG), and recommender systems. At this scale, <a href="https://www.vastdata.com/blog/powering-enterprise-ai-high-velocity-vector-search-sql">vector indexing on CPUs can take days or even weeks</a>, slowing experimentation and making large index updates operationally expensive. This is especially challenging for workloads that require regular index rebuilds, such as those driven by frequent data updates or frequent embedding model updates. For example, several ecommerce teams refresh their product catalog nightly and train their own embedding models, each requiring an index rebuild.</p><p></p><p>CPU-based indexing can also degrade search performance when indexing and search happen simultaneously, especially in online systems, like ecommerce platforms and ad-serving pipelines. This happens because index builds consume significant CPU resources, leaving fewer cycles available for queries and driving up search latency and latency variance. Even after indexing completes, latency can remain high due to index fragmentation. Customers often run a forced merge to consolidate segments, but it can be slow on CPUs, further delaying recovery to the expected search latency by hours to days.</p><p></p><p><a href="https://www.elastic.co/elasticsearch/vector-database">Elasticsearch</a> has <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/gpu-vector-indexing">introduced GPU-accelerated vector indexing</a> powered by <a href="https://developer.nvidia.com/cuvs">NVIDIA cuVS</a>. In this blog, we show how Elasticsearch achieves up to 7x faster vector indexing throughput by offloading hierarchical navigable small world (HNSW) index construction to GPUs, indexing 138 million vectors in under 10 minutes using eight NVIDIA RTX PRO 6000 GPUs.</p><p></p><h2>Indexing over 100M Vectors in under 10 minutes</h2><p>Figure 1 shows vector indexing throughput on 138 million 1,024-dimensional vectors from the <a href="https://github.com/elastic/rally-tracks/tree/master/msmarco-v2-vector">MS MARCO</a> dataset, representing roughly 4 TB of multimodal documents (Figure 1). The benchmark was run with <a href="https://github.com/elastic/rally">Rally</a>, Elasticsearch’s macro-benchmarking framework. The test ran on a server with 8 NVIDIA RTX Pro 6000 GPUs with a 2 socket AMD EPYC 9555. GPUs were turned on for the GPU test and turned off for the CPU test. The result was that Elasticsearch achieved 7x higher indexing throughput, lowering time to index 138M vectors from 1 hour on CPUs to under 10 minutes on GPUs.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt05167c135e5e0727/6ab4e86f7a4851d33df9803f/vector-indexing-throughput.png" alt="Bar chart of vector indexing throughput: 281K docs/s on GPU versus 38K docs/s on CPU for 138M vectors" /><p></p><h3>Does GPU vector indexing change recall?</h3><p>Index acceleration isn’t helpful if indexes built on GPU have lower search quality compared to CPUs. The throughput-versus-recall curves for GPU-built and CPU-built indexes overlap, showing that GPU indexing delivers the same recall profile as Elasticsearch’s CPU-based indexing path, with minimal accuracy or search-time performance penalty (Figure 2).</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltdbb6f3deb0bcaec9/6ab4e8ebafd558dce24a0a33/qps-vs-recall.png" alt="Chart showing QPS versus recall curves overlapping for GPU-built and CPU-built vector indexes at k=100" /><p></p><h3>Search latency under concurrent indexing load</h3><p>Offloading indexing to the GPU improved search latency on the CPU by 6x when running indexing and search simultaneously (Figure 3) because indexing on GPUs frees CPU resources for faster search latency. This allows a single Elasticsearch cluster to support real-time ingestion and low-latency retrieval at the same time, without forcing a tradeoff between freshness and search performance.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltd34d2464499b03e4/6ab4e9637c5f023f526f3c67/p90-search-latency.png" alt="Bar chart of p90 search latency during indexing: 7 ms with a GPU-built index versus 45 ms on CPU" /><p></p><h3>Force merge time for HNSW indexes on GPU versus CPU</h3><p>Additionally, GPU-accelerated indexing in Elasticsearch reduced forced merge time from roughly four hours to five minutes, helping restore low-latency search in real-time (Figure 4). Note: Force merge was used to merge each shard into four segments.</p><p></p><p>Here are the full results for GPU versus CPU vector indexing in Elasticsearch:</p><p></p><p><strong>Metric</strong></p><p><strong>CPU</strong></p><p><strong>GPU</strong></p><p><strong>Change</strong></p><p>Index build time, 138 million vectors</p><p>~1 hour</p><p>Under 10 min</p><p>7x throughput</p><p>p90 search latency under indexing load</p><p>Baseline</p><p>6x lower</p><p>6x</p><p>Force merge, four segments per shard</p><p>~4 hours</p><p>~5 min</p><p>~48x</p><p>Recall at target</p><p>95%</p><p>95%</p><p>Unchanged</p><p></p><p></p><h2>What GPU vector indexing means for Elasticsearch users</h2><p>Elasticsearch’s integration of NVIDIA cuVS brings GPU-accelerated vector indexing to production AI search with over 84% reduction in search latency to deliver faster ingest with a 640% increase in indexing throughput with a lower CPU overhead and limited infrastructure tradeoffs. This unlocks a new class of applications, such as faster RAG pipelines, semantic search, multimodal retrieval, and fraud detection at scale.</p><p></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt54b3129de4981654/6ab4e9c7bf5efbe04ad38b76/force-merge-time.png" alt="Bar chart of force merge time for a 138M vector index: 5 minutes on GPU versus 249 minutes on CPU" /><p></p><h2>Get started with GPU vector indexing in Elasticsearch</h2><p>Install Elasticsearch with GPU-accelerated vector indexing by following <a href="https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/gpu-vector-indexing">these instructions</a>. Replicate the above benchmarks by running<a href="https://github.com/elastic/rally-tracks/tree/master/msmarco-v2-vector"> rally msmarco track</a> with parameters in <a href="https://github.com/elastic/rally-tracks/issues/1188">the GitHub issue</a>. Learn more about GPU-accelerated vector indexing, search, and preprocessing by visiting<a href="https://github.com/NVIDIA/cuvs"> NVIDIA cuVS</a>. </p><h2>Acknowledgments</h2><p>The authors would like to thank Chris Hegarty, Gilad Gal, and Mayya Sharipova from Elastic, as well as Nathan Stephens from NVIDIA, for their contributions to this article.</p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/vector-indexing-gpu-elasticsearch-cuvs</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/vector-indexing-gpu-elasticsearch-cuvs</guid>
    <category><![CDATA[Vector Database]]></category>
    <category><![CDATA[Inside Elastic]]></category>
    <category><![CDATA[Index Data]]></category>
    <dc:creator><![CDATA[Bao Tong,Manas Singh,Corey Nolet]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt2415fbdb68d01a49/6ab4e2a39a8267823dc47d59/NVIDIA.jpg" length="0" type="image/jpeg"/>
    <pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate>
  </item>
  <item>
    <title><![CDATA[Up to 12x Faster Vector Indexing in Elasticsearch with NVIDIA cuVS: GPU-acceleration Chapter 2]]></title>
    <description><![CDATA[Discover how Elasticsearch achieves nearly 12x higher indexing throughput with GPU-accelerated vector indexing and NVIDIA cuVS.]]></description>
    <content:encoded><![CDATA[<p>Earlier this year, Elastic announced the <a href="https://ir.elastic.co/news/news-details/2025/Elastic-Brings-Enterprise-Data-to-NVIDIA-AI-Factories/default.aspx">collaboration</a> with NVIDIA to bring GPU acceleration to Elasticsearch, integrating with <a href="https://developer.nvidia.com/cuvs">NVIDIA cuVS</a>—as detailed in a <a href="https://www.nvidia.com/en-us/on-demand/session/gtc25-S71286/">session at NVIDIA GTC</a> and various <a href="https://www.elastic.co/search-labs/blog/gpu-accelerated-vector-search-elasticsearch-nvidia">blogs</a>. This post is an update on the co-engineering effort with the NVIDIA vector search team.</p><h2>Recap</h2><p>First, let’s bring you up to speed. Elasticsearch has established itself as a powerful vector database, offering a rich set of features and strong performance for large-scale similarity search. With capabilities such as scalar quantization, Better Binary Quantization (<a href="https://www.elastic.co/search-labs/blog/better-binary-quantization-lucene-elasticsearch">BBQ</a>), <a href="https://www.elastic.co/blog/accelerating-vector-search-simd-instructions">SIMD</a> vector operations, and more disk-efficient algorithms like <a href="https://www.elastic.co/search-labs/blog/diskbbq-elasticsearch-introduction">DiskBBQ</a>, it already provides efficient and flexible options for managing vector workloads.</p><p>By integrating NVIDIA cuVS as a callable module for vector search tasks, we aim to deliver significant gains in vector indexing performance and efficiency to better support large-scale vector workloads.</p><h2>The challenge</h2><p>One of the toughest challenges in building a high-performance vector database is constructing the vector index - the <a href="https://arxiv.org/abs/1603.09320">HNSW</a> graph. Index building quickly becomes dominated by millions or even billions of arithmetic operations as every vector is compared against many others. In addition, index lifecycle operations, such as compaction and merges, can further increase the overall compute overhead of indexing. As data volumes and associated vector embeddings grow exponentially, accelerated computing GPUs, built for massive parallelism and high-throughput math, are ideally positioned to handle these workloads.</p><h2>Enter the Elasticsearch-GPU Plugin</h2><p><a href="https://developer.nvidia.com/cuvs">NVIDIA cuVS</a> is an open-source CUDA-X library for GPU-accelerated vector search and data clustering that enables fast index building and embedding retrieval for AI and recommendation workloads.</p><p>Elasticsearch uses cuVS through <a href="https://mvnrepository.com/artifact/com.nvidia.cuvs/cuvs-java">cuvs-java</a>, an open-source library developed by the community and maintained by NVIDIA. The cuvs-java library is lightweight and builds on the <a href="https://docs.nvidia.com/cuvs/api-reference/c-api-core-c-api">cuVS C API</a> using <a href="https://openjdk.org/projects/panama/">Panama</a> Foreign Function to expose cuVS features in an idiomatic Java way, while remaining modern and performant.</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltc7fd7361099da05a/6a17e920be608670af00477f/5f6daa1eb07f704a6707d9e6b7ccb81d0abaa8c9-566x419.png" alt="How Elasticsearch works with NVIDIA cuVS, CPU and GPU indexing" /><p>The cuvs-java library is integrated into a <a href="https://github.com/elastic/elasticsearch/pull/135545">new Elasticsearch plugin</a>; therefore, vector indexing on the GPU can occur on the same Elasticsearch node and process, without the need to provision any external code or hardware. During index building, if the cuVS library is installed and a GPU is present and configured, Elasticsearch will use the GPU to accelerate the vector indexing process. The vectors are given to the GPU, which constructs a <a href="https://arxiv.org/abs/2308.15136">CAGRA</a> graph. This graph is then converted to the HNSW format, making it immediately available for vector search on the CPU. The final format of the built graph is the same as what would be built on the CPU; this allows Elasticsearch to leverage GPUs for high-throughput vector indexing when the underlying hardware supports it, while freeing CPU power for other tasks (concurrent search, data processing, etc.).</p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt485f55f29d6df5c4/6a17e922be6086dcf3004785/3ea255bd9bfd7983f78143c5eba999d2149d72be-671x356.png" alt="" /><h2>Index build acceleration</h2><p>As part of integrating GPU acceleration into Elasticsearch, several enhancements were made to cuvs-java, focusing on efficient data input/output and function invocation. A key enhancement is the use of <a href="https://github.com/rapidsai/cuvs/blob/2cf5fa7666d703dccbe655f8214656b0952bb69b/java/cuvs-java/src/main/java/com/nvidia/cuvs/CuVSMatrix.java">cuVSMatrix</a> to transparently model vectors, whether they reside on the Java heap, off-heap, or in GPU memory. This enables data to move efficiently between memory and the GPU, avoiding unnecessary copies of potentially billions of vectors.</p><p>Thanks to this underlying zero-copy abstraction, both transferring to GPU memory and retrieving the graph can occur directly. During indexing, vectors are first buffered in memory on the Java heap, then sent to the GPU to construct the CAGRA graph. The graph is subsequently retrieved from the GPU, converted into HNSW format, and persisted to disk.</p><p>At merge time, the vectors are already stored on disk, bypassing the Java heap entirely. Index files are memory-mapped, and data is transferred directly into GPU memory. The design also easily accommodates different bit-widths, such as float32 or int8, and naturally extends to other quantization schemes.</p><h2>Drumroll…so, how does it perform?</h2><p>Before we get into the numbers, a bit of context is helpful. Segment merging in Elasticsearch typically runs automatically in the background during indexing, which makes it difficult to benchmark in isolation. To obtain reproducible results, we used force-merge to explicitly trigger segment merging in a controlled experiment. Since force-merge performs the same underlying merge operations as background merging, its performance serves as a useful indicator of expected improvements, even though the exact gains may differ in real-world indexing workloads.</p><p>Now, let’s see the numbers.</p><p>Our initial benchmark results are very promising. We ran the benchmark on an AWS <code>g6.4xlarge</code> instance with locally attached NVMe storage. A single node of Elasticsearch was configured to use the default, optimal number of indexing threads (8 - one for each physical core), and to disable <a href="https://www.elastic.co/docs/reference/elasticsearch/index-settings/merge">merge throttling</a> (which is less applicable with fast NVMe disks).</p><p>For the dataset, we used 2.6 million vectors with 1,536 dimensions from the <a href="https://github.com/elastic/rally-tracks/blob/master/openai_vector/README.md">OpenAI Rally vector track</a>, encoded as <a href="https://github.com/elastic/elasticsearch/pull/137072">base64 strings</a>, and indexed as float32 <em>hnsw</em>. In all scenarios, the constructed graphs achieve recall levels of up to 95%. Here’s what we found:</p><ul><li><p><strong>Indexing Throughput:</strong> By moving graph construction to the GPU during in-memory buffer flushes, we increase throughput by ~12x.</p></li><li><p><strong>Force-merge:</strong> After indexing completes, the GPU continues to accelerate segment merging, speeding up the force-merge phase by ~7x.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/bltfea4ee13b5a3b10d/6a17e923e9ea879c6aa9c616/f60ea9ee5996e456f393ffd195ee7eada6e5a7c2-948x387.png" alt="" /><ul><li><p><strong>CPU usage:</strong> Offloading graph construction to the GPU significantly reduces both average and peak CPU utilization. The graphs below illustrate CPU usage during indexing and merging, highlighting how much lower it is when these operations run on the GPU. Lower CPU utilization during GPU indexing frees up CPU cycles that can be redirected to improve search performance.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt80ff9c53f9b6884a/6a17e925445de9ee4b4d0187/5e680a5fc41700a877f3d8b2e5ce18ebd3f37a0b-1600x562.png" alt="" /><ul><li><p><strong>Recall:</strong> Accuracy remains effectively the same between CPU and GPU runs, with the GPU-built graph reaching marginally higher recall.</p></li></ul><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt5cbe084eca27b8e4/6a17e926faa913317093c8b7/48a2b7758606bd321712b7d8378cd2640e652a4e-1384x544.png" alt="" /><h2>Comparing along another dimension: Price</h2><p>The earlier comparison intentionally used identical hardware, with the only difference being whether the GPU was used during indexing. That setup is useful for isolating raw compute effects, but we can also look at the comparison from a cost perspective.</p><p>At roughly the same hourly price as the GPU-accelerated configuration, one can provision a CPU-only setup with approximately twice the comparable CPU and memory resources: 32 vCPUs (AMD EPYC) and 64 GB of RAM, allowing to double the number of indexing threads to 16.</p><p>To keep the comparison fair and consistent, we ran this CPU-only experiment on an AWS g6.8xlarge instance, with the GPU explicitly disabled. This allowed us to hold all other hardware characteristics constant while evaluating the cost–performance trade-off of GPU acceleration versus CPU-only indexing.</p><p>The more powerful CPU instance does show improved performance compared to the benchmarks in the above section, as you would expect. However, when we compare this more powerful CPU instance against the original GPU-accelerated results, the GPU still delivers substantial performance gains: <strong>~5x</strong> improvement in indexing throughput, and <strong>~6x </strong>in force merge, all while building graphs that achieve recall levels of up to <strong>95%.</strong></p><img src="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt94b5eb6f95ba307d/6a17e928abe0f255d4dfea35/8ffa58cae3ad175ef2932a351aeef4c34a1407b9-948x394.png" alt="" /><h2>Conclusion</h2><p>In end-to-end scenarios, GPU acceleration with NVIDIA cuVS delivers nearly a 12x improvement in indexing throughput and a 7x decrease in force-merge latency, with significantly lower CPU utilization. This shows that vector indexing and merge workloads benefit significantly from GPU acceleration. On a cost-adjusted comparison, GPU acceleration continues to yield substantial performance gains, with approximately 5x higher indexing throughput and 6x faster force-merge operations.</p><p>GPU-accelerated vector indexing is currently planned for Tech Preview in Elasticsearch 9.3, which is scheduled to be released early in 2026.</p><p>Stay tuned for more.</p>]]></content:encoded>
    <link>https://www.elastic.co/search-labs/blog/elasticsearch-gpu-accelerated-vector-indexing-nvidia</link>
    <guid isPermaLink="true">https://www.elastic.co/search-labs/blog/elasticsearch-gpu-accelerated-vector-indexing-nvidia</guid>
    <category><![CDATA[Vector Database]]></category>
    <dc:creator><![CDATA[Chris Hegarty,Hemant Malik,Corey Nolet,Manas Singh,Mithun Radhakrishnan,Mayya Sharipova,Lorenzo Dematte,Ben Frederickson]]></dc:creator>
    <enclosure url="https://static-www.elastic.co/v3/assets/bltefdd0b53724fa2ce/blt1248d51633bd75d9/6a17e92ae9ea8714b3a9c61a/08f7469a4daaf67b7c5999585aae179b6680c78d-896x746.png" length="0" type="image/png"/>
    <pubDate>Wed, 03 Dec 2025 00:00:00 GMT</pubDate>
  </item>
  </channel>
</rss>