Suryansh Gupta

dblp:311/9940 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Indexing and storage engines · 33% Information retrieval · 33% Database system architecture and tuning · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.912025
Cost-Effective, Low Latency Vector Search with Azure Cosmos DB · Proc. VLDB Endow. 2025
Indexing and storage engines
vector index
0.912025
Cost-Effective, Low Latency Vector Search with Azure Cosmos DB · Proc. VLDB Endow. 2025
Distributed systems
distributed database
0.312025
Cost-Effective, Low Latency Vector Search with Azure Cosmos DB · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

partitioning · 1.7DiskANN · 1.7
YearPublicationVenuePosition
2025 Cost-Effective, Low Latency Vector Search with Azure Cosmos DB
abstract
Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43× and 12× lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases.
Nitish Upreti, Harsha Vardhan Simhadri, Hari Sudan Sundar, Krishnan Sundaram, Samer Boshra, Balachandar Perumalswamy, Shivam Atri, Martin Chisholm, Revti Raman Singh, Greg Yang, Tamara Hass, Nitesh Dudhey, Subramanyam Pattipaka, Mark Hildebrand, Magdalen Dobson, Jack Moffitt, Naren Datha, Suryansh Gupta, Ravishankar Krishnaswamy, Hemeswari Varada, Sudhanshu Barthwal, Ritika Mor, James Codella, Shaun Cooper, Kevin Pilch, Simon Moreno, Aayush Kataria, Neil Deshpande, Amar Sagare, Dinesh Billa, Zishan Fu, Vipul Vishal
Proc. VLDB Endow.19