Shaun Cooper

dblp:77/4236 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Indexing and storage engines · 33% Information retrieval · 33% Database system architecture and tuning · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.912025
Cost-Effective, Low Latency Vector Search with Azure Cosmos DB · Proc. VLDB Endow. 2025
Indexing and storage engines
vector index
0.912025
Cost-Effective, Low Latency Vector Search with Azure Cosmos DB · Proc. VLDB Endow. 2025
Distributed systems
distributed database
0.312025
Cost-Effective, Low Latency Vector Search with Azure Cosmos DB · Proc. VLDB Endow. 2025
Query processing and optimization › join processing › join algorithms
hash join
0.011998
Hash Joins and Hash Teams in Microsoft SQL Server · VLDB 1998
Query processing and optimization
join processing
0.011998
Hash Joins and Hash Teams in Microsoft SQL Server · VLDB 1998

Methods — techniques the papers use, named apart from their topics

partitioning · 1.7DiskANN · 1.7
YearPublicationVenuePosition
2025 Cost-Effective, Low Latency Vector Search with Azure Cosmos DB
abstract
Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43× and 12× lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases.
Nitish Upreti, Harsha Vardhan Simhadri, Hari Sudan Sundar, Krishnan Sundaram, Samer Boshra, Balachandar Perumalswamy, Shivam Atri, Martin Chisholm, Revti Raman Singh, Greg Yang, Tamara Hass, Nitesh Dudhey, Subramanyam Pattipaka, Mark Hildebrand, Magdalen Dobson, Jack Moffitt, Naren Datha, Suryansh Gupta, Ravishankar Krishnaswamy, Hemeswari Varada, Sudhanshu Barthwal, Ritika Mor, James Codella, Shaun Cooper, Kevin Pilch, Simon Moreno, Aayush Kataria, Neil Deshpande, Amar Sagare, Dinesh Billa, Zishan Fu, Vipul Vishal
Proc. VLDB Endow.27
2005 Fast, Accurate Microarchitecture Simulation Using Statistical Phase Detection
abstract
Simulation-based microarchitecture research is often hindered by the slow speed of simulators. In this work, we propose a novel statistical technique to identify highly representative unique behaviors or phases in a benchmark based on its IPC (instructions committed per cycle) trace. By simulating the timing of only the unique phases, the cycle-accurate simulation time for the SPEC suite is reduced from 5 months to 5 days, with a significant retention of the original dynamic behavior. Evaluation across many processor configurations within the same architecture family shows that the algorithm is robust. A cost function is provided that enables users to easily optimize the parameters of the algorithm for either simulation speed or accuracy depending on preference. A new measure is introduced to quantify the ability of a simulation speedup technique to retain behavior realized in the original workload. Unlike a first order statistic such as mean value, the newly introduced measure captures important differences in dynamic behavior between the complete and the sampled simulations
Ram Srinivasan, Jeanine E. Cook, Shaun Cooper
ISPASS3
1998 Hash Joins and Hash Teams in Microsoft SQL Server
Goetz Graefe, Ross Bunker, Shaun Cooper
VLDB3
1995 Configuring a large LAN for TCP/IP, Appletalk, and IPX
abstract
This paper discusses a particular network configuration plan used at New Mexico State University which enables us to effectively manage IP, Appletalk, and IPX. The network addressing scheme is based on an IP configuration of the network. From this configuration we derive Appletalk and IPX configurations. In this manner, we configure the network only once, and benefit from this configuration each time that we implement another network protocol. This configuration has helped us manage a network supporting 5000 hosts in 100 buildings at five geographically distinct locations. We believe that our network addressing scheme is simple and, hence, easy to administer. As we demonstrate in this paper, our scheme is easily applicable to other protocols and, therefore, can be a benefit to network managers.
Shaun Cooper, Patricia J. Teller
LCN1
1988 Analogical Reasoning and Proof Discovery
Bishop Brock, Shaun Cooper, William Pierce
CADE2