Leonardo Kuffó

dblp:363/8018 · also Leonardo Xavier Kuffó Rivero · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-3575-0528ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Bang for the Buck: Vector Search on Cloud CPUs
abstract
Vector databases have emerged as a new type of systems that support efficient querying of high-dimensional vectors.Many of these offer their database as a service in the cloud.However, the variety of available CPUs and the lack of vector search benchmarks across CPUs make it difficult for users to choose one.In this study, we show that CPU microarchitectures available in the cloud perform significantly differently across vector search scenarios.For instance, in an IVF index on float32 vectors, AMD's Zen4 gives almost 3x more queries per second (QPS) compared to Intel's Sapphire Rapids, but for HNSW indexes, the tables turn.However, when looking at the number of queries per dollar (QP$), Graviton3 is the best option for most indexes and quantization settings, even over Graviton4 (Table 1).With this work, we hope to guide users in getting the best "bang for the buck" when deploying vector search systems.
Leonardo Kuffó, Peter Boncz
DaMoN1
2025 PDX: A Data Layout for Vector Similarity Search
abstract
We propose Partition Dimensions Across (PDX), a data layout for vectors (e.g., embeddings) that, similar to PAX [6], stores multiple vectors in one block, using a vertical layout for the dimensions (Figure 1). PDX accelerates exact and approximate similarity search thanks to its dimension-by-dimension search strategy that operates on multiple-vectors-at-a-time in tight loops. It beats SIMD-optimized distance kernels on standard horizontal vector storage (avg 40% faster), only relying on scalar code that gets auto-vectorized. We combined the PDX layout with recent dimension-pruning algorithms ADSampling [19] and BSA [52] that accelerate approximate vector search. We found that these algorithms on the horizontal vector layout can lose to SIMD-optimized linear scans, even if they are SIMD-optimized. However, when used on PDX, their benefit is restored to 2-7x. We find that search on PDX is especially fast if a limited number of dimensions has to be scanned fully, which is what the dimension-pruning approaches do. We finally introduce PDX-BOND, an even more flexible dimension-pruning strategy, with good performance on exact search and reasonable performance on approximate search. Unlike previous pruning algorithms, it can work on vector data ''as-is'' without preprocessing; making it attractive for vector databases with frequent updates.
Leonardo Kuffó, Elena Krippner, Peter Boncz
Proc. ACM Manag. Data1
2023 ALP: Adaptive Lossless floating-Point Compression
abstract
IEEE 754 doubles do not exactly represent most real values, introducing rounding errors in computations and [de]serialization to text. These rounding errors inhibit the use of existing lightweight compression schemes such as Delta and Frame Of Reference (FOR), but recently new schemes were proposed: Gorilla, Chimp128, PseudoDecimals (PDE), Elf and Patas. However, their compression ratios are not better than those of general-purpose compressors such as Zstd; while [de]compression is much slower than Delta and FOR. We propose and evaluate ALP, that significantly improves these previous schemes in both speed and compression ratio (Figure 1). We created ALP after carefully studying the datasets used to evaluate the previous schemes. To obtain speed, ALP is designed to fit vectorized execution. This turned out to be key for also improving the compression ratio, as we found in-vector commonalities to create compression opportunities. ALP is an adaptive scheme that uses a strongly enhanced version of PseudoDecimals [31] to losslessly encode doubles as integers if they originated as decimals, and otherwise uses vectorized compression of the doubles' front bits. Its high speeds stem from our implementation in scalar code that auto-vectorizes, using building blocks provided by our FastLanes library [6], and an efficient two-stage compression algorithm that first samples row-groups and then vectors.
Azim Afroozeh, Leonardo Kuffó, Peter Boncz
Proc. ACM Manag. Data2
2018 Know your customer: Detection of Customer Experience (CX) in Social Platforms using Text Categorization
abstract
Customers nowadays are one online post away from their stores, specially when it comes to post-shopping experiences. This translates to large amounts of text messages to evaluate and process for big brands that aim to maintain a good quality of service as well as a digital channel of communication for their customers. Automating the understanding of this text data poses questions such as how large the corpus should be and which are the best algorithms to discriminate whether a social media post is related or not to customer experience (CX). In order to help answering these questions, first, we get hold of posts from three different platforms: Foursquare (77K) , Twitter (153K) and Facebook (2.2M). Such posts are directed to brands ranked in the ForeSee CX Index and the Forrester CX Index rankings. Second, we build a binary classifier using different algorithms to identify customer experience posts on a social platform. The accuracy of the best performing setting is 86.4% for Facebook and 91.2% for Twitter. Third, we explore the effect of increasing the number of training samples, and how a plateau is reached after 5K posts. Finally, we conduct experiments using different combinations of n-grams as features for the text mining process. As a result we observe that uni-grams and bi-grams are the best combination when we need to choose features for a classifier discriminating customer experience social media posts on Twitter and a combination of up to four-grams on Facebook.
Leonardo Kuffó, Carmen Vaca, Edgar Izquierdo, Juan Carlos Bustamante 0003
IEEE BigData1