Fabien André

dblp:146/7818 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Storage systems · 34% Processor architecture and microarchitecture · 22% Cloud and datacenter computing · 17%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search
nearest neighbor search
0.722021
Quicker ADC : Unlocking the Hidden Potential of Product Quantization With SIMD · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan · Proc. VLDB Endow. 2015
Information retrieval › similarity search › vector quantization
product quantization
0.512021
Quicker ADC : Unlocking the Hidden Potential of Product Quantization With SIMD · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Processor architecture and microarchitecture
SIMD
0.512021
Quicker ADC : Unlocking the Hidden Potential of Product Quantization With SIMD · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Parallel and multicore computing › parallel computing › parallel optimization
SIMD optimization
0.212015
Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan · Proc. VLDB Endow. 2015
Storage systems
archival storage
0.212014
Archiving cold data in warehouses with clustered network coding · EuroSys 2014
Storage systems › storage reliability
erasure coding
0.212014
Archiving cold data in warehouses with clustered network coding · EuroSys 2014
Storage systems › storage reliability › erasure coding
network coding
0.212014
Archiving cold data in warehouses with clustered network coding · EuroSys 2014
Storage systems
storage reliability
0.212014
Archiving cold data in warehouses with clustered network coding · EuroSys 2014
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.112021
Quicker ADC : Unlocking the Hidden Potential of Product Quantization With SIMD · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Operating systems › network stack
kernel networking
0.112018
Don't share, Don't lock: Large-scale Software Connection Tracking with Krononat · USENIX ATC 2018
Memory systems › data locality
cache locality
0.112015
Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan · Proc. VLDB Endow. 2015
Cloud and datacenter computing
datacenter storage
0.112014
Archiving cold data in warehouses with clustered network coding · EuroSys 2014

Methods — techniques the papers use, named apart from their topics

product quantization · 1.4SIMD · 1.4lookup table · 1.0
YearPublicationVenuePosition
2021 Quicker ADC : Unlocking the Hidden Potential of Product Quantization With SIMD
abstract
Efficient Nearest Neighbor (NN) search in high-dimensional spaces is a foundation of many multimedia retrieval systems. A common approach is to rely on Product Quantization, which allows the storage of large vector databases in memory and efficient distance computations. Yet, implementations of nearest neighbor search with Product Quantization have their performance limited by the many memory accesses they perform. Following this observation, André et al. proposed Quick ADC with up to 6× faster implementations of PQ m×4 product quantizers (PQ) leveraging specific SIMD instructions. Quicker ADC is a generalization of Quick ADC not limited to PQ m×4 codes and supporting AVX-512, the latest revision of SIMD instruction set. In doing so, Quicker ADC faces the challenge of using efficiently 5,6 and 7-bit shuffles that do not align to computer bytes or words. To this end, we introduce (i) irregular product quantizers combining sub-quantizers of different granularity and (ii) split tables allowing lookup tables larger than registers. We evaluate Quicker ADC with multiple indexes including Inverted Multi-Indexes and IVF HNSW and show that it outperforms the reference optimized implementations (i.e., FAISS and polysemous codes) for numerous configurations. Finally, we release an open-source fork of FAISS enhanced with Quicker ADC.
Fabien André, Anne-Marie Kermarrec, Nicolas Le Scouarnec
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Don't share, Don't lock: Large-scale Software Connection Tracking with Krononat
Fabien André, Stéphane Gouache, Nicolas Le Scouarnec, Antoine Monsifrot
USENIX ATC1
2017 Accelerated Nearest Neighbor Search with Quick ADC
abstract
Efficient Nearest Neighbor (NN) search in high-dimensional spaces is a foundation of many multimedia retrieval systems. Because it offers low responses times, Product Quantization (PQ) is a popular solution. PQ compresses high-dimensional vectors into short codes using several sub-quantizers, which enables in-RAM storage of large databases. This allows fast answers to NN queries, without accessing the SSD or HDD. The key feature of PQ is that it can compute distances between short codes and high-dimensional vectors using cache-resident lookup tables. The efficiency of this technique, named Asymmetric Distance Computation (ADC), remains limited because it performs many cache accesses.
Fabien André, Anne-Marie Kermarrec, Nicolas Le Scouarnec
ICMR1
2015 Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan
abstract
Nearest Neighbor (NN) search in high dimension is an important feature in many applications (e.g., image retrieval, multimedia databases). Product Quantization (PQ) is a widely used solution which offers high performance, i.e., low response time while preserving a high accuracy. PQ represents high-dimensional vectors (e.g., image descriptors) by compact codes. Hence, very large databases can be stored in memory, allowing NN queries without resorting to slow I/O operations. PQ computes distances to neighbors using cache-resident lookup tables, thus its performance remains limited by (i) the many cache accesses that the algorithm requires, and (ii) its inability to leverage SIMD instructions available on modern CPUs. In this paper, we advocate that cache locality is not sufficient for efficiency. To address these limitations, we design a novel algorithm, PQ Fast Scan, that transforms the cache-resident lookup tables into small tables, sized to fit SIMD registers. This transformation allows (i) in-register lookups in place of cache accesses and (ii) an efficient SIMD implementation. PQ Fast Scan has the exact same accuracy as PQ, while having 4 to 6 times lower response time (e.g., for 25 million vectors, scan time is reduced from 74ms to 13ms).
Fabien André, Anne-Marie Kermarrec, Nicolas Le Scouarnec
Proc. VLDB Endow.1
2014 Archiving cold data in warehouses with clustered network coding
abstract
Modern storage systems now typically combine plain replication and erasure codes to reliably store large amount of data in datacenters. Plain replication allows a fast access to popular data, while erasure codes, e.g., Reed-Solomon codes, provide a storage-efficient alternative for archiving less popular data. Although erasure codes are now increasingly employed in real systems, they experience high overhead during maintenance, i.e., upon failures, typically requiring files to be decoded before being encoded again to repair the encoded blocks stored at the faulty node.
Fabien André, Anne-Marie Kermarrec, Erwan Le Merrer, Nicolas Le Scouarnec, Gilles Straub, Alexandre van Kempen
EuroSys1