VLDB 2026 Research / reviewers in the wild / expert
Penghao Chen
dblp:348/4362
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0006-1223-7787ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 83% Data mining · 13% Machine learning and data management · 4% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search |
2.5 | 3 | 2025 | LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN Search · WWW 2025 Towards Accurate Distance Estimation for Distribution-Aware c-ANN Search · ICDE 2025 Efficient Data-aware Distance Comparison Operations for High-Dimensional Approximate Nearest Neighbor Search · Proc. VLDB Endow. 2024 |
Information retrieval › similarity search
nearest neighbor search |
1.6 | 2 | 2025 | LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN Search · WWW 2025 Efficient Data-aware Distance Comparison Operations for High-Dimensional Approximate Nearest Neighbor Search · Proc. VLDB Endow. 2024 |
Information retrieval › hashing › hashing for nearest neighbor search
locality-sensitive hashing |
0.9 | 1 | 2025 | Towards Accurate Distance Estimation for Distribution-Aware c-ANN Search · ICDE 2025 |
Data mining
distance estimation |
0.8 | 1 | 2024 | Efficient Data-aware Distance Comparison Operations for High-Dimensional Approximate Nearest Neighbor Search · Proc. VLDB Endow. 2024 |
Algorithms and data structures › similarity search
high-dimensional similarity search |
0.3 | 1 | 2025 | Towards Accurate Distance Estimation for Distribution-Aware c-ANN Search · ICDE 2025 |
Methods — techniques the papers use, named apart from their topics
unbiased distance estimation · 1.7data distribution modeling · 1.7learning-based redundancy strategy · 0.9learned probing model · 0.9hypothesis testing · 0.8dimensionality reduction · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Accurate Distance Estimation for Distribution-Aware c-ANN SearchabstractLocality sensitive hashing (LSH) is a representative approach for nearest neighbor (NN) search in high-dimensional spaces, which is able to answer c-approximate NN (c-ANN) queries in sublinear time with constant probability. Existing advanced LSH methods leverage a plurality of novel techniques such as query-aware dynamic bucketing, virtual rehashing, and efficient indexing to achieve state-of-the-art performance. However, they rely on similar random LSH functions, which provides distance estimations that are irrelevant to the given data distribution. Therefore, the quality of the searched candi-dates is suboptimal. In this study, we reformulate the c-ANN query from the perspective of data distribution. Specifically, we propose a novel distribution-aware c-ANN query, which can guarantee the quality of searched results from the query distribution perspective. We introduce an accurately unbiased distance estimator into LSH methods, which can provide more precise distance estimations by modeling the data distribution. We also conduct rigorous theoretical analysis to prove that our methods can correctly answer the distribution-aware c-ANN query with at least a constant probability. Experiments on seven real datasets with different sizes and dimensionalities indicate that the proposed method can achieve better performance than existing LSH methods in terms of efficiency and effectiveness. Liwei Deng 0001, Penghao Chen, Ximu Zeng, Yuchen Fang 0001, Jin Chen 0008, Yan Zhao 0008 |
ICDE | 2 |
| 2025 | LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN SearchabstractApproximate nearest neighbor search is fundamental in information retrieval. Previous partition-based methods enhance search efficiency by probing partial partitions, yet they face two common issues. In the query phase, a common strategy is to probe partitions based on the distance ranks of a query to partition centroids, which inevitably probes irrelevant partitions as it ignores data distribution. In the partition construction phase, all partition-based methods face the boundary problem that separates a query's nearest neighbors to multiple partitions, resulting in a long-tailed kNN distribution and degrading the optimal nprobe (i.e., the number of probing partitions). To address this gap, we propose LIRA, a LearnIng-based queRy-aware pArtition framework. Specifically, we propose a probing model to directly probe the partitions containing the kNN of a query, which can reduce probing waste and allow for query-aware probing with nprobe individually. Moreover, we incorporate the probing model into a learning-based redundancy strategy to mitigate the adverse impact of the long-tailed kNN distribution on search efficiency. Extensive experiments on real-world vector datasets demonstrate the superiority of LIRA in the trade-off among accuracy, latency, and query fan-out. The codes are available at https://github.com/SimoneZeng/LIRA-ANN-search. Ximu Zeng, Liwei Deng 0001, Penghao Chen, Xu Chen 0023, Han Su 0001, Kai Zheng 0001 |
WWW | 3 |
| 2024 | Efficient Data-aware Distance Comparison Operations for High-Dimensional Approximate Nearest Neighbor SearchabstractHigh-dimensional approximate K nearest neighbor search (AKNN) is a fundamental task for various applications, including information retrieval. Most existing algorithms for AKNN can be decomposed into two main components, i.e., candidate generation and distance comparison operations (DCOs). While different methods have unique ways of generating candidates, they all share the same DCO process. In this study, we focus on accelerating the process of DCOs that dominates the time cost in most existing AKNN algorithms. To achieve this, we propose an Data-Aware Distance Estimation approach, called DADE , which approximates the exact distance in a lower-dimensional space. We theoretically prove that the distance estimation in DADE is unbiased in terms of data distribution. Furthermore, we propose an optimized estimation based on the unbiased distance estimation formulation. In addition, we propose a hypothesis testing approach to adaptively determine the number of dimensions needed to estimate the exact distance with sufficient confidence. We integrate DADE into widely-used AKNN search algorithms, e.g., IVF and HNSW , and conduct extensive experiments to demonstrate the superiority. Liwei Deng 0001, Penghao Chen, Ximu Zeng, Tianfu Wang 0002, Yan Zhao 0008, Kai Zheng 0001 |
Proc. VLDB Endow. | 2 |