VLDB 2026 Research / reviewers in the wild / expert
Patrick H. Chen
dblp:222/2938
· DBLP profile ↗
11ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-6247-6317ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attend to Fragments: How Key Information Affects Large Language Models for Factual Inconsistency DetectionabstractAs large language models (LLMs) continue to advance, a key challenge remains their tendency to hallucinate, generating fluent yet inconsistent content that lacks factual grounding. Natural language inference (NLI)-based methods, which determine whether one statement can be logically inferred from another, are widely considered the most effective for detecting input-output inconsistencies in LLMs. However, several fundamental questions, such as whether LLMs can identify relevant information to make correct factual inconsistency detections and how different arrangements of the source document affect reasoning, are not discussed in prior studies. To bridge this research gap, we design a new benchmark, KIFI, which comprises 1032 carefully selected instances from the TRUE and ScreenEval datasets, with key information annotated. Using KIFI, we show that LLMs frequently fail to use the appropriate information to make correct decisions. In addition, we find that LLMs tend to make predictions by overemphasizing certain keywords or fragments, a new phenomenon we term "Attend to Fragments". We further introduce a novel token-based permutation method to identify untrustworthy inconsistencies. Experiments show that filtering out these instances improves the overall correlation by 1.3% on the standard TRUE benchmark. The project is available at https://github.com/VibeHPC/attend-to-fragments Xindi Guo, Patrick H. Chen |
SIGIR | 3 |
| 2026 | An 8-Way Taxonomy for Multimodal Disinformation and Detection Benchmark
Shuhan Cui, Ruimin Chu, Hanrui Wang 0005, Patrick H. Chen, Ching-Chun Chang, Isao Echizen |
WWW | 4 |
| 2023 | FINGER: Fast Inference for Graph-based Approximate Nearest Neighbor SearchabstractApproximate K-Nearest Neighbor Search (AKNNS) has now become ubiquitous in modern applications, such as a fast search procedure with two-tower deep learning models. Graph-based methods for AKNNS in particular have received great attention due to their superior performance. These methods rely on greedy graph search to traverse the data points as embedding vectors in a database. Under this greedy search scheme, we make a key observation: many distance computations do not influence search updates so that these computations can be approximated without hurting performance. As a result, we propose FINGER, a fast inference method for efficient graph search in AKNNS. FINGER approximates the distance function by estimating angles between neighboring residual vectors. The approximated distance can be used to bypass unnecessary computations for faster searches. Empirically, when it comes to speeding up the inference of HNSW, which is one of the most popular graph-based AKNNS methods, FINGER significantly outperforms existing acceleration approaches and conventional libraries by 20 to 60 across different benchmark datasets. Patrick H. Chen, Wei-Cheng Chang, Jyun-Yu Jiang, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui Hsieh |
WWW | 1 |
| 2022 | ELIAS: End-to-End Learning to Index and Search in Large Output SpacesabstractExtreme multi-label classification (XMC) is a popular framework for solving many real-world problems that require accurate prediction from a very large number of potential output choices. A popular approach for dealing with the large label space is to arrange the labels into a shallow tree-based index and then learn an ML model to efficiently search this index via beam search. Existing methods initialize the tree index by clustering the label space into a few mutually exclusive clusters based on pre-defined features and keep it fixed throughout the training procedure. This approach results in a sub-optimal indexing structure over the label space and limits the search performance to the quality of choices made during the initialization of the index. In this paper, we propose a novel method ELIAS which relaxes the tree-based index to a specialized weighted graph-based index which is learned end-to-end with the final task objective. More specifically, ELIAS models the discrete cluster-to-label assignments in the existing tree-based index as soft learnable parameters that are learned jointly with the rest of the ML model. ELIAS achieves state-of-the-art performance on several large-scale extreme classification benchmarks with millions of labels. In particular, ELIAS can be up to 2.5% better at precision@$1$ and up to 4% better at recall@$100$ than existing XMC methods. A PyTorch implementation of ELIAS along with other resources is available at https://github.com/nilesh2797/ELIAS. Nilesh Gupta, Patrick H. Chen, Hsiang-Fu Yu, Cho-Jui Hsieh, Inderjit S. Dhillon |
NeurIPS | 2 |
| 2021 | DRONE: Data-aware Low-rank Compression for Large NLP ModelsabstractThe representations learned by large-scale NLP models such as BERT have been widely used in various tasks. However, the increasing model size of the pre-trained models also brings efficiency challenges, including inference speed and model size when deploying models on mobile devices. Specifically, most operations in BERT consist of matrix multiplications. These matrices are not low-rank and thus canonical matrix decomposition could not find an efficient approximation. In this paper, we observe that the learned representation of each layer lies in a low-dimensional space. Based on this observation, we propose DRONE (data-aware low-rank compression), a provably optimal low-rank decomposition of weight matrices, which has a simple closed form solution that can be efficiently computed. DRONE can be applied to both fully connected and self-attention layers appearing in the BERT model. In addition to compressing standard models, out method can also be used on distilled BERT models to further improve compression rate. Experimental results show that DRONE is able to improve both model size and inference speed with limited loss in accuracy. Specifically, DRONE alone achieves 1.92x speedup on the MRPC task with only 1.5% loss in accuracy, and when DRONE is combined with distillation, it further achieves over 12.3x speedup on various natural language inference tasks. Patrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui Hsieh |
NeurIPS | 1 |
| 2020 | Sign-OPT: A Query-Efficient Hard-label Adversarial Attack
Minhao Cheng, Simranjit Singh 0003, Patrick H. Chen, Sijia Liu 0001, Cho-Jui Hsieh |
ICLR | 3 |
| 2020 | Clustering and Constructing User Coresets to Accelerate Large-scale Top-K Recommender SystemsabstractTop-K recommender systems aim to generate few but satisfactory personalized recommendations for various practical applications, such as item recommendation for e-commerce and link prediction for social networks. However, the numbers of users and items can be enormous, thereby leading to myriad potential recommendations as well as the bottleneck in evaluating and ranking all possibilities. Existing Maximum Inner Product Search (MIPS) based methods treat the item ranking problem for each user independently and the relationship between users has not been explored. In this paper, we propose a novel model for clustering and navigating for top-K recommenders (CANTOR) to expedite the computation of top-K recommendations based on latent factor models. A clustering-based framework is first presented to leverage user relationships to partition users into affinity groups, each of which contains users with similar preferences. CANTOR then derives a coreset of representative vectors for each affinity group by constructing a set cover with a theoretically guaranteed difference to user latent vectors. Using these representative vectors in the coreset, approximate nearest neighbor search is then applied to obtain a small set of candidate items for each affinity group to be used when computing recommendations for each user in the affinity group. This approach can significantly reduce the computation without compromising the quality of the recommendations. Extensive experiments are conducted on six publicly available large-scale real-world datasets for item recommendation and personalized link prediction. The experimental results demonstrate that CANTOR significantly speeds up matrix factorization models with high precision. For instance, CANTOR can achieve 355.1x speedup for inferring recommendations in a million-user network with 99.5% [email protected] to the original system while the state-of-the-art method can only obtain 93.7x speedup with 99.0% [email protected] Jyun-Yu Jiang, Patrick H. Chen, Cho-Jui Hsieh, Wei Wang 0010 |
WWW | 2 |
| 2019 | MulCode: A Multiplicative Multi-way Model for Compressing Neural Language ModelabstractYukun Ma, Patrick H. Chen, Cho-Jui Hsieh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Patrick H. Chen, Cho-Jui Hsieh |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Learning to Screen for Fast Softmax Inference on Large Vocabulary Neural Networks
Patrick H. Chen, Si Si, Sanjiv Kumar, Yang Li 0058, Cho-Jui Hsieh |
ICLR (Poster) | 1 |
| 2019 | Efficient Contextual Representation Learning With Continuous OutputsabstractContextual representation models have achieved great success in improving various downstream natural language processing tasks. However, these language-model-based encoders are difficult to train due to their large parameter size and high computational complexity. By carefully examining the training procedure, we observe that the softmax layer, which predicts a distribution of the target word, often induces significant overhead, especially when the vocabulary size is large. Therefore, we revisit the design of the output layer and consider directly predicting the pre-trained embedding of the target word for a given context. When applied to ELMo, the proposed approach achieves a 4-fold speedup and eliminates 80% trainable parameters while achieving competitive performance on downstream tasks. Further analysis shows that the approach maintains the speed advantage under various settings, even when the sentence encoder is scaled up. Liunian Harold Li, Patrick H. Chen, Cho-Jui Hsieh, Kai-Wei Chang 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model ShrinkingabstractModel compression is essential for serving large deep neural nets on devices with limited resources or applications that require real-time responses. For advanced NLP problems, a neural language model usually consists of recurrent layers (e.g., using LSTM cells), an embedding matrix for representing input tokens, and a softmax layer for generating output tokens. For problems with a very large vocabulary size, the embedding and the softmax matrices can account for more than half of the model size. For instance, the bigLSTM model achieves state-of-the-art performance on the One-Billion-Word (OBW) dataset with around 800k vocabulary, and its word embedding and softmax matrices use more than 6GBytes space, and are responsible for over 90\% of the model parameters. In this paper, we propose GroupReduce, a novel compression method for neural language models, based on vocabulary-partition (block) based low-rank matrix approximation and the inherent frequency distribution of tokens (the power-law distribution of words). We start by grouping words into $c$ blocks based on their frequency, and then refine the clustering iteratively by constructing weighted low-rank approximation for each block, where the weights are based the frequencies of the words in the block. The experimental results show our method can significantly outperform traditional compression methods such as low-rank approximation and pruning. On the OBW dataset, our method achieved 6.6x compression rate for the embedding and softmax matrices, and when combined with quantization, our method can achieve 26x compression rate without losing prediction accuracy. Patrick H. Chen, Si Si, Yang Li 0058, Ciprian Chelba, Cho-Jui Hsieh |
NeurIPS | 1 |