Minghan Li 0002

dblp:214/2450-2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2024
0009-0007-8972-7714ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Unifying Multimodal Retrieval via Document Screenshot Embedding
abstract
In the real world, documents are organized in different formats and varied modalities.Traditional retrieval pipelines require tailored document parsing techniques and content extraction modules to prepare input for indexing.This process is tedious, prone to errors, and has information loss.To this end, we propose Document Screenshot Embedding (DSE), a novel retrieval paradigm that regards document screenshots as a unified input format, which does not require any content extraction preprocess and preserves all the information in a document (e.g., text, image and layout).DSE leverages a large vision-language model to directly encode document screenshots into dense representations for retrieval.To evaluate our method, we first craft the dataset of Wiki-SS, a 1.3M Wikipedia web page screenshots as the corpus to answer the questions from the Natural Questions dataset.In such a text-intensive document retrieval setting, DSE shows competitive effectiveness compared to other text retrieval methods relying on parsing.For example, DSE outperforms BM25 by 17 points in top-1 retrieval accuracy.Additionally, in a mixed-modality task of slide retrieval, DSE significantly outperforms OCR text retrieval methods by over 15 points in [email protected] experiments show that DSE is an effective document retrieval paradigm for diverse types of documents.Model checkpoints, code, and Wiki-SS collection are released at http://tevatron.ai.
Xueguang Ma, Sheng-Chieh Lin, Minghan Li 0002, Wenhu Chen, Jimmy Lin
EMNLP3
2024 Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
abstract
Large language models (LLMs) often hallucinate and lack the ability to provide attribution for their generations. Semi-parametric LMs, such as kNN-LM, approach these limitations by refining the output of an LM for a given prompt using its nearest neighbor matches in a non-parametric data store. However, these models often exhibit slow inference speeds and produce non-fluent texts. In this paper, we introduce Nearest Neighbor Speculative Decoding (NEST), a novel semi-parametric language modeling approach that is capable of incorporating real-world text spans of arbitrary length into the LM generations and providing attribution to their sources. NEST performs token-level retrieval at each inference step to compute a semi-parametric mixture distribution and identify promising span continuations in a corpus. It then uses an approximate speculative decoding procedure that accepts a prefix of the retrieved span or generates a new token. NEST significantly enhances the generation quality and attribution rate of the base LM across a variety of knowledge-intensive tasks, surpassing the conventional kNN-LM method and performing competitively with in-context retrieval augmentation. In addition, NEST substantially improves the generation speed, achieving a 1.8x speedup in inference time when applied to Llama-2-Chat 70B. Code will be released at https://github.com/facebookresearch/NEST/tree/main.
Minghan Li 0002, Xilun Chen 0002, Ari Holtzman, Beidi Chen, Jimmy Lin, Scott Yih, Xi Victoria Lin
NeurIPS1
2024 Can Query Expansion Improve Generalization of Strong Cross-Encoder Rankers?
abstract
Query expansion has been widely used to improve the search results of first-stage retrievers, yet its influence on second-stage, cross-encoder rankers remains under-explored. A recent study shows that current expansion techniques benefit weaker models but harm stronger rankers. In this paper, we re-examine this conclusion and raise the following question: Can query expansion improve generalization of strong cross-encoder rankers? To answer this question, we first apply popular query expansion methods to different cross-encoder rankers and verify the deteriorated zero-shot effectiveness. We identify two vital steps in the experiment: high-quality keyword generation and minimally-disruptive query modification. We show that it is possible to improve the generalization of a strong neural ranker, by generating keywords through a reasoning chain and aggregating the ranking results of each expanded query via self-consistency, reciprocal rank weighting, and fusion. Experiments on BEIR and TREC Deep Learning 2019/2020 show that the nDCG@10 scores of both MonoT5 and RankT5 following these steps are improved, which points out a direction for applying query expansion to strong cross-encoder rankers.
Minghan Li 0002, Honglei Zhuang, Kai Hui 0001, Zhen Qin 0001, Jimmy Lin, Rolf Jagerman, Xuanhui Wang, Michael Bendersky
SIGIR1
2023 CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval
abstract
Minghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Minghan Li 0002, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Scott Yih, Xilun Chen 0002
ACL (1)1
2023 SLIM: Sparsified Late Interaction for Multi-Vector Retrieval with Inverted Indexes
abstract
This paper introduces Sparsified Late Interaction for Multi-vector (SLIM) retrieval with inverted indexes. Multi-vector retrieval methods have demonstrated their effectiveness on various retrieval datasets, and among them, ColBERT is the most established method based on the late interaction of contextualized token embeddings of pre-trained language models. However, efficient ColBERT implementations require complex engineering and cannot take advantage of off-the-shelf search libraries, impeding their practical use. To address this issue, SLIM first maps each contextualized token vector to a sparse, high-dimensional lexical space before performing late interaction between these sparse token embeddings. We then introduce an efficient two-stage retrieval architecture that includes inverted index retrieval followed by a score refinement module to approximate the sparsified late interaction, which is fully compatible with off-the-shelf lexical search libraries such as Lucene. SLIM achieves competitive accuracy on MS MARCO Passages and BEIR compared to ColBERT while being much smaller and faster on CPUs. To our knowledge, we are the first to explore using sparse token representations for multi-vector retrieval. Source code and data are integrated into the Pyserini IR toolkit.
Minghan Li 0002, Sheng-Chieh Lin, Xueguang Ma, Jimmy Lin
SIGIR1
2023 Aggretriever: A Simple Approach to Aggregate Textual Representations for Robust Dense Passage Retrieval
abstract
Abstract Pre-trained language models have been successful in many knowledge-intensive NLP tasks. However, recent work has shown that models such as BERT are not “structurally ready” to aggregate textual information into a [CLS] vector for dense passage retrieval (DPR). This “lack of readiness” results from the gap between language model pre-training and DPR fine-tuning. Previous solutions call for computationally expensive techniques such as hard negative mining, cross-encoder distillation, and further pre-training to learn a robust DPR model. In this work, we instead propose to fully exploit knowledge in a pre-trained language model for DPR by aggregating the contextualized token embeddings into a dense vector, which we call agg★. By concatenating vectors from the [CLS] token and agg★, our Aggretriever model substantially improves the effectiveness of dense retrieval models on both in-domain and zero-shot evaluations without introducing substantial training overhead. Code is available at https://github.com/castorini/dhr.
Sheng-Chieh Lin, Minghan Li 0002, Jimmy Lin
Trans. Assoc. Comput. Linguistics2
2022 Another Look at DPR: Reproduction of Training and Replication of Retrieval
Xueguang Ma, Ronak Pradeep, Minghan Li 0002, Jimmy Lin
ECIR (1)4
2022 Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking
abstract
In information retrieval (IR), candidate set pruning has been commonly used to speed up twostage relevance ranking.However, such an approach lacks accurate error control and often trades accuracy against computational efficiency in an empirical fashion, missing theoretical guarantees.In this paper, we propose the concept of certified error control of candidate set pruning for relevance ranking, which means that the test error after pruning is guaranteed to be controlled under a user-specified threshold with high probability.Both in-domain and outof-domain experiments show that our method successfully prunes the first-stage retrieved candidate sets to improve the second-stage reranking speed while satisfying the pre-specified accuracy constraints in both settings.For example, on MS MARCO Passage v1, our method reduces the average candidate set size from 1000 to 27, increasing reranking speed by about 37 times, while keeping MRR@10 greater than a pre-specified value of 0.38 with about 90% empirical coverage.In contrast, empirical baselines fail to meet such requirements.
Minghan Li 0002, Xinyu Zhang 0018, Ji Xin, Hongyang Zhang 0001, Jimmy Lin
EMNLP1
2021 Simple and Effective Unsupervised Redundancy Elimination to Compress Dense Vectors for Passage Retrieval
abstract
Recent work has shown that dense passage retrieval techniques achieve better ranking accuracy in open-domain question answering compared to sparse retrieval techniques such as BM25, but at the cost of large space and memory requirements.In this paper, we analyze the redundancy present in encoded dense vectors and show that the default dimension of 768 is unnecessarily large.To improve space efficiency, we propose a simple unsupervised compression pipeline that consists of principal component analysis (PCA), product quantization, and hybrid search.We further investigate other supervised baselines and find surprisingly that unsupervised PCA outperforms them in some settings.We perform extensive experiments on five question answering datasets and demonstrate that our best pipeline achieves good accuracy-space trade-offs, for example, 48× compression with less than 3% drop in top-100 retrieval accuracy on average or 96× compression with less than 4% drop.Code and data are available at http: //pyserini.io/.
Xueguang Ma, Minghan Li 0002, Ji Xin, Jimmy Lin
EMNLP (1)2