VLDB 2026 Research / reviewers in the wild / expert
Yibin Lei
dblp:321/0804
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-9558-5548ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 33% Language models and text generation · 27% Machine translation · 19% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 100% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
1.9 | 2 | 2026 | Making Large Language Models Efficient Dense Retrievers · ACL (1) 2026 Neural Lexical Search with Learned Sparse Retrieval · SIGIR 2025 |
Natural language and speech › Language models and text generation
large language model |
1.1 | 2 | 2026 | Meta-Task Prompting Elicits Embeddings from Large Language Models · ACL (1) 2024 Making Large Language Models Efficient Dense Retrievers · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | Making Large Language Models Efficient Dense Retrievers · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression
pruning |
1.0 | 1 | 2026 | Making Large Language Models Efficient Dense Retrievers · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
1.0 | 1 | 2026 | Making Large Language Models Efficient Dense Retrievers · ACL (1) 2026 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.0 | 1 | 2026 | Making Large Language Models Efficient Dense Retrievers · ACL (1) 2026 |
Natural language and speech › Language models and text generation
decoding |
0.9 | 1 | 2025 | Calibrating Translation Decoding with Quality Estimation on LLMs · NeurIPS 2025 |
Natural language and speech › Machine translation › neural machine translation
large language model translation |
0.9 | 1 | 2025 | Calibrating Translation Decoding with Quality Estimation on LLMs · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
text embedding |
0.9 | 1 | 2025 | Enhancing Lexicon-Based Text Embeddings with Large Language Models · ACL (1) 2025 |
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation |
0.9 | 1 | 2025 | Calibrating Translation Decoding with Quality Estimation on LLMs · NeurIPS 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models · EMNLP 2025 |
Information retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models · EMNLP 2025 |
Information retrieval › retrieval models › sparse retrieval
learned sparse retrieval |
0.9 | 1 | 2025 | Neural Lexical Search with Learned Sparse Retrieval · SIGIR 2025 |
Information retrieval › image retrieval
visual document retrieval |
0.9 | 1 | 2025 | SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models · EMNLP 2025 |
Information retrieval › document processing › document analysis › document representation
text embedding |
0.8 | 1 | 2024 | Meta-Task Prompting Elicits Embeddings from Large Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model
large language model representation |
0.3 | 1 | 2025 | Enhancing Lexicon-Based Text Embeddings with Large Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.3 | 1 | 2025 | Calibrating Translation Decoding with Quality Estimation on LLMs · NeurIPS 2025 |
Information retrieval › retrieval models
neural retrieval |
0.3 | 1 | 2025 | Neural Lexical Search with Learned Sparse Retrieval · SIGIR 2025 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.2 | 1 | 2024 | Meta-Task Prompting Elicits Embeddings from Large Language Models · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
retrieval-specific fine-tuning · 2.0layer redundancy analysis · 2.0MLP compression · 2.0zero-shot retrieval · 1.7vision-language model captioning · 1.7text embedding · 1.7token embedding clustering · 0.9pooling strategies · 0.9pearson correlation optimization · 0.9neural network · 0.9bidirectional attention · 0.9prompt engineering · 0.8meta-task prompting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Making Large Language Models Efficient Dense RetrieversabstractRecent work has shown that directly fine-tuning large language models (LLMs) for dense retrieval yields strong performance, but their substantial parameter counts make them computationally inefficient.While prior studies have revealed significant layer redundancy in LLMs for generative tasks, it remains unclear whether similar redundancy exists when these models are adapted for retrieval tasks, which require encoding entire sequences into fixed representations rather than generating tokens iteratively.To this end, we conduct a comprehensive analysis of layer redundancy in LLM-based dense retrievers.We find that, in contrast to generative settings, MLP layers are substantially more prunable, while attention layers remain critical for semantic aggregation.Building on this insight, we propose EffiR, a framework for developing efficient retrievers that performs large-scale MLP compression through a coarseto-fine strategy (coarse-grained depth reduction followed by fine-grained width reduction), combined with retrieval-specific fine-tuning.Across diverse BEIR datasets and LLM backbones, EffiR achieves substantial reductions in model size and inference cost while preserving the performance of full-size models. 1 * Equal contribution 1 Our code and models are available at https://github. com/ Yibin Lei, Shwai He, Andrew Yates |
ACL (1) | 1 |
| 2026 | Neural Lexical Search with Learned Sparse Retrieval
Andrew Yates, Carlos Eduardo Rosar Kós Lassance, Cosimo Rulli, Eugene Yang 0001, Sean MacAvaney, Siddharth A. K. Singh, Thong Nguyen 0004, Yibin Lei |
ECIR (4) | 8 |
| 2026 | Improving zero-shot translation with the navigation ability-enhanced language tags
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibin Lei, Yibing Zhan, Weifeng Liu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Enhancing Lexicon-Based Text Embeddings with Large Language ModelsabstractRecent large language models (LLMs) have demonstrated exceptional performance on general-purpose text embedding tasks.While dense embeddings have dominated related research, we introduce the first lexicon-based embeddings (LENS) leveraging LLMs that achieve competitive performance on these tasks.LENS consolidates the vocabulary space through token embedding clustering to handle the issue of token redundancy in LLM vocabularies.To further improve performance, we investigate bidirectional attention and various pooling strategies.Specifically, LENS simplifies lexical matching with redundant vocabularies by assigning each dimension to a specific token cluster, where semantically similar tokens are grouped together.Extensive experiments demonstrate that LENS outperforms dense embeddings on the Massive Text Embedding Benchmark (MTEB), delivering compact representations with dimensionality comparable to dense counterparts.Furthermore, LENS inherently supports efficient embedding dimension pruning without any specialized objectives like Matryoshka Representation Learning.Notably, combining LENS with dense embeddings achieves state-of-the-art performance on the retrieval subset of MTEB (i.e., BEIR). 1 Yibin Lei, Tao Shen 0001, Yu Cao 0014, Andrew Yates |
ACL (1) | 1 |
| 2025 | SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language ModelsabstractVisual Document Retrieval (VDR) typically operates as text-to-image retrieval using specialized bi-encoders trained to directly embed document images.We revisit a zero-shot generateand-encode pipeline: a vision-language model first produces a detailed textual description of each document image, which is then embedded by a standard text encoder.On the ViDoRe-v2 benchmark, the method reaches 63.4% nDCG@5, surpassing the strongest specialised multi-vector visual document encoder.It also scales better to large collections and offers broader multilingual coverage.Analysis shows that modern vision-language models capture complex textual and visual cues with sufficient granularity to act as a reusable semantic proxy.By offloading modality alignment to pretrained vision-language models, our approach removes the need for computationally intensive text-image contrastive training and establishes a strong zero-shot baseline for future VDR systems.Our code is available for reproduction at: thongnt99/serval Thong Nguyen 0004, Yibin Lei, Jia-Huei Ju, Andrew Yates |
EMNLP | 2 |
| 2025 | Calibrating Translation Decoding with Quality Estimation on LLMsabstractNeural machine translation (NMT) systems typically employ maximum *a posteriori* (MAP) decoding to select the highest-scoring translation from the distribution. However, recent evidence highlights the inadequacy of MAP decoding, often resulting in low-quality or even pathological hypotheses as the decoding objective is only weakly aligned with real-world translation quality. This paper proposes to directly calibrate hypothesis likelihood with translation quality from a distributional view by directly optimizing their Pearson correlation, thereby enhancing decoding effectiveness. With our method, translation with large language models (LLMs) improves substantially after limited training (2K instances per direction). This improvement is orthogonal to those achieved through supervised fine-tuning, leading to substantial gains across a broad range of metrics and human evaluations. This holds even when applied to top-performing translation-specialized LLMs fine-tuned on high-quality translation data, such as Tower, or when compared to recent preference optimization methods, like CPO. Moreover, the calibrated translation likelihood can directly serve as a strong proxy for translation quality, closely approximating or even surpassing some state-of-the-art translation quality estimation models, like CometKiwi.
Lastly, our in-depth analysis demonstrates that calibration enhances the effectiveness of MAP decoding, thereby enabling greater efficiency in real-world deployment. The resulting state-of-the-art translation model, which covers 10 languages, along with the accompanying code and human evaluation data, has been released: https://github.com/moore3930/calibrating-llm-mt. Yibin Lei, Christof Monz |
NeurIPS | 2 |
| 2025 | Neural Lexical Search with Learned Sparse RetrievalabstractLearned Sparse Retrieval (LSR) techniques use neural machinery to represent queries and documents as learned bags of words. In contrast with other neural retrieval techniques, such as generative retrieval and dense retrieval, LSR has been shown to be a remarkably robust, transferable, and efficient family of methods for retrieving high-quality search results. This half-day tutorial aims to provide an extensive overview of LSR, ranging from its fundamentals to the latest emerging techniques. By the end of the tutorial, attendees will be familiar with the important design decisions of an LSR system, know how to apply them to text and other modalities, and understand the latest techniques for retrieving with them efficiently. Website: https://lsr-tutorial.github.io Andrew Yates, Carlos Eduardo Rosar Kós Lassance, Cosimo Rulli, Eugene Yang 0001, Sean MacAvaney, Siddharth A. K. Singh, Thong Nguyen 0004, Yibin Lei |
SIGIR | 8 |
| 2024 | Meta-Task Prompting Elicits Embeddings from Large Language ModelsabstractWe introduce a new unsupervised text embedding method, Meta-Task Prompting with Explicit One-Word Limitation (MetaEOL), for generating high-quality sentence embeddings from Large Language Models (LLMs) without the need for model fine-tuning.Leveraging meta-task prompting, MetaEOL guides LLMs to produce embeddings through a series of carefully designed prompts that address multiple representational aspects.Our comprehensive experiments demonstrate that embeddings averaged from various meta-tasks are versatile embeddings that yield competitive performance on Semantic Textual Similarity (STS) benchmarks and excel in downstream tasks, surpassing contrastive-trained models.Our findings suggest a new scaling law, offering a versatile and resource-efficient approach for embedding generation across diverse scenarios.1 Yibin Lei, Tianyi Zhou 0001, Tao Shen 0001, Yu Cao 0014, Chongyang Tao, Andrew Yates |
ACL (1) | 1 |