VLDB 2026 Research / reviewers in the wild / expert
Yeqin Zhang
dblp:160/9920
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 67% Language models and text generation · 24% Question answering and dialogue systems · 5% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
1.0 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Machine learning › Representation and self-supervised learning › text embedding
text representation learning |
1.0 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
unsupervised sentence embeddings |
1.0 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Natural language and speech › Language models and text generation
LLM agents |
0.8 | 1 | 2024 | Retrospex: Language Agent Meets Offline Reinforcement Learning Critic · EMNLP 2024 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.8 | 1 | 2024 | Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization · AAAI 2024 |
Information retrieval › retrieval models
retrieval model training |
0.8 | 1 | 2024 | Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization · AAAI 2024 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.3 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.2 | 1 | 2024 | Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization · AAAI 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game tree search |
0.2 | 1 | 2015 | TDS+: Improving Temperature Discovery Search · AAAI 2015 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 2.5noise contrastive estimation · 1.5hard negative sampling · 1.5masked next-token prediction · 1.0context compression · 1.0offline reinforcement learning · 0.8action rescoring · 0.8monte carlo tree search · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text RepresentationabstractText representation plays a critical role in tasks like clustering, retrieval, and other downstream applications. With the emergence of large language models (LLMs), there is increasing interest in harnessing their capabilities for this purpose. However, most of the LLMs are inherently causal and optimized for next-token prediction, making them suboptimal for producing holistic representations. To address this, recent studies introduced pretext tasks to adapt LLMs for text representation. Most of these tasks, however, rely on token-level prediction objectives, such as the masked next-token prediction (MNTP) used in LLM2Vec. In this work, we explore the untapped potential of context compression as a pretext task for unsupervised adaptation of LLMs. During compression pre-training, the model learns to generate compact memory tokens, which substitute the whole context for downstream sequence prediction. Experiments demonstrate that a well-designed compression objective can significantly enhance LLM-based text representations, outperforming models trained with token-level pretext tasks. Further improvements through contrastive learning produce a strong representation model (LLM2Comp) that outperforms contemporary LLM-based text encoders on a wide range of tasks while being more sample-efficient, requiring significantly less training data. Yeqin Zhang, Yizheng Zhao, Binxing Jiao, Daxin Jiang, Ruihang Miao, Cam-Tu Nguyen |
AAAI | 1 |
| 2024 | Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence RegularizationabstractIn open-domain Question Answering (QA), dense text retrieval is crucial for finding relevant passages to generate answers. Typically, contrastive learning is used to train a retrieval model, which maps passages and queries to the same semantic space, making similar ones closer and dissimilar ones further apart. However, training such a system is challenging due to the false negative problem, where relevant passages may be missed during data annotation. Hard negative sampling, commonly used to improve contrastive learning, can introduce more noise in training. This is because hard negatives are those close to a given query, and thus more likely to be false negatives. To address this, we propose a novel contrastive confidence regularizer for Noise Contrastive Estimation (NCE) loss, a commonly used contrastive loss. Our analysis shows that the regularizer helps make the dense retrieval model more robust against false negatives with a theoretical guarantee. Additionally, we propose a model-agnostic method to filter out noisy negative passages in the dataset, improving any downstream dense retrieval models. Through experiments on three datasets, we demonstrate that our method achieves better retrieval performance in comparison to existing state-of-the-art dense retrieval systems. Shiqi Wang 0003, Yeqin Zhang, Cam-Tu Nguyen |
AAAI | 2 |
| 2024 | Retrospex: Language Agent Meets Offline Reinforcement Learning CriticabstractLarge Language Models (LLMs) possess extensive knowledge and commonsense reasoning capabilities, making them valuable for creating powerful agents.However, existing LLM agent frameworks have not fully utilized past experiences for improvement.This work introduces a new LLM-based agent framework called Retrospex , which addresses this challenge by analyzing past experiences in depth.Unlike previous approaches, Retrospex does not directly integrate experiences into the LLM's context.Instead, it combines the LLM's action likelihood with action values estimated by a Reinforcement Learning (RL) Critic, which is trained on past experiences through an offline "retrospection" process.Additionally, Retrospex employs a dynamic action rescoring mechanism that increases the importance of experience-based values for tasks that require more interaction with the environment.We evaluate Retrospex in ScienceWorld, ALFWorld and Webshop environments, demonstrating its advantages over strong, contemporary baselines 1 . Yufei Xiang, Yiqun Shen, Yeqin Zhang, Cam-Tu Nguyen |
EMNLP | 3 |
| 2023 | Coarse-To-Fine Knowledge Selection for Document Grounded DialogsabstractMulti-document grounded dialogue systems (DGDS) belong to a class of conversational agents that answer users’ requests by finding supporting knowledge from a collection of documents. Most previous studies aim to improve the knowledge retrieval model or propose more effective ways to incorporate external knowledge into a parametric generation model. These methods, however, focus on retrieving knowledge from mono-granularity language units (e.g. passages, sentences, or spans in documents), which is not enough to effectively and efficiently capture precise knowledge in long documents. This paper proposes Re3G, which aims to optimize both coarse-grained knowledge retrieval and fine-grained knowledge extraction in a unified framework. Specifically, the former efficiently finds relevant passages in a retrieval-and-reranking process, whereas the latter effectively extracts finer-grain spans within those passages to incorporate into a parametric answer generation model (BART, T5). Experiments on DialDoc Shared Task demonstrate the effectiveness of our method. Yeqin Zhang, Haomin Fu, Cheng Fu 0003, Haiyang Yu 0003, Yongbin Li 0001, Cam-Tu Nguyen |
ICASSP | 1 |
| 2023 | Long Short-Term Planning for Conversational Recommendation Systems
Hongguang Shi, Yeqin Zhang, Xubin Li, Cam-Tu Nguyen |
ICONIP (6) | 4 |
| 2015 | TDS+: Improving Temperature Discovery Search
Yeqin Zhang, Martin Müller 0003 |
AAAI | 1 |