VLDB 2026 Research / reviewers in the wild / expert
Yuwei Zhang 0001
dblp:95/8351-1
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-6910-8130ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge InjectionabstractYuwei Zhang, Wenhao Yu, Shangbin Feng, Yifan Zhu, Letian Peng, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuwei Zhang 0001, Wenhao Yu 0002, Shangbin Feng, Letian Peng, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang |
ACL (1) | 1 |
| 2025 | Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit ReasoningabstractThere has been a surge in the use of large language models (LLM) conversational agents to generate responses based on long-term history from multiple sessions.However, existing long-term open-domain dialogue datasets lack complex, real-world personalization and fail to capture implicit reasoning-where relevant information is embedded in subtle, syntactic, or semantically distant connections rather than explicit statements.In such cases, traditional retrieval methods fail to capture relevant context, and long-context modeling also becomes inefficient due to numerous complicated personarelated details.To address this gap, we introduce IMPLEXCONV, a large-scale long-term dataset with 2,500 examples, each containing approximately 100 conversation sessions, designed to study implicit reasoning in personalized dialogues.Additionally, we propose TAC-ITREE, a novel hierarchical tree framework that structures conversation history into multiple levels of summarization.Instead of brute-force searching all data, TACITREE enables an efficient, level-based retrieval process where models refine their search by progressively selecting relevant details.Our experiments demonstrate that TACITREE significantly improves the ability of LLMs to reason over long-term conversations with implicit contextual dependencies.I love playing sports like basketball and swimming.I'm going on a sports stadium tour!Aug 15, 2024 I broke my leg in a car accident... Xintong Li 0001, Jalend Bantupalli, Ria Dharmani, Yuwei Zhang 0001, Jingbo Shang |
EMNLP | 4 |
| 2025 | Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingabstractRetrieval augmented generation (RAG) combines the generative abilities of large language models (LLMs) with external knowledge sources to provide more accurate and up-to-date responses. Recent RAG advancements focus on improving retrieval outcomes through iterative LLM refinement or self-critique capabilities acquired through additional instruction tuning of LLMs. In this work, we introduce Speculative RAG - a framework that leverages a larger generalist LM to efficiently verify multiple RAG drafts produced in parallel by a smaller, distilled specialist LM. Each draft is generated from a distinct subset of retrieved documents, offering diverse perspectives on the evidence while reducing input token counts per draft. This approach enhances comprehension of each subset and mitigates potential position bias over long context. Our method accelerates RAG by delegating drafting to the smaller specialist LM, with the larger generalist LM performing a single verification pass over the drafts. Extensive experiments demonstrate that Speculative RAG achieves state-of-the-art performance with reduced latency on TriviaQA, MuSiQue, PopQA, PubHealth, and ARC-Challenge benchmarks. It notably enhances accuracy by up to 12.97% while reducing latency by 50.83% compared to conventional RAG systems on PubHealth. Zilong Wang 0002, Zifeng Wang 0002, Long T. Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang 0001, Anush Mattapalli, Ankur Taly, Jingbo Shang, Chen-Yu Lee, Tomas Pfister |
ICLR | 7 |
| 2025 | LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryabstractRecent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory capabilities in sustained interactions remain underexplored. We introduce LongMemEval, a comprehensive benchmark designed to evaluate five core long-term memory abilities of chat assistants: information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. With 500 meticulously curated questions embedded within freely scalable user-assistant chat histories, LongMemEval presents a significant challenge to existing long-term memory systems, with commercial chat assistants and long-context LLMs showing a 30% accuracy drop on memorizing information across sustained interactions. We then present a unified framework that breaks down the long-term memory design into three stages: indexing, retrieval, and reading. Built upon key experimental insights, we propose several memory design optimizations including session decomposition for value granularity, fact-augmented key expansion for indexing, and time-aware query expansion for refining the search scope. Extensive experiments show that these optimizations greatly improve both memory recall and downstream question answering on LongMemEval. Overall, our study provides valuable resources and guidance for advancing the long-term memory capabilities of LLM-based chat assistants, paving the way toward more personalized and reliable conversational AI. Our benchmark and code are publicly available at https://github.com/xiaowu0162/LongMemEval. Di Wu 0054, Wenhao Yu 0002, Yuwei Zhang 0001, Kai-Wei Chang 0001, Dong Yu 0001 |
ICLR | 4 |
| 2025 | RADAR: Benchmarking Language Models on Imperfect Tabular DataabstractLanguage models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness—the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies—remains underexplored. These artifacts are especially common in real-world tabular data and, if mishandled, can significantly compromise the validity of analytical conclusions. To address this gap, we present RADAR, a benchmark for systematically evaluating data-aware reasoning on tabular data. We develop a framework to simulate data artifacts via programmatic perturbations to enable targeted evaluation of model behavior. RADAR comprises 2,980 table-query pairs, grounded in real-world data spanning 9 domains and 5 data artifact types. In addition to evaluating artifact handling, RADAR systematically varies table size to study how reasoning performance holds when increasing table size. Our evaluation reveals that, despite decent performance on tables without data artifacts, frontier models degrade significantly when data artifacts are introduced, exposing critical gaps in their capacity for robust, data-aware analysis. Designed to be flexible and extensible, RADAR supports diverse perturbation types and controllable table sizes, offering a valuable resource for advancing tabular reasoning. Ken Gu, Zhihan Zhang 0002, Kate Lin, Yuwei Zhang 0001, Akshay Paruchuri, Hong Yu 0001, Mehran Kazemi, Kumar Ayush, A. Ali Heydari, Maxwell A. Xu, Yun Liu 0013, Ming-Zher Poh, Yuzhe Yang 0003, Mark Malhotra, Shwetak N. Patel, Hamid Palangi, Xuhai Xu, Daniel McDuff, Tim Althoff, Xin Liu 0034 |
NeurIPS | 4 |
| 2025 | SensorLM: Learning the Language of Wearable SensorsabstractWe present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks. Code is available at https://github.com/Google-Health/consumer-health-research/tree/main/sensorlm. Yuwei Zhang 0001, Kumar Ayush, Siyuan Qiao, A. Ali Heydari, Girish Narayanswamy, Maxwell A. Xu, Ahmed Metwally 0002, Jinhua Xu, Jake Garrison, Xuhai Xu, Tim Althoff, Yun Liu 0013, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Cecilia Mascolo, Xin Liu 0034, Daniel McDuff, Yuzhe Yang 0003 |
NeurIPS | 1 |
| 2024 | Answer is All You Need: Instruction-following Text Embedding via Answering the QuestionabstractLetian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa, Gaowen Liu, Zihan Wang, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Letian Peng, Yuwei Zhang 0001, Zilong Wang 0002, Jayanth Srinivasa, Gaowen Liu, Zihan Wang 0001, Jingbo Shang |
ACL (1) | 2 |
| 2024 | Towards Open Respiratory Acoustic Foundation Models: Pretraining and BenchmarkingabstractRespiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Generalizable respiratory acoustic foundation models pretrained with unlabeled data would offer appealing advantages and possibly unlock this impasse. However, given the safety-critical nature of healthcare applications, it is pivotal to also ensure openness and replicability for any proposed foundation model solution. To this end, we introduce OPERA, an OPEn Respiratory Acoustic foundation model pretraining and benchmarking system, as the first approach answering this need. We curate large-scale respiratory audio datasets ($\sim$136K samples, over 400 hours), pretrain three pioneering foundation models, and build a benchmark consisting of 19 downstream respiratory health tasks for evaluation. Our pretrained models demonstrate superior performance (against existing acoustic models pretrained with general audio on 16 out of 19 tasks) and generalizability (to unseen datasets and new respiratory audio modalities). This highlights the great promise of respiratory acoustic foundation models and encourages more studies using OPERA as an open resource to accelerate research on respiratory audio for health. The system is accessible from https://github.com/evelyn0414/OPERA. Yuwei Zhang 0001, Tong Xia, Jing Han 0010, Yu Wu 0021, Georgios Rizos, Yang Liu 0101, Mohammed Mosuily, Jagmohan Chauhan, Cecilia Mascolo |
NeurIPS | 1 |
| 2023 | ClusterLLM: Large Language Models as a Guide for Text ClusteringabstractWe introduce CLUSTERLLM, a novel text clustering framework that leverages feedback from an instruction-tuned large language model, such as ChatGPT.Compared with traditional unsupervised methods that builds upon "small" embedders, CLUSTERLLM exhibits two intriguing advantages: (1) it enjoys the emergent capability of LLM even if its embeddings are inaccessible; and (2) it understands the user's preference on clustering through textual instruction and/or a few annotated data.First, we prompt ChatGPT for insights on clustering perspective by constructing hard triplet questions , where A, B and C are similar data points that belong to different clusters according to small embedder.We empirically show that this strategy is both effective for fine-tuning small embedder and cost-efficient to query ChatGPT.Second, we prompt ChatGPT for helps on clustering granularity by carefully designed pairwise questions , and tune the granularity from cluster hierarchies that is the most consistent with the ChatGPT answers.Extensive experiments on 14 datasets show that CLUSTERLLM consistently improves clustering quality, at an average cost of ∼$0.6 1 per dataset.The code will be available at https: //github.com/zhang-yu-wei/ClusterLLM. Yuwei Zhang 0001, Zihan Wang 0001, Jingbo Shang |
EMNLP | 1 |
| 2023 | Toward Unsupervised Realistic Visual Question AnsweringabstractThe problem of realistic VQA (RVQA), where a model has to reject unanswerable questions (UQs) and answer answerable ones (AQs), is studied. We first point out 2 drawbacks in current RVQA research, where (1) datasets contain too many unchallenging UQs and (2) a large number of annotated UQs are required for training. To resolve the first drawback, we propose a new testing dataset, RGQA, which combines AQs from an existing VQA dataset with around 29K human-annotated UQs. These UQs consist of both fine-grained and coarse-grained image-question pairs generated with 2 approaches: CLIP-based and Perturbation-based. To address the second drawback, we introduce an unsupervised training approach. This combines pseudo UQs obtained by randomly pairing images and questions, with an RoI Mixup procedure to generate more fine-grained pseudo UQs, and model ensembling to regularize model confidence. Experiments show that using pseudo UQs significantly outperforms RVQA baselines. RoI Mixup and model ensembling further increase the gain. Finally, human evaluation reveals a performance gap between humans and models, showing that more RVQA research is needed. Code and dataset is released on https://github.com/chihhuiho/RGQA. Yuwei Zhang 0001, Chih-Hui Ho, Nuno Vasconcelos |
ICCV | 1 |
| 2022 | New Intent Discovery with Pre-training and Contrastive LearningabstractNew intent discovery aims to uncover novel intent categories from user utterances to expand the set of supported intent classes.It is a critical task for the development and service expansion of a practical dialogue system.Despite its importance, this problem remains under-explored in the literature.Existing approaches typically rely on a large amount of labeled utterances and employ pseudo-labeling methods for representation learning and clustering, which are label-intensive, inefficient, and inaccurate.In this paper, we provide new solutions to two important research questions for new intent discovery: (1) how to learn semantic utterance representations and (2) how to better cluster utterances.Particularly, we first propose a multi-task pre-training strategy to leverage rich unlabeled data along with external labeled data for representation learning.Then, we design a new contrastive loss to exploit self-supervisory signals in unlabeled data for clustering.Extensive experiments on three intent recognition benchmarks demonstrate the high effectiveness of our proposed method, which outperforms state-of-the-art methods by a large margin in both unsupervised and semi-supervised scenarios.The source code will be available at https://github. com/ Yuwei Zhang 0001, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu 0003, Albert Y. S. Lam |
ACL (1) | 1 |
| 2022 | Fine-tuning Pre-trained Language Models for Few-shot Intent Detection: Supervised Pre-training and IsotropizationabstractHaode Zhang, Haowen Liang, Yuwei Zhang, Li-Ming Zhan, Xiao-Ming Wu, Xiaolei Lu, Albert Lam. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Haode Zhang, Haowen Liang, Yuwei Zhang 0001, Li-Ming Zhan, Xiao-Ming Wu 0003, Xiaolei Lu, Albert Y. S. Lam |
NAACL-HLT | 3 |