VLDB 2026 Research / reviewers in the wild / expert
Wenhan Liu
dblp:192/1039
· DBLP profile ↗
18ranked-venue papers
11as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReasonRank: Empowering Passage Ranking with Strong Reasoning AbilityabstractLarge Language Model (LLM) based listwise ranking has shown superior performance in many passage ranking tasks.With the development of Large Reasoning Models (LRMs), many studies have demonstrated that step-bystep reasoning during test-time helps improve listwise ranking performance.However, due to the scarcity of reasoning-intensive training data, existing rerankers perform poorly in many complex ranking scenarios, and the ranking ability of reasoning-intensive rerankers remains largely underdeveloped.In this paper, we first propose an automated reasoning-intensive training data synthesis framework, which sources training queries and passages from diverse domains and applies DeepSeek-R1 to generate high-quality training labels.To empower the listwise reranker with strong reasoning ability, we further propose a two-stage training approach, which includes a cold-start supervised fine-tuning (SFT) stage and a reinforcement learning (RL) stage.During the RL stage, we design a novel multi-view ranking reward tailored to the multi-turn nature of listwise ranking.Extensive experiments demonstrate that our trained reasoning-intensive reranker Rea-sonRank outperforms existing baselines significantly and also achieves much lower latency than the pointwise reranker. Wenhan Liu, Xinyu Ma 0001, Weiwei Sun 0001, Yutao Zhu 0001, Yuchen Li 0006, Dawei Yin 0001, Zhicheng Dou |
ACL (1) | 1 |
| 2026 | TriFusNet: A Triple-Fusion Convolutional Network for Multivariate Time Series ClassificationabstractMultivariate time series classification (MTSC) plays a critical role in a wide range of real-world applications, such as healthcare, finance, and industrial monitoring. This paper proposes a triple-fusion network (TriFusNet), a novel convolutional network designed to address the challenges of MTSC. TriFusNet employs a specialized architecture that captures both variable-specific features and features shared across variables through the parallel use of standard, depth-wise, and shared-kernel convolutions. A hierarchical triple-fusion strategy is introduced to enhance representation learning across three stages: input-level fusion transforms raw variables, intermediate-level fusion integrates heterogeneous features, and output-level fusion improves decision robustness. Extensive experiments on 26 benchmark datasets show that TriFusNet outperforms 15 competitive baselines, achieving the best average rank (4.3077) with a Win/Draw/Loss of 5/4/17. The effectiveness of its architectural design and parameter settings is empirically validated, and a qualitative theoretical discussion is conducted to support the proposed fusion strategy. These results highlight TriFusNet's strong potential for real-world applications involving complex and high-dimensional time series data. Wenhan Liu, Shurong Pan, Sheng Chang 0003, Qijun Huang, Nan Jiang 0013 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | DemoRank: Selecting Effective Demonstrations for Large Language Models in Ranking TaskabstractLarge Language Models (LLMs) have been proven to have strong zero-shot passage ranking capabilities. In-context learning effectively enhances LLM performance by providing few-shot demonstrations, opening avenues for further improving LLM’s ranking ability. However, existing studies usually retrieve the most similar demonstrations to the input, ignoring the demonstration dependencies and diversity, which is insufficient to inspire the LLM for assessing the current query-passage relevance. In this article, we propose a framework named DemoRank, which selects few-shot demonstrations by performing a novel dependency-aware reranking of the retrieved demonstrations. Considering the dependency, combining top-ranked demonstrations yields better results. Nevertheless, generating the training samples for such a dependency-aware demonstration reranker faces two challenges: (1) the traditional demonstration ranked list assumes demonstration independence, which cannot be used to train our reranker, and (2) obtaining the optimal demonstration ranked list from the retrieved set is NP-hard and inefficient. To overcome these challenges, we propose an approach to construct a kind of dependency-aware training samples efficiently and design a list-pairwise training approach for the optimization of the demonstration reranker. We conduct extensive experiments on a series of passage ranking datasets, and the results demonstrate the superior performance of our proposed DemoRank framework under various scenarios. Our code is publicly available at https://github.com/8421bcd/demorank . Wenhan Liu, Yutao Zhu 0001, Zhicheng Dou, Yujia Zhou 0002 |
ACM Trans. Inf. Syst. | 1 |
| 2026 | Large Language Models for Information Retrieval: A SurveyabstractAs a primary means of information acquisition, information retrieval (IR) systems, such as search engines, have integrated themselves into our daily lives. These systems also serve as components of dialogue, question-answering, and recommender systems. The trajectory of IR has evolved dynamically from its origins in term-based methods to its integration with advanced neural models. While the neural models excel at capturing complex contextual signals and semantic nuances, they still face challenges such as data scarcity, interpretability, and the generation of contextually plausible yet potentially inaccurate responses. This evolution requires a combination of traditional methods (such as term-based sparse retrieval methods with rapid response) and modern neural architectures (such as language models with powerful language understanding capacity). Meanwhile, the emergence of large language models (LLMs) has revolutionized natural language processing due to their remarkable language understanding, generation, and reasoning abilities. Consequently, recent research has sought to leverage LLMs to improve IR systems. Given the rapid evolution of this research trajectory, it is necessary to consolidate existing methodologies and provide nuanced insights through a comprehensive overview. In this survey, we delve into the confluence of LLMs and IR systems, including crucial aspects such as query rewriters, retrievers, rerankers, readers, and search agents. Yutao Zhu 0001, Huaying Yuan, Shuting Wang 0002, Jiongnan Liu 0001, Wenhan Liu, Chenlong Deng, Haonan Chen 0005, Zheng Liu 0011, Zhicheng Dou, Ji-Rong Wen |
ACM Trans. Inf. Syst. | 5 |
| 2025 | Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language ModelsabstractWenhan Liu, Xinyu Ma, Yutao Zhu, Ziliang Zhao, Shuaiqiang Wang, Dawei Yin, Zhicheng Dou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wenhan Liu, Xinyu Ma 0001, Yutao Zhu 0001, Ziliang Zhao 0001, Shuaiqiang Wang, Dawei Yin 0001, Zhicheng Dou |
ACL (1) | 1 |
| 2025 | Direct Lead Assignment: A Simple and Scalable Contrastive Learning Method for ECG and Its IoMT ApplicationsabstractNowadays, applying deep learning (DL) to electrocardiogram (ECG) analysis has become a significant topic in intelligent healthcare. DL models heavily rely on large-scale labeled ECGs in supervised learning, while labeling ECGs is a costly and time-consuming process. This article proposes a simple self-supervised learning (SSL) method to pretrain models using unlabeled ECGs, improving model performances in a low-data regime. It is termed direct lead assignment (DLA). In pretraining, DLA employs multilead and single-lead encoders to interact between global and lead-specific representations. The pretrained encoders can constitute scalable models, which can be deployed on Internet of Medical Things (IoMT) devices for ECG monitoring with different leads. According to the experiments, DLA outperforms existing SSL methods for ECGs and reduces the reliance on labels by$4\times $–$8\times $. In other words, DLA can make the model obtain better results when labeled data are scarce than the one trained from scratch, as the pretraining helps the model recognize critical ECG patterns in a low-data regime. For IoMT applications, the models are deployed on a Raspberry Pi 2 W with an ARM Cortex-A53 processor. It can run the models in real time. The maximum running time is about 388 ms/10-s record. Thus, DLA shows excellent potential for automatic ECG analysis based on IoMT, aiding cardiologists in diagnosing cardiovascular diseases. Wenhan Liu, Shurong Pan, Sheng Chang 0003, Qijun Huang, Nan Jiang 0013 |
IEEE Internet Things J. | 1 |
| 2024 | Mining Exploratory Queries for Conversational SearchabstractUsers' queries are usually vague, and their search intents tend to be ambiguous, thereby needing search clarification to clarify users' current intent by asking a clarifying question and providing several clickable sub-intent items as clarification options. However, in addition to drilling down the current query, users may also have exploratory needs that diverge from their current intent. For example, a user searching for the query "Cartier women watches'' may also potentially want to explore some parallel information by issuing queries such as "Rolex women watches'' or "Cartier women bracelets'', named exploratory queries in this paper. These exploratory needs are common during the search process yet cannot be satisfied by current search clarification approaches which typically stick to the sub-intents of the query. This paper focuses on mining exploratory queries as additional options to meet users' exploratory needs in conversational search systems. Specifically, we first design a rule-based model that generates exploratory queries based on the current query's top retrieved documents. Then, we propose using the data generated by the rule-based model to train a neural generation model through multi-task learning for further generalization. Finally, we borrow the in-context learning ability of the large language model to generate exploratory queries based on prompt engineering. We constructed an evaluation dataset based on human annotations and conduct an extensive set of experiments. The results show that our proposed methods generate higher-quality exploratory queries compared with several baselines. Wenhan Liu, Ziliang Zhao 0001, Yutao Zhu 0001, Zhicheng Dou |
WWW | 1 |
| 2024 | DDDG: A dual bi-directional knowledge distillation method with generative self-supervised pre-training and its hardware implementation on SoC for ECG
Huaicheng Zhang, Wenhan Liu, Qianxi Guo, Jiguang Shi, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
Expert Syst. Appl. | 2 |
| 2024 | SIGxCL: A Signal-Image-Graph Cross-Modal Contrastive Learning Framework for CVD Diagnosis Based on Internet of Medical ThingsabstractRecently, contrastive learning (CL) has garnered wide interest because it enables unsupervised pretraining to alleviate conventional deep learning methods’ strong reliance on artificial labels. While CL-based methods have been applied to cardiovascular disease (CVD) diagnosis with noninvasive electrocardiogram (ECG), most of these methods are limited within the 1-D signal modality and primarily focus on temporal features like amplitude and time sequence. The morphological features derived from clinically significant image-like ECGs are ignored. Furthermore, the relationships among different leads are neglected as well, describing the activities and interaction of various heart regions that are essential in CVD diagnosis and lesion localization. To address these limitations, this work proposes a novel cross-modal CL framework named signal–image–graph cross-modal contrastive learning (SIGxCL), which represents and jointly analyzes ECGs in signal, image, and graph modalities. Crucial for CL, modality-specific transformations are introduced for ECGs in the three modalities. SIGxCL enables signal–image–graph correspondence by maximizing the agreement of the accordant cross-modal ECGs in the invariant space. Consequently, SIGxCL could capture and leverage temporal, morphological, and spatially physiological features simultaneously. Compared to random initial and conventional supervised methods, SIGxCL achieves remarkable enhancements. Considering the best performances of the existing CL-based methods, SIGxCL outperforms them by up to 4.72%, 9.41%, and 4.31% across three data sets. SIGxCL is designed to be compatible with Internet of Medical Things (IoMT) and can be deployed on resource-limited portable devices. The deployment includes pretraining, online/offline tuning, and real-time inference modules. In conclusion, SIGxCL demonstrates superior performance and provides a promising approach for real-time IoMT-based diagnosis. Huaicheng Zhang, Wenhan Liu, Zhoutong Li, Jiguang Shi, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
IEEE Internet Things J. | 2 |
| 2024 | An adaptive threshold-based semi-supervised learning method for cardiovascular disease detection
Jiguang Shi, Zhoutong Li, Wenhan Liu, Huaicheng Zhang, Deyu Luo, Yue Ge, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
Inf. Sci. | 3 |
| 2024 | Honest-GE: 2-step heuristic optimization and node-level embedding empower spatial-temporal graph model for ECG
Huaicheng Zhang, Wenhan Liu, Deyu Luo, Jiguang Shi, Qianxi Guo, Yue Ge, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
Inf. Sci. | 2 |
| 2024 | How to personalize and whether to personalize? Candidate documents decide
Wenhan Liu, Yujia Zhou 0002, Yutao Zhu 0001, Zhicheng Dou |
Knowl. Inf. Syst. | 1 |
| 2024 | Defending Batch-Level Label Inference and Replacement Attacks in Vertical Federated LearningabstractIn a vertical federated learning (VFL) scenario where features and models are split into different parties, it has been shown that sample-level gradient information can be exploited to deduce crucial label information that should be kept secret. An immediate defense strategy is to protect sample-level messages communicated with Homomorphic Encryption (HE), exposing only batch-averaged local gradients to each party. In this paper, we show that even with HE-protected communication, private labels can still be reconstructed with high accuracy by gradient inversion attack, contrary to the common belief that batch-averaged information is safe to share under encryption. We then show that backdoor attack can also be conducted by directly replacing encrypted communicated messages without decryption. To tackle these attacks, we propose a novel defense method, Confusional AutoEncoder (termedCAE), which is based on autoencoder and entropy regularization to disguise true labels. To further defend attackers with sufficient prior label knowledge, we introduce DiscreteSGD-enhanced CAE (termedDCAE), and show that DCAE significantly boosts the main task accuracy than other known methods when defending various label inference attacks. Tianyuan Zou, Yang Liu 0165, Yan Kang 0001, Wenhan Liu, Yuanqin He, Zhihao Yi, Qiang Yang 0001, Ya-Qin Zhang |
IEEE Trans. Big Data | 4 |
| 2024 | ST-ReGE: A Novel Spatial-Temporal Residual Graph Convolutional Network for CVDabstractRecently, deep learning (DL) has enabled rapid advancements in electrocardiogram (ECG)-based automatic cardiovascular disease (CVD) diagnosis. Multi-lead ECG signals have lead systems based on the potential differences between electrodes placed on the limbs and the chest. When applying DL models, ECG signals are usually treated as synchronized signals arranged in Euclidean space, which is the abstraction and generalization of real space. However, conventional DL models typically merely focus on temporal features when analyzing Euclidean data. These approaches ignore the spatial relationships of different leads, which are physiologically significant and useful for CVD diagnosis because different leads represent activities of specific heart regions. These relationships derived from spatial distributions of electrodes can be conveniently created in non-Euclidean data, making multi-lead ECGs better conform to their nature. Considering graph convolutional network (GCN) adept at analyzing non-Euclidean data, a novel spatial-temporal residual GCN for CVD diagnosis is proposed in this work. ECG signals are firstly divided into single-channel patches and transferred into nodes, which will be connected by spatial-temporal connections. The proposed model employs residual GCN blocks and feed-forward networks to alleviate over-smoothing and over-fitting. Moreover, residual connections and patch dividing enable the capture of global and detailed spatial-temporal features. Experimental results reveal that the proposed model achieves at least a 5.85% and 6.80% increase inF1over other state-of-the-art algorithms with similar parameters and computations in both PTB-XL and Chapman databases. It indicates that the proposed model provides a promising avenue for intelligent diagnosis with limited computing resources. Huaicheng Zhang, Wenhan Liu, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Dense lead contrast for self-supervised representation learning of multilead electrocardiograms
Wenhan Liu, Zhoutong Li, Huaicheng Zhang, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
Inf. Sci. | 1 |
| 2022 | Lead Separation and Combination: A Novel Unsupervised 12-Lead ECG Feature Learning Framework for Internet of Medical ThingsabstractThe development of healthcare industry, especially Internet of Medical Things (IoMT), has generated considerable unlabeled electrocardiogram (ECG) signals. This article proposes a new unsupervised feature learning method for these unlabeled 12-lead ECGs, a type of 12-channel 1-D time series. Based on contrastive predictive coding (CPC), it considers the characteristics of 12-lead ECGs and develops novel lead-separation CPC (LSCPC) and lead-combination CPC (LCCPC). Specifically, LSCPC captures intralead features for each lead, while LCCPC combines all the leads and explores interlead relationships. Furthermore, a fusion model of LSCPC and LCCPC generates final representations. The Physikalisch-Technische Bundesanstalt (PTB)-XL database that contains 21837 12-lead records is used for unsupervised feature learning. Using learned features, linear classifiers are trained to accomplish the downstream tasks. 448 ECG records from 148 myocardial infarction (MI) and 52 healthy control subjects of the PTB database are used for MI detection. 6877 records from the CPSC-2018 database are used for atrial fibrillation (AF) detection, including 918 normal records, 1098 AF records, and 4861 other records. Using fivefold cross-validation, our model achieves 90.38% and 73.27% accuracy in MI and AF detection, respectively. Compared with existing models, it improves the performances by at least 2.32% for MI detection and 3.99% for AF detection. The model has been deployed on a lightweight embedded system (800-MHz ARM processor, 1-GB RAM). The maximum latency is only 465.34 ms, which can satisfy the real-time constraints. Overall, all the results have demonstrated the potential of our method for real-world healthcare, and lightweight IoMT applications. Wenhan Liu, Qianxi Guo, Xinwei Gao, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang |
IEEE Internet Things J. | 1 |
| 2020 | MFB-CBRNN: A Hybrid Network for MI Detection Using 12-Lead ECGsabstractThis paper proposes a novel hybrid network named multiple-feature-branch convolutional bidirectional recurrent neural network (MFB-CBRNN) for myocardial infarction (MI) detection using 12-lead ECGs. The model efficiently combines convolutional neural network-based and recurrent neural network-based structures. Each feature branch consists of several one-dimensional convolutional and pooling layers, corresponding to a certain lead. All the feature branches are independent from each other, which are utilized to learn the diverse features from different leads. Moreover, a bidirectional long short term memory network is employed to summarize all the feature branches. Its good ability of feature aggregation has been proved by the experiments. Furthermore, the paper develops a novel optimization method, lead random mask (LRM), to alleviate overfitting and implement an implicit ensemble like dropout. The model with LRM can achieve a more accurate MI detection. Class-based and subject-based fivefold cross validations are both carried out using Physikalisch-Technische Bundesanstalt diagnostic database. Totally, there are 148 MI and 52 healthy control subjects involved in the experiments. The MFB-CBRNN achieves an overall accuracy of 99.90% in class-based experiments, and an overall accuracy of 93.08% in subject-based experiments. Compared with other related studies, our algorithm achieves a comparable or even better result on MI detection. Therefore, the MFB-CBRNN has a good generalization capacity and is suitable for MI detection using 12-lead ECGs. It has a potential to assist the real-world MI diagnostics and reduce the burden of cardiologists. Wenhan Liu, Qijun Huang, Sheng Chang 0003, Hao Wang 0046, Jin He 0002 |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | Real-Time Multilead Convolutional Neural Network for Myocardial Infarction DetectionabstractIn this paper, a novel algorithm based on a convolutional neural network (CNN) is proposed for myocardial infarction detection via multilead electrocardiogram (ECG). A beat segmentation algorithm utilizing multilead ECG is designed to obtain multilead beats, and fuzzy information granulation is adopted for preprocessing. Then, the beats are input into our multilead-CNN (ML-CNN), a novel model that includes sub two-dimensional (2-D) convolutional layers and lead asymmetric pooling (LAP) layers. As different leads represent various angles of the same heart, LAP can capture multiscale features of different leads, exploiting the individual characteristics of each lead. In addition, sub 2-D convolution can utilize the holistic characters of all the leads. It uses 1-D kernels shared among the different leads to generate local optimal features. These strategies make the ML-CNN suitable for multilead ECG processing. To evaluate our algorithm, actual ECG datasets from the PTB diagnostic database are used. The sensitivity of our algorithm is 95.40%, the specificity is 97.37%, and the accuracy is 96.00% in the experiments. Targeting lightweight mobile healthcare applications, real-time analyses are performed on both MATLAB and ARM Cortex-A9 platforms. The average processing times for each heartbeat are approximately 17.10 and 26.75 ms, respectively, which indicate that this method has good potential for mobile healthcare applications. Wenhan Liu, Mengxin Zhang, Qijun Huang, Sheng Chang 0003, Hao Wang 0046, Jin He 0002 |
IEEE J. Biomed. Health Informatics | 1 |