VLDB 2026 Research / reviewers in the wild / expert
Zehan Li
dblp:287/5139
· DBLP profile ↗
20ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Boundary Token Graph for Zero-Shot Relation Triplet Extraction Involving Discontinuous EntitiesabstractZero-Shot Relation Triplet Extraction (ZSRTE) aims to extract head-tail entity pairs and their corresponding relations from sentences, where the relations available during inference are not seen during training. Existing methods typically assume that entities are continuous; however, in practice, entities can be discontinuous, which poses challenges to these approaches. To address this issue, we are the first to discuss and study the ZSRTE task involving discontinuous entities, and propose an innovative BoG framework, which is based on our proposed Boundary Token Graph structure. This method first predicts and adds edges between boundary tokens of (dis)continuous entities to construct a token graph, and then innovatively transforms the relation triplet extraction task into a process of finding paths in the graph. Additionally, we design a Boundary Token-Aware Prompt for each relation to further enhance the interaction between boundary tokens and relation semantics. Experimental results on four ZSRTE datasets—with or without discontinuous entities—consistently demonstrate that our method outperforms previous approaches, achieving state-of-the-art results. Kailun Lyu, Zehan Li, Fu Zhang 0001, Jingwei Cheng |
AAAI | 2 |
| 2026 | Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic SystemsabstractGuqin (古 琴) Jianzi (減 字) is an open and freely compositional tablature system that encodes performance actions rather than acoustic outcomes.Its automatic recognition remains largely unexplored, as conventional OCR assumes a closed and enumerable glyph set and struggles with Jianzi's unbounded composition and manuscript-level variability.We introduce Zero-shot Jianzi Recognition, which formulates Jianzi recognition as visionto-sequence prediction of canonical component sequences under a zero-shot split.To enable scalable supervision, we construct Synthetic-JZ from aligned online composition metadata.We then synthesize manuscriptlike training images via component-wise style recomposition and manuscript-domain noise modeling, and fine-tune a VLM for end-toend component sequence recognition.At inference time, a lightweight legality-guided correction module re-ranks decoding candidates, suppressing structural hallucinations without modifying the backbone.Experiments on two benchmarks show that our method achieves 63.02% sequence accuracy on Real-JZ, our manually annotated realworld Jianzi benchmark, surpassing Gemini-3-Pro by 35.11%.This result highlights the feasibility of reliable automated Jianzi recognition and its potential for large-scale digitization of historical Guqin Jianzi Pu manuscripts. Zehan Li, Fu Zhang 0001, Jingwei Cheng |
ACL (1) | 1 |
| 2025 | Re-Cent: A Relation-Centric Framework for Joint Zero-Shot Relation Triplet ExtractionabstractZero-shot Relation Triplet Extraction (ZSRTE) aims to extract triplets from the context where the relation patterns are unseen during training. Due to the inherent challenges of the ZSRTE task, existing extractive ZSRTE methods often decompose it into named entity recognition and relation classification, which overlooks the interdependence of two tasks and may introduce error propagation. Motivated by the intuition that crucial entity attributes might be implicit in the relation labels, we propose a Relation-Centric joint ZSRTE method named Re-Cent. This approach uses minimal information, specifically unseen relation labels, to extract triplets in one go through a unified model. We develop two span-based extractors to identify the subjects and objects corresponding to relation labels, forming span-pairs. Additionally, we introduce a relation-based correction mechanism that further refines the triplets by calculating the relevance between span-pairs and relation labels. Experiments demonstrate that Re-Cent achieves state-of-the-art performance with fewer parameters and does not rely on synthetic data or manual labor. Zehan Li, Fu Zhang 0001, Kailun Lyu, Jingwei Cheng, Tianyue Peng |
COLING | 1 |
| 2025 | CE-DA: Custom Embedding and Dynamic Aggregation for Zero-Shot Relation ExtractionabstractZero-shot Relation Extraction (ZSRE) aims to predict novel relations from sentences with given entity pairs, where the relations have not been encountered during training. Prototypebased methods, which achieve ZSRE by aligning the sentence representation and the relation prototype representation, have shown great potential. However, most existing works focus solely on improving the quality of prototype representations, neglecting sentence representations and lacking interaction between different types of relation side information. In this paper, we propose a novel ZSRE framework named CE-DA, which includes two modules: Custom Embedding and Dynamic Aggregation. We employ a two-stage approach to obtain customized embeddings of sentences. In the first stage, we train a sentence encoder through unsupervised contrastive learning, and in the second stage, we highlight the potential relations between entities in sentences using carefully designed entity emphasis prompts to further enhance sentence representations. Additionally, our dynamic aggregation method assigns different weights to different types of relation side information through a learnable network to enhance the quality of relation prototype representations. In contrast to traditional methods that treat the importance of all side information equally, our dynamic aggregation method further strengthen the interaction between different types of relation side information. Our method demonstrates competitive performance across various metrics on two ZSRE datasets. Fu Zhang 0001, Zehan Li, Jingwei Cheng |
COLING | 3 |
| 2025 | Frame First, Then Extract: A Frame-Semantic Reasoning Pipeline for Zero-Shot Relation Triplet ExtractionabstractLarge Language Models (LLMs) have shown impressive capabilities in language understanding and generation, leading to growing interest in zero-shot relation triplet extraction (Ze-roRTE), a task that aims to extract triplets for unseen relations without annotated data.However, existing methods typically depend on costly fine-tuning and lack the structured semantic guidance required for accurate and interpretable extraction.To overcome these limitations, we propose FrameRTE, a novel Ze-roRTE framework that adopts a "frame first, then extract" paradigm.Rather than extracting triplets directly, FrameRTE first constructs high-quality Relation Semantic Frames (RSFs) through a unified pipeline that integrates frame retrieval, synthesis, and enhancement.These RSFs serve as structured and interpretable knowledge scaffolds that guide frozen LLMs in the extraction process.Building upon these RSFs, we further introduce a human-inspired three-stage reasoning pipeline consisting of semantic frame evocation, frame-guided triplet extraction, and core frame elements validation to achieve semantically constrained extraction.Experiments demonstrate that FrameRTE achieves competitive zero-shot performance on multiple benchmarks.Moreover, the RSFs we construct serve as high-quality semantic resources that can enhance other extraction methods, showcasing the synergy between linguistic knowledge and foundation models. Frame ( Work )An Agent expends effort towards achieving a Goal.Alternatively, a Salient_entity involved in the Goal can be expressed in place of a Goal expression.Definition Agent: The Agent puts effort into reaching Goal. Core Frame Elements Goal:The Goal is what the Agent expends effort to achieve.Salient_entity: An entity that is centrally involved in the Goal that the Agent is attempting to acheive.Circumstances, Degree, Zehan Li, Fu Zhang 0001, Jingwei Cheng, Tianyue Peng |
EMNLP | 1 |
| 2025 | Privacy Text Clustering Method Based on Burst Feature of WordsabstractABSTRACT Real‐time detection of privacy‐relevant events in social media faces two fundamental challenges: (1) cluster instability caused by sparse and noisy text data, which leads to center drift; and (2) poor event discernibility in traditional online clustering methods. These limitations severely impair effective privacy monitoring in dynamic social media environments. To address these challenges, we propose an innovative edge intelligence‐driven framework that integrates adaptive burst word detection using wavelet‐based signal analysis; spectral clustering of identified burst words to establish stable event anchors; and real‐time incremental text clustering centered around these fixed anchors. We conduct a comprehensive evaluation on a dataset of 116 million COVID‐19‐related tweets and obtain the following results: Burst word identification accuracy of 86.28%; cluster purity of 0.875 (37% improvement over the baseline method); throughput of 3000 tweets per minute; and 78% reduction of irrelevant content through effective noise filtering. The key advantages of our approach include: Addressing the persistent cluster drift problem via burst anchoring centers; enabling efficient distributed processing via edge intelligence architecture; providing a practical and scalable solution for real‐time social media monitoring; and establishing a new paradigm for privacy‐aware event detection systems. Zehan Li, Hangyu Hu |
Concurr. Comput. Pract. Exp. | 2 |
| 2024 | ProCQA: A Large-scale Community-based Programming Question Answering Dataset for Code SearchabstractRetrieval-based code question answering seeks to match user queries in natural language to relevant code snippets. Previous approaches typically rely on pretraining models using crafted bi-modal and uni-modal datasets to align text and code representations. In this paper, we introduce ProCQA, a large-scale programming question answering dataset extracted from the StackOverflow community, offering naturally structured mixed-modal QA pairs. To validate its effectiveness, we propose a modality-agnostic contrastive pre-training approach to improve the alignment of text and code representations of current code language models. Compared to previous models that primarily employ bimodal and unimodal pairs extracted from CodeSearchNet for pre-training, our model exhibits significant performance improvements across a wide range of code retrieval benchmarks. Zehan Li, Jianfei Zhang 0003, Chuantao Yin, Yuanxin Ouyang, Wenge Rong |
LREC/COLING | 1 |
| 2024 | Improving large language models for clinical named entity recognition via prompt engineeringabstractIMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets. OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance. MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, and tests from clinical notes in the MTSamples corpus, following the 2010 i2b2 concept extraction shared task, and (2) to identify nervous system disorder-related adverse events from safety reports in the vaccine adverse event reporting system (VAERS). To improve the GPT models' performance, we developed a clinical task-specific prompt framework that includes (1) baseline prompts with task description and format specification, (2) annotation guideline-based prompts, (3) error analysis-based instructions, and (4) annotated samples for few-shot learning. We assessed each prompt's effectiveness and compared the models to BioClinicalBERT. RESULTS: Using baseline prompts, GPT-3.5 and GPT-4 achieved relaxed F1 scores of 0.634, 0.804 for MTSamples and 0.301, 0.593 for VAERS. Additional prompt components consistently improved model performance. When all 4 components were used, GPT-3.5 and GPT-4 achieved relaxed F1 socres of 0.794, 0.861 for MTSamples and 0.676, 0.736 for VAERS, demonstrating the effectiveness of our prompt framework. Although these results trail BioClinicalBERT (F1 of 0.901 for the MTSamples dataset and 0.802 for the VAERS), it is very promising considering few training samples are needed. DISCUSSION: The study's findings suggest a promising direction in leveraging LLMs for clinical NER tasks. However, while the performance of GPT models improved with task-specific prompts, there's a need for further development and refinement. LLMs like GPT-4 show potential in achieving close performance to state-of-the-art models like BioClinicalBERT, but they still require careful prompt engineering and understanding of task-specific knowledge. The study also underscores the importance of evaluation schemas that accurately reflect the capabilities and performance of LLMs in clinical settings. CONCLUSION: While direct application of GPT models to clinical NER tasks falls short of optimal performance, our task-specific prompt framework, incorporating medical knowledge and training samples, significantly enhances GPT models' feasibility for potential clinical applications. Qingyu Chen 0001, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou 0003, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | RefAI: a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarizationabstractOBJECTIVES: Precise literature recommendation and summarization are crucial for biomedical professionals. While the latest iteration of generative pretrained transformer (GPT) incorporates 2 distinct modes-real-time search and pretrained model utilization-it encounters challenges in dealing with these tasks. Specifically, the real-time search can pinpoint some relevant articles but occasionally provides fabricated papers, whereas the pretrained model excels in generating well-structured summaries but struggles to cite specific sources. In response, this study introduces RefAI, an innovative retrieval-augmented generative tool designed to synergize the strengths of large language models (LLMs) while overcoming their limitations. MATERIALS AND METHODS: RefAI utilized PubMed for systematic literature retrieval, employed a novel multivariable algorithm for article recommendation, and leveraged GPT-4 turbo for summarization. Ten queries under 2 prevalent topics ("cancer immunotherapy and target therapy" and "LLMs in medicine") were chosen as use cases and 3 established counterparts (ChatGPT-4, ScholarAI, and Gemini) as our baselines. The evaluation was conducted by 10 domain experts through standard statistical analyses for performance comparison. RESULTS: The overall performance of RefAI surpassed that of the baselines across 5 evaluated dimensions-relevance and quality for literature recommendation, accuracy, comprehensiveness, and reference integration for summarization, with the majority exhibiting statistically significant improvements (P-values <.05). DISCUSSION: RefAI demonstrated substantial improvements in literature recommendation and summarization over existing tools, addressing issues like fabricated papers, metadata inaccuracies, restricted recommendations, and poor reference integration. CONCLUSION: By augmenting LLM with external resources and a novel ranking algorithm, RefAI is uniquely capable of recommending high-quality literature and generating well-structured summaries, holding the potential to meet the critical needs of biomedical professionals in navigating and synthesizing vast amounts of scientific literature. Jeff Zhao, Manqi Li, Yifang Dang, Evan Yu, Jianfu Li, Zenan Sun, Usama Hussein, Jianguo Wen, Ahmed M. Abdelhameed, Junhua Mai, Shenduo Li, Yue Yu 0012, Xinyue Hu 0002, Daowei Yang, Jingna Feng, Zehan Li, Jianping He 0002, Tiehang Duan, Yanyan Lou, Fang Li 0011, Cui Tao |
J. Am. Medical Informatics Assoc. | 17 |
| 2024 | Artificial intelligence-powered pharmacovigilance: A review of machine and deep learning in clinical text-based adverse drug event detection for benchmark datasets
Zehan Li, Zenan Sun, Fang Li 0011, Susan H. Fenton, Hua Xu 0001, Cui Tao |
J. Biomed. Informatics | 3 |
| 2024 | Factorized and progressive knowledge distillation for CTC-based ASR models
Sanli Tian, Zehan Li, Zhaobiao Lyv, Gaofeng Cheng, Ta Li, Qingwei Zhao |
Speech Commun. | 2 |
| 2023 | Text Representation Distillation via Information Bottleneck PrincipleabstractPre-trained language models (PLMs) have recently shown great success in text representation field.However, the high computational cost and high-dimensional representation of PLMs pose significant challenges for practical applications.To make models more accessible, an effective method is to distill large models into smaller representation models.In order to relieve the issue of performance degradation after distillation, we propose a novel Knowledge Distillation method called IBKD.This approach is motivated by the Information Bottleneck principle and aims to maximize the mutual information between the final representation of the teacher and student model, while simultaneously reducing the mutual information between the student model's representation and the input data.This enables the student model to preserve important learned information while avoiding unnecessary information, thus reducing the risk of over-fitting.Empirical studies on two main downstream applications of text representation (Semantic Textual Similarity and Dense Retrieval tasks) demonstrate the effectiveness of our proposed approach 1 . Yanzhao Zhang, Dingkun Long, Zehan Li, Pengjun Xie |
EMNLP | 3 |
| 2023 | An Application of Quantum Mechanics to Attention Methods in Computer VisionabstractThis work proposes the quantum-state-based mapping (QSM) for machine learning. QSM uses wave functions that describe microscopic particle systems as mappings. By QSM, original inputs or features extracted by neural networks are processed as quantum states to train wave function parameters. QSM has a low computational cost, almost no additional parameters, and is easy to integrate with other modules. We demonstrate the simplest form of the wave function as a mapping, that is, when a one-dimensional particle is in an infinitely deep potential well, in combination with advanced attention modules. Experiments show that QSM significantly improves the feature recalibration ability of attention module in transfer learning tasks. Then, we tried to analyze the effectiveness of QSM. This work indicates that QSM has an important application value in interdisciplinary machine learning. Yihao Luo, Zehan Li, Wenbo An |
ICASSP | 4 |
| 2023 | Applications of Quantum Embedding in Computer Vision
Zehan Li, Wenbo An |
ICONIP (11) | 6 |
| 2023 | QCA-Net: Quantum-based Channel Attention for Deep Neural NetworksabstractThe channel attention mechanism, which adaptively recalibrates each channel's weight, can enhance the performance of deep neural networks. Most channel attention modules use simple pooling operations to aggregate spatial information. The drawback is the incapability to express complex global spatial information effectively. In this paper, we propose a Quantum-based Channel Attention (QCA), which only involves a handful of parameters but brings apparent performance gain. Using quantum mechanics analogy, we utilize wave functions describing microscopic particles to generate complex global spatial information. In addition, the QCA module has no convolutional layer, making it suitable for integration with various network architectures, including transformer and multilayer perceptron (MLP). We evaluate QCA through experiments on ImageNet-1K, and we also demonstrated the effect of QCA in combination with pre-training networks on small downstream transfer learning tasks. Zehan Li, Wenbo An |
IJCNN | 3 |
| 2023 | A feature engineering method for machine learning inspired by quantum mechanicsabstractThis work proposes a quantum-state-based feature engineering (QSFE) method for machine learning. QSFE uses wave functions that describe microscopic particle systems as mappings. By QSFE, original inputs or features extracted by neural networks are processed as quantum states to train wave function parameters. The experiments demonstrate that QSFE can improve the feature recalibration ability in deep neural networks. QSFE has a low computational cost, almost no additional parameters, and is easy to integrate with other modules. This work unfolds two effectiveness of QSFE: firstly, QSFE can enhance the expression ability of the model and make full use of the features extracted from the previous network; secondly, QSFE can extract complex spatial and temporal interactions, following the self-organization theory. The validations on various machine learning tasks, including a classical self-organization model, indicate that QSFE is valuable in interdisciplinary machine learning applications. Zehan Li, Wenbo An |
IJCNN | 3 |
| 2023 | QEA-Net: Quantum-Effects-based Attention Networks
Zehan Li, Wenbo An |
PRCV (3) | 6 |
| 2022 | Effects of Information Presentation Modalities on Antibiotic Reassessment Decision-Making in PICU: A Comparison Study
Zehan Li, Jules Bergmann, James C. Fackler, Harold P. Lehmann |
AMIA | 1 |
| 2022 | Improving Streaming End-to-End ASR on Transformer-based Causal Models with Encoder States Revision StrategiesabstractThere is often a trade-off between performance and latency in streaming automatic speech recognition (ASR).Traditional methods such as look-ahead and chunk-based methods, usually require information from future frames to advance recognition accuracy, which incurs inevitable latency even if the computation is fast enough.A causal model that computes without any future frames can avoid this latency, but its performance is significantly worse than traditional methods.In this paper, we propose corresponding revision strategies to improve the causal model.Firstly, we introduce a real-time encoder states revision strategy to modify previous states.Encoder forward computation starts once the data is received and revises the previous encoder states after several frames, which is no need to wait for any right context.Furthermore, a CTC spike position alignment decoding algorithm is designed to reduce time costs brought by the proposed revision strategy.Experiments are all conducted on Librispeech datasets.Fine-tuning on the CTC-based wav2vec2.0model, our best method can achieve 3.7/9.2WERs on test-clean/other sets and brings 45% relative improvement for causal models, which is also competitive with the chunkbased methods and the knowledge distillation methods. Zehan Li, Haoran Miao, Keqi Deng, Gaofeng Cheng, Sanli Tian, Ta Li, Yonghong Yan 0002 |
INTERSPEECH | 1 |
| 2022 | Knowledge Distillation For CTC-based Speech Recognition Via Consistent Acoustic Representation Learning
Sanli Tian, Keqi Deng, Zehan Li, Lingxuan Ye, Gaofeng Cheng, Ta Li, Yonghong Yan 0002 |
INTERSPEECH | 3 |