VLDB 2026 Research / reviewers in the wild / expert
Zequn Liu
dblp:219/8666
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 79% Question answering and dialogue systems · 8% Transfer learning and domain adaptation · 8% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 58% Data mining · 21% Knowledge graphs · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 54% Bioinformatics and computational biology · 46% | |
| Theoretical computer science
1 paper |
Combinatorics and discrete mathematics · 100% |
Topics — the 14 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
1.0 | 1 | 2026 | SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models · ACL (1) 2026 |
Medical and health informatics › medical report generation
radiology report generation |
1.0 | 1 | 2026 | EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language Models · SIGIR 2026 |
Information retrieval
retrieval-augmented generation |
1.0 | 1 | 2026 | EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language Models · SIGIR 2026 |
Natural language and speech › Language models and text generation
masked language modeling |
0.9 | 1 | 2025 | ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models · ICML 2025 |
Bioinformatics and computational biology › molecular informatics
molecular representation learning |
0.9 | 1 | 2025 | SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision · ICLR 2025 |
Data mining › structured data mining › graph mining
heterogeneous information network |
0.6 | 1 | 2022 | MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information Networks · EMNLP 2022 |
Knowledge graphs
knowledge graph embedding |
0.6 | 1 | 2022 | MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information Networks · EMNLP 2022 |
Natural language and speech › Language models and text generation › text generation › conditional text generation
definition generation |
0.5 | 1 | 2021 | Graphine: A Dataset for Graph-aware Terminology Definition Generation · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.4 | 1 | 2020 | Learning to Customize Model Structures for Few-shot Dialogue Generation Tasks · ACL 2020 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.4 | 1 | 2020 | Learning to Customize Model Structures for Few-shot Dialogue Generation Tasks · ACL 2020 |
Combinatorics and discrete mathematics › matrix theory
semi-tensor product |
0.3 | 1 | 2018 | From STP to game-based control · Sci. China Inf. Sci. 2018 |
Information retrieval
multimodal retrieval |
0.3 | 1 | 2026 | EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language Models · SIGIR 2026 |
Information retrieval › cross-modal retrieval
vision-language retrieval |
0.3 | 1 | 2026 | EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language Models · SIGIR 2026 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.3 | 1 | 2025 | SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 2.0structured triplet alignment · 2.0dense embedding similarity · 2.0fragment-level supervision · 1.7edit-based pre-training · 1.7automatic evaluation generation · 1.0text infilling · 0.6pre-trained language model · 0.6transformer · 0.5graph representation learning · 0.5graph neural network · 0.5private module · 0.4meta-learning · 0.4gating module · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsabstractYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan, Bin Feng, Yingce Xia, Shufang Xie, Kaili Liu, Bohan Wu, Qi Shi, Haoran Li, Beier Xiao, Zhiping Xiao, Xiao Luo, Weizhi Zhang, Philip S. Yu, Zequn Liu, Ming Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yiyang Gu, Junyu Luo 0002, Ye Yuan 0016, Yingce Xia, Shufang Xie 0003, Kaili Liu, Bohan Wu, Haoran Li 0003, Beier Xiao, Zhiping Xiao 0001, Xiao Luo 0001, Weizhi Zhang 0001, Philip S. Yu, Zequn Liu, Ming Zhang 0004 |
ACL (1) | 17 |
| 2026 | EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language ModelsabstractRetrieval-augmented generation (RAG) is widely adopted for radiology report generation with medical vision-language models, leveraging external reports as linguistic references. However, existing RAG methods rely primarily on dense embedding similarity, which may retrieve reports that are semantically related yet clinically inconsistent with respect to presence or laterality constraints. Such inconsistencies are often propagated into generation, resulting in contradictory or unsupported findings. We propose an evidence-guided retrieval-augmented framework EviRAG that decomposes retrieval into structured and unstructured alignment levels. First, we induce structured clinical triplets from both query and database cases through targeted visual interrogation, projecting images into a shared evidence space. Triplet-level alignment enforces explicit agreement over presence and laterality variables, yielding a clinically admissible candidate set via structural ranking. Within this constrained space, we perform semantic alignment in a shared multimodal embedding space to capture nuanced descriptive correspondence. The top-ranked reports and query image are jointly fed into a medical vision-language model for report generation. Comprehensive experiments on radiology report generation benchmarks show that EviRAG substantially reduces clinical inconsistencies compared to strong medical vision-language baselines. The source code is available at https://github.com/liamgu06/EviRAG. Yiyang Gu, Jiayue Fan, Kaili Liu, Bohan Wu, Binqi Chen, Zequn Liu, Zhiping Xiao 0001, Rongcheng Tu, Xiao Luo 0001, Ming Zhang 0004 |
SIGIR | 6 |
| 2026 | MAP-MIL: Dual-branch collaborative learning of mask enhancement and pseudo-bag generation for whole slide image classification
Zequn Liu, Liangkuan Zhu, Yining Xie, Jiayi Ma 0001 |
Expert Syst. Appl. | 1 |
| 2026 | M2PL-GAN: Multi-View Multi-Level Pathology Semantic Perception Learning for H&E-to-IHC Virtual StainingabstractImmunohistochemistry (IHC) staining is crucial for determining tumor subtypes, obtaining protein expression information, and developing personalized treatment plans. But compared with hematoxylin and eosin (H&E) staining, IHC staining is more complex and expensive. With the advancement of deep learning, converting H&E stained images into IHC stained images has gradually emerged as a solution for obtaining IHC staining. However, current virtual staining processes suffer from difficulties in aligning pathological semantic features, posing significant challenges for network training, which poses significant challenges for network training. To solve these issues, we propose a multi-view multi-level pathology semantic perception learning method for H&E-to-IHC virtual staining (M2PL-GAN). Unlike prior approaches, M2PL-GAN introduces a comprehensive semantic learning paradigm from three views: structural contextual relations, feature distribution, and topology-aware fine-grained semantics. These correspond to the Context-aware Correlation Mechanism (CACM), the Local-aware Distribution Alignment Mechanism (LDAM), and the Graph- aware Bidirectional Contrastive Learning Mechanism (GBCLM) respectively. Among them, CACM enhances contextual consistency by establishing semantic correlations between virtual and real IHC images at local scales. LDAM ensures alignment of semantic feature distributions between virtual and real IHC images, mitigating semantic shifts caused by HE-IHC staining. GBCLM leverages graph neural network to capture topology-aware semantic representations and optimizes semantic feature alignment through bidirectional contrastive learning. Extensive experiments on both public and private datasets demonstrate that our method outperforms state-of-the-art approaches in both quantitative metrics and qualitative evaluations. Our code is available in https://github.com/Pikachu-one/M2PL-GAN. Zequn Liu, Liangkuan Zhu, Yining Xie, Xiaoqing Hu, Haochen Qi, Jiayi Ma 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | SMI-Editor: Edit-based SMILES Language Model with Fragment-level SupervisionabstractSMILES, a crucial textual representation of molecular structures, has garnered significant attention as a foundation for pre-trained language models (LMs). However, most existing pre-trained SMILES LMs focus solely on the single-token level supervision during pre-training, failing to fully leverage the substructural information of molecules. This limitation makes the pre-training task overly simplistic, preventing the models from capturing richer molecular semantic information. Moreover, during pre-training, these SMILES LMs only process corrupted SMILES inputs, never encountering any valid SMILES, which leads to a train-inference mismatch. To address these challenges, we propose SMI-Editor, a novel edit-based pre-trained SMILES LM. SMI-Editor disrupts substructures within a molecule at random and feeds the resulting SMILES back into the model, which then attempts to restore the original SMILES through an editing process. This approach not only introduces fragment-level training signals, but also enables the use of valid SMILES as inputs, allowing the model to learn how to reconstruct complete molecules from these incomplete structures. As a result, the model demonstrates improved scalability and an enhanced ability to capture fragment-level molecular information. Experimental results show that SMI-Editor achieves state-of-the-art performance across multiple downstream molecular tasks, and even outperforming several 3D molecular representation models. Kangjie Zheng, Siyue Liang, Zequn Liu, Wei Ju 0001, Zhiping Xiao 0001, Ming Zhang 0004 |
ICLR | 5 |
| 2025 | ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
Kangjie Zheng, Siyue Liang, Zequn Liu, Wei Ju 0001, Zhiping Xiao 0001, Ming Zhang 0004 |
ICML | 5 |
| 2025 | SD-MIL: Multiple instance learning with dual perception of scale and distance information fusion for whole slide image classification
Yining Xie, Zequn Liu, Wei Zhang 0259, Jiayi Ma 0001 |
Expert Syst. Appl. | 2 |
| 2024 | A Comprehensive Survey on Deep Graph Representation Learning
Wei Ju 0001, Zheng Fang 0007, Yiyang Gu, Zequn Liu, Qingqing Long, Ziyue Qiao, Yifang Qin, Jianhao Shen, Zhiping Xiao 0001, Jingyang Yuan, Yusheng Zhao, Yifan Wang 0014, Xiao Luo 0001, Ming Zhang 0004 |
Neural Networks | 4 |
| 2023 | Few-shot Molecular Property Prediction via Hierarchically Structured Learning on Relation Graphs
Wei Ju 0001, Zequn Liu, Yifang Qin, Zhihui Guo, Xiao Luo 0001, Ming Zhang 0004 |
Neural Networks | 2 |
| 2022 | MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information NetworksabstractHeterogeneous Information Network (HIN) is essential to study complicated networks containing multiple edge types and node types.Meta-path, a sequence of node types and edge types, is the core technique to embed HINs.Since manually curating meta-paths is timeconsuming, there is a pressing need to develop automated meta-path generation approaches.Existing meta-path generation approaches cannot fully exploit the rich textual information in HINs, such as node names and edge type names.To address this problem, we propose MetaFill, a text-infilling-based approach for meta-path generation.The key idea of MetaFill is to formulate meta-path identification problem as a word sequence infilling problem, which can be advanced by Pretrained Language Models (PLMs).We observed the superior performance of MetaFill against existing meta-path generation methods and graph embedding methods that do not leverage meta-paths in both link prediction and node classification on two real-world HIN datasets.We further demonstrated how MetaFill can accurately classify edges in the zero-shot setting, where existing approaches cannot generate any meta-paths.MetaFill exploits PLMs to generate meta-paths for graph embedding, opening up new avenues for language model applications in graph analysis. Zequn Liu, Kefei Duan, Ming Zhang 0004, Sheng Wang 0012 |
EMNLP | 1 |
| 2021 | Graphine: A Dataset for Graph-aware Terminology Definition GenerationabstractPrecisely defining the terminology is the first step in scientific communication.Developing neural text generation models for definition generation can circumvent the laborintensity curation, further accelerating scientific discovery.Unfortunately, the lack of large-scale terminology definition dataset hinders the process toward definition generation.In this paper, we present a large-scale terminology definition dataset Graphine covering 2,010,648 terminology definition pairs, spanning 227 biomedical subdisciplines.Terminologies in each subdiscipline further form a directed acyclic graph, opening up new avenues for developing graph-aware text generation models.We then proposed a novel graphaware definition generation model Graphex that integrates transformer with graph neural network.Our model outperforms existing text generation models by exploiting the graph structure of terminologies.We further demonstrated how Graphine can be used to evaluate pretrained language models, compare graph representation learning methods and predict sentence granularity.We envision Graphine to be a unique resource for definition generation and many other NLP tasks in biomedicine. 1 Zequn Liu, Shukai Wang, Yiyang Gu, Ming Zhang 0004, Sheng Wang 0012 |
EMNLP (1) | 1 |
| 2020 | Learning to Customize Model Structures for Few-shot Dialogue Generation TasksabstractTraining the generative models with minimal corpus is one of the critical challenges for building open-domain dialogue systems.Existing methods tend to use the meta-learning framework which pre-trains the parameters on all non-target tasks then fine-tunes on the target task.However, fine-tuning distinguishes tasks from the parameter perspective but ignores the model-structure perspective, resulting in similar dialogue models for different tasks.In this paper, we propose an algorithm that can customize a unique dialogue model for each task in the few-shot setting.In our approach, each dialogue model consists of a shared module, a gating module, and a private module.The first two modules are shared among all the tasks, while the third one will differentiate into different network structures to better capture the characteristics of the corresponding task.The extensive experiments on two datasets show that our method outperforms all the baselines in terms of task consistency, response quality, and diversity. Yiping Song, Zequn Liu, Wei Bi, Rui Yan 0001, Ming Zhang 0004 |
ACL | 2 |
| 2020 | Multi-task Learning via Adaptation to Similar Tasks for Mortality Prediction of Diverse Rare Diseases
Luchen Liu, Zequn Liu, Haoxian Wu, Zichang Wang, Jianhao Shen, Yiping Song, Ming Zhang 0004 |
AMIA | 2 |
| 2018 | From STP to game-based control
Daizhan Cheng, Hongsheng Qi, Zequn Liu |
Sci. China Inf. Sci. | 3 |