Zhizhe Liu

dblp:205/7730 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-7571-4161ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BiKT: Unleashing the Potential of GNNs via Bi-Directional Knowledge Transfer
abstract
Based on the message-passing paradigm, there has been an amount of research proposing diverse and impressive feature propagation mechanisms to improve the performance of GNNs. However, less focus has been put on feature transformation, another major operation of the message-passing framework. In this paper, we first empirically investigate the performance of the feature transformation operation in several typical GNNs. Unexpectedly, we notice that GNNs do not completely free up the power of the inherent feature transformation operation. By this observation, we propose the Bi-directional Knowledge Transfer (BiKT), a plug-and-play approach to unleash the potential of the feature transformation operations without modifying the original architecture. Taking the feature transformation operation as a derived representation learning model that shares parameters with the original GNN, the direct prediction by this model provides a topological-agnostic knowledge feedback that can further instruct the learning of GNN and the feature transformations therein. On this basis, BiKT not only allows us to acquire knowledge from both the GNN and its derived model but also promotes each other by injecting the knowledge into the other. In addition, a theoretical analysis is further provided to demonstrate that BiKT improves the generalization bound of the GNNs from the perspective of domain adaptation. An extensive group of experiments on up to 7 datasets with 5 typical GNNs demonstrates that BiKT brings up to 0.5% - 4% performance gain over the original GNN, which means a boosted GNN is obtained. Meanwhile, the derived model also shows a powerful performance to compete with or even surpass the original GNN, enabling us to flexibly apply it independently to some other specific downstream tasks.
Shuai Zheng 0005, Zhizhe Liu, Zhenfeng Zhu, Xingxing Zhang 0001, Jianxin Li 0002, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Controllable Traffic Simulation through LLM-Guided Hierarchical Reasoning and Refinement
abstract
Evaluating autonomous driving systems in complex and diverse traffic scenarios through controllable simulation is essential to ensure their safety and reliability. However, existing traffic simulation methods face challenges in their controllability. To address this, we propose a novel diffusion-based and LLM-enhanced traffic simulation framework. Our approach incorporates a high-level understanding module and a low-level refinement module, which systematically examines the hierarchical structure of traffic elements, guides LLMs to thoroughly analyze traffic scenario descriptions step by step, and refines the generation by self-reflection, enhancing their understanding of complex situations. Furthermore, we propose a Frenet-frame-based cost function framework that provides LLMs with geometrically meaningful quantities, improving their grasp of spatial relationships in a scenario and enabling more accurate cost function generation. Experiments on the Waymo Open Motion Dataset (WOMD) demonstrate that our method can handle more intricate descriptions and generate a broader range of scenarios in a controllable manner.
Leheng Li, Haotian Lin 0006, Zhizhe Liu, Jianqiang Wang 0003
IROS6
2024 Node-Oriented Spectral Filtering for Graph Neural Networks
abstract
Graph neural networks (GNNs) have shown remarkable performance on homophilic graph data while being far less impressive when handling non-homophilic graph data due to the inherent low-pass filtering property of GNNs. In general, since real-world graphs are often complex mixtures of diverse subgraph patterns, learning a universal spectral filter on the graph from the global perspective as in most current works may still suffer from great difficulty in adapting to the variation of local patterns. On the basis of the theoretical analysis of local patterns, we rethink the existing spectral filtering methods and propose theNode-oriented spectralFiltering forGraphNeuralNetwork (namely NFGNN). By estimating the node-oriented spectral filter for each node, NFGNN is provided with the capability of precise local node positioning via the generalized translated operator, thus discriminating the variations of local homophily patterns adaptively. Meanwhile, the utilization of re-parameterization brings a good trade-off between global consistency and local sensibility for learning the node-oriented spectral filters. Furthermore, we theoretically analyze the localization property of NFGNN, demonstrating that the signal after adaptive filtering is still positioned around the corresponding node. Extensive experimental results demonstrate that the proposed NFGNN achieves more favorable performance.
Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Youru Li, Yao Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Graph meets probabilistic generation model: A new perspective for graph disentanglement
Zouzhang Peng, Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Jian Cheng 0001, Honghui Dong, Yao Zhao 0001
Pattern Recognit.4
2024 The Devil Is in the Boundary: Boundary-Enhanced Polyp Segmentation
abstract
Due to the various appearance of the polyps and the tiny contrast between the polyp area and its surrounding background, accurate polyp segmentation has become a challenging task. To tackle this issue, we introduce a boundary-enhanced framework for polyp segmentation, called the Focused on Boundary Segmentation (FoBS) framework, that leverages multi-level collaboration among sample, feature, and optimization. It places greater emphasis on the polyp boundary to improve the accuracy of segmentation. Firstly, a boundary-aware mixup method is designed to improve the model’s awareness of the boundary. More importantly, we propose deformable laplacian-based feature refining to explicitly strengthen the representation ability of the boundary features. It employs a deformable Laplacian refinement function to capture discriminative information from a deformable perceptual field, thereby improving its ability to adapt to boundary variations. In addition, we introduce the self-adjusting refinement coefficient learning that enables adaptive control over the refinement strength at each location. Furthermore, we develop a location-sensitive compensation criterion that assigns more importance to the degraded feature after feature refinement during optimization. Extensive quantitative and qualitative experiments on four polyp benchmarks demonstrate the effectiveness of our method for automatic polyp segmentation. Our code is available at https://github.com/TFboys-lzz/ FoBS.
Zhizhe Liu, Shuai Zheng 0005, Xiaoyi Sun, Zhenfeng Zhu, Xuebing Yang, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 From Observation to Concept: A Flexible Multi-View Paradigm for Medical Report Generation
abstract
Automated radiology report generation aims to generate accurate and radiologist-like descriptions for the patient's images, which can greatly relieve the workload of radiologists. However, due to the data bias and long report problems, medical report generation has been a challenging task. In this article, we propose aFlexibleMulti-viewParadigm (FMVP) for medical report generation in a novel observation-to-concept manner. It first makes some medical observations automatically or with the help of a radiologist on the patient's image to obtain patient-related priori knowledge, just as radiologists do in practice. Furthermore, to bridge the gap betweenpretrainandgenerationphases, the hierarchical alignment is proposed to jointly conduct the implicit alignment between region-tag and the explicit global alignment of the image-report pair. Finally, a compatible decoder towards decoding the fused multi-view knowledge is proposed to capture more complementary information for the report generation, which breaks the traditional entrenched decoding mechanism guided by visual information. Extensive quantitative and qualitative experiments on the public MIMIC-CXR and IU-Xray datasets show that our model achieves competitive performance compared to state-of-the-art methods.
Zhizhe Liu, Zhenfeng Zhu, Shuai Zheng 0005, Kunlun He, Yao Zhao 0001
IEEE Trans. Multim.1
2023 Sylvester Equation Induced Collaborative Representation Learning for Recommendation
abstract
For an actual recommendation system, it generally involves a variety of heterogeneous interactive relationships, such as the typical user-user (U2U), item-item (I2I), and user-item (U2I) interaction relationships. With the application of graph neural networks (GNNs) in embedding various interactive relations, recommendation technology has made gratifying progress in recent years, which benefits lot from its powerful ability in relation modeling. However, most of the existing GNN-based methods fail to collaboratively explore the above heterogeneous multiple interactive relationships, including the internal correlations among multiple relationships and the intrinsic association behind different relationships. As a consequence, the user's personalized preference for the items to be recommended will not be well captured. In this paper, we propose aSylvester equation inducedCollaborativeRepresentationLearning framework (S-CRL) for recommendation system by utilizing the heterogeneous multiple interactive relationships. In particular, we ingeniously define a novel Sylvester equation to associate tactfully the multiple heterogeneous relations together. From the perspective of rating propagation, such Sylvester equation is shown theoretically to be the optimal solution of a local structure sensitive rating propagation function. Additionally, to seek more expressive embeddings about user and item, a layer-wise attention is introduced to aggregate the multi-hop information from U2U and I2I graphs, respectively, so as to promote the aggregation with the corresponding embeddings from the U2I interaction graph. Extensive experiments on three real-world datasets verify that our model achieves more favorable performance over currently representative methods.
Xingyuan Li 0002, Zhenfeng Zhu, Shuai Zheng 0005, Zhizhe Liu, Youru Li, Deqiang Kong, Yao Zhao 0001
IEEE Trans. Knowl. Data Eng.4
2022 SGT: Scene Graph-Guided Transformer for Surgical Report Generation
Shuai Zheng 0005, Zhizhe Liu, Youru Li, Zhenfeng Zhu, Yao Zhao 0001
MICCAI (8)3
2022 Attention-Enhanced Disentangled Representation Learning for Unsupervised Domain Adaptation in Cardiac Segmentation
Xiaoyi Sun, Zhizhe Liu, Shuai Zheng 0005, Zhenfeng Zhu, Yao Zhao 0001
MICCAI (8)2
2022 The More, The Better? Active Silencing of Non-Positive Transfer for Efficient Multi-Domain Few-Shot Classification
abstract
Few-shot classification refers to recognizing several novel classes given only a few labeled samples. Many recent methods try to gain an adaptation benefit by learning prior knowledge from more base training domains, aka. multi-domain few-shot classification. However, with extensive empirical evidence, we find more is not always better: current models do not necessarily benefit from pre-training on more base classes and domains, since the pre-trained knowledge might be non-positive for a downstream task. In this work, we hypothesize that such redundant pre-training can be avoided without compromising the downstream performance. Inspired by the selective activating/silencing mechanism in the biological memory system, which enables the brain to learn a new concept from a few experiences both quickly and accurately, we propose to actively silence those redundant base classes and domains for efficient multi-domain few-shot classification. Then, a novel data-driven approach named Active Silencing with hierarchical Subset Selection (AS3) is developed to address two problems: 1) finding a subset of base classes that adequately represent novel classes for efficient positive transfer; and 2) finding a subset of base learners (i.e., domains) with confident accurate prediction in a new domain. Both problems are formulated as distance-based sparse subset selection. We extensively evaluate AS3 on the recent META-DATASET benchmark as well as MNIST, CIFAR10, and CIFAR100, where AS3 achieves over 100% acceleration while maintaining or even improving accuracy. Our code and Appendix are available at https://github.com/indussky8/AS3.
Xingxing Zhang 0001, Zhizhe Liu, Weikai Yang, Jun Zhu 0001
ACM Multimedia2
2022 MFHI: Taking Modality-Free Human Identification as Zero-Shot Learning
abstract
Human identification is an important topic in event detection, person tracking, and public security. There have been numerous methods proposed for human identification, such as face identification, person re-identification, and gait identification. Typically, existing methods predominantly classify a queried image to a specific identity in an image gallery set (I2I). This is seriously limited for the scenario where only a textual description of the query or an attribute gallery set is available in a wide range of video surveillance applications (A2IorI2A). However, very few efforts have been devoted towards modality-free identification, i.e., identifying a query in a gallery set in a scalable way. In this work, we take an initial attempt, and formulate such a novelModality-FreeHumanIdentification (named MFHI) task as a generic zero-shot learning model in a scalable way. Meanwhile, it is capable of bridging the visual and semantic modalities by learning a discriminative prototype of each identity. In addition, the semantics-guided spatial attention is enforced on visual modality to obtain interpretable representations with both high global category-level and local attribute-level discrimination. Finally, we design and conduct an extensive group of experiments on two common challenging identification tasks, including face identification and person re-identification, demonstrating that our method outperforms a wide variety of state-of-the-art methods on modality-free human identification.
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Margin Preserving Self-Paced Contrastive Learning Towards Domain Adaptation for Medical Image Segmentation
abstract
To bridge the gap between the source and target domains in unsupervised domain adaptation (UDA), the most common strategy puts focus on matching the marginal distributions in the feature space through adversarial learning. However, such category-agnostic global alignment lacks of exploiting the class-level joint distributions, causing the aligned distribution less discriminative. To address this issue, we propose in this paper a novel margin preserving self-paced contrastive Learning (MPSCL) model for cross-modal medical image segmentation. Unlike the conventional construction of contrastive pairs in contrastive learning, the domain-adaptive category prototypes are utilized to constitute the positive and negative sample pairs. With the guidance of progressively refined semantic prototypes, a novel margin preserving contrastive loss is proposed to boost the discriminability of embedded representation space. To enhance the supervision for contrastive learning, more informative pseudo-labels are generated in target domain in a self-paced way, thus benefiting the category-aware distribution alignment for UDA. Furthermore, the domain-invariant representations are learned through joint contrastive learning between the two domains. Extensive experiments on cross-modal cardiac segmentation tasks demonstrate that MPSCL significantly improves semantic segmentation performance, and outperforms a wide variety of state-of-the-art methods by a large margin.
Zhizhe Liu, Zhenfeng Zhu, Shuai Zheng 0005, Yang Liu 0235, Yao Zhao 0001
IEEE J. Biomed. Health Informatics1
2022 Multi-Modal Graph Learning for Disease Prediction
abstract
Benefiting from the powerful expressive capability of graphs, graph-based approaches have been popularly applied to handle multi-modal medical data and achieved impressive performance in various biomedical applications. For disease prediction tasks, most existing graph-based methods tend to define the graph manually based on specified modality (e.g., demographic information), and then integrated other modalities to obtain the patient representation by Graph Representation Learning (GRL). However, constructing an appropriate graph in advance is not a simple matter for these methods. Meanwhile, the complex correlation between modalities is ignored. These factors inevitably yield the inadequacy of providing sufficient information about the patient's condition for a reliable diagnosis. To this end, we propose an end-to-end Multi-modal Graph Learning framework (MMGL) for disease prediction with multi-modality. To effectively exploit the rich information across multi-modality associated with the disease, modality-aware representation learning is proposed to aggregate the features of each modality by leveraging the correlation and complementarity between the modalities. Furthermore, instead of defining the graph manually, the latent graph structure is captured through an effective way of adaptive graph learning. It could be jointly optimized with the prediction model, thus revealing the intrinsic connections among samples. Our model is also applicable to the scenario of inductive learning for those unseen data. An extensive group of experiments on two disease prediction tasks demonstrates that the proposed MMGL achieves more favorable performance. The code of MMGL is available at https://github.com/SsGood/MMGL.
Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Yang Liu 0235, Yao Zhao 0001
IEEE Trans. Medical Imaging3
2021 CETransformer: Casual Effect Estimation via Transformer Based Representation Learning
Shuai Zheng 0005, Zhizhe Liu, Zhenfeng Zhu
PRCV (4)3
2020 Distribution-Induced Bidirectional Generative Adversarial Network for Graph Representation Learning
abstract
Graph representation learning aims to encode all nodes of a graph into low-dimensional vectors that will serve as input of many computer vision tasks. However, most existing algorithms ignore the existence of inherent data distribution and even noises. This may significantly increase the phenomenon of over-fitting and deteriorate the testing accuracy. In this paper, we propose a Distribution-induced Bidirectional Generative Adversarial Network (named DBGAN) for graph representation learning. Instead of the widely used Gaussian assumption, the prior distribution of latent representation in our DBGAN is estimated in a structure-aware way, which implicitly bridges the graph and content spaces by prototype learning. Thus discriminative and robust representations are generated for all nodes. Furthermore, to improve their generalization ability while preserving representation ability, the sample-level and distribution-level consistency are well balanced via a bidirectional adversarial learning framework. An extensive group of experiments is then carefully designed and presented, demonstrating that our DBGAN obtains remarkably more favorable trade-off between representation and robustness, and meanwhile is dimension-efficient, over currently available alternatives in various tasks.
Shuai Zheng 0005, Zhenfeng Zhu, Xingxing Zhang 0001, Zhizhe Liu, Jian Cheng 0001, Yao Zhao 0001
CVPR4
2020 Convolutional prototype learning for zero-shot recognition
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001
Image Vis. Comput.1