Zhengyi Liu

dblp:155/4838 · DBLP profile ↗
← Back
40ranked-venue papers
26as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 20 first-author · 26 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GTAvatar: Efficient garment transfer for aligned avatars simultaneously reconstructed from monocular videos
Xianyong Fang, Baofeng Zhou, Linbo Wang 0001, Zhengyi Liu
Comput. Aided Geom. Des.5
2026 Multi-view simulation for robust polyp segmentation via cross-gated decoding and soft-attention fusion
Linbo Wang 0001, Jinxian Qiu, Zhengyi Liu, Xianyong Fang, Shaohua Wan 0001
Eng. Appl. Artif. Intell.4
2026 Text-Driven Medical Image Segmentation With LLM Semantic Bridge and LLM Prompt Bridge
abstract
Text-driven medical image segmentation aims to accurately segment pathological regions in medical images based on textual descriptions. Existing methods face two major challenges: (a) The significant modality heterogeneity between textual and visual features leads to inefficient cross-modal feature alignment; (b) The insufficient utilization of medical shared knowledge restricts semantic understanding. To address these challenges, two large language model (LLM) bridges are constructed. LLM semantic bridge leverages the sequential modeling capability of a frozen LLM to reorganize visual features into semantically coherent units that possess linguistic logic, thereby effectively bridging vision and language. The LLM prompt bridge appends learnable prompts, which encode medical shared knowledge from the LLM, to text embeddings, thereby effectively bridging case-specificity and medical consensus knowledge. Experimental results show the predominant performance due to LLM participation.https://github.com/liuzywen/MedBridge
Zhengyi Liu, Jiali Wu, Xianyong Fang, Linbo Wang 0001
IEEE Signal Process. Lett.1
2025 Embedding Space Decomposition Meets Invertible Networks: A New Paradigm for Unpaired Low-Light Enhancement
Linbo Wang 0001, Zhuo Yan, Zhengyi Liu, Xianyong Fang, Ping Li 0016
CGI (3)3
2025 Enhanced Context-Guided Aggregation Network for Remote Sensing Change Detection
abstract
Detecting changes in remote sensing imagery is crucial for applications including environmental surveillance, urban growth assessment, and disaster response planning. However, accurate identification of change regions remains challenging due to environmental factors, complex structures, and image quality variations. To enhance semantic understanding and retain local details in remote sensing change detection, we propose a novel framework called the Enhanced Context-Guided Aggregation Network (ECGA-Net). It includes three main parts: a Siamesestyle encoder, a difference-aware fusion module, and a decoder. We design an Enhanced Context Aggregation (ECA) module to fuse local and global contextual information, thereby enhancing feature representation capabilities. Additionally, we introduce a Dual Attention Fusion (DAF) module to extract difference features and suppress background interference, further improving the discriminative ability for change regions. Experiments show that ECGA-Net achieves an F1 score of 91.76% and an IoU of 84.62% on the LEVIR-CD dataset, and an F1 score of$\mathbf{9 2. 9 3} \%$and an IoU of$\mathbf{8 6. 6 1}$% on the WHU-CD dataset.
Zhengyi Liu, Ruibin Zhao, Xingsheng Lu
CW2
2025 Discrepancy Correction Reweight Aggregation Federated Learning for Incomplete Multimodal of Brain Tumor Segmentation
Zhengyi Liu, Weikai Shi, Linbo Wang 0001, Xianyong Fang
PRCV (14)1
2025 Incomplete Multi-modal Brain Tumor Segmentation via Multi-expert Collaboration
Zhengyi Liu, Jinhai Yu, Xianyong Fang, Linbo Wang 0001
PRCV (14)1
2025 Unified-Modal Salient Object Detection via Adaptive Prompt Learning
abstract
Existing single-modal and multi-modal salient object detection (SOD) methods focus on designing specific architectures tailored for their respective tasks. However, developing completely different models for different tasks leads to labor and time consumption, as well as high computational and practical deployment costs. In this paper, we attempt to address both single-modal and multi-modal SOD in a unified framework called UniSOD, which fully exploits the overlapping prior knowledge between different tasks. Nevertheless, assigning appropriate strategies to modality variable inputs is challenging. To this end, UniSOD learns modality-aware prompts with task-specific hints through adaptive prompt learning, which are seamlessly plugged into the proposed pre-trained baseline SOD model to handle corresponding tasks, while only requiring few learnable parameters compared to training the entire model from scratch. In particular, each modality-aware prompt is solely generated from a homogeneous switchable prompt generation (SPG) block, which adaptively performs structural switching based on single-modal and multi-modal inputs without manual intervention, ensuring that the framework can effectively handle diverse input cases (e.g., RGB-only, RGB-D, RGB-T) with a unified approach. Through end-to-end joint training, UniSOD achieves ovrall competitive performance on 14 benchmark datasets, demonstrating its ability to efficiently unify single-modal and multi-modal SOD tasks. Code has been available athttps://github.com/Angknpng/UniSOD
Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Zhengyi Liu, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 DiffRSD: Diffusion-Based and Integrity-Aware RGB-D Rail Surface Defect Inspection
Zhengyi Liu, Junnan Zhou, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001
IEEE Trans. Intell. Transp. Syst.1
2025 SSFam: Scribble Supervised Salient Object Detection Family
abstract
Scribble supervised salient object detection (SSSOD) constructs segmentation ability of attractive objects from surroundings under the supervision of sparse scribble labels. For the better segmentation, depth and thermal infrared modalities serve as the supplement to RGB images in the complex scenes. Existing methods specifically design various feature extraction and multi-modal fusion strategies for RGB, RGB-Depth, RGB-Thermal, and Visual-Depth-Thermal image input respectively, leading to similar model flood. As the recently proposed Segment Anything Model (SAM) possesses extraordinary segmentation and prompt interactive capability, we propose an SSSOD family based on SAM, namedSSFam, for the combination input with different modalities. Firstly, different modal-aware modulators are designed to attain modal-specific knowledge which cooperates with modal-agnostic information extracted from the frozen SAM encoder for the better feature ensemble. Secondly, a siamese decoder is tailored to bridge the gap between the training with scribble prompt and the testing with no prompt for the stronger decoding ability. Our model demonstrates the remarkable performance among combinations of different modalities and refreshes the highest level of scribble supervised methods and comes close to the ones of fully supervised methods.
Zhengyi Liu, Sheng Deng, Linbo Wang 0001, Xianyong Fang, Bin Tang 0003
IEEE Trans. Multim.1
2025 SemanticAvatar: human surface reconstruction based on semantically consistent biplane features
Baofeng Zhou, Xianyong Fang, Linbo Wang 0001, Zhengyi Liu
Vis. Comput.4
2024 Towards Finer Human Reconstruction for Single RGB-D Images
Renlong Dai, Linbo Wang 0001, Zhengyi Liu, Xianyong Fang
CGI (2)5
2024 LFSamba: Marry SAM With Mamba for Light Field Salient Object Detection
abstract
A light field camera can reconstruct 3D scenes using captured multi-focus images that contain rich spatial geometric information, enhancing applications in stereoscopic photography, virtual reality, and robotic vision. In this work, a state-of-the-art salient object detection model for multi-focus light field images, called LFSamba, is introduced to emphasize four main insights: (a) Efficient feature extraction, where SAM is used to extract modality-aware discriminative features; (b) Inter-slice relation modeling, leveraging Mamba to capture long-range dependencies across multiple focal slices, thus extracting implicit depth cues; (c) Inter-modal relation modeling, utilizing Mamba to integrate all-focus and multi-focus images, enabling mutual enhancement; (d) Weakly supervised learning capability, developing a scribble annotation dataset from an existing pixel-level mask dataset, establishing the first scribble-supervised baseline for light field salient object detection.
Zhengyi Liu, Longzhen Wang, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001
IEEE Signal Process. Lett.1
2023 Scribble-Supervised RGB-T Salient Object Detection
abstract
Salient object detection segments attractive objects in scenes. RGB and thermal modalities provide complementary information and scribble annotations alleviate large amounts of human labor. Based on the above facts, we propose a scribble-supervised RGB-T salient object detection model. By a four-step solution (expansion, prediction, aggregation, and supervision), label-sparse challenge of scribble-supervised method is solved. To expand scribble annotations, we collect the superpixels that foreground scribbles pass through in RGB and thermal images, respectively. The expanded multi-modal labels provide the coarse object boundary. To further polish the expanded labels, we propose a prediction module to alleviate the sharpness of boundary. To play the complementary roles of two modalities, we combine the two into aggregated pseudo labels. Supervised by scribble annotations and pseudo labels, our model achieves the state-of-the-art performance on the relabeled RGBT-S dataset. Furthermore, the model is applied to RGB-D and video scribble-supervised applications, achieving consistently excellent performance.1
Zhengyi Liu, Xiaoshen Huang, Xianyong Fang, Linbo Wang 0001, Bin Tang 0003
ICME1
2023 Sub-Band Based Attention for Robust Polyp Segmentation
abstract
This article proposes a novel spectral domain based solution to the challenging polyp segmentation. The main contribution is based on an interesting finding of the significant existence of the middle frequency sub-band during the CNN process. Consequently, a Sub-Band based Attention (SBA) module is proposed, which uniformly adopts either the high or middle sub-bands of the encoder features to boost the decoder features and thus concretely improve the feature discrimination. A strong encoder supplying informative sub-bands is also very important, while we highly value the local-and-global information enriched CNN features. Therefore, a Transformer Attended Convolution (TAC) module as the main encoder block is introduced. It takes the Transformer features to boost the CNN features with stronger long-range object contexts. The combination of SBA and TAC leads to a novel polyp segmentation framework, SBA-Net. It adopts TAC to effectively obtain encoded features which also input to SBA, so that efficient sub-bands based attention maps can be generated for progressively decoding the bottleneck features. Consequently, SBA-Net can achieve the robust polyp segmentation, as the experimental results demonstrate.
Xianyong Fang, Yuqing Shi, Linbo Wang 0001, Zhengyi Liu
IJCAI5
2023 Local Consensus Enhanced Siamese Network with Reciprocal Loss for Two-view Correspondence Learning
abstract
Recent studies of two-view correspondence learning usually establish an end-to-end network to jointly predict correspondence reliability and relative pose. We improve such a framework from two aspects. First, we propose a Local Feature Consensus (LFC) plugin block to augment the features of existing models. Given a correspondence feature, the block augments its neighboring features with mutual neighborhood consensus and aggregates them to produce an enhanced feature. As inliers obey a uniform cross-view transformation and share more consistent learned features than outliers, feature consensus strengthens inlier correlation and suppresses outlier distraction, which makes output features more discriminative for classifying inliers/outliers. Second, existing approaches supervise network training with the ground truth correspondences and essential matrix projecting one image to the other for an input image pair, without considering the information from the reverse mapping. We extend existing models to a Siamese network with a reciprocal loss that exploits the supervision of mutual projection, which considerably promotes the matching performance without introducing additional model parameters. Building upon MSA-Net [30], we implement the two proposals and experimentally achieve state-of-the-art performance on benchmark datasets.
Linbo Wang 0001, Xianyong Fang, Zhengyi Liu, Chenjie Cao, Yanwei Fu 0001
ACM Multimedia4
2023 Fine Back Surfaces Oriented Human Reconstruction for Single RGB-D Images
abstract
Abstract Current single RGB‐D image based human surface reconstruction methods generally take both the RGB images and the captured frontal depth maps together so that the 3D cues from the frontal surfaces can help infer the full surface geometries. However, we observe that the back surfaces can often be quite different from the frontal surfaces and, therefore, current methods can mess the recovery process by adopting such 3D cues, especially for the unseen back surfaces. We need to do the back surface inference without the frontal depth map. Consequently, a novel human reconstruction framework is proposed, so that human models with fine geometric details, especially for the back surfaces, can be obtained. In this approach, a progressive estimation method is introduced to effectively recover the unseen back depth maps. The coarse back depth maps are recovered by the parametric models of the subjects, with the fine ones further obtained by the normal‐maps conditioned GAN. This framework also includes a cross‐attention based denoising method for the frontal depth maps. This method adopts the cross attention between the features of the last two layers encoded from the frontal depth maps and thus suppresses the noise for fine depth maps by the attentions of features from the low‐noise and globally‐structured highest layer. Experimental results show the efficacies of the proposed ideas.
Xianyong Fang, Jinshen He, Linbo Wang 0001, Zhengyi Liu
Comput. Graph. Forum5
2023 Parallel matters: Efficient polyp segmentation with parallel structured feature augmentation modules
abstract
Abstract The large variations of polyp sizes and shapes and the close resemblances of polyps to their surroundings call for features with long‐range information in rich scales and strong discrimination. This article proposes two parallel structured modules for building those features. One is the Transformer Inception module (TI) which applies Transformers with different reception fields in parallel to input features and thus enriches them with more long‐range information in more scales. The other is the Local‐Detail Augmentation module (LDA) which applies the spatial and channel attentions in parallel to each block and thus locally augments the features from two complementary dimensions for more object details. Integrating TI and LDA, a new Transformer encoder based framework, Parallel‐Enhanced Network (PENet), is proposed, where LDA is specifically adopted twice in a coarse‐to‐fine way for accurate prediction. PENet is efficient in segmenting polyps with different sizes and shapes without the interference from the background tissues. Experimental comparisons with state‐of‐the‐arts methods show its merits.
Xianyong Fang, Kaibing Wang, Yuqing Shi, Linbo Wang 0001, Enming Zhang, Zhengyi Liu
IET Image Process.7
2023 Dilated high-resolution network driven RGB-T multi-modal crowd counting
Zhengyi Liu, Yacheng Tan, Bin Tang 0003
Signal Process. Image Commun.1
2023 LFTransNet: Light Field Salient Object Detection via a Learnable Weight Descriptor
abstract
Light Field Salient Object Detection (LF SOD) aims to segment the visually distinctive objects out of surroundings. Since light field images provide a multi-focus stack (many focal slices in different depth levels) and an all-focus image for the same scene, they record comprehensive but redundant information. Existing methods exploit the useful cue by long short-term memory with attention mechanism, 3D convolution, and graph learning. However, the importance of intra-slice and inter-slice in the focal stack is not well investigated. In the paper, we propose a learnable weight descriptor to simultaneously exploit different weights in slice, spatial region, and channel dimensions, and therefore propose an LF SOD method based on the learnable descriptor. The method extracts slice features and all-focus features from a weight-shared backbone and another backbone, respectively. A transformer decoder is used to learn the weight descriptor which both emphasizes the importance of each slice (inter-slice) and discriminates the spatial and channel importance of each slice (intra-slice). The learnt descriptor serves as the weight to make slice features attend to important slices, regions, and channels. Furthermore, we propose the hierarchical multi-modal fusion which aggregates high-layer features by modelling the long-range dependency to fully excavate common salient semantics and combines low-layer features by spatial constraint to eliminate the blurring effect of slice features. The experimental result exceeds the state-of-the-art methods at least 25% in terms of mean absolute error evaluation metric. It demonstrates a significant improvement in LF SOD performance via the designed learnable weight descriptor.https://github.com/liuzywen/LFTransNet
Zhengyi Liu, Linbo Wang 0001, Xianyong Fang, Bin Tang 0003
IEEE Trans. Circuits Syst. Video Technol.1
2023 HRTransNet: HRFormer-Driven Two-Modality Salient Object Detection
abstract
The High-Resolution Transformer (HRFormer) can maintain high-resolution representation and share global receptive fields. It is friendly towards salient object detection (SOD) in which the input and output have the same resolution. However, two critical problems need to be solved for two-modality SOD. One problem is two-modality fusion. The other problem is the HRFormer output’s fusion. To address the first problem, a supplementary modality is injected into the primary modality by using global optimization and an attention mechanism to select and purify the modality at the input level. To solve the second problem, a dual-direction short connection fusion module is used to optimize the output features of HRFormer, thereby enhancing the detailed representation of objects at the output level. The proposed model, named HRTransNet, first introduces an auxiliary stream for feature extraction of supplementary modality. Then, features are injected into the primary modality at the beginning of each multi-resolution branch. Next, HRFormer is applied to achieve forwarding propagation. Finally, all the output features with different resolutions are aggregated by intra-feature and inter-feature interactive transformers. Application of the proposed model results in impressive improvement for driving two-modality SOD tasks, e.g., RGB-D, RGB-T, and light field SOD.https://github.com/liuzywen/HRTransNet
Bin Tang 0003, Zhengyi Liu, Yacheng Tan
IEEE Trans. Circuits Syst. Video Technol.2
2022 RGB-T Multi-Modal Crowd Counting Based on Transformer
Zhengyi Liu, Yacheng Tan
BMVC1
2022 Boosting Camouflaged Object Detection with Dual-Task Interactive Transformer
abstract
Camouflaged object detection intends to discover the concealed objects hidden in the surroundings. Existing methods follow the bio-inspired framework, which first locates the object and second refines the boundary. We argue that the discovery of camouflaged objects depends on the recurrent search for the object and the boundary. The recurrent processing makes the human tired and helpless, but it is just the advantage of the transformer with global search ability. Therefore, a dual-task interactive transformer is proposed to detect both accurate position of the camouflaged object and its detailed boundary. The boundary feature is considered as Query to improve the camouflaged object detection, and meanwhile the object feature is considered as Query to improve the boundary detection. The camouflaged object detection and the boundary detection are fully interacted by multi-head self-attention. Besides, to obtain the initial object feature and boundary feature, transformer-based backbones are adopted to extract the foreground and background. The foreground is just object, while foreground minus background is considered as boundary. Here, the boundary feature can be obtained from blurry boundary region of the foreground and background. Supervised by the object, the background and the boundary ground truth, the proposed model achieves state-of-the-art performance in public datasets. https://github.com/liuzywen/COD
Zhengyi Liu, Yacheng Tan
ICPR1
2022 BGRDNet: RGB-D salient object detection with a bidirectional gated recurrent decoding network
Zhengyi Liu, Yacheng Tan
Multim. Tools Appl.1
2022 Global-guided cross-reference network for co-salient object detection
Zhengyi Liu, Yun Xiao 0003
Mach. Vis. Appl.1
2022 AGRFNet: Two-stage cross-modal and multi-level attention gated recurrent fusion network for RGB-D saliency detection
Zhengyi Liu, Yacheng Tan, Yun Xiao 0003
Signal Process. Image Commun.1
2022 SwinNet: Swin Transformer Drives Edge-Aware RGB-D and RGB-T Salient Object Detection
abstract
Convolutional neural networks (CNNs) are good at extracting contexture features within certain receptive fields, while transformers can model the global long-range dependency features. By absorbing the advantage of transformer and the merit of CNN, Swin Transformer shows strong feature representation ability. Based on it, we propose a cross-modality fusion model,SwinNet, for RGB-D and RGB-T salient object detection. It is driven by Swin Transformer to extract the hierarchical features, boosted by attention mechanism to bridge the gap between two modalities, and guided by edge information to sharp the contour of salient object. To be specific, two-stream Swin Transformer encoder first extracts multi-modality features, and then spatial alignment and channel re-calibration module is presented to optimize intra-level cross-modality features. To clarify the fuzzy boundary, edge-guided decoder achieves inter-level cross-modality fusion under the guidance of edge features. The proposed model outperforms the state-of-the-art models on RGB-D and RGB-T datasets, showing that it provides more insight into the cross-modality complementarity task.
Zhengyi Liu, Yacheng Tan, Yun Xiao 0003
IEEE Trans. Circuits Syst. Video Technol.1
2021 TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
abstract
Salient object detection is the pixel-level dense prediction task which can highlight the prominent object in the scene. Recently U-Net framework is widely used, and continuous convolution and pooling operations generate multi-level features which are complementary with each other. In view of the more contribution of high-level features for the performance, we propose a triplet transformer embedding module to enhance them by learning long-range dependencies across layers. It is the first to use three transformer encoders with shared weights to enhance multi-level features. By further designing scale adjustment module to process the input, devising three-stream decoder to process the output and attaching depth features to color features for the multi-modal fusion, the proposed triplet transformer embedding network (TriTransNet) achieves the state-of-the-art performance in RGB-D salient object detection, and pushes the performance to a new level. Experimental results demonstrate the effectiveness of the proposed modules and the competition of TriTransNet.
Zhengyi Liu, Zhengzheng Tu, Yun Xiao 0003, Bin Tang 0003
ACM Multimedia1
2021 A cross-modal edge-guided salient object detection for RGB-D image
Zhengyi Liu, Kaixun Wang
Neurocomputing1
2021 Multi-level progressive parallel attention guided salient object detection for RGB-D images
Zhengyi Liu, Quntao Duan, Peng Zhao 0010
Vis. Comput.1
2020 Attentive Part-aware Networks for Partial Person Re- identification
abstract
Partial person re-identification (re-ID) refers to re-identify a person through occluded images. It suffers from two major challenges, i.e., insufficient training data and incomplete probe image. In this paper, we introduce a part-aware learning method for partial person re-identification. On the one hand, we adopt data augmentation operation to enrich the training data and improve the robustness of the model. On the other hand, we intuitively find that the partial person images usually have fixed percentages of parts, therefore, in partial person re-ID task, the probe image could be cropped from the pictures and divided into several different partial types following fixed ratios. Based on the cropped images, we propose the Cropping Type Consistency (CTC) loss to classify the cropping types of partial images. Moreover, in order to help the network better fit the generated and cropped data, we incorporate the Block Attention Mechanism (BAM) into the framework for attentive learning. To enhance the retrieval performance in the inference stage, we implement cropping on gallery images according to the predicted types of probe partial images. Through calculating feature distances between the partial image and the cropped holistic gallery images, the model can recognize the right person from the gallery. To validate the effectiveness of our approach, we conduct extensive experiments on the partial re- ID benchmarks and achieve state-of-the-art performance.
Lijuan Huo, Chunfeng Song, Zhengyi Liu, Zhaoxiang Zhang 0001
ICPR3
2020 Deep layer guided network for salient object detection
Zhengyi Liu, Quanlong Li
Neurocomputing1
2020 A cross-modal adaptive gated fusion generative adversarial network for RGB-D salient object detection
Zhengyi Liu, Wei Zhang 0098, Peng Zhao 0010
Neurocomputing1
2020 Salient object detection for RGB-D images by generative adversarial network
Zhengyi Liu, Jiting Tang, Peng Zhao 0010
Multim. Tools Appl.1
2020 Salient object detection via hybrid upsampling and hybrid loss computing
Zhengyi Liu, Jiting Tang, Peng Zhao 0010
Vis. Comput.1
2020 Robust salient object detection for RGB images
Zhengyi Liu, Jiting Tang, Peng Zhao 0010
Vis. Comput.1
2019 Salient object detection for RGB-D image by single stream recurrent convolution neural network
Zhengyi Liu, Quntao Duan, Wei Zhang 0098, Peng Zhao 0010
Neurocomputing1
2019 RGB-D image saliency detection from 3D perspective
Zhengyi Liu, Tengfei Song
Multim. Tools Appl.1
2018 Co-saliency Detection for RGBD Images Based on Multi-constraint Superpixels Matching and Co-cellular Automata
Zhengyi Liu
PRCV (1)1
2014 A feature-emphasized clustering method for 2D vector field
abstract
Large-scale vector data produce the vector field clustering in flow visualization. To emphasize essential flow features, a new clustering method for 2D vector fields is proposed in this paper. With this method, the vector field is firstly initialized as a cluster, which is then iteratively divided into a hierarchy of clusters. During the iteration, clusters are segmented with streamlines instead of straight lines. This change enables it to emphasize flow features, since streamlines are consistent with flow behaviors, and clusters shaped by streamlines are aligned to the underlying flow. It is easy to capture flow patterns and features from resulting clusters. Moreover, our method improves representative vectors of clusters, leading to a more efficient approximation to the original field. Test results show that it is superior to other similar methods in terms of preserving flow features and approximating vector fields.
Mengyuan Guan, Zhengyi Liu
SMC4