Jiangyun Li

dblp:228/2717 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0003-2288-7901ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
YearPublicationVenuePosition
2026 BeltTear-seg: A lightweight model for belt tear segmentation with multi-scale feature squeeze attention and enhanced classification decision
Hebin Zhou, Yuqi Kong, Li Liu 0026, Jiangyun Li
Eng. Appl. Artif. Intell.5
2026 Generalized referring expression segmentation driven by instance-oriented queries
Jiangyun Li, Zhaokun Wen, Yisi Zhang, Wenxuan Wang 0002, Yuanxiu Cai, Xingjian He, Jing Liu 0001
Pattern Recognit.1
2026 3D ear reconstruction with Neural-Powell optimized SOP and curvature guidance
abstract
To address the limited precision in three-dimensional (3D) ear reconstruction, this study proposes a keypoint-driven hybrid optimization method for end-to-end reconstruction, achieving high-fidelity 3D ear model from single-view image. We propose the Neural-Powell optimized SOP(scaled orthographic projection), which integrates three strategies: least squares linear solution(3DMM), Powell’s gradient-free nonlinear optimization, and neural network-based nonlinear refinement. This approach enables high-precision estimation of SOP parameters. Building on this, we develop a curvature-guided personalized 3D reconstruction module that employs Gaussian curvature-guided Laplacian smoothing to preserve geometric details while ensuring surface continuity. Additionally, a visibility-aware adaptive texture mapping algorithm is proposed that uses surface normal analysis and bilinear interpolation to achieve high-fidelity texture reconstruction. Experimental results demonstrate that, compared to 3D Morphable Models (3DMM), the Neural-Powell optimized SOP reduces normalized root mean square error (RMSE) and normalized mean error (NME) by 80.7% and 81.2% respectively.And our method exhibits significant advantages in personalized feature representation and geometric detail preservation,offering a novel technical pathway to tackle the limited availability of 3D ear datasets.
He-Bin Zhou, Li Liu 0026, Jiangyun Li
Pattern Recognit. Lett.3
2026 Deep Compression on Segment Anything Model for Efficient Industrial Manufacturing
abstract
Segment Anything Model (SAM) is a popular vision foundation model that can segment data from any domain. Benefiting from its outstanding generalization ability, SAM has been widely adopted in many industrial scenarios. However, as SAM is built upon a heavy Vision Transformer (ViT), it suffers from memory-hungry and low latency, which restricts the deployment on edge devices. In this paper, we systematically explore how to compress SAM effectively, making it feasible to adapt edge devices with limited calculation abilities. Specifically, our method consists of three aspects: weight initialization, knowledge distillation, and model quantization. It is notable that all three aspects are not simply inherited from previous methods, but tactfully designed based on the teacher-student learning paradigm, considering the task-attributes of SAM pre-training. Firstly, we design a weight initialization method for fully using the pre-training knowledge implicitly contained in the teacher’s parameter space. Secondly, based on the weight initialization, we design a novel distillation method tailored to SAM pre-training, focusing on learning the semantic differences among areas. Lastly, we perform quantization on our distilled models. Unlike the previous method, we used both the teacher and the student to calibrate our model in the quantization process. We conduct systematic experiments on various teacher-student network pairs to validate the broad effectiveness of our method. By applying our method, our target models achieve over 64.5× speed increase compared to the original SAM. Core code is available at: https://github.com/ZG-ZZ/DC-SAM.
Yang Zheng 0002, Jie Liu 0028, Qing Li 0015, Jiangyun Li, Zhenghao Xi
IEEE Trans Autom. Sci. Eng.4
2026 MGD-SAM2: Multi-View Guided Detail-Enhanced Segment Anything Model 2 for High-Resolution Class-Agnostic Segmentation
abstract
Segment Anything Models (SAMs), as vision foundation models, have demonstrated remarkable performance across various image analysis tasks. Despite their strong generalization capabilities, SAMs encounter challenges in fine-grained detail segmentation for high-resolution class-independent segmentation (HRCS), due to the limitations in the direct processing of high-resolution inputs and low-resolution mask predictions, and the reliance on accurate manual prompts. To address these limitations, we propose MGD-SAM2, which integrates SAM2 with multi-view feature interaction between a global image and local patches to achieve precise segmentation. MGD-SAM2 incorporates the pre-trained SAM2 with four novel modules: the Multi-view Perception Adapter (MPAdapter), the Multi-view Complementary Enhancement Module (MCEM), the Hierarchical Multi-view Interaction Module (HMIM), and the Detail Refinement Module (DRM). Specifically, we first introduce MPAdapter to adapt the SAM2 encoder for enhanced local-global perception required by HRCS images. Then, MCEM and HMIM are proposed to further enrich the local texture and global semantics by aggregating multi-view features within and across multi-scales. Finally, DRM is designed to generate gradually restored high-resolution mask predictions by jointly leveraging mask features and original images, compensating for the detail loss caused by directly upsampling low-resolution prediction maps. Experimental results demonstrate the superior performance and strong generalization of our model on multiple high-resolution and normal-resolution datasets, while requiring only 8.8% trainable parameters. Code is available at https://github.com/sevenshr/MGD-SAM2.
Haoran Shen, Peixian Zhuang, Jiahao Kou, Yuxin Zeng, Haoying Xu, Jiangyun Li
IEEE Trans. Circuits Syst. Video Technol.6
2025 BeltDiff: Diffusion-based self-labeled generation system for conveyor belt damage detection
Peixian Zhuang, Yuanxiu Cai, Xianchao Zheng, Fuheng Xiao, Jiangyun Li
Eng. Appl. Artif. Intell.6
2025 Enhancing Semantic Information Representation in Multi-View Geo-Localization through Dual-Branch Network with Feature Consistency Enhancement and Multi-Level Feature Mining
abstract
ABSTRACT Metric learning is fundamental to multi‐view geo‐localization, as it aims to establish a distance metric that minimizes the feature space distance between similar data points while maximizing the separation between dissimilar ones. However, in Siamese networks employed for metric learning, individual branches may exhibit discrepancies in their interpretation of semantic information from input data, resulting in semantically inconsistent feature representations. To address this issue, a method is designed to enhance significant region consistency within multi‐view spaces by integrating feature consistency enhancement (FCE) and multi‐level feature mining (MLFM) techniques into a dual‐branch network. The FCE method emphasizes critical components of the input data, ensuring feature consistency between the two branches. Additionally, the MLFM mechanism facilitates feature integration across multiple levels, thereby enabling a more comprehensive extraction of semantic information. This approach enhances semantic understanding and promotes feature consistency across branches. The proposed method achieves AP values of 82.38% for drone‐to‐satellite and 77.36% for satellite‐to‐drone image matching. Notably, the method maintains computational efficiency without significantly affecting inference time. Additionally, improvements are observed in R@1, R@5 and R@10 metrics. The experimental results show that integrating FCE and MLFM into the dual‐branch network improves semantic representation and outperforms existing methods.
Yang Zheng 0002, Qing Li 0015, Jiangyun Li, Zhenghao Xi, Jie Liu 0028
IET Image Process.3
2025 State space models meet transformers for hyperspectral image classification
Xuefei Shi, Yisi Zhang, Zhaokun Wen, Wenxuan Wang 0002, Jiangyun Li
Signal Process.7
2025 DBMGNet: A Dual-Branch Mamba-GCN Network for Hyperspectral Image Classification
abstract
In hyperspectral image (HSI) classification, convolutional neural networks (CNNs) excel at local feature modeling but are limited to Euclidean space. Transformers offer long-range dependency modeling but suffer from high computational complexity. In contrast, graph convolutional networks (GCNs) can process information in non-Euclidean space, compensating for the limitations of CNNs. Meanwhile, the state space model Mamba, thanks to its linear complexity and strong long-range dependency modeling, shows great potential to offer an alternative to Transformers for HSI classification. To address the limitations of CNNs and Transformers while exploiting the potential of Mamba, we propose a dual-branch hybrid architecture named DBMGNet that combines Mamba with GCN for the HSI classification. In the Mamba branch, we design Band Selection Enhanced Bidirectional Mamba (BSEBM), which leverages Mamba’s long-range dependency modeling and sequential modeling capabilities to process spatial-spectral information. In the GCN branch, we apply reparameterized Chebyshev graph convolution to model similarity dependencies in non-Euclidean space, along with designing an adjacency matrix based on the intrinsic characteristics of HSIs. Extensive experiments demonstrate that our DBMGNet achieves the state-of-the-art performance of HSI classification against thirteen mainstream approaches. The code for this work will be available at https://github.com/Wanghao00pro/DBMGNet.
Hao Wang 0215, Peixian Zhuang, Jiangyun Li
IEEE Trans. Geosci. Remote. Sens.4
2024 Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype Enhancement
abstract
The Few-Shot Segmentation (FSS) aims to accomplish the novel class segmentation task with a few annotated images. Current FSS research based on meta-learning focuses on designing a complex interaction mechanism between the query and support feature. However, unlike humans who can rapidly learn new things from limited samples, the existing approach relies solely on fixed feature matching to tackle new tasks, lacking adaptability. In this paper, we propose a novel framework based on the adapter mechanism, namely Adaptive FSS, which can efficiently adapt the existing FSS model to the novel classes. In detail, we design the Prototype Adaptive Module (PAM), which utilizes accurate category information provided by the support set to derive class prototypes, enhancing class-specific information in the multi-stage representation. In addition, our approach is compatible with diverse FSS methods with different backbones by simply inserting PAM between the layers of the encoder. Experiments demonstrate that our method effectively improves the performance of the FSS models (e.g., MSANet, HDMNet, FPTrans, and DCAMA) and achieves new state-of-the-art (SOTA) results (i.e., 72.4% and 79.1% mIoU on PASCAL-5i 1-shot and 5-shot settings, 52.7% and 60.0% mIoU on COCO-20i 1-shot and 5-shot settings). Our code is available at https://github.com/jingw193/AdaptiveFSS.
Jing Wang 0222, Jiangyun Li, Chen Chen 0001, Yisi Zhang, Haoran Shen
AAAI2
2024 Med-DANet V2: A Flexible Dynamic Architecture for Efficient Medical Volumetric Segmentation
abstract
Recent works have shown that the computational efficiency of 3D medical image (e.g. CT and MRI) segmentation can be impressively improved by dynamic inference based on slice-wise complexity. As a pioneering work, a dynamic architecture network for medical volumetric segmentation (i.e. Med-DANet [44]) has achieved a favorable accuracy and efficiency trade-off by dynamically selecting a suitable 2D candidate model from the pre-defined model bank for different slices. However, the issues of incomplete data analysis, high training costs, and the two-stage pipeline in Med-DANet require further improvement. To this end, this paper further explores a unified formulation of the dynamic inference framework from the perspective of both the data itself and the model structure. For each slice of the input volume, our proposed method dynamically selects an important foreground region for segmentation based on the policy generated by our Decision Network and Crop Position Network. Besides, we propose to insert a stage-wise quantization selector to the employed segmentation model (e.g. U-Net) for dynamic architecture adapting. Extensive experiments on BraTS 2019 and 2020 show that our method achieves comparable or better performance than previous state-of-the-art methods with much less model complexity. Compared with previous methods Med-DANet and TransBTS with dynamic and static architecture respectively, our framework improves the model efficiency by up to nearly 4.1 and 17.3 times with comparable segmentation results on BraTS 2019. Code will be available at https://github.com/Rubics-Xuan/Med-DANet.
Haoran Shen, Wenxuan Wang 0002, Chen Chen 0001, Jing Liu 0001, Jiangyun Li
WACV7
2024 FreMIM: Fourier Transform Meets Masked Image Modeling for Medical Image Segmentation
abstract
The research community has witnessed the powerful potential of self-supervised Masked Image Modeling (MIM), which enables the models capable of learning visual representation from unlabeled data. In this paper, to incorporate both the crucial global structural information and local details for dense prediction tasks, we alter the perspective to the frequency domain and present a new MIM-based framework named FreMIM for self-supervised pre-training to better accomplish medical image segmentation tasks. Based on the observations that the detailed structural information mainly lies in the high-frequency components and the high-level semantics are abundant in the low-frequency counterparts, we further incorporate multi-stage supervision to guide the representation learning during the pre-training phase. Extensive experiments on three benchmark datasets show the superior advantage of our FreMIM over previous state-of-the-art MIM methods. Compared with various baselines trained from scratch, our FreMIM could consistently bring considerable improvements to model performance. The code will be publicly available at https://github.com/jingw193/FreMIM.
Wenxuan Wang 0002, Jing Wang 0222, Chen Chen 0001, Jianbo Jiao, Yuanxiu Cai, Jiangyun Li
WACV7
2024 Decomposition-Estimation-Reconstruction: An Automatic and Accurate Neuron Extraction Paradigm
abstract
The extraction of spatiotemporal neuron activity from calcium imaging videos plays a crucial role in unraveling the coding properties of neurons. While existing neuron extraction approaches have shown promising results, disturbing and scattering background and unused depth still impede their performance. To address these limitations, we develop an automatic and accurate neuron extraction paradigm, dubbed as decomposition–estimation–reconstruction (DER), consisting of D-procedure, E-procedure, and R-procedure. Specifically, the D-procedure first decomposes the raw data into a low-rank background and a sparse neuron signal, and regularizes$L_{0}$-norm priors of intensity and gradient of the neuron signal to suppress blurring and artifact effects. Then, the E-procedure estimates the depth-dependent transmission of the neuron signal based on its bright and dark channel priors. The R-procedure finally integrates the depth estimation of the neuron signal as a content-importance weight into a constrained non-negative matrix decomposition framework, which facilitates accurate neuron locations to boost the quality of extracted neurons. These three procedures are coupled in a cascade manner, where the former copes with calcium imaging data to facilitate the subsequent one. Comprehensive experiments on neuron extraction from calcium imaging videos demonstrate the superiority of our DER paradigm in both qualitative results and quantitative assessments over state-of-the-art methods.
Peixian Zhuang, Jiangyun Li, Qing Li 0015, Sam Kwong
IEEE Trans. Cybern.2
2024 CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation
abstract
Referring image segmentation (RIS) is a fundamental vision-language task that intends to segment a desired object from an image based on a given natural language expression. Due to the essentially distinct data properties between image and text, most of existing methods either introduce complex designs towards fine-grained vision-language alignment or lack required dense alignment, resulting in scalability issues or mis-segmentation problems such as over- or under-segmentation. To achieve effective and efficient fine-grained feature alignment in the RIS task, we explore the potential of masked multimodal modeling coupled with self-distillation and propose a novel cross-modality masked self-distillation framework named CM-MaskSD, in which our method inherits the transferred knowledge of image-text semantic alignment from CLIP model to realize fine-grained patch-word feature alignment for better segmentation accuracy. Moreover, our CM-MaskSD framework can considerably boost model performance in a nearly parameter-free manner, since it shares weights between the main segmentation branch and the introduced masked self-distillation branches, and solely introduces negligible parameters for coordinating the multimodal features. Comprehensive experiments on three benchmark datasets (ie RefCOCO, RefCOCO+, G-Ref) for the RIS task convincingly demonstrate the superiority of our proposed framework over previous state-of-the-art methods.
Wenxuan Wang 0002, Xingjian He, Yisi Zhang, Longteng Guo, Jiangyun Li, Jing Liu 0001
IEEE Trans. Multim.6
2023 Positive-negative equal contrastive loss for semantic segmentation
Jing Wang 0222, Jiangyun Li, Lingfei Xuan, Wenxuan Wang 0002
Neurocomputing2
2022 Med-DANet: Dynamic Architecture Network for Efficient Medical Volumetric Segmentation
Wenxuan Wang 0002, Chen Chen 0001, Jing Wang 0222, Sen Zha, Yan Zhang 0141, Jiangyun Li
ECCV (21)6
2022 Attention Guided Global Enhancement and Local Refinement Network for Semantic Segmentation
abstract
The encoder-decoder architecture is widely used as a lightweight semantic segmentation network. However, it struggles with a limited performance compared to a well-designed Dilated-FCN model for two major problems. First, commonly used upsampling methods in the decoder such as interpolation and deconvolution suffer from a local receptive field, unable to encode global contexts. Second, low-level features may bring noises to the network decoder through skip connections for the inadequacy of semantic concepts in early encoder layers. To tackle these challenges, a Global Enhancement Method is proposed to aggregate global information from high-level feature maps and adaptively distribute them to different decoder layers, alleviating the shortage of global contexts in the upsampling process. Besides, aLocal Refinement Module is developed by utilizing the decoder features as the semantic guidance to refine the noisy encoder features before the fusion of these two (the decoder features and the encoder features). Then, the two methods are integrated into a Context Fusion Block, and based on that, a novel Attention guided Global enhancement and Local refinement Network (AGLN) is elaborately designed. Extensive experiments on PASCAL Context, ADE20K, and PASCAL VOC 2012 datasets have demonstrated the effectiveness of the proposed approach. In particular, with a vanilla ResNet-101 backbone, AGLN achieves the state-of-the-art result (56.23% mean IOU) on the PASCAL Context dataset. The code is available at https://github.com/zhasen1996/AGLN.
Jiangyun Li, Sen Zha, Chen Chen 0001, Meng Ding 0001
IEEE Trans. Image Process.1
2021 TransBTS: Multimodal Brain Tumor Segmentation Using Transformer
Wenxuan Wang 0002, Chen Chen 0001, Meng Ding 0001, Sen Zha, Jiangyun Li
MICCAI (1)6
2020 Modeling Local and Global Contexts for Image Captioning
abstract
Image captioning aims to first observe an image, most notably the involved objects that are highly context-dependent, and then depict it with a natural description. However, most of the current models solely use the isolated objects vectors as image representations, ignoring the contexts among them. In this paper, we introduce a Local-Global Context (LGC) network, endowing the independent object features with shortrange perception (local contexts) and long-range dependence (global contexts). LGC network can be viewed as feature refiner, much beneficial to reason the novel objects and verbal words for the caption decoder. The local contexts are modeled with 1-D group convolution on adjacent objects, strengthening the local connections. Still further, self-attention mechanism is utilized to model the global contexts by correlating all the local contexts. Extensive experiments on MSCOCO dataset demonstrate that LGC network can easily plug into almost any neural captioning models and significantly improve the model performance.
Jiangyun Li, Longteng Guo, Jing Liu 0001
ICME2
2019 3D Dilated Multi-fiber Network for Real-Time Brain Tumor Segmentation in MRI
Chen Chen 0001, Meng Ding 0001, Junfeng Zheng, Jiangyun Li
MICCAI (3)5
2018 Collaborative Deconvolutional Neural Networks for Joint Depth Estimation and Semantic Segmentation
abstract
Semantic segmentation and single-view depth estimation are two fundamental problems in computer vision. They exploit the semantic and geometric properties of images, respectively, and are thus complementary in scene understanding. In this paper, we propose a collaborative deconvolutional neural network (C-DCNN) to jointly model these two problems for mutual promotion. The C-DCNN consists of two DCNNs, of which each is for one task. The DCNNs provide a finer resolution reconstruction method and are pretrained with hierarchical supervision. The feature maps from these two DCNNs are integrated via a pointwise bilinear layer, which fuses the semantic and depth information and produces higher order features. Then, the integrated features are fed into two sibling classification layers to simultaneously learn for semantic segmentation and depth estimation. In this way, we combine the semantic and depth features in a unified deep network and jointly train them to benefit each other. Specifically, during network training, we process depth estimation as a classification problem where a soft mapping strategy is proposed to map the continuous depth values into discrete probability distributions and the cross entropy loss is used. Besides, a fully connected conditional random field is also used as postprocessing to further improve the performance of semantic segmentation, where the proximity relations of pixels on position, intensity, and depth are jointly considered. We evaluate our approach on two challenging benchmarks: NYU Depth V2 and SUN RGB-D. It is demonstrated that our approach effectively utilizes these two kinds of information and achieves state-of-the-art results on both the semantic segmentation and depth estimation tasks.
Jing Liu 0001, Yong Li 0034, Jun Fu 0005, Jiangyun Li, Hanqing Lu
IEEE Trans. Neural Networks Learn. Syst.5