VLDB 2026 Research / reviewers in the wild / expert
Yucheng Shu
dblp:150/0715
· DBLP profile ↗
45ranked-venue papers
15as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 20 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Selective and safe structure transfer network for multi-contrast MRI super-resolution
Guoqing Ge, Weisheng Li 0001, Yucheng Shu, Xiaoyu Qiao |
Expert Syst. Appl. | 3 |
| 2026 | Multi-contrast feature cross entanglement network for joint MR image reconstruction and super-resolution
Guoqing Ge, Weisheng Li 0001, Yucheng Shu, Xiaoyu Qiao |
Knowl. Based Syst. | 3 |
| 2026 | DVAP-Reg: Dual-view anatomical prior-driven cross-dimensional registration for spinal surgery navigation
Zhengyang Wu 0002, Wenjie Zheng 0004, Yingjie Hao, Jing Ling, Maodan Nie, Rui Zuo, Minghan Liu, Zegang Shi, Wen Xia, Fayuan Zhou, Zhuojun Cao, Weisheng Li 0001, Guifeng Xia, Yucheng Shu, Chao Zhang 0106 |
Medical Image Anal. | 16 |
| 2026 | Label distribution learning via modeling label correlation on Gaussian components
Xueting Zheng, Chunlong Hu, Chang-Bin Shao, Yucheng Shu |
Multim. Syst. | 4 |
| 2026 | Rethinking normalization strategies and convolutional kernels for multimodal image fusion
Dan He 0010, Guofen Wang, Weisheng Li 0001, Yucheng Shu |
Pattern Recognit. | 4 |
| 2026 | Dual Knowledge Distillation Framework With Class-Adaptive Temperature and TopK Feature Perturbation for Few-Shot Prompt LearningabstractPre-trained vision-language models have shown great potential in few-shot learning. However, existing methods typically employ either KL divergence or feature similarity-based knowledge distillation, and rarely integrate both. Our analysis reveals that a naive simultaneous deployment of these two strategies yields suboptimal results. To address this, we propose a unified dual knowledge distillation framework. This framework is grounded in a theoretical derivation of class-adaptive temperature parameters, effectively resolving the incompatibility between KL divergence and feature similarity approaches. Furthermore, we introduce a top-K feature perturbation technique that targets specific features for more consistent enhancement than traditional noise regularization. Experimental results across 11 diverse benchmarks show that our approach yields consistent performance gains over various baselines. Notably, it improves the harmonic mean (H) by 0.41% to 0.72% and enhances generalization to unseen classes with an accuracy boost of up to 1.41%. Our source code is available at: https://github.com/sydney72380/DKL. Weisheng Li 0001, Yucheng Shu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Parallel Trajectory Constraint Sampling for Solving Universal Medical Inverse ProblemsabstractInverse problems in medical imaging, such as undersampled magnetic resonance imaging (MRI) and sparse-view computed tomography (CT) reconstruction, are essential yet challenging tasks for achieving accurate and reliable diagnostic images. Traditional reconstruction approaches, including iterative optimization algorithms and supervised deep learning methods, often struggle with limited adaptability across imaging protocols, substantial computational requirements, and poor generalization between different imaging modalities. Diffusion-based generative models have recently demonstrated promising results; however, these methods frequently suffer from cumulative estimation errors in their sampling processes, limiting their practical performance and robustness. In this paper, we propose a novel framework called Parallel Trajectory Constrained Sampling (PCS), which substantially enhances image reconstruction quality by explicitly enforcing consistency with the underlying physical measurement process. Specifically, PCS introduces a measurement-domain diffusion model whose reverse stochastic differential equation (SDE) trajectory is analytically determinable, thus obviating the need for a learned score estimator within the measurement domain. Furthermore, a parallel trajectory constraint is formulated to rigorously align the reverse sampling paths of the measurement and image diffusion processes, ensuring strict adherence to the known physical model at every sampling step. The proposed PCS method is flexible and can seamlessly integrate various SDE-based diffusion priors. Extensive experiments on representative inverse problems—including undersampled MRI reconstruction, sparse-view CT reconstruction, and image super-resolution—demonstrate that PCS consistently outperforms existing state-of-the-art diffusion-based reconstruction methods. Although current evaluations focus specifically on MRI and CT modalities, the PCS framework holds considerable promise for broader applicability to other imaging modalities and inverse problems, which we plan to investigate in future studies. Lihong Qiao, Rongxuan Wang, Yucheng Shu, Weisheng Li 0001, Zhanchuan Cai, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | LTOFusion: A Learning-to-Optimize Framework With Flow Matching for Unsupervised Image FusionabstractMultimodal Image Fusion (MMIF) aims to synthesize complementary information from different modalities to generate comprehensive fused images, thereby facilitating downstream applications. Existing methods typically employ deep neural networks to directly construct high-dimensional image-to-image mappings, which is highly challenging, struggling to extract generalizable patterns for various fusion scenarios. Inspired by meta learning, we propose a learning-to-optimize fusion framework, named LTOFusion, which formulates image fusion as a trajectory optimization problem, decoupling the complicated fusion problem into multistage subproblems. Subsequently, a restricted state transition function based on flow matching is designed to compress the prediction space and lead the network to build an image-to-flow mapping and fine-tune the current fusion state. To facilitate model training, we collect intermediate fusion states and utilize a memory-replay strategy, further enhancing the sample diversity and model robustness. In addition, a hybrid loss with respect to intensity, gradient, structure, and local normalized cross-correlation is designed to improve image details and reduce potential artifacts for fusion results. Experimental results demonstrate that the proposed method achieves the state-of-the-art performance across multiple fusion tasks and downstream applications without requiring fine-tuning. The code is available at https://github.com/HeDan-11/LTOFusion. Dan He 0010, Guofen Wang, Yucheng Shu, Weisheng Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Mask-Guided Proxy Mining Network for Few-Shot Medical Image SegmentationabstractFew-shot medical image segmentation (FSMIS) has attracted increasing attention as a promising technique for solving medical image segmentation tasks by relying on only a small amount of labeled data from new classes. Current FSMIS methods typically employ pixel-level semantic correlations between support-query image pairs to guide the segmentation of query images. However, the class information gap between support and query images may induce severe mismatches, leading to semantic ambiguity between foreground and background pixels. To address this issue, we propose a novel mask-guided proxy mining network (MPMNet), which mines a set of representative reference features (termed proxies) from support and query images to rectify foreground-background ambiguity. Specifically, to eliminate false pairwise matches caused by excessive intra-class variations, we design a mask-guided proxy mining module to adaptively learn representative proxies that can perceive visual differences between objects with different scales and shapes. Moreover, we integrate a hierarchical prior generation module and a context-aware feature enrichment module into MPMNet to obtain multi-scale information and enhance the discriminability of features. With these well-designed components and structures, our MPMNet can effectively overcome the adverse effects of false pixel matches by establishing proxy-level semantic correlations. Extensive experiments on three standard medical segmentation benchmarks demonstrate that our MPMNet significantly outperforms previous state-of-the-art methods, with a mean gain of 2.71% in DSC across all datasets. The code is available at: https://github.com/donglongzi/MPMNet. Wendong Huang, Jinwu Hu, Yongchao Wang 0004, Xiuli Bi, Yucheng Shu, Xuezong Yang, Bin Xiao 0002 |
IEEE Trans. Image Process. | 5 |
| 2026 | Multi-View Chest X-Ray Vision-Language Pre-Training via Semantic-Aware Masked Language Modeling and High-Order AlignmentabstractChest X-Ray Vision-Language pretraining (VLP) leverages large-scale radiograph-report pairs to develop joint image-text representations, demonstrating significant potential for medical image diagnosis. However, existing VLP approaches often overlook the multi-view nature of chest X-Rays, and some multi-view methods apply uniform feature fusion, neglecting view-key semantic contributions. Moreover, random cross-modal Masked Language Modeling (MLM) fails to facilitate effective interactions, impeding representation alignment. Additionally, global alignment in VLP may lead to the false-negative problem. To address these limitations, we propose a novel medical VLP framework comprising three core components. First, a Key Semantics-enhanced Multi-view MLM module aggregates pathology-relevant patches across views, providing semantically rich supervision for MLM. A local semantics enhancing approach, which identifies and aggregates pathology-relevant key patches across views to guide MLM. Second, a Frontal-Lateral Alignment module extracts view-specific pathological features, ensuring semantic consistency and preserving critical information during aggregation. This module independently extracts pathological features from both views to preserve view-specific information while ensuring semantic consistency, which mitigates the loss of crucial information during aggregation. Third, a High-order Semantic Alignment approach mitigates false-negative issues by aligning features with semantically consistent clusters, enhancing global alignment through prototype-level semantics. Extensive experiments across seven public datasets demonstrate that our framework outperforms state-of-the-art methods in four downstream tasks, validating its efficacy. The code is available at https://github.com/sajiutea/F-L. Lihong Qiao, Jingya Gong, Yucheng Shu, Lifang Zhou, Baobin Li, Weisheng Li 0001, Bai Ying Lei |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Self-Geometry-Guided Direct Pose Regression Based on Dual Perspective Fusion for 2D-3D Cross Dimensional Spinal Surgery Navigationabstract2D-3D cross-dimensional registration for spinal surgery navigation, which aims to achieve real-time visual navigation of preoperative 3D vertebrae based on intraoperative 2D fluoroscopy images, faces significant challenges due to semantic and dimensional gaps. Traditional 2D-3D registration methods often require fine adjustment steps and have low computational efficiency. In this paper, we propose a self-geometry-guided direct regression method based on dual perspective images. Firstly, an effective mechanism for unifying the dual view coordinate system was proposed. Secondly, a novel feature extraction module based on a face-graph convolutional network (F-GCN) is proposed to effectively extract 3D vertebra posture features. Finally, a posture direct regression network guided by self-vertebral geometry based on 2D-3D fusion features was constructed. Experimental results show that our method has made significant progress in solving the problem of 2D-3D cross-dimensional registration for spinal surgery navigation. Jing Ling, Zhengyang Wu 0002, Weisheng Li 0001, Chao Zhang 0106, Yucheng Shu |
ICASSP | 6 |
| 2025 | DGMIR: Dual-Guided Multimodal Medical Image Registration Based on Multi-view Augmentation and On-Site Modality Removal
Gao Le, Yucheng Shu, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
MICCAI (1) | 2 |
| 2025 | Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-trainingabstractCurrent medical vision-language pre-training models primarily follow two paradigms: report-supervised cross-modal alignment pre-training and reconstruction-based self-supervised pre-training. The former enhances the discriminative power of representations, while the latter facilitates fine-grained representation learning. However, naively combining these two paradigms inherits their inherent limitations: reconstruction-based methods treat all image patches equally during reconstruction, failing to effectively capture critical pathological details-since disease-related regions typically occupy only a small fraction of the image. Meanwhile, alignment-based methods suffer from suboptimal representations due to the presence of false negatives. To address these challenges, we propose a novel pre-training framework that integrates two key components: Pathology-Aware Reconstruction (PAR) and Discriminative Knowledge-Boosted Alignment (DKBA). Through a cascaded training strategy, our framework effectively combines the strengths of both paradigms while mitigating their inherent limitations. During the reconstruction pre-training stage, PAR incorporates pathology-aware priors to enhance the model's ability to capture fine-grained pathological details. In the alignment pre-training stage, DKBA leverages a medical knowledge graph as external supervision to improve cross-modal clustering alignment, thereby reducing the negative impact of false negatives. Extensive experiments on diverse downstream medical imaging tasks including image classification, object detection, and semantic segmentation, demonstrate the superior generalization capabilities of our method. Our code is publicly available at https://github.com/Felix1118/PADKB. Lihong Qiao, Shiyi Gao, Yucheng Shu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ACM Multimedia | 3 |
| 2025 | The Overlooked Matters: Revisiting Background, Prototype, and Activation in Few-Shot Medical Image Segmentation
Yucheng Shu, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2025 | Rethinking the CNN and transformer for deformable image registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Guofen Wang, Bin Xiao 0002 |
Expert Syst. Appl. | 3 |
| 2025 | A method for identifying urban road intersections considering the merging of parallel lanesabstractRoad intersections are crucial nodes in urban networks, where transportation lanes converge and socioeconomic activity is concentrated. Although methods to identify road intersections using raster maps, satellite images and trace data have been explored, challenges in accuracy and consistency remain. This paper proposes a method for identifying intersections based on OpenStreetMap data, which records networks at the lane level. Unlike geometric line intersections in OpenStreetMap, the identified intersections ‘summarize’ parallel lanes and incorporate branch information, such as counts and orientations. The proposed method uses a vector-raster fusion approach to initially locate intersections, followed by hierarchical geometric configuration to improve accuracy and extract branch data. Experimental results show that the method effectively handles complex road networks in various cities, accurately identifying intersections and their branches. Experiments conducted on OpenStreetMap data from 7 cities yielded over 97% precision and 99% recall, outperforming the state-of-the-art methods. Additionally, lane synthesis at intersections achieved 98.84% precision and 96.24% recall. Urban characteristics can be quantitatively analyzed based on the identified road intersections. For instance, the proportion of four-way road intersections in New York is 52.2%, whereas it is 8.9% in London, which may be attributed to the differing urban histories of these cities. Yucheng Shu, Dajiang Wang, Renyu Chen, Anxun Ren, Songshan Yue, Yongning Wen |
Int. J. Geogr. Inf. Sci. | 2 |
| 2025 | Cardiac cavity segmentation review in the past decade: Methods and future perspectives
Weisheng Li 0001, Yucheng Shu, Yidong Peng, Bin Xiao 0002 |
Neurocomputing | 3 |
| 2025 | 3D point cloud semantic segmentation based on visual guidance and feature enhancement
Yucheng Shu, Lihong Qiao, Zhengyang Wu 0002, Jing Ling, Jiang Wu 0006, Weisheng Li 0001 |
Multim. Syst. | 2 |
| 2025 | Fast Sampling of Diffusion Models for Accelerated MRI Using Dual Manifold ConstraintsabstractDiffusion models show great potential in solving inverse problems, including MRI reconstruction. With its unique characteristics, medical imaging demands both efficiency and accuracy in the reconstruction process. However, existing MRI reconstruction methods based on diffusion models often fall short of fully leveraging the available measurements during sampling. Consequently, these methods suffer from compromised reconstruction quality and elevated bias, especially when dealing with large acceleration factors. In response to these challenges, we propose Dual Manifold Constraints (DMC), a fast MRI reconstruction method based on diffusion models. We treat the sampling process as a combination of denoising and adding noise processes, and we constrain these two processes using both pristine measurements and their noisy counterparts to adapt to the geometry of diffusion. It’s worth noting that we propose a method to estimate the noisy measurement that satisfies the sub-sampling process to maintain the current data manifold when performing data consistency constraints. Experimental results show that our method outperforms the latest diffusion-based methods regarding both reconstruction speed and accuracy, and exhibits strong out-of-distribution generalization performance. Lihong Qiao, Rongxuan Wang, Yucheng Shu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001, Zhanchuan Cai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Whole Heart Segmentation Based on 3D Contour-Guided Multi-Head Attention Network From CT and MRI ImagesabstractHeart image segmentation is a critical task in medical image processing, which is crucial for the diagnosis and treatment planning of cardiovascular diseases. It helps doctors understand patients' cardiac anatomy and functional status more comprehensively and lays the foundation for personalized medicine and precision medicine research. Addressing the current challenges of rough surfaces on the entire heart, incomplete segmentation of heart substructures, and the lack of structured prediction of pulmonary arteries due to artifacts, scale diversity, uneven intensity, and boundary ambiguity in cardiac computed tomography (CT) and magnetic resonance imaging (MRI) images, we propose a whole heart segmentation algorithm based on 3D contour guided network. The proposed algorithm achieves robust whole heart segmentation results and has few network structure parameters. To enhance the consistency of features extracted by the codec, we propose a 3D codec information integration module to focus on task-related areas. In the final stage of information integration, features of different scales are combined. A 3D contour attention module enhances the perception of the heart's structure and shape. Contour prediction results from the initial stage, generating a low-resolution voxel of the entire heart with contour details. The second stage builds upon the initial phase of secondary learning to achieve multi-label segmentation results. The proposed algorithm achieved average Dice scores of 0.905 and 0.865 for the CT and MRI modalities, respectively, in 40 cases. Weisheng Li 0001, Yidong Peng, Yucheng Shu |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Learning from Inside: Self-driven Intra-modality Siamese Knowledge Generation and Inter-modality Alignment for Chest X-rays Vision-Language Pre-trainingabstractSince pathology occupies only a small portion of an X-ray, which means that a large portion of the information may be irrelevant to the paired radiology report, the Chest X-rays Report Understanding (CRU) task focuses on how to utilize small regions of the case to improve the performance of medical VLP. However, existing studies have neglected the fine-grained false negative samples of medical visual representations, resulting in their poor performance in CRU scenarios, which we attribute this to the fine-grained feature collapse problem. To address this issue, we propose an intra-modality siamese knowledge generation and inter-modality alignment framework, termed Chest X-rays Report Understanding Framework(CRUF). CRUF leverages the siamese knowledge in image-text pairs as guiding signals to distinguish fine-grained false negative and negative samples within the modality, and further narrows the distance between false negative and positive samples between modalities, accurately aligning the case regions of each image with the corresponding medical terms. Experimental results on multiple downstream medical image datasets covering tasks such as image classification, object detection, and semantic segmentation demonstrate the stability and outstanding performance of our framework. Code is available at https://github.com/cl-red/CRUF. Lihong Qiao, Yucheng Shu, Xiao Luan, Bin Xiao 0002 |
BIBM | 3 |
| 2024 | Re3adapter: Efficient Parameter Fing-Tuning with Triple Reparameterization for Adapter without Inference LatencyabstractWith the rise of large-scale model applications, leveraging these models as the base network for efficient transfer learning has garnered increasing attention. Currently, parameter-efficient transfer learning methods have made significant improvements in reducing the number of trainable parameters but introduce latency during inference. In this study, we propose an enhanced adaptation of the adapter using a reparameterization technique, revamping the activating layers into linear layers. This modification retains the high-dimensional fine-tuning capability of the adapter for visual tasks while avoiding additional inference latency. We name this plug-and-play module the Re3adapter, which optimizes the model with only 0.26% of the parameters and introduces no inference latency. Experimental results demonstrate its clear advantages in traditional classification and medical tasks. Lihong Qiao, Rui Wang 0173, Yucheng Shu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 3 |
| 2024 | Focal-Guided Multi-Consistency for Unsupervised Partial-to-Partial Point Cloud RegistrationabstractPoint Cloud Registration (PCR) is fundamental for the automatic perception of our space. With the rapid development of deep neural network, the community has swiftly adapted to this data-driven technique, and achieved promising performances. However, most existing learning-based methods attempt to conduct PCR within specific ideal experimental settings, in which the ground truth transformations are accessible and most of the data points have one-to-one correspondences. But in real-world scenarios, the GT transformations are often unknown, and point clouds may only share partially overlapped regions. It leads us to a challenging yet practical issue: How to perform Partial-to-Partial (PtP) Point Cloud Registration without pre-acquired supervisions? In this paper, we aim to tackle both challenges under a unified framework. To achieve this, we propose a novel Focal Anchor Generator to emulate the human perceptual process, particularly focusing on the mutual cloud parts. On top of it, a set of Multi-Consistency constraints are introduced to equip our model with the unsupervised learning ability, which is highly applicable. Extensive experiments have demonstrated the distinctive quality of our proposed framework. We believe this work will broaden the scope of PCR research and enhance the applicative potential of PCR algorithms. (The project code has been released on github.com/chengxiaojin/FGMC-UPCR). Yucheng Shu, Longjin Cheng, Bin Xiao 0002, Lihong Qiao, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 1 |
| 2024 | C3T: Contrastive Consistency Cross-Network Learning for Semi-Supervised Semantic SegmentationabstractSemi-supervised image semantic segmentation, a vital but challenging task in multimedia applications, aims to accurately classify pixels with limited labeled data. Traditional approaches in this domain often grapple with the confirmation bias problem, where models, influenced by their own predictions, become prone to replicating errors. To address this critical issue, our research introduces a cross-network-crossview consistency learning framework. This novel paradigm significantly reduce the confirmation bias through diversifying the learning perspectives. Integral to our approach are two components: a pseudo-label validation and filtering mechanism, and a cross-contrastive learning module within the feature domain. These elements work in synergy to not only amplify the accuracy of the model but also its robustness against varied data scenarios. Extensive evaluations, conducted across multiple datasets, clearly demonstrate the effectiveness of our method. In comparison to existing state-of-the-art models, our approach exhibits marked improvements, especially in the challenging contexts of semisupervised image semantic segmentation. The code is available at https://github.com/Sstar2orchid/C3T. Yucheng Shu, Jiaxin Xie, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 1 |
| 2024 | ShiftMorph: A Fast and Robust Convolutional Neural Network for 3D Deformable Medical Image Registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Bin Xiao 0002 |
ACM Multimedia | 3 |
| 2024 | CMRVAE: Contrastive margin-restrained variational auto-encoder for class-separated domain adaptation in cardiac segmentation
Lihong Qiao, Rui Wang 0173, Yucheng Shu, Bin Xiao 0002, Xidong Xu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001, Bai Ying Lei |
Knowl. Based Syst. | 3 |
| 2024 | Boosting Robust Multi-Focus Image Fusion With Frequency Mask and Hyperdimensional ComputingabstractMulti-focus image fusion (MFIF) creates an image from different source images with various sensors or optical settings as the devices can’t focus all objects at different distances. Most of the MFIF methods have several limitations in encoder enough features from the images and the result are not robust. To overcome the primary issue, we present a robust fusion algorithm based on the Frequency mask and the Hyperdimensional computing. We propose the Frequency Mask Filter (FMF) to get the narrow-band signals by encoding the frequency domain vector through the mask filter in the frequency domain. The Hyperdimensional encoder uses monogenic mapping, in which the multi-modulation features (MMF) such as the frequency, phase and amplitude are dynamically selected to obtain robust focus maps. Generated by multiscale monogenic representations of each image, the narrow-band image are mapped to hypervector encoding. Hyperdimensional encoder shows the energetic and structural information and leads to robust fusion results. Our proposed method is far superior to the existing MFIF method in terms of both objective evaluation metrics and visual effects on three publicly available datasets.Additionally, our proposed method requires only 0.88 seconds and has a parameter count of 0.13 million for multi-focus image fusion. Lihong Qiao, Shixin Wu, Bin Xiao 0002, Yucheng Shu, Xiao Luan, Sicheng Lu, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | SRPA: ScribbleMatch and Reliable-Guided Pixel Alignment for Scribble-Supervised Medical Image SegmentationabstractMedical image segmentation is a critical task in the field of medical image analysis. Recently, there has been increasing attention on scribble-supervised medical image segmentation due to its simplicity for label generation. However, the performance of a scribble-based task highly relies on the quality of learning from inconsistent annotations and the classification of pixels in high-entropy regions. In this paper, we propose a novel framework called SRPA that combines both ScribbleMatch and Reliable-Guided Pixel Alignment to enhance the performance of the scribble-based task. The ScribbleMatch technique utilizes the pseudo label incorporated from two different weakly perturbed views of the same image to supervise a strongly perturbed view, which assists in boosting the quality of the shape information learning as the scribble-based task lacks the prototypes to consistently capture shape prior to model training. The Reliable-Guided Pixel Alignment technique employs reliable pixels selected by contrasting two weakly perturbed views, which serve as the standard for blurred pixels in strongly perturbed images to maximally align. This ensures the reliable classification of high-entropy pixels. Our method is evaluated on the public ACDC and MSCMRseg datasets, and the results demonstrate that our approach surpasses current scribble-supervised segmentation methods. Code will be available at https://github.com/RheinSXY/SRPA. Tingjie Liu, Lihong Qiao, Yucheng Shu, Weisheng Li 0001, Xinbo Gao 0001 |
BIBM | 3 |
| 2023 | Non-rigid Medical Image Registration Based on Unsupervised Self-driven Prior FusionabstractDeformable image registration is a basic building block in intelligent bioinformatical analysis and biomedicine systems. With the rapid development of deep learning, the community has witnessed a great leap via this effective data-driven technique. Recently, the Vision Transformer, famous by its long-range modeling ability, has been successfully used in the field of medical image registration. However, the existing ViT based techniques are deemed to have certain limitations. Firstly, these methods often transplanted the transformer module directly into the networks, while did not dive deeper to explore its compatibility to practical registration tasks. Moreover, self-attention’s relatively rigid all-to-all patching strategy may cause undesirable discontinuity effect to the spatial calculation. To address these issues, we propose a novel medical image registration framework based on an efficient image prior learning and fusion mechanism. Unlike the existing prior-based registration methods, our model is capable of learning task-specific saliency priors, without the need of hand-crafted features, or heavy-loaded auxiliary tasks, or pre-acquired expensive annotations. Then, followed by a multi-scale patch embedding module, the self-driven saliency prior is integrated into a ViT block with an active feature fusion mechanism, to further expand our network’s structural learning abilities. Extensive experiments on multiple data sets have demonstrated the superior quality of the proposed framework. We believe this plug-and-play model will bring about more application potentials to the community (Project webpage: https://github.com/raincity212/SPF-Net). Yucheng Shu, Xuxuan Guan, Bin Xiao 0002, Lihong Qiao, Weisheng Li 0001, Xinbo Gao 0001 |
BIBM | 1 |
| 2023 | QDRJL: Quaternion dynamic representation with joint learning neural network for heart sound signal abnormal detectionabstractAt present, deep learning based heart sound diagnosis algorithms are mostly complex and large models for high accuracy, which are difficult to deploy on mobile devices due to the high number of parameters and large computational cost. The current mainstream approach for processing heart sound signals involves utilizing their Mel-frequency cepstral coefficients (MFCC) features. However, most existing methods have overlooked the multi-channel characteristics of MFCC. To address this issue, we propose a Quaternion Dynamic Representation with Joint Learning (QDRJL) neural network for learning MFCC multi-channel features. Our proposed approach combines quaternion dynamic convolution with dynamic weighting and the Quaternion Interior Learning Block (QILB). Finally, we present a global and energy joint learning branch for jointly learning MFCC features. The success of the proposed quaternion network depends on its ability to utilize the internal relations between quaternion-valued input features and the definition of the dynamic weight variables in the augmented quaternion domain. We assessed various state-of-the-art classification algorithms for detecting heart sounds and found that our proposed classifier achieved an accuracy of up to 97.2%, outperforming existing models. Our experimental evaluation, using the 2016 PhysioNet/CinC Challenge dataset, revealed that our model could reduce the number of network parameters to 25% due to quaternion properties. Lihong Qiao, Bin Xiao 0002, Yucheng Shu, Yuhang Shi, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2023 | Cross-Mix Monitoring for Medical Image Segmentation With Limited SupervisionabstractImage segmentation is a fundamental building block of automatic medical applications. It has been greatly improved since the emergence of deep neural networks. However, deep-learning based models often require a large number of manual annotations, which has seriously hindered its practical usage. To alleviate this problem, numerous works were proposed by utilizing unlabeled data based on semi-supervised frameworks. Recently, the Mean-Teacher (MT) model has been successfully applied in many scenarios due to its effective learning strategy. Nevertheless, the existing MT model still have certain limitations. Firstly, various sorts of perturbations are often added to the training data to gain extra generalization ability through consistency training. However, if the variation is too weak, it may cause the Lazy Student Phenomenon, and bring large fluctuations to the learning model. On the contrary, large image perturbations may enlarge the performance gap between the teacher and student. In this case, the student may lose its learning momentum, and more seriously, drag down the overall performance of the whole system. In order to address these issues, we introduce a novel semi-supervised medical image segmentation framework, in which a Cross-Mix Teaching paradigm is proposed to provide extra data flexibility, thus effectively avoid Lazy Student Phenomenon. Moreover, a lightweight Transductive Monitor is applied to server as the bridge that connect the teacher and student for active knowledge distillation. In the light of this cross-network information mixing and transfer mechanism, our method is able to continuously explore the discriminative information contained in unlabeled data. Extensive experiments on challenging medical image data sets demonstrate that our method is able to outperform current state-of-the-art semi-supervised segmentation methods under severe lack of supervision. Yucheng Shu, Hengbo Li, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Partial-to-Partial Point Cloud Registration Based on Multi-Level Semantic-Structural Cognitionabstract3D point cloud registration attempt to establish spatial correspondences between the source point cloud and the target point cloud. It is a fundamental task in computer vision and multimedia applications. Recently, many learning-based methods have been proposed and achieved promising performance. However, in the Partial-to-Partial (PtP) registration problem, the existence of a large number of external points may greatly handicap the effectiveness of these methods. In this paper, we propose to address the PtP issue under a novel multi-task cognition framework. At the global semantic level, we introduce a multi-scale feature exchanging network to actively evaluate the matching credibility. For the local structural learning, an inner-inter attention fusion branch is applied to generate discriminative features. Moreover, we integrate a novel alternating correspondences searching mechanism with a flexible bi-directional dislocation loss to perform robust learning, and a simple yet effective SVD weighting scheme is introduced at the inference stage. Experiment results on two challenging PtP 3D point cloud registration data sets show that our proposed method outperforms all the SOTA methods with higher precision and robustness. Yucheng Shu, Zongzhuang Hou, Bin Xiao 0002, Xiuli Bi, Xiao Luan, Weisheng Li 0001 |
ICME | 1 |
| 2022 | Registration-Is-Evaluation: Robust Point Set Matching With Multigranular Prior AssessmentabstractPoint set registration is one of the challenging tasks in remote sensing image processing and analysis. Its critical step is to find the corresponding relationships between the fixed scene point set and the moving model point set that undergo different sorts of transformations. Existing algorithms primarily utilize different types of prior information to improve the registration performance, such as spatial consistency, local similarity, and uniform outliers. However, due to the lack of active evaluation on the prior and intermediate information during the registration process, these strategies are susceptible to large data transformations. In order to enhance the robustness and accuracy for point set registration, we propose in this article a novel framework, namely Registration-is-Evaluation (RisE). Based on a multigranular probability model, our method exploits and utilizes both prior and posterior information to dynamically evaluate the matching status. What is more, instead of adding an extra uniform prior, we unified the outliers, noise, missing points, and heavily warped points into our registration evaluation model and address them simultaneously. We also apply a novel point set descriptor, called local polar relative geometry (LPRG), to have a more robust local similarity measurement. It adopts the local polar coordinate to perform multiscale pooling and relative geometric computation. Based on our proposed method, the matching relationships and the spatial transformations can be actively evaluated to provide useful contextual guidance for the registration process. Experimental results on multiple data sets show that our algorithm outperforms the state-of-the-art methods, in terms of both accuracy and robustness under large data degradations. Yucheng Shu, Zhenlong Liao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Rubik-Net: Learning Spatial Information via Rotation-Driven Convolutions for Brain SegmentationabstractThe accurate segmentation of brain tissue in Magnetic Resonance Image (MRI) slices is essential for assessing neurological conditions and brain diseases. However, it is challenging to segment MRI slices because of the low contrast between different brain tissues and the partial volume effect. 2-Dimensional (2-D) convolutional networks cannot handle such volumetric image data well because they overlook spatial information between MRI slices. Although 3-Dimensional (3-D) convolutions capture volumetric spatial information, they have not been fully exploited to enhance representative ability of deep networks; moreover, they may lead to overfitting with insufficient training data. In this paper, we propose a novel convolutional mechanism, termed Rubik convolution, to capture multi-dimensional information between MRI slices. Rubik convolution rotates the axis of a set of consecutive slices, enabling 2-D convolution kernels to extract features of each axial plane simultaneously. Next, feature maps are rotated back to fuse multidimensional information by the Max-View-Maps. Furthermore, we propose an efficient 2-D convolutional network, namely Rubik-Net, where the residual connections and the bottleneck structure are used to enhance information transmission and reduce the number of network parameters. The Rubik-Net shows promising results on iSeg2017, iSeg2019, IBSR and BrainWeb datasets in terms of segmentation accuracy. In particular, we achieved the best results in 95th percentile Hausdorff distance and average surface distance in cerebrospinal fluid segmentation on the most challenging iSeg2019 dataset. The experiments indicate that Rubik-Net improves the accuracy and efficiency of medical image segmentation. Moreover, Rubik convolution can be easily embedded into existing 2-D convolutional networks. Xiao Luan, Xinyu zheng, Weisheng Li 0001, Linghui Liu, Yucheng Shu |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Collaborative Learning With a Multi-Branch Framework for Feature EnhancementabstractFeature representation is highly important for many computer vision tasks. A broad range of prior studies have been proposed to strengthen representation ability of architectures via built-in blocks. However, during the forward propagation, the reduction in feature map scales still leads to the lack of representation ability. In this paper, we focus on boosting the representational power of a convolutional network by the multi-branch framework that we term the BranchNet. Each branch is directly supervised by label information to enrich the hierarchy features in BranchNet. Based on this framework, we further propose a collaborative learning loss and a soft target loss to transfer knowledge from deeper layers to shallow layers. BranchNet is an efficient training framework without extra parameters introduced in inference and can be integrated in existing networks, e.g., VGG, ResNet, and DenseNet. We evaluate BranchNet on all of these models and find that our method outperforms the baseline models on the widely-used CIFAR and ImageNet datasets. In particular, on the CIFAR-100 dataset, the classification error of ResNet-164 with BranchNet decreases by 4.51 percent. We also conduct experiments on the representative computer vision tasks of instance segmentation and class activation mapping, further verifying the superiority of BranchNet over the baseline models. Models and code are available athttps://github.com/zyyupup/BranchNet/. Xiao Luan, Weihua Ou, Linghui Liu, Weisheng Li 0001, Yucheng Shu, Hongmin Geng |
IEEE Trans. Multim. | 6 |
| 2021 | Medical Image Registration Based on Uncoupled Learning and Accumulative Enhancement
Yucheng Shu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001 |
MICCAI (4) | 1 |
| 2021 | Medical image segmentation based on active fusion-transduction of multi-stream features
Yucheng Shu, Bin Xiao 0002, Weisheng Li 0001 |
Knowl. Based Syst. | 1 |
| 2020 | AFT-Net: Active Fusion-Transduction for Multi-stream Medical Image SegmentationabstractAs an important building block in automatic medical applications, image segmentation has made a great progress due to the data-driving mechanism of deep architecture. Recently, numerous methods have been proposed to boost the segmentation performance based on U-shape network. However, they often built feature encoders with only one data routine, which have limited the representation ability of the networks. Although some methods applied multiple learning paths to fix this problem, the deep supervision techniques are required to monitor the training status at individual path, which may bring extra burden to practical usage of the algorithm. Additionally, under these frameworks, the semantic gap between different paths may interfere with model's learning performance, and the potential transduction ability of skip connections still needs further investigation. To address these issues, we introduce a novel medical image segmentation framework, namely AFT-Net, in which an attention-based data fusion model is proposed to effectively cooperate with multi-stream encoder. By progressively accumulating the features from different paths, our method can establish meaningful connections between structural and semantic features, while keeping an integral and flexible layout without deeply customized supervisions. Extensive experiments on two medical image data sets demonstrate that our method is able to acquire image features with both diversity and quality, thereby outperforms current state-of-the-art segmentation methods. Yucheng Shu, Bin Xiao 0002, Xiao Luan, Linghui Liu, Chunlong Hu |
ICTAI | 1 |
| 2019 | Robust Point Set Registration with Mixture Re-Weighting Based on Relative Geometric StructuresabstractPoint set registration is one of the challenging tasks in computer vision. One critical step is to find the corresponding relationship between the model point set and the scene point set. Existing registration algorithms primarily utilize the information of global and local shape, yet neglect the credibility of corresponding relation, therefore they may lead to the insufficient estimation of spatial transformation. To tackle this problem, we firstly adopt a relative polar coordinate system, it performs spatial pooling operation and further divides the feature extraction region into sub-areas with different scales. Then, based on the Relative Average Distance (RAD) and the Relative Average Offset Angle (RAOA), we propose multi-granular MRGS descriptor to extract visual structures of the point set. The similarity between the model point set and scene point set is then represented by the Gaussian Mixture Model, where the weights can be dynamically adjusted during the process of registration. Finally, we apply the robust mixture re-weighting to reduce the impact of false corresponding pairs and reinforce the weight of correct matching points. Experimental results on synthetic data and medical image data not only show that our method outperform state-of-the-art methods, but also demonstrate the robustness of our method when the non-grid transformation of point sets suffers from deformations, noises and outliers. Yucheng Shu, Zhenlong Liao |
ICTAI | 1 |
| 2019 | LVC-Net: Medical Image Segmentation with Noisy Label Based on Local Visual Cues
Yucheng Shu, Weisheng Li 0001 |
MICCAI (6) | 1 |
| 2018 | Multi-task Micro-expression Recognition Combining Deep and Handcrafted FeaturesabstractMicro-expression recognition is a challenging problem due to its short duration and low intensity. Most previous work on micro-expression mainly used the handcrafted features. Recently, deep learning methods were also employed for some difficult face recognition tasks. This paper presents a new framework to recognize micro-expression by combining handcrafted features and deep features. The employed handcrafted feature is called Local Gabor Binary Pattern from Three Orthogonal Panels (LGBP-TOP) feature. LGBP-TOP combines spatial and temporal analysis to encode the local facial movements. The employed deep feature is based on the Convolutional Neural Network (CNN) model trained on the micro-expression dataset. And then, the sparse multi-task learning framework with adaptive penalty term is employed to remove the irrelevant information from the combined LGBP-TOP and CNN features. The experimental evaluation is performed on two widely used micro-expression databases. The results demonstrate that the proposed approach achieves a competitive performance compared with other popular micro-expression recognition methods. Chunlong Hu, Dengbiao Jiang, Yucheng Shu |
ICPR | 5 |
| 2016 | Discriminative transform of receptive field patterns for feature representation
Yucheng Shu, Tianjiang Wang, Guangpu Shao, Chunlong Hu |
Multim. Tools Appl. | 1 |
| 2014 | DTRF: A physiologically motivated method for image descriptionabstractExtensive neurophysiological studies have shown that the receptive field plays a significant role in the human visual system. It has various kinds of properties such as orientation-selectivity, correlativity, etc. Motivated by these structural and functional properties, we propose in this paper a novel local image descriptor namely the Discriminative Transform of Receptive Fields (DTRF). Specifically, Receptive Field Patterns (RFP) are defined around each sample pixel and then divided into two kinds of components: RFP-Surround and RFP-Center. The RFP-Surround serves as the basic feature structure, which is extracted based on Local Annular Discrete Cosine Transform (LADCT) algorithm. The RFP-Center is used to pool these local features to simulate the correlative property of receptive field. Experimental results on the standard Oxford data set demonstrate the superiority of DTR-F over the state-of-the-art descriptors under various types of image transformations such as rotation and scaling changes, viewpoint changes, image blurring, JPEG compression, illumination changes, and image noise. Yucheng Shu, Tianjiang Wang, Guangpu Shao, Fang Liu 0011, Qi Feng 0003 |
ICIP | 1 |
| 2014 | Fuzzy c-means clustering with a new regularization term for image segmentationabstractWe present a new fuzzy c-means algorithm for image segmentation by introducing a novel spatially constrained Student's t-distribution and a new regularization term. Firstly, considering that conventional distribution models lack spatial information and the multivariate Student's t-distribution is heavily tailed, we propose a new way to incorporate spatial information between neighboring pixels into the Student's t-distribution based on Markov random field (MRF) in order to enhance robustness. Secondly, the new regularization term, inspired by the geodesic active contour (GAC) with a strong ability in capturing boundary, can preserve the details of edges and further enhance its robustness to noise and outliers by capitalizing on the local context information and edge information. Finally, in comparison to other Markov random fields that are complex and computationally expensive, the parameters are easily optimized with the EM algorithm in our proposed method. The proposed algorithm demonstrates the robustness and effectiveness, compared with other state-of-the-art methods on synthetic and real images. Guangpu Shao, Junbin Gao, Tianjiang Wang, Fang Liu 0011, Yucheng Shu |
IJCNN | 5 |
| 2014 | Robust Differential Circle Patterns based on fuzzy membership-pooling: A novel local image descriptor
Yucheng Shu, Tianjiang Wang, Guangpu Shao, Fang Liu 0011, Qi Feng 0003 |
Neurocomputing | 1 |