Weisheng Li 0001

dblp:10/7821-1 · DBLP profile ↗
← Back
225ranked-venue papers
8as first author
179since 2021 · last 2026
0000-0002-9033-8245ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 94 · 6 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 90 · 1 first-author · 77 since 2021Applied, interdisciplinary, general and emerging computing · 40 · 1 first-author · 35 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Computer networks · 4 · 3 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RIFD-DETR: Rotation-Invariant Face Detection with DETR and Polar-Aware Landmarks
Hathai Kaewkorn, Lifang Zhou, Weisheng Li 0001
FG3
2026 Selective and safe structure transfer network for multi-contrast MRI super-resolution
Guoqing Ge, Weisheng Li 0001, Yucheng Shu, Xiaoyu Qiao
Expert Syst. Appl.2
2026 TumorAL: Evidence-aware active learning for 3D tumor segmentation
Hongyi Wang 0006, Jiaxu Leng, Yue Zhao 0012, Weikai Li 0003, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing6
2026 MMP-YOLO: A multi-branch defect detection model based on rich gradient information
abstract
In the manufacturing process, the diversity of products and the complexity of the production environment pose severe challenges to the detection of small defects, which can lead to serious missed detections and false positives. To address these issues, this paper proposes a multi-branch defect detection model based on rich gradient information (MMP-YOLO), which significantly improves the performance of detecting defective objects. Specifically, we design three innovative modules integrated into MMP-YOLO. (1) The Multi-Level Gradient Lightweight Deep Network (MGLD) module processes multi-gradient information through a deep network integrated with large kernel convolution, ensuring accurate transmission of original input information and efficient feature extraction of small objects. (2) The Multi-Scale Function Complementary Upsampling (MFCU) module exploits the complementarity between high-resolution and low-resolution features and introduces transposed convolution and dilated convolution to further reduce information loss further. (3) The Parallel Task-Related Feature Selection (PTFS) module selectively suppresses background interference through a combination of global and local information. Extensive experiments on multiple datasets demonstrate that MMP-YOLO outperforms other state-of-the-art methods in reducing information loss and minimizing background noise interference.
Weisheng Li 0001, Shaoze Wang
Neurocomputing2
2026 DENAS: Differential evolution neural architecture search for prediction of diabetic retinopathy
Hangjiang Liu, Liming Xu, Jie Shao 0001, Weisheng Li 0001
Neurocomputing6
2026 EEL-RIS: Explain everything on localization via randomized input sampling
Jian-Xun Mi, Shenglei Shi, Lin Luo 0008, Weisheng Li 0001
Inf. Sci.4
2026 Multi-contrast feature cross entanglement network for joint MR image reconstruction and super-resolution
Guoqing Ge, Weisheng Li 0001, Yucheng Shu, Xiaoyu Qiao
Knowl. Based Syst.2
2026 DVAP-Reg: Dual-view anatomical prior-driven cross-dimensional registration for spinal surgery navigation
Zhengyang Wu 0002, Wenjie Zheng 0004, Yingjie Hao, Jing Ling, Maodan Nie, Rui Zuo, Minghan Liu, Zegang Shi, Wen Xia, Fayuan Zhou, Zhuojun Cao, Weisheng Li 0001, Guifeng Xia, Yucheng Shu, Chao Zhang 0106
Medical Image Anal.13
2026 SODAS: Second-order optimization differential architecture search for diabetic retinopathy prediction
Liming Xu, Jie Shao 0001, Jiancheng Lv 0001, Weisheng Li 0001
Neural Networks6
2026 HiViTrack: Hierarchical vision transformer with efficient target-prompt update for visual object tracking
Bailian Xie, Zongyi Xu, Weisheng Li 0001, Xinbo Gao 0001
Pattern Recognit.6
2026 Rethinking normalization strategies and convolutional kernels for multimodal image fusion
Dan He 0010, Guofen Wang, Weisheng Li 0001, Yucheng Shu
Pattern Recognit.3
2026 Text-Guided Modality Fusion for Brain Tumor Segmentation
Xiao Luan, Xiongfeng Huang, Linghui Liu, Weisheng Li 0001, Xinbo Gao 0001
Pattern Recognit.4
2026 Multi-view subspace clustering via tensor nuclear norm factorization
Tinghe Yan, Qiang Guo 0003, Jian-Xun Mi, Weisheng Li 0001
Pattern Recognit.4
2026 Harmonize task divergence with MUNet: Bridging semantic gaps between segmentation and classification for medical images
Laquan Li, Weisheng Li 0001, Shenhai Zheng
Pattern Recognit.3
2026 Dual Knowledge Distillation Framework With Class-Adaptive Temperature and TopK Feature Perturbation for Few-Shot Prompt Learning
abstract
Pre-trained vision-language models have shown great potential in few-shot learning. However, existing methods typically employ either KL divergence or feature similarity-based knowledge distillation, and rarely integrate both. Our analysis reveals that a naive simultaneous deployment of these two strategies yields suboptimal results. To address this, we propose a unified dual knowledge distillation framework. This framework is grounded in a theoretical derivation of class-adaptive temperature parameters, effectively resolving the incompatibility between KL divergence and feature similarity approaches. Furthermore, we introduce a top-K feature perturbation technique that targets specific features for more consistent enhancement than traditional noise regularization. Experimental results across 11 diverse benchmarks show that our approach yields consistent performance gains over various baselines. Notably, it improves the harmonic mean (H) by 0.41% to 0.72% and enhances generalization to unseen classes with an accuracy boost of up to 1.41%. Our source code is available at: https://github.com/sydney72380/DKL.
Weisheng Li 0001, Yucheng Shu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Diff-AEPNet: Facial Aesthetic Enhancement and Prediction Network Based on Differential Average Aesthetic Perceptions
abstract
With the advent of the intelligent era, increasing attention has been given to facial aesthetics. While the academic community has achieved notable progress in facial aesthetic research, current efforts predominantly concentrate on two isolated subtasks: aesthetic evaluation and enhancement. Crucially, the intrinsic correlation between these tasks and their integration within a unified framework remain underexplored. To bridge this gap, this paper proposes a facial aesthetic enhancement and prediction network based on differential average aesthetic perceptions (Diff-AEPNet) that synergistically combines facial aesthetic enhancement with prediction. The proposed framework implements a four-stage architecture: (1) a transformer module learns latent code beautification trajectories to guide preenhancement feature modification; (2) a dual-stream encoder extracts and contrasts pre/postbeautification features to refine evaluation accuracy; (3) a lightweight network generates attention-guided image mask for image fusion; and (4) a deghosting block eliminates fusion artifacts through residual learning. The experimental results demonstrate that the model achieves a favorable beautification effect in the enhancement task and exhibits better generalization performance across datasets in the evaluation task than existing aesthetic evaluation models do.
Weisheng Li 0001, Bin Xiao 0002, Yong Wang 0009
IEEE Trans. Circuits Syst. Video Technol.2
2026 HandJoKe: Joint-Guided Keypoint Denoising Transformer for Depth-Based 3D Hand Pose Estimation
abstract
Existing depth-based 3D hand pose estimation methods typically estimate hand joints from either 2D depth images or 3D point clouds, whereas the approaches that fuse multimodal data remain underexplored. Furthermore, previous methods often struggle to learn geometric-facilitated features and precise joint correlations, especially for occluded hands, due to the lack of explicit prior guidance and insufficient cross-dimensional interaction. By taking advantage of multi-modal fusion, cross-dimensional interaction, and prior guidance, we propose a novel joint-guided keypoint denoising Transformer (named HandJoKe) to achieve more precise hand pose estimation, which can iteratively estimate hand poses based on keypoint features from both 2D depth images and 3D point clouds under explicit joint guidance within only several denoising steps. Rather than directly applying existing multi-modal fusion to perform redundant interactions among many background pixels and irrelevant points, HandJoKe focuses on modeling correlations and capturing dependencies among local informative hand regions (i.e., keypoints), thus attaining higher learning capability with lower computation redundancy. Moreover, a novel joint-guided denoising estimation strategy is introduced to adequately fuse cross-modal keypoint features under explicit joint guidance, achieving geometric-facilitated cross-modal keypoint interaction in both 2D and 3D spaces. The effectiveness of joint guidance can be further strengthened through iterative denoising, since it can subsequently update cross-modal keypoint features based on previous denoised hand poses and thus can help better locate confused joints, especially for occluded hands. Extensive experiments show that HandJoKe has achieved state-of-the-art performance on four public challenging benchmarks, including single-hand datasets NYU and ICVL, and hand-object datasets DexYCB and HO3D.
Ji Gan, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Parallel Trajectory Constraint Sampling for Solving Universal Medical Inverse Problems
abstract
Inverse problems in medical imaging, such as undersampled magnetic resonance imaging (MRI) and sparse-view computed tomography (CT) reconstruction, are essential yet challenging tasks for achieving accurate and reliable diagnostic images. Traditional reconstruction approaches, including iterative optimization algorithms and supervised deep learning methods, often struggle with limited adaptability across imaging protocols, substantial computational requirements, and poor generalization between different imaging modalities. Diffusion-based generative models have recently demonstrated promising results; however, these methods frequently suffer from cumulative estimation errors in their sampling processes, limiting their practical performance and robustness. In this paper, we propose a novel framework called Parallel Trajectory Constrained Sampling (PCS), which substantially enhances image reconstruction quality by explicitly enforcing consistency with the underlying physical measurement process. Specifically, PCS introduces a measurement-domain diffusion model whose reverse stochastic differential equation (SDE) trajectory is analytically determinable, thus obviating the need for a learned score estimator within the measurement domain. Furthermore, a parallel trajectory constraint is formulated to rigorously align the reverse sampling paths of the measurement and image diffusion processes, ensuring strict adherence to the known physical model at every sampling step. The proposed PCS method is flexible and can seamlessly integrate various SDE-based diffusion priors. Extensive experiments on representative inverse problems—including undersampled MRI reconstruction, sparse-view CT reconstruction, and image super-resolution—demonstrate that PCS consistently outperforms existing state-of-the-art diffusion-based reconstruction methods. Although current evaluations focus specifically on MRI and CT modalities, the PCS framework holds considerable promise for broader applicability to other imaging modalities and inverse problems, which we plan to investigate in future studies.
Lihong Qiao, Rongxuan Wang, Yucheng Shu, Weisheng Li 0001, Zhanchuan Cai, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 C-GAN: Medical Image Steganography Based on Convergent GANs With Localization
abstract
Image steganography aims to hide secret message into cover image in an imperceptible and undetectable way, and only allows the informed receivers to decode stego image. Generative adversarial nets have been proved to be promising against other generative models, and some recently propose to use GANs to hide secret message to reach image steganography. However, it is still facing low embedding capacity, high detectability and poor convergence. To hand the pitfalls, we propose a novel medical image steganography method with convergent divergence measurement to achieve large capacity and undetectable hiding. Specifically, generator, extractor and discriminator are jointed into end-to-end framework where generator yields visually and detectably indistinguishable steganography image from which extractor recovers diagnose report while discriminator tries to distinguish the steganographic and original images. Then, we design Zero-centered Wasserstein distance to achieve controllable and stable training. It can be proved that the proposed method with the defined Zero-centered Wasserstein can converge to a local equilibrium with finite discriminator updates per generator updates. Besides, local regularization which can be also proved to be effective for accelerating convergence is imposed on generator to improve embedding capacity and achieve homeomorphism manifold mapping in low-dimension latent space. Extensive experiments on Open-I, LGK and COV-CTR medical dataset show that the proposed method outperforms recent state-of-the-art methods in capacity, detectability and convergence rate.
Liming Xu, Bochuan Zheng, Weisheng Li 0001
IEEE Trans. Dependable Secur. Comput.4
2026 LTOFusion: A Learning-to-Optimize Framework With Flow Matching for Unsupervised Image Fusion
abstract
Multimodal Image Fusion (MMIF) aims to synthesize complementary information from different modalities to generate comprehensive fused images, thereby facilitating downstream applications. Existing methods typically employ deep neural networks to directly construct high-dimensional image-to-image mappings, which is highly challenging, struggling to extract generalizable patterns for various fusion scenarios. Inspired by meta learning, we propose a learning-to-optimize fusion framework, named LTOFusion, which formulates image fusion as a trajectory optimization problem, decoupling the complicated fusion problem into multistage subproblems. Subsequently, a restricted state transition function based on flow matching is designed to compress the prediction space and lead the network to build an image-to-flow mapping and fine-tune the current fusion state. To facilitate model training, we collect intermediate fusion states and utilize a memory-replay strategy, further enhancing the sample diversity and model robustness. In addition, a hybrid loss with respect to intensity, gradient, structure, and local normalized cross-correlation is designed to improve image details and reduce potential artifacts for fusion results. Experimental results demonstrate that the proposed method achieves the state-of-the-art performance across multiple fusion tasks and downstream applications without requiring fine-tuning. The code is available at https://github.com/HeDan-11/LTOFusion.
Dan He 0010, Guofen Wang, Yucheng Shu, Weisheng Li 0001
IEEE Trans. Image Process.6
2026 Topology-Guided Semantic Face Center Estimation for Rotation-Invariant Face Detection
abstract
Face detection accuracy significantly decreases under rotational variations, including in-plane (RIP) and out-of-plane (ROP) rotations. ROP is particularly problematic due to its impact on landmark distortion, which leads to inaccurate face center localization. Meanwhile, many existing rotation-invariant models are primarily designed to handle RIP, they often fail under ROP because they lack the ability to capture semantic and topological relationships. Moreover, existing datasets frequently suffer from unreliable landmark annotations caused by imperfect ground truth labeling, the absence of precise center annotations, and imbalanced data across different rotation angles. To address these challenges, we propose a topology-guided semantic face center estimation method that leverages graph-based landmark relationships to preserve structural integrity under both RIP and ROP. Additionally, we construct a rotation-aware face dataset with accurate face center annotations and balanced rotational diversity to support training under extreme pose conditions. Next, we introduce a Hybrid-ViT model that fuses CNN spatial features with transformer-based global context and employ a center-guided module for robust landmark localization under extreme rotations. In order to evaluate center quality, we further design a hybrid metric that combines topological geometry with semantic perception for a more comprehensive evaluation of face center accuracy. Finally, experimental results demonstrate that our method outperforms state-of-the-art models in cross-dataset evaluations. Code: https://github.com/Catster111/TCE_RIFD.
Hathai Kaewkorn, Lifang Zhou, Weisheng Li 0001, Chengjiang Long
IEEE Trans. Image Process.3
2026 GarmentRec: Towards Individual Garment Reconstruction From a Monocular Human Image
abstract
Reconstructing high-quality garment models from monocular images is important as it provides a practical and effective solution for human digitization and virtual try-on etc. Recent implicit function-based garment reconstruction methods recover free-form geometry but struggle to reconstruct individual garment meshes from human images and tend to produce disembodied limbs or degenerate shapes for novel views. In contrast, explicit parametric garment template models can be utilised to construct separate meshes and constrain the shape reconstruction robustly. However, this limits the reconstruction of garment details and shape variations, such as the wrinkles and pockets etc. To address this problem, in this paper, we introduce a novel explicit garment template that is designed for both closed and open garment topology. Powered by our new garment template, we further propose a detailed garment reconstruction method based on a monocular view that can process both the closed and open types for shape recovery. To capture those challenging parts with unknown geometry and topology, we predict displacement maps on the parameterization domain for the target garment from the monocular image and elaborate it to the 3D garment surface via the UV coordinates, achieving realistic details on the 3D garment shape. Extensive experiments demonstrate the accuracy and robustness of our method and show that realistic details like garment wrinkles and pockets can be faithfully recovered in an explicit way. The code and dataset are available at https://github.com/worryDes/GarmentRec.
Zongyi Xu, Shiyang Cheng 0001, Wang Fei, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.6
2026 SAFusion: Scenario-Adaptive Network for Multimodal Medical Image Fusion
abstract
Multimodal medical image fusion aims to integrate complementary information from different modalities to support clinical diagnosis and treatment. Although deep learning has significantly advanced this field, existing methods often overlook the differences between various fusion scenarios, making a single network inadequate for diverse fusion requirements. Therefore, we propose a novel scenario-adaptive fusion network. The network employs a two-stage training process. In the first stage, an autoencoder is trained for multiscale feature extraction and image reconstruction. In the second stage, the autoencoder parameters are frozen, and a Fusion Layer is trained to achieve multimodal feature integration. The Fusion Layer consists of a Scenario-Specific Fusion Module and a Scenario-General Fusion Module. The former uses a mixture-of-experts model to customize fusion strategies for different scenarios to optimize the fusion process. The latter employs a dual-path fusion structure based on standard convolution and deformable convolution gating mechanisms to achieve general feature fusion across multi-scenario. Compared to eleven state-of-the-art methods, our method demonstrates superior information integration and visual consistency, offering a flexible and efficient solution for various fusion scenarios.
Weisheng Li 0001, Pengtao Jia, Dan He 0010, Guofen Wang
IEEE J. Biomed. Health Informatics1
2026 Multi-View Chest X-Ray Vision-Language Pre-Training via Semantic-Aware Masked Language Modeling and High-Order Alignment
abstract
Chest X-Ray Vision-Language pretraining (VLP) leverages large-scale radiograph-report pairs to develop joint image-text representations, demonstrating significant potential for medical image diagnosis. However, existing VLP approaches often overlook the multi-view nature of chest X-Rays, and some multi-view methods apply uniform feature fusion, neglecting view-key semantic contributions. Moreover, random cross-modal Masked Language Modeling (MLM) fails to facilitate effective interactions, impeding representation alignment. Additionally, global alignment in VLP may lead to the false-negative problem. To address these limitations, we propose a novel medical VLP framework comprising three core components. First, a Key Semantics-enhanced Multi-view MLM module aggregates pathology-relevant patches across views, providing semantically rich supervision for MLM. A local semantics enhancing approach, which identifies and aggregates pathology-relevant key patches across views to guide MLM. Second, a Frontal-Lateral Alignment module extracts view-specific pathological features, ensuring semantic consistency and preserving critical information during aggregation. This module independently extracts pathological features from both views to preserve view-specific information while ensuring semantic consistency, which mitigates the loss of crucial information during aggregation. Third, a High-order Semantic Alignment approach mitigates false-negative issues by aligning features with semantically consistent clusters, enhancing global alignment through prototype-level semantics. Extensive experiments across seven public datasets demonstrate that our framework outperforms state-of-the-art methods in four downstream tasks, validating its efficacy. The code is available at https://github.com/sajiutea/F-L.
Lihong Qiao, Jingya Gong, Yucheng Shu, Lifang Zhou, Baobin Li, Weisheng Li 0001, Bai Ying Lei
IEEE Trans. Medical Imaging7
2026 Mutually Guided Fusion Learning for Collaborative Camouflaged Object Segmentation
abstract
Collaborative camouflaged object segmentation (CoCOS) is a challenging task, focusing on identifying objects that blend closely with their backgrounds by jointly processing intraclass images. Existing methods fail to fully leverage the shared features (e.g., shape, texture, and contour) from these intraclass images, which leads to poor segmentation performance in relatively complex scenarios. To address this issue, we propose a novel mutually guided fusion refinement network (MFRNet), which improves the model performance by more effectively collaborating and optimizing the shared information. Specifically, it includes feature encoding, single-image branch feature enhancement, multiimage branch feature enhancement, and mutual guidance. After the feature encoding step, we design the graph convolution self-attention (GCS) and spatial context exploration (SCE) modules to enhance multilevel features of the single-image and multiimage branches, respectively. Moreover, we propose a mutual guidance fusion (MGF) module to utilize cross-scene image information for mutual guidance and progressive refinement, enhancing intraclass collaboration for improving target feature distinction. Extensive experimental results demonstrate that our MFRNet significantly outperforms existing CoCOS methods, achieving a mean E-measure score of 0.846 on the CoCOD8K dataset. Our code will be published at https://github.com/another-u/MFRNet.
Chen Li 0048, Xiao Luan, Linghui Liu, Yanzhao Su, Yule Fu, Weisheng Li 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
abstract
Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations.
Xiuli Bi, Bo Liu 0047, Xiaodong Cun, Yong Zhang 0034, Weisheng Li 0001, Bin Xiao 0002
AAAI6
2025 Beyond Probability Guided Categorization: A Subspace Projection-Aware Approach for Medical Image Segmentation
abstract
Most current medical image segmentation (MIS) methods focus on the performance of models while overlooking the interpretability of models. For existing segmentation models, the number of channels of the output feature maps is equal to the number of segmentation categories. The category of each pixel is determined by the channel with the highest probability, yet the reasons for this decision are unclear. Inspired by the principle that samples from the same category are spatially closer, we propose a novel subspace projection method to enhance both the performance and interpretability of MIS models. Our approach replaces the output layer of baseline models with a new convolutional layer that enhances channel-wise representational power, and concatenates the channel features of each pixel into feature vectors to integrate channel-level information. Considering the sparse structure of feature vectors, an overcomplete dictionary learning is applied to extract the most discriminative features. These feature vectors are projected into category subspaces, with pixel classification determined based on the smallest projection distance. To guarantee that feature vectors within the same category subspace are grouped closely and to promote nonoverlapping feature learning across subspaces, we introduce the distance loss and the subspace orthogonality loss. Our method is evaluated on ten popular MIS models using two publicly available datasets. Experimental results show that our approach outperforms baseline models in terms of both segmentation accuracy and the interpretability of model decision.
Xiao Luan, Yule Fu, Linghui Liu, Chen Li 0048, Weisheng Li 0001
BIBM5
2025 Towards Universal AI-Generated Image Detection by Variational Information Bottleneck Network
abstract
The rapid advancement of generative models has significantly improved the quality of generated images. Mean-while, it challenges information authenticity and credibility. Current generated image detection methods based on large-scale pre-trained multimodal models have achieved impressive results. Although these models provide abundant features, the authentication task-related features are often submerged. Consequently, those authentication task-irrelated features cause models to learn superficial biases, thereby harming their generalization performance across different model genera (e.g., GANs and Diffusion Models). To this end, we proposed VIB-Net, which uses Variational Information Bottlenecks to enforce authentication task-related feature learning. We tested and analyzed the proposed method and existing methods on samples generated by 17 different generative models. Compared to SOTA methods, VIB-Net achieved a 5.55% improvement in mAP and a 9.33% increase in accuracy. Notably, in generalization tests on unseen generative models from different series, VIB-Net improved mAP by 12.48% and accuracy by 23.59% over SOTA methods. The code is available at https://github.com/oceanzhf/VIBAIGCDetect.
Qinghui He, Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002
CVPR4
2025 Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling
abstract
The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images. To tackle these issues, a novel Bidirectional Reference Image Quality Assessment (BRIQA) method is proposed, centering on leveraging bidirectional reference images and content-quality correlation modeling. First, triplets of content-aligned low-quality and content-non-aligned high-quality reference images are generated using two easily accessible approaches. To prevent the extraction of redundant information, two feature extractors pretrained through unsupervised contrastive learning are utilized to independently extract content and quality features for the triplet images. Then, an attention-mixer is introduced to further mine quality difference information and enhance content feature. Finally, a content-quality correlation modeler is proposed to model the relationship between quality differences and visual contents. Experimental results on benchmark datasets demonstrate that the BRIQA outperforms existing state-of-the-art methods.
Bo Hu 0008, Wenzhi Chen, Chunyi Li 0001, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001
ICASSP5
2025 Self-Geometry-Guided Direct Pose Regression Based on Dual Perspective Fusion for 2D-3D Cross Dimensional Spinal Surgery Navigation
abstract
2D-3D cross-dimensional registration for spinal surgery navigation, which aims to achieve real-time visual navigation of preoperative 3D vertebrae based on intraoperative 2D fluoroscopy images, faces significant challenges due to semantic and dimensional gaps. Traditional 2D-3D registration methods often require fine adjustment steps and have low computational efficiency. In this paper, we propose a self-geometry-guided direct regression method based on dual perspective images. Firstly, an effective mechanism for unifying the dual view coordinate system was proposed. Secondly, a novel feature extraction module based on a face-graph convolutional network (F-GCN) is proposed to effectively extract 3D vertebra posture features. Finally, a posture direct regression network guided by self-vertebral geometry based on 2D-3D fusion features was constructed. Experimental results show that our method has made significant progress in solving the problem of 2D-3D cross-dimensional registration for spinal surgery navigation.
Jing Ling, Zhengyang Wu 0002, Weisheng Li 0001, Chao Zhang 0106, Yucheng Shu
ICASSP4
2025 Stacking U-Nets in U-shape: Redesigning the Information Flow in Model-based Networks for MRI Reconstruction
abstract
Model-based networks have shown convincing performance in MRI reconstruction. However, the unrolled cascades within the networks are constrained to solely obtain information from the preceding counterpart, resulting in potential error accumulation. Moreover, the linear structure fails to address the challenge of recovering fine-grained details. To tackle these problems, we propose to redesign the information flow in model-based networks. Our method features a large U-shaped network, where the nodes are built with unrolled cascades and U-Net-based regularizers. We design an input-level integration module to help the cascades acquire information from adjacent and skip-connected counterparts, building robust mappings to the target. We further design a coarse-to-fine feature-level integration module, aiming at guiding the network to progressively recover fine details. Intermediate reconstructions produced by subnetworks of different scales are integrated, enabling the extraction of complementary information to enhance the final performance. Compared with cutting-edge methods on different datasets, our method exhibits superior performances.
Xiaoyu Qiao, Weisheng Li 0001, Bin Xiao 0002
ICASSP2
2025 Subsampling Decomposition based k-Space Refinement for Accelerated MRI Reconstruction
abstract
In accelerated MRI reconstruction problem, directly recovering all the missing k-space data from undersampled measurements is highly ill-posed and often leads to suboptimal performance. To address the problem, we propose a novel deep unfolding network (DUN) with subsampling decomposition (SD) based k-space refinement to mitigate the ill-posedness. Our method employs a parallel network architecture with a primary branch unfolded by gradient descent-inspired optimization process (GD-PB) for reconstruction. Additionally, we introduce an SD-based auxiliary branch (SD-AB) that decompose the inverse problem into moderately corrupted subproblems. We design a novel subsampling mask predictor that captures both global and local spatial correlations in k-space, ensuring the SD-AB effectively preserves the most well-reconstructed subsets as reliable region (RR). The RR in SD-AB is used to periodically refine the intermediate outputs of the GD-PB, achieving improved accuracy. Experimental results reveal that our method significantly outperforms conventional and SD-based DUN techniques, achieving superior PSNR and SSIM results compared with cutting-edge methods.
Xiaoyu Qiao, Weisheng Li 0001, Bin Xiao 0002
ICASSP2
2025 Learning Preconditioners in Gates-controlled Deep Unfolding Networks based on Quasi-Newton Methods For Accelerated MRI Reconstruction
abstract
Deep unfolding networks (DUNs) have made significant progress in MRI reconstruction, successfully tackling the problem of prolonged imaging time. However, the ill-conditioned nature of MRI reconstruction often causes slow convergence in iterative optimization, potentially compromising the performance of DUNs. In this study we propose a preconditioned and gates-controlled DUN (PGDUN) to address these challenges. Our approach starts with optimizing the step size of proximal gradient descent (PGD) through a preconditioner. To improve flexibility and adaptability, we relax the constrains on quasi-newton-based optimization procedure. We design ConvLSTM-based modules, where the gate units automatically preserve necessary long- and short-term information, facilitating the learning of optimized variables and their combinations. Furthermore, we design gate units to modulate the features fed to regularizers across different iterations, boosting their robustness against potential accumulated errors. Evaluations using PSNR and SSIM metrics reveal that our approach outperforms existing state-of-the-art methods, achieving superior reconstruction results across various sequences.
Xiaoyu Qiao, Weisheng Li 0001, Bin Xiao 0002
ICASSP2
2025 ALCReg: Active Label Correction for Partial Point Cloud Registration
abstract
Deep point cloud registration methods encounter challenges due to partial overlaps and are heavily reliant on labeled data. In this paper, we propose ALCReg, an active label correction method for partial point cloud registration learning. ALCReg utilises a multimodal approach to generate pseudo labels, mitigating the cold-start issue in active learning. To ensure the diversity and representativeness of selected samples, we propose an inlier ratio based query strategy for manual correction. Furthermore, an innovative self-correction mechanism based on consistency is introduced, allowing the model to refine pseudo labels autonomously and further improve model performance. Experimental results on the 3DMatch and 3DLoMatch datasets demonstrate that ALCReg achieves comparable performance with the fully-supervised registration methods, even with only 5% of labeled samples, making it the first active learning method tailored for partial point cloud registration. Code is available at https://github.com/Jiang0903/ALCReg.
Zongyi Xu, Xinqi Jiang, Shanshan Zhao 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
ICME6
2025 DGMIR: Dual-Guided Multimodal Medical Image Registration Based on Multi-view Augmentation and On-Site Modality Removal
Gao Le, Yucheng Shu, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
MICCAI (1)6
2025 Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-training
abstract
Current medical vision-language pre-training models primarily follow two paradigms: report-supervised cross-modal alignment pre-training and reconstruction-based self-supervised pre-training. The former enhances the discriminative power of representations, while the latter facilitates fine-grained representation learning. However, naively combining these two paradigms inherits their inherent limitations: reconstruction-based methods treat all image patches equally during reconstruction, failing to effectively capture critical pathological details-since disease-related regions typically occupy only a small fraction of the image. Meanwhile, alignment-based methods suffer from suboptimal representations due to the presence of false negatives. To address these challenges, we propose a novel pre-training framework that integrates two key components: Pathology-Aware Reconstruction (PAR) and Discriminative Knowledge-Boosted Alignment (DKBA). Through a cascaded training strategy, our framework effectively combines the strengths of both paradigms while mitigating their inherent limitations. During the reconstruction pre-training stage, PAR incorporates pathology-aware priors to enhance the model's ability to capture fine-grained pathological details. In the alignment pre-training stage, DKBA leverages a medical knowledge graph as external supervision to improve cross-modal clustering alignment, thereby reducing the negative impact of false negatives. Extensive experiments on diverse downstream medical imaging tasks including image classification, object detection, and semantic segmentation, demonstrate the superior generalization capabilities of our method. Our code is publicly available at https://github.com/Felix1118/PADKB.
Lihong Qiao, Shiyi Gao, Yucheng Shu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
ACM Multimedia5
2025 The Overlooked Matters: Revisiting Background, Prototype, and Activation in Few-Shot Medical Image Segmentation
Yucheng Shu, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
ACM Multimedia6
2025 A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
abstract
While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditions, leading to significant performance drops in drone-view scenarios.To bridge this gap, we introduce A2Seek (Aerial Anomaly Seek), a large-scale, reasoning-centric benchmark dataset for aerial anomaly understanding. This dataset covers various scenarios and environmental conditions, providing high-resolution real-world aerial videos with detailed annotations, including anomaly categories, frame-level timestamps, region-level bounding boxes, and natural language explanations for causal reasoning. Building on this dataset, we propose A2Seek-R1, a novel reasoning framework that generalizes R1-style strategies to aerial anomaly understanding, enabling a deeper understanding of “Where” anomalies occur and “Why” they happen in aerial frames.To this end, A2Seek-R1 first employs a graph-of-thought (GoT)-guided supervised fine-tuning approach to activate the model's latent reasoning capabilities on A2Seek. Then, we introduce Aerial Group Relative Policy Optimization (A-GRPO) to design rule-based reward functions tailored to aerial scenarios. Furthermore, we propose a novel “seeking” mechanism that simulates UAV flight behavior by directing the model's attention to informative regions.Extensive experiments demonstrate that A2Seek-R1 achieves up to a 22.04\% improvement in AP for prediction accuracy and a 13.9\% gain in mIoU for anomaly localization, exhibiting strong generalization across complex environments and out-of-distribution scenarios. Our dataset and code are released at https://2-mo.github.io/A2Seek/.
Mengjingcheng Mo, Xinyang Tong, Mingpi Tan, Jiaxu Leng, Jiankang Zheng, Haosheng Chen 0001, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001
NeurIPS9
2025 A spatiotemporal fusion method based on two-stream high temporal sensitive convolutional neural network
Dajiang Lei, Qianwei Zhu, Weisheng Li 0001
Appl. Intell.6
2025 FMCA-Net: A feature secondary multiplexing and dilated convolutional attention polyp segmentation network based on pyramid vision transformer
Weisheng Li 0001, Xiaolong Nie, Zhaopeng Huang, Guofeng Zeng
Expert Syst. Appl.1
2025 Rethinking the CNN and transformer for deformable image registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Guofen Wang, Bin Xiao 0002
Expert Syst. Appl.2
2025 Transfer morphological features for segmentation with few labels on fluorescent mitochondria images
abstract
Abstract Automated segmentation of mitochondria is crucial for statistical analysis in biological research. Existing segmentation techniques often face challenges with fluorescence images. Handcrafted methods have poor segmentation results while deep learning‐based methods lack the labeled mitochondrial data. However, although the number of labeled mitochondrial images is limited, the unlabeled fluorescent data is easy to obtain. The authors aim to leverage a large amount of unlabeled data to learn mitochondrial morphological features. The approach begins with self‐supervised learning from a vast set of unlabeled images through masked image modeling. This technique involves presenting images with randomly masked patches, prompting the model to predict the content of these masked areas. By doing so, the model learns the distinctive features of mitochondria. In the subsequent phase, the trained encoder is transferred to the segmentation task, replacing the original reconstruction decoder with the Segformer segmentation decoder. The model is then fine‐tuned using a small labeled dataset. By reconstructing mitochondria in the masked regions, the model learns features more effectively on unlabeled samples, and improves segmentation performance even with limited labeled data. Empirical results validate the effectiveness of the approach, showing an 11.8% improvement in Intersection over Union metrics compared to existing fluorescence mitochondrial segmentation techniques.
Junchao Fan, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002, Xiaoshuai Huang
IET Image Process.4
2025 CS-CoLBP: Cross-Scale Co-occurrence Local Binary Pattern for Image Classification
Bin Xiao 0002, Danyu Shi, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
Int. J. Comput. Vis.4
2025 Cardiac cavity segmentation review in the past decade: Methods and future perspectives
Weisheng Li 0001, Yucheng Shu, Yidong Peng, Bin Xiao 0002
Neurocomputing2
2025 MCU-Net: A multi-prior collaborative deep unfolding network with gates-controlled spatial attention for accelerated MRI reconstruction
Xiaoyu Qiao, Weisheng Li 0001, Guofen Wang
Neurocomputing2
2025 A multi-granularity facial aesthetic evaluation model based on image-text modality
Yong Wang 0009, Weisheng Li 0001, Bin Xiao 0002
Knowl. Based Syst.3
2025 YOLOCS: Object detection based on dense channel compression for feature spatial solidification
Weisheng Li 0001, Yujuan Tan, LinLin Shen, Jing Yu 0026, Haojie Fu
Knowl. Based Syst.2
2025 Dual-modality visual feature flow for medical report generation
abstract
Medical report generation, a cross-modal task of generating medical text information, aiming to provide professional descriptions of medical images in clinical language. Despite some methods have made progress, there are still some limitations, including insufficient focus on lesion areas, omission of internal edge features, and difficulty in aligning cross-modal data. To address these issues, we propose Dual-Modality Visual Feature Flow (DMVF) for medical report generation. Firstly, we introduce region-level features based on grid-level features to enhance the method's ability to identify lesions and key areas. Then, we enhance two types of feature flows based on their attributes to prevent the loss of key information, respectively. Finally, we align visual mappings from different visual feature with report textual embeddings through a feature fusion module to perform cross-modal learning. Extensive experiments conducted on four benchmark datasets demonstrate that our approach outperforms the state-of-the-art methods in both natural language generation and clinical efficacy metrics.
Quan Tang 0006, Liming Xu, Yongheng Wang, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001
Medical Image Anal.7
2025 Knowledge-driven multi-graph convolutional network for brain network analysis and potential biomarker discovery
Jianhua Gong, Weisheng Li 0001, Zhuoya Yang
Medical Image Anal.3
2025 3D point cloud semantic segmentation based on visual guidance and feature enhancement
Yucheng Shu, Lihong Qiao, Zhengyang Wu 0002, Jing Ling, Jiang Wu 0006, Weisheng Li 0001
Multim. Syst.7
2025 RP-Net: A Robust Polar Transformation Network for rotation-invariant face detection
Hathai Kaewkorn, Lifang Zhou, Weisheng Li 0001
Pattern Recognit.3
2025 S2Reg: Structure-semantics collaborative point cloud registration
Zongyi Xu, Xinqi Jiang, Shiyang Cheng 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
Pattern Recognit.6
2025 Improving the sparse coding model via hybrid Gaussian priors
Jian-Xun Mi, Weisheng Li 0001, Guofen Wang, Bin Xiao 0002
Pattern Recognit.3
2025 Weak-prior dual cognitive attention lightweight network for abdominal multi-organ segmentation
Shenhai Zheng, Jianfei Li, Haiguo Zhao, Weisheng Li 0001
Pattern Recognit.4
2025 CSSNet: A 3D medical image segmentation network based on compressed sparse dual-branch structure
Xiao Luan, Yule Fu, Linghui Liu, Weisheng Li 0001
Pattern Recognit. Lett.4
2025 SR-LBSCC: Super resolution based screen content image compression at low bitrate
Shenhai Zheng, Weisheng Li 0001
Pattern Recognit. Lett.5
2025 CF3d: Category fused 3D point cloud retrieval
Zongyi Xu, Ruicheng Zhang, Zuo Li, Shiyang Cheng 0001, Huiyu Zhou 0001, Weisheng Li 0001, Xinbo Gao 0001
Signal Process.6
2025 Reversible Feature Learning for Brain Tumor Segmentation With Incomplete Modalities
abstract
Accurate brain tumor segmentation is vital for clinical diagnosis and treatment. Due to motion artifacts and image damage, it is challenging to obtain accurate segmentation results of brain tumors in the presence of incomplete MRI modalities. We propose a multimodal reversible feature learning method to tackle this problem. This method can fully explore the potential feature similarities and complementarities between MRI modalities. To compensate the information of missing modalities, we propose a reversible feature interaction module. It explores information similarity among existing modalities as priors, which are used to reconstruct features of missing modalities at the feature level. With enhanced discriminative information, the model suppresses the noise in the missing modalities. Furthermore, we propose a dual-scale attention module to enhance the detail restoration and reconstruction accuracy of images. Comparison results on the BRATS challenge datasets show the superiority of our method over current popular methods. The code used in this study is available athttps://github.com/fybgogogo/reverse.
Yanbing Fan, Linghui Liu, Xiao Luan, Weisheng Li 0001
IEEE Signal Process. Lett.4
2025 Dual-Branch Network for No-Reference Super-Resolution Image Quality Assessment
abstract
No-reference super-resolution image quality assessment (SR-IQA) has become an critical technique for optimizing SR algorithms, the key challenge is how to comprehensively learn visual related features of SR image. Existing methods ignore the context information and feature correlation. To tackle this problem, this letter proposes a dual-branch network for no-reference super-resolution image quality assessment (DBSRNet). First, dual-branch feature extraction module is designed, where residual network and receptive field block net are combined to learn multi-scale local features, stacked vision transformer blocks are utilized to learn global features. Then, correlations between dual-branch features are learned and fused based on self-attention mechanism structure, final predicted score is obtained by adaptive feature pooling strategy. Finally, experimental results show that DBSRNet significantly outperforms State-of-the-Art methods in terms of prediction accuracy on all SR-IQA datasets.
Fan Yang 0159, Weisheng Li 0001
IEEE Signal Process. Lett.4
2025 Feature Transductive Distribution Optimization for Few-Shot Image Classification
abstract
Few-shot learning (FSL) requires vision models to quickly adapt to brand-new classification tasks with changing task distributions in the presence of limited annotated samples. However, the learned model is susceptible to overfitting and may fail to identify effective classification boundaries due to the biased distribution resulting from a limited number of training samples. Moreover, if the support samples from different classes in the new task are in close proximity, this may lead to fuzzy or even biased class decision boundaries. To address the issues, we propose a generation-based Feature Transductive Distribution Optimization (FTDO) in our research. Specifically, we calibrate the distribution of novel classes by utilizing high-confidence unlabeled query samples from these novel classes, together with the statistics of similar base classes, to generate a sufficient number of virtual training samples. In addition, we introduce a task commonality removal and discriminability enhancement module, which eliminates commonality from all features in the task along the task-commonality direction, and reinforces the retained discriminative features through a channel transformation function. Our method can be implemented using off-the-shelf pre-trained feature extractors and classification models, without requiring additional parameters. Experiments conducted on four few-shot classification datasets substantiate the superiority of our proposed method.
Xinyan Jiang, Weisheng Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Fast Sampling of Diffusion Models for Accelerated MRI Using Dual Manifold Constraints
abstract
Diffusion models show great potential in solving inverse problems, including MRI reconstruction. With its unique characteristics, medical imaging demands both efficiency and accuracy in the reconstruction process. However, existing MRI reconstruction methods based on diffusion models often fall short of fully leveraging the available measurements during sampling. Consequently, these methods suffer from compromised reconstruction quality and elevated bias, especially when dealing with large acceleration factors. In response to these challenges, we propose Dual Manifold Constraints (DMC), a fast MRI reconstruction method based on diffusion models. We treat the sampling process as a combination of denoising and adding noise processes, and we constrain these two processes using both pristine measurements and their noisy counterparts to adapt to the geometry of diffusion. It’s worth noting that we propose a method to estimate the noisy measurement that satisfies the sub-sampling process to maintain the current data manifold when performing data consistency constraints. Experimental results show that our method outperforms the latest diffusion-based methods regarding both reconstruction speed and accuracy, and exhibits strong out-of-distribution generalization performance.
Lihong Qiao, Rongxuan Wang, Yucheng Shu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001, Zhanchuan Cai
IEEE Trans. Circuits Syst. Video Technol.5
2025 ABIBE: Adaptive Building Information-Based Extraction From Remote Sensing Imagery Using Vision-Language Models
abstract
Building extraction from remote sensing imagery is essential for urban planning, population monitoring, and emergency response. However, traditional methods often struggle with accurate building delineation due to complex architectural features and varying imaging conditions. In this paper, we propose ABIBE, an Adaptive Building Information-Based Extraction framework that leverages Vision-Language Models (VLMs) for high-precision building extraction from remote sensing imagery. Our approach introduces two key innovations: (1) a hierarchical feature transfer mechanism that selectively extracts and adapts visual representations from the JanusPro vision-language model, and (2) a dynamic multi-scale feature fusion technique that adaptively integrates these features with spatial information through cross-modal attention mechanisms. Extensive experiments on the WHU Building Dataset demonstrate that our method significantly outperforms state-of-the-art approaches, achieving 91.39% IoU and 95.50% F1 score. We also provide comprehensive performance analysis including computational complexity, runtime efficiency, and resource requirements to evaluate the practical applicability of the proposed framework. The ABIBE framework provides a promising solution for high-precision building extraction from remote sensing imagery, with particular advantages in areas with complex architectural structures and high building density.
Yongtao Deng, Dajiang Lei, Yidong Peng, Weisheng Li 0001, Liping Zhang 0012
IEEE Trans. Geosci. Remote. Sens.4
2025 GAPL-SegNet: Geometry-Aware Prototype Learning for Few-Shot Building Segmentation in Remote Sensing Imagery
abstract
Few-shot building segmentation in remote sensing imagery remains challenging due to limited annotated data and complex geometric structures inherent in building footprints. Traditional prototype-based methods ignore rich geometric priors of buildings, leading to suboptimal feature representations and poor generalization across different geographic regions. We propose GAPL-SegNet, a novel architecture that integrates Geometry-Aware Prototype Learning with adaptive feature aggregation for superior few-shot performance. Our approach introduces several key innovations: (1) a geometric feature extractor that captures building-specific structural patterns including edges, corners, and rectangularity through dedicated detection branches; (2) a geometry-aware prototype learning framework that leverages spatial importance weighting for more discriminative prototype construction based on geometric significance; (3) a progressive adaptive training strategy with dynamic geometric loss weighting that ensures effective integration of geometric priors; and (4) systematic cross-dataset analysis quantifying domain adaptation challenges. Extensive experiments demonstrate strong performance across multiple datasets and scenarios. On the WHU Building dataset with 100 training samples, GAPL-SegNet achieves 85.75±0.43% IoU (95% CI: [85.22, 86.29]) based on five independent runs, with notable stability (CV < 0.5%), outperforming the best baseline by 3.67% IoU. Cross-dataset evaluations reveal significant challenges from resolution differences: WHU to Inria achieves 51.68% IoU (41.4% with resolution alignment) and Inria to WHU reaches 53.71% IoU (33.6% with resolution alignment). Multi-resolution analysis demonstrates that spatial resolution differences (0.3m to 1.2m) cause up to 46% IoU degradation, identifying resolution adaptation as the primary domain shift challenge that exceeds the scope of geometric priors alone. Computational efficiency analysis reveals that our method achieves a favorable performance-efficiency balance with 22.5M parameters and 5.3ms inference time. In building segmentation’s specific few-shot setting, geometric priors bring better boundary consistency and deployment efficiency compared to general-purpose foundation models.
Yongtao Deng, Dajiang Lei, Liping Zhang 0012, Yidong Peng, Weisheng Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Contrastive Learning Guided Fusion Network for Brain CT and MRI
abstract
Medical image fusion technology provides professionals with more detailed and precise diagnostic information. This paper introduces a new efficient CT and MRI fusion network, CLGFusion, based on a contrastive learning-guided network. CLGFusion includes two encoding branches at the feature encoding stage, enabling them to interact and learn from each other. The approach begins with training a single-view encoder to predict the feature representation of an image from varied augmented views. Simultaneously, the multi-view encoder is improved using the exponential moving average of the single-view encoder. Contrastive learning is integrated into medical image fusion by creating a feature contrast space without constructing negative samples. This feature contrast space cleverly uses the information of the difference in the feature product of the source image and its corresponding augmented image. It continuously guides the network to constantly optimize its fusion effect by combining the method of structural similarity loss, to achieve more accurate and efficient image fusion. This approach represents an end-to-end unsupervised fusion model. Experimental validation shows that our proposed method demonstrates performance comparable to state-of-the-art techniques in both subjective evaluation and objective metrics.
Weisheng Li 0001, Bin Xiao 0002, Guofen Wang, Dan He 0010, Xiaoyu Qiao
IEEE J. Biomed. Health Informatics2
2025 Whole Heart Segmentation Based on 3D Contour-Guided Multi-Head Attention Network From CT and MRI Images
abstract
Heart image segmentation is a critical task in medical image processing, which is crucial for the diagnosis and treatment planning of cardiovascular diseases. It helps doctors understand patients' cardiac anatomy and functional status more comprehensively and lays the foundation for personalized medicine and precision medicine research. Addressing the current challenges of rough surfaces on the entire heart, incomplete segmentation of heart substructures, and the lack of structured prediction of pulmonary arteries due to artifacts, scale diversity, uneven intensity, and boundary ambiguity in cardiac computed tomography (CT) and magnetic resonance imaging (MRI) images, we propose a whole heart segmentation algorithm based on 3D contour guided network. The proposed algorithm achieves robust whole heart segmentation results and has few network structure parameters. To enhance the consistency of features extracted by the codec, we propose a 3D codec information integration module to focus on task-related areas. In the final stage of information integration, features of different scales are combined. A 3D contour attention module enhances the perception of the heart's structure and shape. Contour prediction results from the initial stage, generating a low-resolution voxel of the entire heart with contour details. The second stage builds upon the initial phase of secondary learning to achieve multi-label segmentation results. The proposed algorithm achieved average Dice scores of 0.905 and 0.865 for the CT and MRI modalities, respectively, in 40 cases.
Weisheng Li 0001, Yidong Peng, Yucheng Shu
IEEE J. Biomed. Health Informatics2
2025 Asymmetric Adaptive Heterogeneous Network for Multi-Modality Medical Image Segmentation
abstract
Existing studies of multi-modality medical image segmentation tend to aggregate all modalities without discrimination and employ multiple symmetric encoders or decoders for feature extraction and fusion. They often overlook the different contributions to visual representation and intelligent decisions among multi-modality images. Motivated by this discovery, this paper proposes an asymmetric adaptive heterogeneous network for multi-modality image feature extraction with modality discrimination and adaptive fusion. For feature extraction, it uses a heterogeneous two-stream asymmetric feature-bridging network to extract complementary features from auxiliary multi-modality and leading single-modality images, respectively. For feature adaptive fusion, the proposed Transformer-CNN Feature Alignment and Fusion (T-CFAF) module enhances the leading single-modality information, and the Cross-Modality Heterogeneous Graph Fusion (CMHGF) module further fuses multi-modality features at a high-level semantic layer adaptively. Comparative evaluation with ten segmentation models on six datasets demonstrates significant efficiency gains as well as highly competitive segmentation accuracy. (Our code is publicly available at https://github.com/joker-527/AAHN).
Shenhai Zheng, Chaohui Yang, Weisheng Li 0001, Xinbo Gao 0001, Yue Zhao 0012
IEEE Trans. Medical Imaging5
2025 DM-FNet: Unified Multimodal Medical Image Fusion via Diffusion Process-Trained Encoder-Decoder
abstract
Multimodal medical image fusion (MMIF) extracts the most meaningful information from multiple source images, enabling a more comprehensive and accurate diagnosis. Achieving high-quality fusion results requires a careful balance of brightness, color, contrast, and detail; this ensures that the fused images effectively display relevant anatomical structures and reflect the functional status of the tissues. However, existing MMIF methods have limited capacity to capture detailed features during conventional training and suffer from insufficient cross-modal feature interaction, leading to suboptimal fused image quality. To address these issues, this study proposes a two-stage diffusion model-based fusion network (DM-FNet) to achieve unified MMIF. In Stage I, a diffusion process trains UNet for image reconstruction. UNet captures detailed information through progressive denoising and represents multilevel data, providing a rich set of feature representations for the subsequent fusion network. In Stage II, noisy images at various steps are input into the fusion network to enhance the model's feature recognition capability. Three key fusion modules are also integrated to process medical images from different modalities adaptively. Ultimately, the robust network structure and a hybrid loss function are integrated to harmonize the fused image's brightness, color, contrast, and detail, enhancing its quality and information density. The experimental results across various medical image types demonstrate that the proposed method performs exceptionally well regarding objective evaluation metrics. The fused image preserves appropriate brightness, a comprehensive distribution of radioactive tracers, rich textures, and clear edges. The code is available athttps://github.com/HeDan-11/DM-FNet.
Dan He 0010, Weisheng Li 0001, Guofen Wang
IEEE Trans. Multim.2
2025 MDFA: A Quantitative Framework for the Analysis of Multimodal Facial Esthetics
abstract
In the era of big data, the problem of facial beauty prediction (FBP) has been addressed using a combination of deep learning and esthetics based on data and models. Most existing methods are based on 2-D unimodal information processing. Owing to the high cost of 3-D data acquisition equipment, studies on the use of multimodal features of 2-D and 3-D for esthetic evaluation are scarce. Moreover, most existing methods are based on self-built 3-D datasets, which are limited to practical application scenarios of 2-D facial images. This study proposed a label distribution-based multimodal facial esthetic analysis framework (LDMFE). The LDMFE performed facial esthetic evaluation by combining 2-D and 3-D information following the process used by the human brain to conduct the 3-D esthetic evaluation. FBP was performed by extracting facial depth structure information using a depth information extraction network, DIENet, which comprises a facial structure perception layer (FSP-Layer) and an attention decision block (AD-Block). Furthermore, to ensure a high degree of agreement between the predicted label distribution of the network and the true distribution, a simple and efficient distribution measurement loss function called ${\mathcal {L}}_{\text {WD}}$ was proposed. Compared with the label distribution-based FBP loss and the latest FBP loss, ${\mathcal {L}}_{\text {WD}}$ was more stable and effective. The performance of LDMFE was evaluated using three datasets. The experimental results demonstrate that the LDMFE exhibits state-of-the-art performance.
Weisheng Li 0001, Bin Xiao 0002, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Multi-scale Consistency Deep Lifelong Cross-modal Hashing
abstract
Deep cross-modal hashing methods provide effective and efficient solutions for large-scale cross-modal retrieval. However, existing cross-modal hashing methods fail to capture the dynamic changes of real-world data, and suffer from serious performance degradation when retrieving streaming data. In this paper, we propose a novel hashing method to achieve accurate cross-modal retrieval under continuous and streaming scenarios. Specifically, regularization-based lifelong learning module is introduced to balance plasticity for learning new knowledge and stability for maintaining old knowledge, and update incremental hash codes without retraining cumulative data. Then, multi-scale consistency network which employs multi-scale feature fusion module to extract fine-grained features among multi-scale modalities is introduced to learn multi-level semantic representations with consistency. Additionally, modality alignment with variational information bottleneck is designed to remove irrelevant information and obtain unified representation, which can be proved to be effective to yield high-quality hash code with new and old knowledge. Extensive experiments show that ours gains the advanced performance and the better adaptability to continuous and streaming environments.
Liming Xu, Jie Shao 0001, Weisheng Li 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Focus Stacking with High Fidelity and Superior Visual Effects
abstract
Focus stacking is a technique in computational photography, and it synthesizes a single all-in-focus image from different focal plane images. It is difficult for previous works to produce a high-quality all-in-focus image that meets two goals: high-fidelity to its source images and good visual effects without defects or abnormalities. This paper proposes a novel method based on optical imaging process analysis and modeling. Based on a foreground segmentation - diffusion elimination architecture, the foreground segmentation makes most of the areas in full-focus images heritage information from the source images to achieve high fidelity; diffusion elimination models the physical imaging process and is specially used to solve the transition region (TR) problem that is a long-term neglected issue and degrades visual effects of synthesized images. Based on extensive experiments on simulated dataset, existing realistic dataset and our proposed BetaFusion dataset, the results show that our proposed method can generate high-quality all-in-focus images by achieving two goals simultaneously, especially can successfully solve the TR problem and eliminate the visual effect degradation of synthesized images caused by the TR problem.
Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002
AAAI4
2024 Using My Artistic Style? You Must Obtain My Authorization
Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002
ECCV (86)3
2024 Facial Aesthetic Enhancement Network for Asian Faces Based on Differential Facial Aesthetic Activations
abstract
In this paper, we addressed facial aesthetic enhancement (FAE). Although existing methods have made great progress, the beautified images generated by them are highly prone to poor beautification, which limits their application to real-world scenes. To tackle this problem, we proposed a new method called the facial aesthetic enhancement network for Asian faces based on differential facial aesthetic activations (Diff-FANet), which comprises three important modules: aesthetic average difference perception block (ADP), aesthetic difference evaluation block (ADE), and aesthetic fusion optimization block (AFO). ADP learns the transformation of the latent code of an image before and after beautification. The ADE learns the features of an enhanced image, which guides image fusion. The AFO was used to eliminate ghosting. To evaluate the effectiveness of Diff-FANet, we utilized the wedding dataset for training and the SCUT-FBP5500 and Asian face datasets for testing. The results of experiments revealed that Diff-FANet achieved excellent results.
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002
ICASSP2
2024 CT and MRI Fusion with Anisotropic Guided Filtering
abstract
The combination of CT and MRI can provide more accurate images of lesions, yielding a significantly higher diagnostic value compared to single-modality pathological images. However, in CT-MRI fusion, preserving the gray-scale distribution of the source image while avoiding ‘detail halos’ poses a challenge. Therefore, we propose the utilization of anisotropic guided filtering (AnisGF), which exhibits excellent edge-preservation properties, to address structural inconsistencies in regions between the two modalities. The local neighborhood variance is utilized for optimizing the weight to achieve maximum diffusion, and subsequently decomposing the source image based on this criterion. A pre-trained convolutional neural network (CNN) is employed to accomplish the mapping from the source image to the weight map, while AnisGF is utilized for maintaining local consistency between them. The efficacy of this novel image fusion algorithm in preserving intricate details without compromising has been demonstrated through a combination of qualitative and quantitative experiments.
Weisheng Li 0001, Guofen Wang, Xiaoyu Qiao
ICASSP2
2024 Fast Intra Mode Prediction Algorithms for SCBS in VVC SCC
abstract
Versatile Video Coding (VVC) now supports Screen Content Coding (SCC) by integrating two efficient coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the numerous modes and the Quad-Tree Plus Multi-Type Tree (QTMT) structure inherent to VVC contribute to a very high coding complexity. To effectively reduce the computational complexity of VVC SCC, we propose a fast Intra mode prediction algorithm for VVC SCC. More specifically, we first use the difference of minimum Sum of Absolute Transformed Differences (SATD) value of four Directional Modes (DMs) of Intra and the SATD value of the IBC-merge mode to determine whether to early skip Intra checking. Subsequently, we use a decision tree to determine whether to early terminate the checking after block differential pulse coded modulation (BDPCM). Finally, we employ a decision tree to determine whether to early skip multiple transform selection (MTS) and low frequency non-separable transform (LFNST) checking. The results demonstrate that our algorithm achieves an average encoding time reduction of 34.34% with a negligible Bjøntegaard delta bitrate increase of 0.46%.
Yishen Deng, Weisheng Li 0001, Xin Lu 0001, Frédéric Dufaux, Bo Hang, Ce Zhu
ICASSP3
2024 Window-Based Convolutional Sparse Coding: Towards A Unified Framework
abstract
Sparse Coding (SC) and Convolution Sparse Coding (CSC) are two widely studied sparse methods in computer vision and signal processing. SC encodes the image patches independently, however fails to utilize the correlation among them. CSC adopts a convolution operator to connect the overlapping patches but in an inflexible manner. In this paper, a novel integrated framework for the two sparse models is proposed, wherein the local correlations among patches are controllable by manipulating a window function. Moreover, the inherent border effect of a convolution model is mitigated with a carefully designed weight function. It can be demonstrated that both SC and CSC are two distinct implementations of this framework. Consequently, our unified framework provides a balanced solution by addressing the strengths and limitations of both SC and CSC. Extensive experimental results are presented to demonstrate the superiority and effectiveness of the proposed method for image inpainting tasks.
Jian-Xun Mi, Guofen Wang, Weisheng Li 0001
ICASSP4
2024 Re3adapter: Efficient Parameter Fing-Tuning with Triple Reparameterization for Adapter without Inference Latency
abstract
With the rise of large-scale model applications, leveraging these models as the base network for efficient transfer learning has garnered increasing attention. Currently, parameter-efficient transfer learning methods have made significant improvements in reducing the number of trainable parameters but introduce latency during inference. In this study, we propose an enhanced adaptation of the adapter using a reparameterization technique, revamping the activating layers into linear layers. This modification retains the high-dimensional fine-tuning capability of the adapter for visual tasks while avoiding additional inference latency. We name this plug-and-play module the Re3adapter, which optimizes the model with only 0.26% of the parameters and introduces no inference latency. Experimental results demonstrate its clear advantages in traditional classification and medical tasks.
Lihong Qiao, Rui Wang 0173, Yucheng Shu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001
ICME6
2024 Focal-Guided Multi-Consistency for Unsupervised Partial-to-Partial Point Cloud Registration
abstract
Point Cloud Registration (PCR) is fundamental for the automatic perception of our space. With the rapid development of deep neural network, the community has swiftly adapted to this data-driven technique, and achieved promising performances. However, most existing learning-based methods attempt to conduct PCR within specific ideal experimental settings, in which the ground truth transformations are accessible and most of the data points have one-to-one correspondences. But in real-world scenarios, the GT transformations are often unknown, and point clouds may only share partially overlapped regions. It leads us to a challenging yet practical issue: How to perform Partial-to-Partial (PtP) Point Cloud Registration without pre-acquired supervisions? In this paper, we aim to tackle both challenges under a unified framework. To achieve this, we propose a novel Focal Anchor Generator to emulate the human perceptual process, particularly focusing on the mutual cloud parts. On top of it, a set of Multi-Consistency constraints are introduced to equip our model with the unsupervised learning ability, which is highly applicable. Extensive experiments have demonstrated the distinctive quality of our proposed framework. We believe this work will broaden the scope of PCR research and enhance the applicative potential of PCR algorithms. (The project code has been released on github.com/chengxiaojin/FGMC-UPCR).
Yucheng Shu, Longjin Cheng, Bin Xiao 0002, Lihong Qiao, Weisheng Li 0001, Xinbo Gao 0001
ICME5
2024 C3T: Contrastive Consistency Cross-Network Learning for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised image semantic segmentation, a vital but challenging task in multimedia applications, aims to accurately classify pixels with limited labeled data. Traditional approaches in this domain often grapple with the confirmation bias problem, where models, influenced by their own predictions, become prone to replicating errors. To address this critical issue, our research introduces a cross-network-crossview consistency learning framework. This novel paradigm significantly reduce the confirmation bias through diversifying the learning perspectives. Integral to our approach are two components: a pseudo-label validation and filtering mechanism, and a cross-contrastive learning module within the feature domain. These elements work in synergy to not only amplify the accuracy of the model but also its robustness against varied data scenarios. Extensive evaluations, conducted across multiple datasets, clearly demonstrate the effectiveness of our method. In comparison to existing state-of-the-art models, our approach exhibits marked improvements, especially in the challenging contexts of semisupervised image semantic segmentation. The code is available at https://github.com/Sstar2orchid/C3T.
Yucheng Shu, Jiaxin Xie, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
ICME5
2024 PriFU: Capturing Task-Relevant Information Without Adversarial Learning
abstract
As machine learning advances, machine learning as a service (MLaaS) in the cloud brings convenience to human lives but also privacy risks, as powerful neural networks used for generation, classification or other tasks can also become privacy snoopers. This motivates privacy preservation in the inference phase. Many approaches for preserving privacy in the inference phase introduce multi-objective functions, training models to remove specific private information from users' uploaded data. Although effective, these adversarial learning-based approaches suffer not only from convergence difficulties, but also from limited generalization beyond the specific privacy for which they are trained. To address these issues, we propose a method for privacy preservation in the inference phase by removing task-irrelevant information, which requires no knowledge of the privacy attacks nor introduction of adversarial learning. Specifically, we introduce a metric to distinguish task-irrelevant information from task-relevant information, and achieve more efficient metric estimation to remove task-irrelevant features. The experiments demonstrate the potential of our method in several tasks. Our code will be available at: https://github.com/iwhoyoung/PriFU.
Xiuli Bi, Bo Liu 0047, Weisheng Li 0001, Pamela C. Cosman, Bin Xiao 0002
ACM Multimedia4
2024 MTSNet: Joint Feature Adaptation and Enhancement for Text-Guided Multi-view Martian Terrain Segmentation
Xuefeng Rao, Xinbo Gao 0001, Weisheng Li 0001, Zijian Min
ACM Multimedia4
2024 ShiftMorph: A Fast and Robust Convolutional Neural Network for 3D Deformable Medical Image Registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Bin Xiao 0002
ACM Multimedia2
2024 Siam2C: Siamese visual segmentation and tracking with classification-rank loss and classification-aware
Bang Jun Lei, Qishuai Ding, Weisheng Li 0001, Hao Tian 0002, Lifang Zhou
Appl. Intell.3
2024 Robust and adaptive subspace learning for fast hyperspectral image denoising
Yue Wu 0007, Weisheng Li 0001
Appl. Intell.2
2024 Retrieval-and-alignment based large-scale indoor point cloud semantic segmentation
Zongyi Xu, Xiaoshui Huang, Yangfu Wang, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
Sci. China Inf. Sci.6
2024 A novel time-delay neural grey model and its applications
Dajiang Lei, Liping Zhang 0012, Qun Liu 0005, Weisheng Li 0001
Expert Syst. Appl.5
2024 Blind image quality index with high-level Semantic Guidance and low-level fine-grained Representation
Bo Hu 0008, Leida Li, Ke Gu 0001, Shuaijian Wang, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing6
2024 Multi-branch progressive embedding network for crowd counting
Lifang Zhou, Songlin Rao, Weisheng Li 0001, Bo Hu 0008
Image Vis. Comput.3
2024 CMRVAE: Contrastive margin-restrained variational auto-encoder for class-separated domain adaptation in cardiac segmentation
Lihong Qiao, Rui Wang 0173, Yucheng Shu, Bin Xiao 0002, Xidong Xu, Baobin Li, Weisheng Li 0001, Xinbo Gao 0001, Bai Ying Lei
Knowl. Based Syst.8
2024 An elastic competitive and discriminative collaborative representation method for image classification
Jian-Xun Mi, Shijie Yin, Weisheng Li 0001
Neural Networks4
2024 D-Net: A dual-encoder network for image splicing forgery detection and localization
Bo Liu 0047, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001, Xinbo Gao 0001
Pattern Recognit.5
2024 Dual Graph Neural Networks for Dynamic Users' Behavior Prediction on Social Networking Services
abstract
Social network services (SNSs) provide platforms where users engage in social link behavior (e.g., predicting social relationships) and consumption behavior. Recent advancements in deep learning for recommendation and link prediction explore the symbiotic relationships between these behaviors, leveraging social influence theory and user homogeneity, i.e., users tend to accept recommendations from social friends and connect with like-minded users. These studies yield positive feedback for users and platforms, fostering practical applications and economic development. While previous works jointly model these behaviors, most studies often overlook the evolution of social relationships and users’ preferences in dynamic scenes and the correlations inside, as well as the higher order information within the social network and preference network (consumption history). To address this, we propose the dynamic graph neural joint behavior prediction model (DGN-JBP). Specifically, we actively disentangle and initialize user embeddings from multiple perspectives to refine information for modeling. Additionally, we design an attentive graph neural network and combine it with gate recurrent units (GRUs) to extract high-order dynamic information. Finally, we design a dual framework and purposefully fuse embeddings to mutually enhance the effectiveness of predictions on two prediction tasks. Extensive experimental results on two real-world datasets clearly demonstrate the effectiveness of our proposed model.
Junwei Li 0011, Le Wu 0001, Yulu Du, Richang Hong, Weisheng Li 0001
IEEE Trans. Comput. Soc. Syst.5
2024 Boosting Robust Multi-Focus Image Fusion With Frequency Mask and Hyperdimensional Computing
abstract
Multi-focus image fusion (MFIF) creates an image from different source images with various sensors or optical settings as the devices can’t focus all objects at different distances. Most of the MFIF methods have several limitations in encoder enough features from the images and the result are not robust. To overcome the primary issue, we present a robust fusion algorithm based on the Frequency mask and the Hyperdimensional computing. We propose the Frequency Mask Filter (FMF) to get the narrow-band signals by encoding the frequency domain vector through the mask filter in the frequency domain. The Hyperdimensional encoder uses monogenic mapping, in which the multi-modulation features (MMF) such as the frequency, phase and amplitude are dynamically selected to obtain robust focus maps. Generated by multiscale monogenic representations of each image, the narrow-band image are mapped to hypervector encoding. Hyperdimensional encoder shows the energetic and structural information and leads to robust fusion results. Our proposed method is far superior to the existing MFIF method in terms of both objective evaluation metrics and visual effects on three publicly available datasets.Additionally, our proposed method requires only 0.88 seconds and has a parameter count of 0.13 million for multi-focus image fusion.
Lihong Qiao, Shixin Wu, Bin Xiao 0002, Yucheng Shu, Xiao Luan, Sicheng Lu, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Learning Discriminative Representations From Cross-Scale Features for Camouflaged Object Detection
abstract
The key that hinders the performance improvement of current camouflaged object detection (COD) models is the lack of discriminability of features at fine granularity. We solve this problem from two complementary perspectives. Firstly, complex scenes result in the discriminative feature representations of camouflaged objects being present at different scales and semantic abstraction levels. Therefore, a mechanism is needed to increase the diversity of features to integrate more information potentially beneficial for COD. Second, appearance similarity between objects and environments will inevitably lead to similarity in features. Enhancing feature diversity alone is not enough to solve the above problems. Therefore, it is necessary to give the model semantic perception capabilities to expand the subtle discrepancies between objects and environments in feature embedding. Inspired by the first point, we propose a cross-scale interaction module (CSIM) that utilizes cross-attention between different scales to enhance the diversity of feature representations. Regarding the second point, the semantic guided feature learning (SGFL) is proposed to promote the model to expand feature discrepancies through explicit supervision. Experiments on four popular COD datasets show that our method outperforms recent SOTA methods. In addition, polyp segmentation experiments show that it is also effective for other COD-like tasks.
Yongchao Wang 0004, Xiuli Bi, Bo Liu 0047, Yang Wei 0002, Weisheng Li 0001, Bin Xiao 0002
IEEE Trans. Circuits Syst. Video Technol.5
2024 Deep Lifelong Cross-Modal Hashing
abstract
Hashing methods have made significant progress in cross-modal retrieval tasks with fast query speed and low storage cost. Among them, deep learning-based hashing achieves better performance on large-scale data due to its excellent extraction and representation ability for nonlinear heterogeneous features. However, there are still two main challenges in catastrophic forgetting when data with new categories arrive continuously, and time-consuming for non-continuous hashing retrieval to retrain for updating. To this end, we, in this paper, propose a novel deep lifelong cross-modal hashing to achieve lifelong hashing retrieval instead of re-training hash function repeatedly when new data arrive. Specifically, we design lifelong learning strategy to update hash functions by directly training the incremental data instead of retraining new hash functions using all the accumulated data, which significantly reduce training time. Then, we propose lifelong hashing loss to enable original hash codes participate in lifelong learning but remain invariant, and further preserve the similarity and dis-similarity among original and incremental hash codes to maintain performance. Additionally, considering distribution heterogeneity when new data arriving continuously, we introduce enhanced-semantic similarity to supervise hash learning, and it has been proven that the similarity improves performance with detailed analysis. Experimental results on benchmark datasets show that our proposed method achieves comparative performance comparing with recent state-of-the-art cross-modal hashing methods, and it yields substantial average increments over 20% in retrieval accuracy and almost reduces over 80% training time when new data arrives continuously.
Liming Xu, Bochuan Zheng, Weisheng Li 0001, Jiancheng Lv 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 CTNet: Contrastive Transformer Network for Polyp Segmentation
abstract
Segmenting polyps from colonoscopy images is very important in clinical practice since it provides valuable information for colorectal cancer. However, polyp segmentation remains a challenging task as polyps have camouflage properties and vary greatly in size. Although many polyp segmentation methods have been recently proposed and produced remarkable results, most of them cannot yield stable results due to the lack of features with distinguishing properties and those with high-level semantic details. Therefore, we proposed a novel polyp segmentation framework called contrastive Transformer network (CTNet), with three key components of contrastive Transformer backbone, self-multiscale interaction module (SMIM), and collection information module (CIM), which has excellent learning and generalization abilities. The long-range dependence and highly structured feature map space obtained by CTNet through contrastive Transformer can effectively localize polyps with camouflage properties. CTNet benefits from the multiscale information and high-resolution feature maps with high-level semantic obtained by SMIM and CIM, respectively, and thus can obtain accurate segmentation results for polyps of different sizes. Without bells and whistles, CTNet yields significant gains of 2.3%, 3.7%, 3.7%, 18.2%, and 10.1% over classical method PraNet on Kvasir-SEG, CVC-ClinicDB, Endoscene, ETIS-LaribPolypDB, and CVC-ColonDB respectively. In addition, CTNet has advantages in camouflaged object detection and defect detection. The code is available at https://github.com/Fhujinwu/CTNet.
Bin Xiao 0002, Jinwu Hu, Weisheng Li 0001, Chi-Man Pun, Xiuli Bi
IEEE Trans. Cybern.3
2024 Nonlocal Deep Unfolding Pansharpening Method Based on Degradation Kernel Estimation
abstract
Variational optimization (VO) pansharpening method proposes a model for optimizing the energy function by delineating the acquisition process of remote sensing images. However, traditional VO methods face challenges in solving the operators. Currently, some model-driven methods alleviate this problem by employing deep unfolding techniques. However, most model-driven approaches do not take into account the unknown and variable degradation process inherent in remote sensing images when solving the optimization model. They typically utilize a fixed Gaussian downsampling operator or a two-layer convolutional neural network module to directly simulate the degradation process during model unfolding, which leads to models that lack sufficient generalization. Therefore, this article proposes a pansharpening method based on deep nonlocal unfolding with degradation kernel estimation. Specifically, we expand the iterative process of solving the energy function into three modules: degradation kernel estimator, image generator, and three-branch prior network. First, the degradation kernel estimator is employed to fit the real degradation process of the image and optimize the generation of an adaptive degradation kernel. Simultaneously, we incorporate local priors, nonlocal spatial priors, and nonlocal spectral priors into the three-branch prior network to adaptively capture local and nonlocal prior features. These prior features are then fused with the generated image through residual connections to approximate the real remote sensing image. Experimental results on datasets from two different types of satellites demonstrate the superiority of our approach.
Dajiang Lei, Genyuan Zhang, Qun Liu 0005, Weisheng Li 0001, Liping Zhang 0012
IEEE Trans. Geosci. Remote. Sens.5
2024 Deep Cross-View Reconstruction GAN Based on Correlated Subspace for Multi-View Transformation
abstract
In scenarios where identifying face information in the visible spectrum (VIS) is challenging due to poor lighting conditions, the use of near-infrared (NIR) and thermal (TH) cameras can provide viable alternatives. However, the unique data distribution of images captured by these cameras compared to VIS images presents challenges in matching face identities. To address these challenges, we propose a novel image transformation framework. The framework includes feature extraction from the input image, followed by a transformation network that generates target domain images with perceptual fidelity. Additionally, a reconstruction network preserves original information by reconstructing the original domain image from the extracted features. By considering the correlation between features from both domains, our framework utilizes paired data obtained from the same individual. We apply this framework to two well-established image-to-image transformation models, pix2pix and CycleGAN, known as CRC-pix2pix and CRC-CycleGAN respectively. The versatility of our approach allows extension to other models based on pix2pix or CycleGAN architectures. Our models generate high-quality images while preserving the identity information of the original face. Performance evaluation on TFW and BUAA NIR-VIS datasets demonstrates the superiority of our models in terms of generated image face matching and evaluation metrics such as SSIM, MSE, PSNR, and LPIPS. Moreover, we introduce the CQUPT-VIS-TH dataset, which enriches the paired dataset with thermal-visual face data capturing various angles and expressions.
Jian-Xun Mi, Junchang He, Weisheng Li 0001
IEEE Trans. Image Process.3
2024 Boundary-Aware Prototype in Semi-Supervised Medical Image Segmentation
abstract
The true label plays an important role in semi-supervised medical image segmentation (SSMIS) because it can provide the most accurate supervision information when the label is limited. The popular SSMIS method trains labeled and unlabeled data separately, and the unlabeled data cannot be directly supervised by the true label. This limits the contribution of labels to model training. Is there an interactive mechanism that can break the separation between two types of data training to maximize the utilization of true labels? Inspired by this, we propose a novel consistency learning framework based on the non-parametric distance metric of boundary-aware prototypes to alleviate this problem. This method combines CNN-based linear classification and nearest neighbor-based non-parametric classification into one framework, encouraging the two segmentation paradigms to have similar predictions for the same input. More importantly, the prototype can be clustered from both labeled and unlabeled data features so that it can be seen as a bridge for interactive training between labeled and unlabeled data. When the prototype-based prediction is supervised by the true label, the supervisory signal can simultaneously affect the feature extraction process of both data. In addition, boundary-aware prototypes can explicitly model the differences in boundaries and centers of adjacent categories, so pixel-prototype contrastive learning is introduced to further improve the discriminability of features and make them more suitable for non-parametric distance measurement. Experiments show that although our method uses a modified lightweight UNet as the backbone, it outperforms the comparison method using a 3D VNet with more parameters.
Yongchao Wang 0004, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2024 CGFTrans: Cross-Modal Global Feature Fusion Transformer for Medical Report Generation
abstract
Medical report generation, as a cross-modal automatic text generation task, can be highly significant both in research and clinical fields. The core is to generate diagnosis reports in clinical language from medical images. However, several limitations persist, including a lack of global information, inadequate cross-modal fusion capabilities, and high computational demands. To address these issues, we propose cross-modal global feature fusion Transformer (CGFTrans) to extract global information meanwhile reduce computational strain. Firstly, we introduce mesh recurrent network to capture inter-layer information at different levels to address the absence of global features. Then, we design feature fusion decoder and define 'mid-fusion' strategy to separately fuse visual and global features with medical report embeddings, which enhances the ability of the cross-modal joint learning. Finally, we integrate shifted window attention into Transformer encoder to alleviate computational pressure and capture pathological information at multiple scales. Extensive experiments conducted on three datasets demonstrate that the proposed method achieves average increments of 2.9%, 1.5%, and 0.7% in terms of the BLEU-1, METEOR and ROUGE-L metrics, respectively. Besides, it achieves average increments -22.4% and 17.3% training time and images throughput, respectively.
Liming Xu, Quan Tang 0006, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001
IEEE J. Biomed. Health Informatics5
2024 Shape-Scale Co-Awareness Network for 3D Brain Tumor Segmentation
abstract
The accurate segmentation of brain tumor is significant in clinical practice. Convolutional Neural Network (CNN)-based methods have made great progress in brain tumor segmentation due to powerful local modeling ability. However, brain tumors are frequently pattern-agnostic, i.e. variable in shape, size and location, which can not be effectively matched by traditional CNN-based methods with local and regular receptive fields. To address the above issues, we propose a shape-scale co-awareness network (S2CA-Net) for brain tumor segmentation, which can efficiently learn shape-aware and scale-aware features simultaneously to enhance pattern-agnostic representations. Primarily, three key components are proposed to accomplish the co-awareness of shape and scale. The Local-Global Scale Mixer (LGSM) decouples the extraction of local and global context by adopting the CNN-Former parallel structure, which contributes to obtaining finer hierarchical features. The Multi-level Context Aggregator (MCA) enriches the scale diversity of input patches by modeling global features across multiple receptive fields. The Multi-Scale Attentive Deformable Convolution (MS-ADC) learns the target deformation based on the multiscale inputs, which motivates the network to enforce feature constraints both in terms of scale and shape for optimal feature matching. Overall, LGSM and MCA focus on enhancing the scale-awareness of the network to cope with the size and location variations, while MS-ADC focuses on capturing deformation information for optimal shape matching. Finally, their effective integration prompts the network to perceive variations in shape and scale simultaneously, which can robustly tackle the variations in patterns of brain tumors. The experimental results on BraTS 2019, BraTS 2020, MSD BTS Task and BraTS2023-MEN show that S2CA-Net has superior overall performance in accuracy and efficiency compared to other state-of-the-art methods. Code: https://github.com/jiangyu945/S2CA-Net.
Lifang Zhou, Weisheng Li 0001, Shenhai Zheng
IEEE Trans. Medical Imaging3
2024 Blind Image Quality Index With Cross-Domain Interaction and Cross-Scale Integration
abstract
With the assistance of Convolutional Neural Networks (CNNs), Image Quality Assessment (IQA) models have made great progress in evaluating both simulated distortion and authentic distortion. However, most of the existing IQA models only learn the features of distorted images, and thus do not make full use of the available feature representation of other domains. Furthermore, the common multi-scale fusion strategies are relatively simple, such as downsampling and concatenating, which further limits the prediction performance. To this end, we propose a novel blind image quality index with cross-domain interaction and cross-scale integration, which is designed based on the combination of CNN and Transformer. First, the hierarchical spatial-domain and gradient-domain representations are obtained through a typical CNN architecture. Then, based on the proposed gradient-query cross-attention, these two types of features are fully interacted in the Cross-Domain Interaction (CDI) module. To represent the distortion information more comprehensively, the Cross-Scale Integration (CSI) module is proposed to combine the information between different scales progressively. Finally, the quality score is obtained through a simple regression module. The experimental results on five public IQA databases of both simulated and authentic scenes show that the proposed model outperforms the compared state-of-the-art metrics. In addition, cross-database experiments show that the proposed model has strong generalization performance.
Bo Hu 0008, Leida Li, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.5
2024 IGReg: Image-Geometry-Assisted Point Cloud Registration via Selective Correlation Fusion
abstract
Point cloud registration suffers from repeated patterns and low geometric structures in indoor scenes. The recent transformer utilises attention mechanism to capture the global correlations in feature space and improves the registration performance. However, for indoor scenarios, global correlation loses its advantages as it cannot distinguish real useful features and noise. To address this problem, we propose an image-geometry-assisted point cloud registration method by integrating image information into point features and selectively fusing the geometric consistency with respect to reliable salient areas. Firstly, an Intra-Image-Geometry fusion module is proposed to integrate the texture and structure information into the point feature space by the cross-attention mechanism. Initial corresponding superpoints are acquired as salient anchors in the source and target. Then, a selective correlation fusion module is designed to embed the correlations between the salient anchors and points. During training, the saliency location and selective correlation fusion modules exchange information iteratively to identify the most reliable salient anchors and achieve effective feature fusion. The obtained distinctive point cloud features allow for accurate correspondence matching, leading to the success of indoor point cloud registration. Extensive experiments are conducted on 3DMatch and 3DLoMatch datasets to demonstrate the outstanding performance of the proposed approach compared to the state-of-the-art, particularly in those geometrically challenging cases such as repetitive patterns and low-geometry regions.
Zongyi Xu, Xinqi Jiang, Changjun Gu, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.7
2023 Self-Supervised Image Local Forgery Detection by JPEG Compression Trace
abstract
For image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression chains, and designed a trace extractor to learn such traces. Then, we utilized the trace extractor as the backbone and trained self-supervised to strengthen the discrimination ability of learned traces. With its benefits, regions with different JPEG compression chains can easily be distinguished within a forged image. Furthermore, our method does not rely on a large amount of training data, and even does not require any forged images for training. Experiments show that the proposed method can detect image local forgery on different datasets without re-training, and keep stable performance over various types of image local forgery.
Xiuli Bi, Wuqing Yan, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
AAAI5
2023 SRPA: ScribbleMatch and Reliable-Guided Pixel Alignment for Scribble-Supervised Medical Image Segmentation
abstract
Medical image segmentation is a critical task in the field of medical image analysis. Recently, there has been increasing attention on scribble-supervised medical image segmentation due to its simplicity for label generation. However, the performance of a scribble-based task highly relies on the quality of learning from inconsistent annotations and the classification of pixels in high-entropy regions. In this paper, we propose a novel framework called SRPA that combines both ScribbleMatch and Reliable-Guided Pixel Alignment to enhance the performance of the scribble-based task. The ScribbleMatch technique utilizes the pseudo label incorporated from two different weakly perturbed views of the same image to supervise a strongly perturbed view, which assists in boosting the quality of the shape information learning as the scribble-based task lacks the prototypes to consistently capture shape prior to model training. The Reliable-Guided Pixel Alignment technique employs reliable pixels selected by contrasting two weakly perturbed views, which serve as the standard for blurred pixels in strongly perturbed images to maximally align. This ensures the reliable classification of high-entropy pixels. Our method is evaluated on the public ACDC and MSCMRseg datasets, and the results demonstrate that our approach surpasses current scribble-supervised segmentation methods. Code will be available at https://github.com/RheinSXY/SRPA.
Tingjie Liu, Lihong Qiao, Yucheng Shu, Weisheng Li 0001, Xinbo Gao 0001
BIBM4
2023 Non-rigid Medical Image Registration Based on Unsupervised Self-driven Prior Fusion
abstract
Deformable image registration is a basic building block in intelligent bioinformatical analysis and biomedicine systems. With the rapid development of deep learning, the community has witnessed a great leap via this effective data-driven technique. Recently, the Vision Transformer, famous by its long-range modeling ability, has been successfully used in the field of medical image registration. However, the existing ViT based techniques are deemed to have certain limitations. Firstly, these methods often transplanted the transformer module directly into the networks, while did not dive deeper to explore its compatibility to practical registration tasks. Moreover, self-attention’s relatively rigid all-to-all patching strategy may cause undesirable discontinuity effect to the spatial calculation. To address these issues, we propose a novel medical image registration framework based on an efficient image prior learning and fusion mechanism. Unlike the existing prior-based registration methods, our model is capable of learning task-specific saliency priors, without the need of hand-crafted features, or heavy-loaded auxiliary tasks, or pre-acquired expensive annotations. Then, followed by a multi-scale patch embedding module, the self-driven saliency prior is integrated into a ViT block with an active feature fusion mechanism, to further expand our network’s structural learning abilities. Extensive experiments on multiple data sets have demonstrated the superior quality of the proposed framework. We believe this plug-and-play model will bring about more application potentials to the community (Project webpage: https://github.com/raincity212/SPF-Net).
Yucheng Shu, Xuxuan Guan, Bin Xiao 0002, Lihong Qiao, Weisheng Li 0001, Xinbo Gao 0001
BIBM5
2023 MCF: Mutual Correction Framework for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning is a promising method for medical image segmentation under limited annotation. However, the model cognitive bias impairs the segmentation performance, especially for edge regions. Furthermore, current mainstream semi-supervised medical image segmentation (SSMIS) methods lack designs to handle model bias. The neural network has a strong learning ability, but the cognitive bias will gradually deepen during the training, and it is difficult to correct itself. We propose a novel mutual correction framework (MCF) to explore network bias correction and improve the performance of SSMIS. Inspired by the plain contrast idea, MCF introduces two different subnets to explore and utilize the discrepancies between subnets to correct cognitive bias of the model. More concretely, a contrastive difference review (CDR) module is proposed to find out inconsistent prediction regions and perform a review training. Additionally, a dynamic competitive pseudo-label generation (DCPLG) module is proposed to evaluate the performance of subnets in real-time, dynamically selecting more reliable pseudo-labels. Experimental results on two medical image databases with different modalities (CT and MRI) show that our method achieves superior performance compared to several state-of-the-art methods. The code will be available at https://github.com/WYC-321/MCF.
Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
CVPR4
2023 DLBD: A Self-Supervised Direct-Learned Binary Descriptor
abstract
For learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed binarization processes. Since their binarizaiion processes are not a component of the network, the learning-based binary descriptor cannot fully utilize the advances of deep learning. To solve this issue, we propose a model-agnostic plugin binary transformation layer (BTL), making the network directly generate binary descriptors. Then, we present the first self-supervised, direct-learned binary descriptor, dubbed DLBD. Furthermore, we propose ultra-wide temperature-scaled crossentropy loss to adjust the distribution of learned descriptors in a larger range. Experiments demonstrate that the proposed BTL can substitute the previous binarization process. Our proposed DLBD outperforms SOTA on different tasks such as image retrieval and classification11Our code is available at: https://github.com/CQUPT-CV/DLBD.
Bin Xiao 0002, Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
CVPR5
2023 A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC
abstract
Due to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a novel Mode Selection-Based Fast Intra Prediction algorithm for SSHVC. We reveal the RD costs of Inter-layer Reference (ILR) mode and Intra mode have a significant difference, and the RD costs of these two modes follow Gaussian distribution. Based on this observation, we propose to apply the classic Gaussian Mixture Model and Expectation Maximization in machine learning to determine whether ILR is the best mode so as to skip the Intra mode. Experimental results demonstrate that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss.
Yu Sun 0003, Weisheng Li 0001, Lele Xie, Xin Lu 0001, Frédéric Dufaux, Ce Zhu
ICASSP3
2023 Fast Learning-Based Split Type Prediction Algorithm for VVC
abstract
As the latest video coding standard, Versatile Video Coding (VVC) is highly efficient at the cost of very high coding complexity, which seriously hinders its widespread application. Therefore, it is very crucial to improve its coding speed. In this paper, we propose a learning-based fast split type (ST) prediction algorithm for VVC using a deep learning approach. We first construct a large-scale database containing sufficient STs with diverse video resolution and content. Next, since the ST distributions of coding units (CUs) of different sizes are significantly distinct, so we separately design neural networks for all different CU sizes. Then, we merge ambiguous STs into four merged classes (MCs) to train models to obtain probabilities of MCs and skip unlikely ones. Experimental results demonstrate that the proposed algorithm can reduce the encoding time of VVC by 67.53% with 1.89% increase in Bjøntegaard delta bit-rate (BDBR) on average.
Liulin Chen, Xin Lu 0001, Frédéric Dufaux, Weisheng Li 0001, Ce Zhu
ICIP5
2023 A Probability-Based All-Zero Block Early Termination Algorithm for QSHVC
abstract
To seamlessly adapt to time-varying network bandwidths, Quality Scalable High-Efficiency Video Coding (QSHVC) is developed. However, its coding process is overwhelmingly complex, and this seriously limits its wide applications in realtime environments. Therefore, it is of great significance to study fast coding algorithms for QSHVC. In this paper, we propose a novel probability-based All-Zero Block (AZB) early termination algorithm for QSHVC. We observe that the generated residual coefficients follow the Laplace distribution if a CU is accurately predicted. Based on this observation, we derive the sum of squared differences-based AZB decision condition. Second, the probability of each coding mode and coding depth being chosen as the best ones are combined with AZBs to derive the probability-based early termination condition. The experimental results show that the proposed algorithm can improve the average coding speed by 74.95% with a 0.26% increase in BDBR.
Xin Lu 0001, Frédéric Dufaux, Qianmin Wang, Weisheng Li 0001, Bo Hang, Ce Zhu
ICIP5
2023 Cross-slice Context Consistency for Semi-supervised 3D Left Atrium Segmentation
abstract
Semi-supervised learning is a promising approach in reducing the requirement to collect large amounts of dense annotations, especially in medical image segmentation. However, most existing semi-supervised 3D medical image segmentation methods tend to ignore the cross-slice context that contains extensive structural information. We believe cross-slice context can help the model capture semantic information complementary to slice context and achieve robust and more accurate segmentation. Therefore, in this paper, we propose a novel cross-slice context consistency framework for 3D left atrium segmentation named CSC2-Net. Our method can effectively utilize unlabeled data by encouraging consistent results between slice segmentation and cross-slice inference segmentation. To achieve this, we design a bidirectional gated context inference module (Bi-GCM) to model cross-slice context and predict slice segmentation without direct slice features. Experiments on a public left atrium (LA) databases show that our method achieves higher performance and outperforms state-of-the-art methods by imposing cross-slice consistency constraint.
Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
ICME4
2023 AEP-GAN: Aesthetic Enhanced Perception Generative Adversarial Network for Asian facial beauty synthesis
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002
Appl. Intell.2
2023 A spatiotemporal fusion method based on interpretable deep networks
Dajiang Lei, Jiayang Tan, Qun Liu 0005, Weisheng Li 0001
Appl. Intell.5
2023 A location-aware siamese network for high-speed visual tracking
Lifang Zhou, Weisheng Li 0001, Jiaxu Leng, Bang Jun Lei, Weibin Yang
Appl. Intell.3
2023 Extraction-and-excitation deep neural network for pansharpening
abstract
Abstract With the recent advances achieved by deep neural networks in image processing applications, researchers have begun exploring deep learning in pansharpening and obtained remarkable results. However, the existing methods are generally limited by their weak feature representation ability, often leading to spectral distortion or spatial blur. To generate high‐quality pansharpened images, this article proposes a novel neural network for pansharpening that includes both feature extraction and excitation mechanisms to consider important features. The neural network is modified with domain knowledge in pansharpening to fully extract spectral and spatial structures, and the proposed method outperforms traditional methods.
Dajiang Lei, Liping Zhang 0012, Weisheng Li 0001
Concurr. Comput. Pract. Exp.4
2023 QDRJL: Quaternion dynamic representation with joint learning neural network for heart sound signal abnormal detection
abstract
At present, deep learning based heart sound diagnosis algorithms are mostly complex and large models for high accuracy, which are difficult to deploy on mobile devices due to the high number of parameters and large computational cost. The current mainstream approach for processing heart sound signals involves utilizing their Mel-frequency cepstral coefficients (MFCC) features. However, most existing methods have overlooked the multi-channel characteristics of MFCC. To address this issue, we propose a Quaternion Dynamic Representation with Joint Learning (QDRJL) neural network for learning MFCC multi-channel features. Our proposed approach combines quaternion dynamic convolution with dynamic weighting and the Quaternion Interior Learning Block (QILB). Finally, we present a global and energy joint learning branch for jointly learning MFCC features. The success of the proposed quaternion network depends on its ability to utilize the internal relations between quaternion-valued input features and the definition of the dynamic weight variables in the augmented quaternion domain. We assessed various state-of-the-art classification algorithms for detecting heart sounds and found that our proposed classifier achieved an accuracy of up to 97.2%, outperforming existing models. Our experimental evaluation, using the 2016 PhysioNet/CinC Challenge dataset, revealed that our model could reduce the number of network parameters to 25% due to quaternion properties.
Lihong Qiao, Bin Xiao 0002, Yucheng Shu, Yuhang Shi, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing7
2023 Deep image captioning: A review of methods, trends and future challenges
Liming Xu, Quan Tang 0006, Jiancheng Lv 0001, Bochuan Zheng, Weisheng Li 0001
Neurocomputing6
2023 A Dual Self-Calibrating Framework for Noninvasive Fetal ECG R-Peak Detection
abstract
Fetal heart rate (fHR) is critical for assessing fetal health and diagnosing disorders, such as fetal distress, congenital heart disease, and intrauterine growth retardation. With the rapid development of the Internet of Medical Things (IoMT), fetal R-peak detection plays an important role in diagnosing heart defects during pregnancy. However, due to the nonlinear mixing of multiple sources in the noninvasive signals and the low signal-to-noise ratio (SNR), it is difficult to obtain accurate R-peak detection result. This article presents a dual self-calibrating system based on a spectral attention kernel independent component analysis (SA-KICA) module and a self-calibrating fetal R-peak detection (SC-FRD) module. SA-KICA is an ICA-based calibration module constructed by the spectral attention mechanism, which was sought from short-time Fourier transform (STFT) and was shipped back to original signal with convolution to achieve perfect maternal electrocardiogram (MECG) separation in high-dimensional linear separable space. Then, a periodic and morphological-based channel selector is designed to select the optimal MECG. After MECG removal, to further improve the performance of fetal R-peak detection, the SC-FRD module is introduced to utilize the interior peak information and self-calibrating strategy, which includes variance-based fetal R-peak seed selection, time-varying coarse prediction, and adaptive probability mask calibration. The proposed framework is a primary attempt to concurrently introduce the nonlinear feature, spectral information, and self-calibrating strategy in the field of fetal ECG processing. The framework achieved excellent performance in fetal R-peak detection accuracy on a simulated data set and two public data sets with varying divergence and richness of resources. The experimental results show that our framework is superior to existing methods and can be used as a potential fetal monitoring method in the application of IoMT. The code is released inhttps://github.com/bfyjr/NI-FECG-Extraction.
Lihong Qiao, Shuai Hu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
IEEE Internet Things J.5
2023 MyoPS: A benchmark of myocardial pathology segmentation combining three-sequence cardiac magnetic resonance images
Lei Li 0020, Fuping Wu, Xinzhe Luo, Carlos Martín-Isla, Shuwei Zhai, Zhen Zhang 0057, Markus J. Ankenbrand, Haochuan Jiang, Linhong Wang, Tewodros Weldebirhan Arega, Elif Altunok, Jun Ma 0016, Xiaoping Yang 0001, Élodie Puybareau, Ilkay Öksüz, Stéphanie Bricq, Weisheng Li 0001, Kumaradevan Punithakumar, Sotirios A. Tsaftaris, Laura Maria Schreiber, Guocai Liu, Yong Xia 0001, Guotai Wang, Sergio Escalera, Xiahai Zhuang
Medical Image Anal.23
2023 Scale-free heterogeneous cycleGAN for defogging from a single image for autonomous driving in fog
Yan Zhang 0002, Zhiping Dan, Shuifa Sun, Jun Wan 0005, Weisheng Li 0001
Neural Comput. Appl.7
2023 Hierarchical neural network with efficient selection inference
Jian-Xun Mi, Ke-Yang Huang, Weisheng Li 0001, Lifang Zhou
Neural Networks4
2023 BH2I-GAN: Bidirectional Hash_code-to-Image Translation using Multi-Generative Multi-Adversarial Nets
Liming Xu, Weisheng Li 0001, Yicai Xie
Pattern Recognit.3
2023 YOLOSA: Object detection based on 2D local feature superimposed self-attention
Weisheng Li 0001
Pattern Recognit. Lett.1
2023 Multi-level dynamic error coding for face recognition with a contaminated single sample per person
Xiao Luan, Linghui Liu, Weisheng Li 0001
Pattern Recognit. Lett.4
2023 IEMask R-CNN: Information-Enhanced Mask R-CNN
abstract
The instance segmentation task is relatively difficult in computer vision, which requires not only high-quality masks but also high-accuracy instance category classification. Mask R-CNN has been proven to be a feasible method. However, due to the Feature Pyramid Network (FPN) structure lack useful channel information, global information and low-level texture information, and mask branch cannot obtain useful local-global information, Mask R-CNN is prevented from obtaining high-quality masks and high-accuracy instance category classification. Therefore, we proposed the Information-enhanced Mask R-CNN, called IEMask R-CNN. In the FPN structure of IEMask R-CNN, the information-enhanced FPN will enhance the useful channel information and the global information of the feature maps to solve the issues that the high-level feature map loses useful channel information and inaccurate of instance category classification, meanwhile the bottom-up path enhancement with adaptive feature fusion will ultilize the precise positioning signal in the lower layer to enhance the feature pyramid. In the mask branch of IEMask R-CNN, an encoding-decoding mask head will strength local-global information to gain a high-quality mask. Without bells and whistles, IEMask R-CNN gains significant gains of about 2.60%, 4.00%, 3.17% over Mask R-CNN on MS COCO2017, Cityscapes and LVIS1.0 benchmarks respectively.
Xiuli Bi, Jinwu Hu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Big Data4
2023 Novel Multi-Feature Fusion Facial Aesthetic Analysis Framework
abstract
Machine learning has been used in facial beauty prediction studies. However, the integrity of facial geometric information is not considered in facial aesthetic feature extraction, and the impact of other facial attributes (expression) on aesthetics. We propose a novel multi-feature fusion facial aesthetic analysis framework (NMFA) to overcome this problem. First, we designed a facial shape feature, which is an intuitive, visual quantitative description, based on B-spline. Second, we designed a representative low-dimensional facial structural feature to establish the theoretical basis of the facial structure, based on facial aesthetic structure and expression recognition theory. Next, we designed texture and holistic features based on Gabor and VGG-face network. Finally, we used a multi-feature fusion strategy to fuse them for aesthetic evaluation. Experiments were conducted on four databases. The results revealed that the proposed method realizes the visualization of facial shape features, enriches geometric information, solves the problem of lack of facial geometric information and difficulty to understand, and achieves excellent performance with fewer parameters.
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002
IEEE Trans. Big Data2
2023 A Versatile Detection Method for Various Contrast Enhancement Manipulations
abstract
Contrast enhancement manipulation is a common method to improve the visual effect of an image. Meanwhile, it can also be considered a type of global image forgery because it changes the image’s visual appearance without alerting its semantics. Moreover, for local image forgery, a tampered image may be composited by images with different contrast enhancement manipulations or post-processed by a contrast enhancement manipulation to conceal the trails of tampering. Therefore, contrast enhancement manipulation detection is critical to global image forgery detection. The existing methods can only detect a particular type of contrast enhancement manipulation, such as gamma correction or histogram equalization. To break this limitation, we propose the zero-gap spans (ZGS) as the fingerprint to explore the traces of contrast enhancement manipulations. Based on ZGS, various contrast enhancement manipulations can be distinguished by a simple classification method at image-level and patch-level; different gamma corrections can be identified, and their gamma value can be estimated. Experimental results indicate that the proposed ZGS-based classification method can achieve and maintain good classification performance under different cases (gamma correction, simple histogram equalization, modified histogram equalization techniques). Meanwhile, ZGS can estimate the gamma value with the mean squared error (MSE) below 0.1156. For the local forgery images, the proposed ZGS also can be utilized to locate the regions with different contrast enhancement manipulations.
Xiuli Bi, Yixuan Shang, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Pareto Refocusing for Drone-View Object Detection
abstract
Drone-view Object Detection (DOD) is a meaningful but challenging task. It hits a bottleneck due to two main reasons: (1) The high proportion of difficult objects (e.g., small objects, occluded objects, etc.) makes the detection performance unsatisfactory. (2) The unevenly distributed objects make detection inefficient. These two factors also lead to a phenomenon, obeying the Pareto principle, that some challenging regions occupying a low area proportion of the image have a significant impact on the final detection while the vanilla regions occupying the major area have a negligible impact due to the limited room for performance improvement. Motivated by the human visual system that naturally attempts to invest unequal energies in things of hierarchical difficulty for recognizing objects effectively, this paper presents a novel Pareto Refocusing Detection (PRDet) network that distinguishes the challenging regions from the vanilla regions under reverse-attention guidance and refocuses the challenging regions with the assistance of the region-specific context. Specifically, we first propose a Reverse-attention Exploration Module (REM) that excavates the potential position of difficult objects by suppressing the features which are salient to the commonly used detector. Then, we propose a Region-specific Context Learning Module (RCLM) that learns to generate specific contexts for strengthening the understanding of challenging regions. It is noteworthy that the specific context is not shared globally but unique for each challenging region with the exploration of spatial and appearance cues. Extensive experiments and comprehensive evaluations on the VisDrone2021-DET and UAVDT datasets demonstrate that the proposed PRDet can effectively improve the detection performance, especially for those difficult objects, outperforming state-of-the-art detectors. Furthermore, our method also achieves significant performance improvements on the DTU-Drone dataset for power inspection.
Jiaxu Leng, Mengjingcheng Mo, Yinghua Zhou, Chenqiang Gao, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Accurate and Robust Eye Center Localization by Deep Voting
abstract
Eye Center Localization (ECL) is one of the most crucial technologies for various computer vision applications, such as eye gazing estimation and eye-tracking. Current conventional implementations consist of two phases, including locating the approximate eye regions and finding the eye center position by extracting the semantic features around the corresponding eye region. However, the combination pipeline results in the ECL accuracy being influenced by not only the environmental factors, such as the variability of photographing angles, illuminations, and the probable occlusions by eyelids or glasses, but also the quality of preceding procedures. Inspired by the ensemble mechanism in machine learning, we formulate ECL problem as a process of end-to-end voting, and the core is to select a set of local descriptors which can capture efficient independent information to vote for eye centers. With the help of deep convolutional neural networks, we are able to determine semantic descriptors around the eye regions. Each descriptor proposes a vote pointing to the corresponding eye center, and all the votes indicate the eye centers finally. The experimental results on the public databases, BioID and GI4E, show that our method achieves 80.3% and 95.2% accuracy, respectively, which outperforms the existing state-of-the-art methods, and the results based on our customized challenging database verify the robustness of our method.
Jian-Xun Mi, Shiyao Yuan, Weisheng Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Correlation Filter Tracker With Sample-Reliability Awareness and Self-Guided Update
abstract
In visual tracking, unreliable samples always exist because of occlusion, illumination variation, motion blur, etc. Existing studies have effectively improved the performance of trackers by enhancing the quality of online samples. However, an underappreciated view is that not all samples are equally essential to model training. In this paper, we propose a Sample-Aware Adaptive Updating (SAAU) strategy which can actively adjust the update formula by sensing the reliability of samples. Specifically, the Sample-Reliability Awareness (SRA) module can quantify sample reliability by calculating three specific indicators, where the Residual Peak-to-Correlation Energy (RPCE) is designed to cooperate with the other two introduced indicators to obtain credit scores on each sample. Besides, the Self-Guided Update (SGU) module provides a tracker with an unfixed learning rate that matches with the reliability label during updating, where our label annotator generates the label. Extensive experiments on several public benchmarks demonstrate the outstanding compatibility of SAAU and the superiority of our tracker (SAAU-CF) over state-of-the-art approaches.
Lifang Zhou, Bang Jun Lei, Weisheng Li 0001, Jiaxu Leng
IEEE Trans. Circuits Syst. Video Technol.4
2023 Pansharpening Method Based on Deep Nonlocal Unfolding
abstract
Although deep neural networks (DNNs) have achieved great success in pansharpening, most of them lack transparency and interpretability. Currently, some DNNs methods utilize deep unfolding techniques to alleviate this problem. However, they do not consider the regularization term separately when solving the energy function that represents the image degradation process, making it difficult to extract complex prior information in the unfolding module. Therefore, this paper proposes a pansharpening method based on deep non-local unfolding. Specifically, we expand the iterative process of solving the energy function into the corresponding neural network modules, making each module have a certain physical meaning. Then, we decouple the prior operator containing the prior knowledge of the remote sensing image and approximate the solution using the network module. Meanwhile, we incorporate local and non-local self-similarity priors into the prior operator and design a two-branch prior module for learning the prior features and contribution weights adaptively. Finally, the fused image is corrected with the learned prior features to approximate the real image. Experimental results on datasets from two different types of satellites demonstrate the superiority of our approach.
Guangyao Shi, Liping Zhang 0012, Weisheng Li 0001, Dajiang Lei
IEEE Trans. Geosci. Remote. Sens.5
2023 A Symmetrical Siamese Network Framework With Contrastive Learning for Pose-Robust Face Recognition
abstract
Face recognition has achieved remarkable success owing to the development of deep learning. However, most of existing face recognition models perform poorly against pose variations. We argue that, it is primarily caused by pose-based long-tailed data - imbalanced distribution of training samples between profile faces and near-frontal faces. Additionally, self-occlusion and nonlinear warping of facial textures caused by large pose variations also increase the difficulty in learning discriminative features of profile faces. In this study, we propose a novel framework called Symmetrical Siamese Network (SSN), which can simultaneously overcome the limitation of pose-based long-tailed data and pose-invariant features learning. Specifically, two sub-modules are proposed in the SSN, i.e., Feature-Consistence Learning sub-Net (FCLN) and Identity-Consistence Learning sub-Net (ICLN). For FCLN, the inputs are all face images on training dataset. Inspired by the contrastive learning, we simulate pose variations of faces and constrain the model to focus on the consistent areas between the original face image and its corresponding virtual pose face images. For ICLN, only profile images are used as inputs, and we propose to adopt Identity Consistence Loss to minimize the intra-class feature variation across different poses. The collaborative learning of two sub-modules guarantees that the parameters of network are updated in a relatively equal probability between near-frontal face images and profile images, so that the pose-based long-tailed problem can be effectively addressed. The proposed SSN shows comparable results over the state-of-the-art methods on several public datasets. In this study, LightCNN is selected as the backbone of SSN, and existing popular networks also can be used into our framework for pose-robust face recognition.
Xiao Luan, Zibiao Ding, Linghui Liu, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2023 HS-Vectors: Heart Sound Embeddings for Abnormal Heart Sound Detection Based on Time-Compressed and Frequency-Expanded TDNN With Dynamic Mask Encoder
abstract
In recent years, auxiliary diagnosis technology for cardiovascular disease based on abnormal heart sound detection has become a research hotspot. Heart sound signals are promising in the preliminary diagnosis of cardiovascular diseases. Previous studies have focused on capturing the local characteristics of heart sounds. In this paper, we investigate a method for mapping heart sound signals with complex patterns to fixed-length feature embedding called HS-Vectors for abnormal heart sound detection. To get the full embedding of the complex heart sound, HS-Vectors are obtained through the Time-Compressed and Frequency-Expanded Time-Delay Neural Network(TCFE-TDNN) and the Dynamic Masked-Attention (DMA) module. HS-Vectors extract and utilize the global and critical heart sound characteristics by masking out irreverent information. Based on the TCFE-TDNN module, the heart sound signal within a certain time is projected into fixed-length embedding. Then, with a learnable mask attention matrix, DMA stats pooling aggregates multi-scale hidden features from different TCFE-TDNN layers and masks out irrelevant frame-level features. Experimental evaluations are performed on a 10-fold cross-validation task using the 2016 PhysioNet/CinC Challenge dataset and the new publicly available pediatric heart sound dataset we collected. Experimental results demonstrate that the proposed method excels the state-of-the-art models in abnormality detection.
Lihong Qiao, Yonghao Gao, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
IEEE J. Biomed. Health Informatics5
2023 Cross-Mix Monitoring for Medical Image Segmentation With Limited Supervision
abstract
Image segmentation is a fundamental building block of automatic medical applications. It has been greatly improved since the emergence of deep neural networks. However, deep-learning based models often require a large number of manual annotations, which has seriously hindered its practical usage. To alleviate this problem, numerous works were proposed by utilizing unlabeled data based on semi-supervised frameworks. Recently, the Mean-Teacher (MT) model has been successfully applied in many scenarios due to its effective learning strategy. Nevertheless, the existing MT model still have certain limitations. Firstly, various sorts of perturbations are often added to the training data to gain extra generalization ability through consistency training. However, if the variation is too weak, it may cause the Lazy Student Phenomenon, and bring large fluctuations to the learning model. On the contrary, large image perturbations may enlarge the performance gap between the teacher and student. In this case, the student may lose its learning momentum, and more seriously, drag down the overall performance of the whole system. In order to address these issues, we introduce a novel semi-supervised medical image segmentation framework, in which a Cross-Mix Teaching paradigm is proposed to provide extra data flexibility, thus effectively avoid Lazy Student Phenomenon. Moreover, a lightweight Transductive Monitor is applied to server as the bridge that connect the teacher and student for active knowledge distillation. In the light of this cross-network information mixing and transfer mechanism, our method is able to continuously explore the discriminative information contained in unlabeled data. Extensive experiments on challenging medical image data sets demonstrate that our method is able to outperform current state-of-the-art semi-supervised segmentation methods under severe lack of supervision.
Yucheng Shu, Hengbo Li, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001
IEEE Trans. Multim.5
2023 MFGAN: Multi-modal Feature-fusion for CT Metal Artifact Reduction Using GANs
abstract
Due to the existence of metallic implants in certain patients, the Computed Tomography (CT) images from these patients are often corrupted by undesirable metal artifacts, which causes severe problem of metal artifact. Although many methods have been proposed to reduce metal artifact, reduction is still challenging and inadequate. Some reduced results are suffering from symptom variance, second artifact, and poor subjective evaluation. To address these, we propose a novel method based on generative adversarial nets (GANs) to reduce metal artifacts. Specifically, we firstly encode interactive information (text) and imaging CT (image) to yield multi-modal feature-fusion representation, which overcomes representative ability limitation of single-modal CT images. The incorporation of interaction information constrains feature generation, which ensures symptom consistency between corrected and target CT. Then, we design an enhancement network to avoid second artifact and enhance edge as well as suppress noise. Besides, three radiology physicians are invited to evaluate the corrected CT image. Experiments show that our method gains significant improvement over other methods. Objectively, ours achieves an average increment of 7.44% PSNR and 6.12% SSIM on two medical image datasets. Subjectively, ours outperforms others in comparison in term of sharpness, resolution, invariance, and acceptability.
Liming Xu, Weisheng Li 0001, Bochuan Zheng
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Multi-Scale Adversarial Learning and Difficult Supervision for Kidney and Kidney Tumor Segmentation
Shenhai Zheng, Qiuyu Sun, Weisheng Li 0001, Laquan Li
BMVC4
2022 Detecting Generated Images by Real Images
Bo Liu 0047, Fan Yang 0159, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
ECCV (14)5
2022 Gaussian Distribution-based Mode Selection for Intra Prediction of Spatial SHVC
abstract
Due to the diversity of terminal devices, Spatial Scalable High Efficiency Video Coding (SSHVC) is an efficient solution to meet this requirement. However, its coding process is very complex, which seriously prevents its wide applications. Therefore, it is very crucial to reduce coding complexity and improve coding speed. In this paper, we propose a Gaussian Distribution-based Mode Selection for Intra Prediction of SSHVC. We show that the rate distortion costs of Inter-layer Reference (ILR) mode and Intra mode are significantly different, and both follow a Gaussian distribution. Based on this discovery, we propose to use a Bayes decision rule to determine whether ILR is the best mode so as to skip Intra mode. Experimental results demonstrate that the proposed algorithm can significantly improve coding speed with negligible coding efficiency losses.
Yu Sun 0003, Weisheng Li 0001, Xin Lu 0001, Frédéric Dufaux
ICIP4
2022 Partial-to-Partial Point Cloud Registration Based on Multi-Level Semantic-Structural Cognition
abstract
3D point cloud registration attempt to establish spatial correspondences between the source point cloud and the target point cloud. It is a fundamental task in computer vision and multimedia applications. Recently, many learning-based methods have been proposed and achieved promising performance. However, in the Partial-to-Partial (PtP) registration problem, the existence of a large number of external points may greatly handicap the effectiveness of these methods. In this paper, we propose to address the PtP issue under a novel multi-task cognition framework. At the global semantic level, we introduce a multi-scale feature exchanging network to actively evaluate the matching credibility. For the local structural learning, an inner-inter attention fusion branch is applied to generate discriminative features. Moreover, we integrate a novel alternating correspondences searching mechanism with a flexible bi-directional dislocation loss to perform robust learning, and a simple yet effective SVD weighting scheme is introduced at the inference stage. Experiment results on two challenging PtP 3D point cloud registration data sets show that our proposed method outperforms all the SOTA methods with higher precision and robustness.
Yucheng Shu, Zongzhuang Hou, Bin Xiao 0002, Xiuli Bi, Xiao Luan, Weisheng Li 0001
ICME6
2022 msFormer: Adaptive Multi-Modality 3D Transformer for Medical Image Segmentation
Jiaxin Tan, Chuangbo Jiang, Laquan Li, Weisheng Li 0001, Shenhai Zheng
PRCV (2)5
2022 L2-Norm Scaled Transformer for 3D Head and Neck Primary Tumors Segmentation in PET-CT
abstract
Head and neck (H&N) cancers are among the most common cancers worldwide (5th leading cancer by incidence). Accurate segmentation of H&N tumors can improve the early diagnosis rate of cancers for timely treatment. H&N tumor segmentation challenge is the equidensity between the tumor and surrounding tissues, which shows low contrast in CT. In contrast, PET images can reflect the distinction between the lesion region and normal tissue through metabolic activity but show low spatial resolution. With the underlying assumption that each modality contains complementary information, we introduce a novel L2-Norm Scaled Transformer (NSTR) multi-modal segmentation method in PET-CT images. The proposed network comprises the Embedding block, L2-Norm Transformer blocks, 3D Deformable down-sampling blocks, and Feature fusion module, which can fully exploit the high sensitivity of PET images to tumors and the anatomical information of CT images. Our method proposes a powerful 3D fusion network that uses a U-shaped structure to exploit complementary features of different models at multiple scales to increase the cubical representations between different modalities. We conducted a comprehensive experimental analysis on the HECKTOR PET-CT dataset. The results indicated NSTR has powerful featured representation capability and surpasses the state-of-the-art H&N tumor segmentation methods in DSC, Jaccard, RVD, and HD95. (our code will be publicly available soon).
Shenhai Zheng, Jiaxin Tan, Chuangbo Jiang, Weisheng Li 0001, Laquan Li
SMC4
2022 HessHist: A Hessian-matrix weighted histogram for image contrast enhancement
abstract
Abstract For image contrast enhancement operation, it is a keypoint to obtain more natural enhanced results and keep more details without distortion. In this paper, a novel image Hessian‐matrix weighted histogram for image contrast enhancement is proposed, which can improve the contrast of smooth regions and simultaneously restrain the contrast of texture regions. In the proposed method, the multi‐scale fractional‐order Hessian‐matrix is firstly utilized to detect and quantify the texture information of the input image, which explores the regions that should be contrasted or should be restrained. Then, the strong texture regions are suppressed by a designed suppress function. Finally, the information on unsuppressed regions and suppressed texture regions will be count by a histogram, which is termed as Hessian‐matrix weighted Histogram (HessHist) in this paper. According to HessHist, the corresponding cumulative distribution function will realize the contrast enhancement operation on the input image. For real‐time application, the integral images are introduced for fast computation of the HessHist. Experimental results show that the proposed HessHist‐based image enhancement algorithm preserves more details of input image without distortion, and is competitive with state‐of‐the‐art image enhancement algorithms in both subjective visual perception and objective evaluation metrics.
Junchao Fan, Xuyang Zong, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001
IET Image Process.6
2022 Multimodal medical image fusion based on multichannel coupled neural P systems and max-cloud models in spectral total variation domain
Guofen Wang, Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002, Jiao Du
Neurocomputing2
2022 Siamese visual tracking with multilayer feature fusion and corner distance IoU loss
Weisheng Li 0001, Junye Zhu
J. Vis. Commun. Image Represent.1
2022 DCNet: Diversity convolutional network for ventricle segmentation on short-axis cardiac magnetic resonance images
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002
Knowl. Based Syst.2
2022 Multigroup spatial shift models for thermal infrared tracking
Weisheng Li 0001, Lanbing Lv, Junye Zhu
Knowl. Based Syst.1
2022 MIA-Net: Multi-information aggregation network combining transformers and convolutional feature learning for polyp segmentation
Weisheng Li 0001, Yinghui Zhao, Linhong Wang
Knowl. Based Syst.1
2022 Recent advancement in haze removal approaches
Hira Khan, Bin Xiao 0002, Weisheng Li 0001, Muhammad Nazeer
Multim. Syst.3
2022 Semantic-refined spatial pyramid network for crowd counting
Lifang Zhou, Peiwen Wang, Weisheng Li 0001, Jiaxu Leng, Bang Jun Lei
Pattern Recognit. Lett.3
2022 Learning Unsupervised Face Normalization Through Frontal View Reconstruction
abstract
Face normalization from large pose is a challenging problem. Many Generative Adversarial Network (GAN) based models can infer frontal view of profile faces, while they require paired faces and pose label. Instead, we focus on frontal face synthesis with unpaired and unlabeled training data. We present a Frontal View Reconstruction based GAN (FVR-GAN) for large pose face normalization and recognition. The generator of FVR-GAN can be considered as a dual-input auto-encoder, where the identity encoder extracts identity features from an identity image and the template encoder extracts contour features from a frontal image. The decoder combines those two kinds of features and synthesizes a corresponding frontal face. To learn face normalization effectively, we incorporate the Frontal View Reconstruction (FVR) operation into training stage. The FVR operation includes self-reconstruction and frontalization mapping. A group of sub-discriminators which receive different facial parts are employed for discrimination. Considering that different face parts have different contributions to discrimination, we introduce a dynamic weighting mechanism to balance the output of sub-discriminators. FVR-GAN can recover high-quality frontal images under arbitrary poses. Experimental results on datasets of Multi-PIE, IJB-A, LFW and CFP demonstrate the efficacy of our model in terms of quality of synthesized images and recognition accuracy.
Xiao Luan, Jiezhong Zheng, Weisheng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 PAM-DenseNet: A Deep Convolutional Neural Network for Computer-Aided COVID-19 Diagnosis
abstract
Currently, several convolutional neural network (CNN)-based methods have been proposed for computer-aided COVID-19 diagnosis based on lung computed tomography (CT) scans. However, the lesions of pneumonia in CT scans have wide variations in appearances, sizes, and locations in the lung regions, and the manifestations of COVID-19 in CT scans are also similar to other types of viral pneumonia, which hinders the further improvement of CNN-based methods. Delineating infection regions manually is a solution to this issue, while excessive workload of physicians during the epidemic makes it difficult for manual delineation. In this article, we propose a CNN called dense connectivity network with parallel attention module (PAM-DenseNet), which can perform well on coarse labels without manually delineated infection regions. The parallel attention module automatically learns to strengthen informative features from both channelwise and spatialwise simultaneously, which can make the network pay more attention to the infection regions without any manual delineation. The dense connectivity structure performs feature maps reuse by introducing direct connections from previous layers to all subsequent layers, which can extract representative features from fewer CT slices. The proposed network is first trained on 3530 lung CT slices selected from 382 COVID-19 lung CT scans, 372 lung CT scans infected by other pneumonia, and 200 normal lung CT scans to obtain a pretrained model for slicewise prediction. We then apply this pretrained model to a CT scans dataset containing 94 COVID-19 CT scans, 93 other pneumonia CT scans, and 93 normal lung scans, and achieve patientwise prediction through a voting mechanism. The experimental results show that the proposed network achieves promising results with an accuracy of 94.29%, a precision of 93.75%, a sensitivity of 95.74%, and a specificity of 96.77%, which is comparable to the methods that are based on manually delineated infection regions.
Bin Xiao 0002, Xiaoming Qiu, Guoyin Wang 0001, Wenbing Zeng, Weisheng Li 0001, Yongjian Nian
IEEE Trans. Cybern.7
2022 NLRNet: An Efficient Nonlocal Attention ResNet for Pansharpening
abstract
Remote sensing images often contain many similar components, such as buildings, roads, and water surfaces, which have similar spectra and spatial structures. Although convolutional neural networks (CNNs) based on residual learning can provide excellent performance in pansharpening, the existing methods do not make full use of intrinsically similar information in images. Moreover, since the convolution operation is focused on the local region, even in a deep network, position-independent global information is difficult to obtain. In this article, an efficient nonlocal attention residual network (NLRNet) is proposed to capture the similar contextual dependencies of all pixels. Specifically, to reduce the difficulty of network training caused by the original nonlocal attention, we propose an efficient nonlocal attention (ENLA) mechanism and employ residual with zero initialization (ReZero) technology to make the signal easy to spread through the network. Furthermore, a spectral aggregation module (SpecAM) is proposed to generate fused images and adjust the corresponding spectral information. The experimental results for the QuickBird and WorldView3 data sets show that the proposed method is competitive with other advanced methods based on quality assessment and visual perception.
Dajiang Lei, Liping Zhang 0012, Weisheng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 MCANet: A Multidimensional Channel Attention Residual Neural Network for Pansharpening
abstract
In the remote sensing image fusion field, fusion methods based on deep learning are the latest techniques in panchromatic sharpening (pansharpening). However, existing pansharpening methods based on neural networks cannot adequately inject the spatial feature information of panchromatic (PAN) images into fusion images, and they do not exploit the feature relationships between spatial locations, such as rows and columns of feature maps. To solve these problems, a multidimensional channel attention residual neural network (MCANet) is proposed in this paper. To preserve the structural information in PAN images, a two-stream detail injection (TSDI) module is proposed, and the local skip connection operation is adopted to mine more spectral and structural information. A multidimensional channel attention (MCA) module is also designed to enable the network to learn the nonlinear mapping relationships between image spatial locations. In addition, a multiscale feature fusion (MSFF) module is designed to improve feature representation in the image fusion process, which is conducive to improving the pansharpening effect. The experimental results on the WorldView-2, GaoFen-2 and QuickBird datasets demonstrate that the proposed method outperforms state-of-the-art methods both visually and quantitatively.
Dajiang Lei, Liping Zhang 0012, Weisheng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Multibranch Feature Extraction and Feature Multiplexing Network for Pansharpening
abstract
With the continuous development of deep neural networks in the visual field, their application to panchromatic sharpening has received increasing attention from researchers; however, the existing panchromatic sharpening methods generally lack the ability to combine knowledge of the panchromatic sharpening field for feature extraction with neural networks, which have certain limitations in feature extraction and the discovery of new features. This article proposes a simple, modular, multibranched feature extraction and reuses network architecture designed not only to support feature reuse but to learn well-expressed new features for use in panchromatic sharpening approaches. In addition, we fused field knowledge of panchromatic sharpening to extract spatial structure information of panchromatic maps through gradient calculators and design structural and spectral compensation to fully extract and preserve the spatial structural and spectral information of images. We conducted experiments on the QuickBird and WorldView-3 satellite data sets, and the experimental results reveal that our proposed method has advantages over the best methods currently available, achieving excellent results not only on objective evaluation metrics, such as full-reference and no-reference metrics, but also on subjective visual evaluation.
Dajiang Lei, Liping Zhang 0012, Weisheng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 MHANet: A Multiscale Hierarchical Pansharpening Method With Adaptive Optimization
abstract
In recent years, the powerful nonlinear modeling capability of convolutional neural networks has led to an increasing number of researchers focusing on deep learning-based pansharpening methods. However, due to the diversity of remote sensing image features and the limitations of the convolution operation, the existing methods are still inadequate in restoring the spatial details of complex remote sensing scenes. Therefore, in this paper, we propose a simple and effective network for pansharpening methods. Specifically, in our hierarchical feature integration architecture, a multi-scale grouping dilated block is designed to adequately capture fine-grained representations of multi-level scale features. At the same time, we propose a spatially self-attention block to adaptively improve the feature extraction process by establishing associations between features. The above blocks are connected in a hierarchical design, with selective reuse of features between layers, and a good ability to explore new levels of features while reusing low-level features. Our experiments with the GaoFen-2 satellite dataset, WorldView-2 satellite dataset, and WorldView-3 satellite dataset show that our proposed method is highly competitive with existing excellent methods in both objective indicator evaluation and subjective visual evaluation.
Dajiang Lei, Liping Zhang 0012, Weisheng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Spatiotemporal Reflectance Fusion via Tensor Sparse Representation
abstract
Tradeoffs between the spatial and temporal resolutions of current satellite instruments limit our ability to conduct high-quality and continuous monitoring of the earth’s surface dynamics. Spatiotemporal image fusion has become increasingly necessary to obtain remote sensing images with high spatiotemporal resolution. However, current learning-based methods concentrate on predicting images only from spatial similarity and neglect spectral correlations of remote sensing images, leading to significant spectral information loss. In this article, we develop a novel nonlocal tensor sparse representation-based semicoupled dictionary learning approach (SCDNTSR) for spatiotemporal fusion. In the SCDNTSR method, the spectral correlation and the spatial similarity of the nonlocal similar cubes are simultaneously exploited through the tensor–tensor product-based tensor sparse representation. Furthermore, the semicoupled mapping prior knowledge of sparse coefficients across the high- and low-spatial resolution (HSR\LSR) image spaces is exploited with the coupled dictionary to constrain the similarity of sparse coefficients to improve the prediction performance. In addition, to capture additional prior spatial information, the SCDNTSR provides a new method to determine the degradation relationship between the target HSR and LSR difference images with the help of the known HSR and LSR difference images. The proposed SCDNTSR method was tested on real datasets at both the Coleambally Irrigation Area study site and the Lower Gwydir Catchment study site. Results show that the proposed method outperforms five state-of-the-art methods, especially in maintaining the spectral information, proving the feasibility of integrating the degradation relationship, spatio-spectral-nonlocal correlation, and semicoupled mapping priors of the multisource data into the proposed model.
Yidong Peng, Weisheng Li 0001, Xiaobo Luo, Jiao Du, Xiayan Zhang, Yi Gan, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Registration-Is-Evaluation: Robust Point Set Matching With Multigranular Prior Assessment
abstract
Point set registration is one of the challenging tasks in remote sensing image processing and analysis. Its critical step is to find the corresponding relationships between the fixed scene point set and the moving model point set that undergo different sorts of transformations. Existing algorithms primarily utilize different types of prior information to improve the registration performance, such as spatial consistency, local similarity, and uniform outliers. However, due to the lack of active evaluation on the prior and intermediate information during the registration process, these strategies are susceptible to large data transformations. In order to enhance the robustness and accuracy for point set registration, we propose in this article a novel framework, namely Registration-is-Evaluation (RisE). Based on a multigranular probability model, our method exploits and utilizes both prior and posterior information to dynamically evaluate the matching status. What is more, instead of adding an extra uniform prior, we unified the outliers, noise, missing points, and heavily warped points into our registration evaluation model and address them simultaneously. We also apply a novel point set descriptor, called local polar relative geometry (LPRG), to have a more robust local similarity measurement. It adopts the local polar coordinate to perform multiscale pooling and relative geometric computation. Based on our proposed method, the matching relationships and the spatial transformations can be actively evaluated to provide useful contextual guidance for the registration process. Experimental results on multiple data sets show that our algorithm outperforms the state-of-the-art methods, in terms of both accuracy and robustness under large data degradations.
Yucheng Shu, Zhenlong Liao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Privacy-Preserving Color Image Feature Extraction by Quaternion Discrete Orthogonal Moments
abstract
To implement image storage and computation in cloud servers without violating users’ privacy, privacy-preserving feature extraction has been a new research interest. The existing works are mainly designed for grayscale images. For color images, they tend to convert them to grayscale images or obtain the results of the combination of single-channel processes. While the capabilities of features extracted from the encrypted color images will be affected if color information and inter-relationship between color channels are ignored. To fully preserve features of color images, we introduce quaternion theory to encode each color image and propose an improved vector homomorphic encryption scheme (IVHE) to encrypt quaternion-based color images. IVHE helps protect image content and keep vector characteristics of color images. Based on IVHE, the framework for feature extraction of privacy-preserving Quaternion Discrete Orthogonal Moments (PPQDOMs) is presented. Theoretical analyses prove that Quaternion Discrete Orthogonal Moments (QDOMs) can be extracted from the encrypted color images by PPQDOMs. Furthermore, we apply three common Discrete Orthogonal Moments to the proposed framework to evaluate its performance. Experimental results demonstrate that the proposed framework can protect color image content and perform well compared to QDOMs in image reconstruction and image recognition.
Xiuli Bi, Chao Shuai, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.5
2022 Multi-Manifold Deep Discriminative Cross-Modal Hashing for Medical Image Retrieval
abstract
Benefitting from the low storage cost and high retrieval efficiency, hash learning has become a widely used retrieval technology to approximate nearest neighbors. Within it, the cross-modal medical hashing has attracted an increasing attention in facilitating efficiently clinical decision. However, there are still two main challenges in weak multi-manifold structure perseveration across multiple modalities and weak discriminability of hash code. Specifically, existing cross-modal hashing methods focus on pairwise relations within two modalities, and ignore underlying multi-manifold structures across over 2 modalities. Then, there is little consideration about discriminability, i.e., any pair of hash codes should be different. In this paper, we propose a novel hashing method named multi-manifold deep discriminative cross-modal hashing (MDDCH) for large-scale medical image retrieval. The key point is multi-modal manifold similarity which integrates multiple sub-manifolds defined on heterogeneous data to preserve correlation among instances, and it can be measured by three-step connection on corresponding hetero-manifold. Then, we propose discriminative item to make each hash code encoded by hash functions be different, which improves discriminative performance of hash code. Besides, we introduce Gaussian-binary Restricted Boltzmann Machine to directly output hash codes without using any continuous relaxation. Experiments on three benchmark datasets (AIBL, Brain and SPLP) show that our proposed MDDCH achieves comparative performance to recent state-of-the-art hashing methods. Additionally, diagnostic evaluation from professional physicians shows that all the retrieved medical images describe the same object and illness as the queried image.
Liming Xu, Bochuan Zheng, Weisheng Li 0001
IEEE Trans. Image Process.4
2022 A Novel Framework With Weighted Decision Map Based on Convolutional Neural Network for Cardiac MR Segmentation
abstract
For diagnosing cardiovascular disease, an accurate segmentation method is needed. There are several unresolved issues in the complex field of cardiac magnetic resonance imaging, some of which have been partially addressed by using deep neural networks. To solve two problems of over-segmentation and under-segmentation of anatomical shapes in the short-axis view from different cardiac magnetic resonance sequences, we propose a novel two-stage framework with a weighted decision map based on convolutional neural networks to segment the myocardium (Myo), left ventricle (LV), and right ventricle (RV) simultaneously. The framework comprises a decision map extractor and a cardiac segmenter. A cascaded U-Net++ is used as a decision map extractor to acquire the decision map that decides the category of each pixel. Cardiac segmenter is a multiscale dual-path feature aggregation network (MDFA-Net) which consists of a densely connected network and an asymmetric encoding and decoding network. The input to the cardiac segmenter is derived from processed original images weighted by the output of the decision map extractor. We conducted experiments on two datasets of multi-sequence cardiac magnetic resonance segmentation challenge 2019 (MS-CMRSeg 2019) and myocardial pathology segmentation challenge 2020 (MyoPS 2020). Test results obtained on MyoPS 2020 show that the average Dice coefficients of the proposed method on the segmentation tasks of Myo, LV and RV are 84.70%, 86.00%, and 86.31%, respectively.
Weisheng Li 0001, Xinbo Gao 0001, Bin Xiao 0002
IEEE J. Biomed. Health Informatics2
2022 Rubik-Net: Learning Spatial Information via Rotation-Driven Convolutions for Brain Segmentation
abstract
The accurate segmentation of brain tissue in Magnetic Resonance Image (MRI) slices is essential for assessing neurological conditions and brain diseases. However, it is challenging to segment MRI slices because of the low contrast between different brain tissues and the partial volume effect. 2-Dimensional (2-D) convolutional networks cannot handle such volumetric image data well because they overlook spatial information between MRI slices. Although 3-Dimensional (3-D) convolutions capture volumetric spatial information, they have not been fully exploited to enhance representative ability of deep networks; moreover, they may lead to overfitting with insufficient training data. In this paper, we propose a novel convolutional mechanism, termed Rubik convolution, to capture multi-dimensional information between MRI slices. Rubik convolution rotates the axis of a set of consecutive slices, enabling 2-D convolution kernels to extract features of each axial plane simultaneously. Next, feature maps are rotated back to fuse multidimensional information by the Max-View-Maps. Furthermore, we propose an efficient 2-D convolutional network, namely Rubik-Net, where the residual connections and the bottleneck structure are used to enhance information transmission and reduce the number of network parameters. The Rubik-Net shows promising results on iSeg2017, iSeg2019, IBSR and BrainWeb datasets in terms of segmentation accuracy. In particular, we achieved the best results in 95th percentile Hausdorff distance and average surface distance in cerebrospinal fluid segmentation on the most challenging iSeg2019 dataset. The experiments indicate that Rubik-Net improves the accuracy and efficiency of medical image segmentation. Moreover, Rubik convolution can be easily embedded into existing 2-D convolutional networks.
Xiao Luan, Xinyu zheng, Weisheng Li 0001, Linghui Liu, Yucheng Shu
IEEE J. Biomed. Health Informatics3
2022 Medical Image Fusion and Denoising Algorithm Based on a Decomposition Model of Hybrid Variation-Sparse Representation
abstract
Medical image fusion technology integrates the contents of medical images of different modalities, thereby assisting users of medical images to better understand their meaning. However, the fusion of medical images corrupted by noise remains a challenge. To solve the existing problems in medical image fusion and denoising algorithms related to excessive blur, unclean denoising, gradient information loss, and color distortion, a novel medical image fusion and denoising algorithm is proposed. First, a new image layer decomposition model based on hybrid variation-sparse representation and weighted Schatten p-norm is proposed. The alternating direction method of multipliers is used to update the structure, detail layer dictionary, and detail layer coefficient map of the input image while denoising. Subsequently, appropriate fusion rules are employed for the structure layers and detail layer coefficient maps. Finally, the fused image is restored using the fused structure layer, detail layer dictionary, and detail layer coefficient maps. A large number of experiments confirm the superiority of the proposed algorithm over other algorithms. The proposed medical image fusion and denoising algorithm can effectively remove noise while retaining the gradient information without color distortion.
Guofen Wang, Weisheng Li 0001, Jiao Du, Bin Xiao 0002, Xinbo Gao 0001
IEEE J. Biomed. Health Informatics2
2022 Collaborative Learning With a Multi-Branch Framework for Feature Enhancement
abstract
Feature representation is highly important for many computer vision tasks. A broad range of prior studies have been proposed to strengthen representation ability of architectures via built-in blocks. However, during the forward propagation, the reduction in feature map scales still leads to the lack of representation ability. In this paper, we focus on boosting the representational power of a convolutional network by the multi-branch framework that we term the BranchNet. Each branch is directly supervised by label information to enrich the hierarchy features in BranchNet. Based on this framework, we further propose a collaborative learning loss and a soft target loss to transfer knowledge from deeper layers to shallow layers. BranchNet is an efficient training framework without extra parameters introduced in inference and can be integrated in existing networks, e.g., VGG, ResNet, and DenseNet. We evaluate BranchNet on all of these models and find that our method outperforms the baseline models on the widely-used CIFAR and ImageNet datasets. In particular, on the CIFAR-100 dataset, the classification error of ResNet-164 with BranchNet decreases by 4.51 percent. We also conduct experiments on the representative computer vision tasks of instance segmentation and class activation mapping, further verifying the superiority of BranchNet over the baseline models. Models and code are available athttps://github.com/zyyupup/BranchNet/.
Xiao Luan, Weihua Ou, Linghui Liu, Weisheng Li 0001, Yucheng Shu, Hongmin Geng
IEEE Trans. Multim.5
2022 IDHashGAN: Deep Hashing With Generative Adversarial Nets for Incomplete Data Retrieval
abstract
Benefiting from low storage costs and high retrieval efficiency, hash learning has been a widely adopted technology for approximating nearest neighbor in large-scale data retrieval. Deep learning to hash greatly improves image retrieval performance by integrating feature learning and hash coding into an end-to-end framework. However, subject to application scope, most existing deep hashing methods only apply to retrieval of complete data and have undesirable results when retrieving incomplete but valuable data. In this paper we propose IDHashGAN, a novel deep hashing model with generative adversarial networks to retrieve incomplete data, in which feature restoration, feature learning and hash coding are integrated into an unified end-to-end framework. The proposed model consists of four key components: (1) reconstructive and generative loss are used to generate continuous feature of incomplete data in generative network; (2) supervised manifold similarity is proposed to improve retrieval accuracy and obtain good user acceptance; (3) adversarial and classified loss are designed to distinguish authenticity and similarity in discriminative network; and (4) encoding and quantization loss are adopted to preserve similarity and control hash quality. Extensive experiments on benchmark datasets show that IDHashGAN is competitive on complete dataset and yields substantial boosts of 70% on incomplete datasets compared to state-of-the-art hashing methods.
Liming Xu, Weisheng Li 0001, Ling Bai
IEEE Trans. Multim.3
2021 Multi-Task Wavelet Corrected Network for Image Splicing Forgery Detection and Localization
abstract
Although the existing image splicing forgery detection networks can achieve a promising performance, most of these networks utilize regular pooling operations (max-pooling and mean-pooling) and a single task strategy, which limits the comprehensiveness and representativeness of the features learned by the networks. In this paper, we propose a multi-task wavelet corrected network (MWC-Net) that can learn more comprehensive and representative features for image splicing forgery detection and localization. MWC-Net exploits wavelet-pooling and wavelet un-pooling to compress and reconstruct the features of splicing forgery images, which can reduce information loss during learning features. Mean-while, MWC-Net implements a multi-task strategy to improve its ability to learn and utilize more comprehensive and representative features. The experimental results demonstrate that MWC-Net outperforms the state-of-the-art methods in splicing forgery detection and localization on four public datasets.
Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001
ICME5
2021 Poolingformer: Long Document Modeling with Pooling Attention
abstract
In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs a larger window to increase receptive fields with pooling attention to reduce both computational cost and memory consumption. We first evaluate Poolingformer on two long sequence QA tasks: the monolingual NQ and the multilingual TyDi QA. Experimental results show that Poolingformer sits atop three official leaderboards measured by F1, outperforming previous state-of-the-art models by 1.9 points (79.8 vs. 77.9) on NQ long answer, 1.9 points (79.5 vs. 77.6) on TyDi QA passage answer, and 1.6 points (67.6 vs. 66.0) on TyDi QA minimal answer. We further evaluate Poolingformer on a long sequence summarization task. Experimental results on the arXiv benchmark continue to demonstrate its superior performance.
Hang Zhang 0029, Yeyun Gong, Yelong Shen, Weisheng Li 0001, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen
ICML4
2021 Medical Image Registration Based on Uncoupled Learning and Accumulative Enhancement
Yucheng Shu, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001
MICCAI (4)5
2021 Neural ordinary differential grey model and its applications
Dajiang Lei, Kaili Wu, Liping Zhang 0012, Weisheng Li 0001, Qun Liu 0005
Expert Syst. Appl.4
2021 Remote sensing image super-resolution using cascade generative adversarial nets
Dongen Guo, Liming Xu, Weisheng Li 0001, Xiaobo Luo
Neurocomputing4
2021 A focus measure in discrete cosine transform domain for multi-focus image fast fusion
Xixi Nie, Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001
Neurocomputing4
2021 A contour-aware feature-merged network for liver segmentation based on shape prior knowledge
Lifang Zhou, Xueyuan Deng, Weisheng Li 0001, Shenhai Zheng, Bang Jun Lei
Neurocomputing3
2021 IoU-guided Siamese region proposal network for real-time visual tracking
Lifang Zhou, Weisheng Li 0001, Jian-Xun Mi, Bang Jun Lei
Neurocomputing3
2021 A Lightweight SE-YOLOv3 Network for Multi-Scale Object Detection in Remote Sensing Imagery
abstract
Current state-of-the-art detectors achieved impressive performance in detection accuracy with the use of deep learning. However, most of such detectors cannot detect objects in real time due to heavy computational cost, which limits their wide application. Although some one-stage detectors are designed to accelerate the detection speed, it is still not satisfied for task in high-resolution remote sensing images. To address this problem, a lightweight one-stage approach based on YOLOv3 is proposed in this paper, which is named Squeeze-and-Excitation YOLOv3 (SE-YOLOv3). The proposed algorithm maintains high efficiency and effectiveness simultaneously. With an aim to reduce the number of parameters and increase the ability of feature description, two customized modules, lightweight feature extraction and attention-aware feature augmentation, are embedded by utilizing global information and suppressing redundancy features, respectively. To meet the scale invariance, a spatial pyramid pooling method is used to aggregate local features. The evaluation experiments on two remote sensing image data sets, DOTA and NWPU VHR-10, reveal that the proposed approach achieves more competitive detection effect with less computational consumption.
Lifang Zhou, Guang Deng, Weisheng Li 0001, Jian-Xun Mi, Bang Jun Lei
Int. J. Pattern Recognit. Artif. Intell.3
2021 DSAGAN: A generative adversarial network based on dual-stream attention mechanism for anatomical and functional image fusion
Jun Fu 0004, Weisheng Li 0001, Jiao Du, Liming Xu
Inf. Sci.2
2021 MDFA-Net: Multiscale dual-path feature aggregation network for cardiac segmentation on multi-sequence cardiac MR
Weisheng Li 0001, Sheng Qin, Linhong Wang
Knowl. Based Syst.2
2021 Medical image segmentation based on active fusion-transduction of multi-stream features
Yucheng Shu, Bin Xiao 0002, Weisheng Li 0001
Knowl. Based Syst.4
2021 Selective Domain-Invariant Feature Alignment Network for Face Anti-Spoofing
abstract
One primary challenge in face anti-spoofing refers to suffering a sharp performance drop in cross-domain scenes, where training and testing images are collected from different datasets. Recent methods have achieved promising results by aligning the features of all images among the available source domains. However, due to significant distribution discrepancies among non-face regions of all images, it is challenging to capture domain-invariant features for these regions. In this paper, we propose a novel Selective Domain-invariant Feature Alignment Network (SDFANet) for cross-domain face anti-spoofing, which aims to seek common feature representations by fully exploring the generalization of different regions of images. Different from previous works that align the whole features directly, the proposed SDFANet leverages multiple domain discriminators with the same architecture to balance the generalization of different regions of the all images. Specifically, we firstly design a multi-grained feature alignment network composed of a local-region and global-image alignment subnetworks to learn more generalized feature space for real faces. Besides, the domain adapter module, which aims to alleviate the large domain discrepancy with the help of the domain attention strategy, is adopted to facilitate the learning of our multi-grained feature alignment network. In addition, a multi-scale attention fusion module is designed in our feature generator to refine the different levels of features effectively. Experimental results show that the proposed SDFANet can greatly improve the generalization ability of face anti-spoofing, and that is superior to the existing methods.
Lifang Zhou, Xinbo Gao 0001, Weisheng Li 0001, Bang Jun Lei, Jiaxu Leng
IEEE Trans. Inf. Forensics Secur.4
2021 2D-LCoLBP: A Learning Two-Dimensional Co-Occurrence Local Binary Pattern for Image Recognition
abstract
The rotation, scale and translation invariance of extracted features have a high significance in image recognition. Local binary pattern (LBP) and LBP-based descriptors have been widely used in image recognition due to feature discrimination and computational efficiency. However, most of the existing LBP-based descriptors have been designed to achieve rotation invariance while fail to achieve scale invariance. Moreover, it is usually difficult to achieve a good trade-off between the feature discrimination and the feature dimension. In this work, a learning 2D co-occurrence LBP termed 2D-LCoLBP is proposed to address these issues. Firstly, a weighted joint histogram is constructed in different neighborhoods and scales of an image to represent the multi-neighborhood and multi-scale LBP (2D-MLBP) and achieve the rotation invariance. A feature learning strategy is then designed to learn the compact and robust descriptor (2D-LCoLBP) from LBP pattern pairs across different scales in the extracted 2D-MLBP to characterize the most stable local structures and achieve the scale invariance, as well as decrease the feature dimension and improve the noise robustness. Finally, a linear SVM classifier is employed for recognition. We applied the proposed 2D-LCoLBP on four image recognition tasks-texture, object, face and food recognition with ten image databases. Experimental results show that 2D-LCoLBP has obviously low feature dimension but outperforms the state-of-the-art LBP-based descriptors in terms of recognition accuracy under noise-free, Gaussian noise and JPEG compression conditions.
Xiuli Bi, Yuan Yuan 0015, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2021 Global-Feature Encoding U-Net (GEU-Net) for Multi-Focus Image Fusion
abstract
The convolutional neural network (CNN)-based multi-focus image fusion methods which learn the focus map from the source images have greatly enhanced fusion performance compared with the traditional methods. However, these methods have not yet reached a satisfactory fusion result, since the convolution operation pays too much attention on the local region and generating the focus map as a local classification (classify each pixel into focus or de-focus classes) problem. In this article, a global-feature encoding U-Net (GEU-Net) is proposed for multi-focus image fusion. In the proposed GEU-Net, the U-Net network is employed for treating the generation of focus map as a global two-class segmentation task, which segments the focused and defocused regions from a global view. For improving the global feature encoding capabilities of U-Net, the global feature pyramid extraction module (GFPE) and global attention connection upsample module (GACU) are introduced to effectively extract and utilize the global semantic and edge information. The perceptual loss is added to the loss function, and a large-scale dataset is constructed for boosting the performance of GEU-Net. Experimental results show that the proposed GEU-Net can achieve superior fusion performance than some state-of-the-art methods in both human visual quality, objective assessment and network complexity.
Bin Xiao 0002, Bocheng Xu, Xiuli Bi, Weisheng Li 0001
IEEE Trans. Image Process.4
2020 Global-Local Mutual Guided Learning for Person Re-identification
Junheng Chen, Xiao Luan, Weisheng Li 0001
PRCV (2)3
2020 Editorial: Deep learning for medical image analysis
Ke Lu 0002, Fei Wang 0001, Ling Shao 0001, Weisheng Li 0001
Neurocomputing4
2020 Multi-granularity generative adversarial nets with reconstructive sampling for image inpainting
Liming Xu, Weisheng Li 0001
Neurocomputing3
2020 Follow the Sound of Children's Heart: A Deep-Learning-Based Computer-Aided Pediatric CHDs Diagnosis System
abstract
Auscultation of heart sounds is a noninvasive and less costly way for congenital heart disease (CHD) diagnosis, especially for pediatric individuals. The deep-learning-based computer-aided heart sound analysis has been widely studied and developed in recent years. In this article, we develop a deep-learning-based computer-aided system for pediatric CHDs diagnosis using two novel lightweight convolution neural networks (CNNs). One key issue of most existing deep-learning-based systems is the scarcity of large-scale data sets for CNN learning. To this end, we collect heart sounds from newborns and children with physicians' annotations to construct a pediatric heart sound data set that contains 528 high-quality recordings (nearly 4 h in total) from 137 subjects. With the constructed data set, deep CNN models can be easily trained as classifiers in computer-aided CHDs diagnosis systems. The experimental results demonstrate the superiority of our proposed methods in terms of diagnosis performance and parameter consumption in the application of Internet of Things.
Bin Xiao 0002, Yunqiu Xu, Xiuli Bi, Weisheng Li 0001, Zhuo Ma 0001
IEEE Internet Things J.4
2020 Three-layer medical image fusion with tensor-based features
Jiao Du, Weisheng Li 0001, Hengliang Tan
Inf. Sci.2
2020 Fractional discrete Tchebyshev moments and their applications in image encryption and watermarking
Bin Xiao 0002, Jiangxia Luo, Xiuli Bi, Weisheng Li 0001, Beijing Chen
Inf. Sci.4
2020 Image splicing forgery detection combining coarse to refined convolutional neural network and adaptive clustering
Bin Xiao 0002, Yang Wei 0002, Xiuli Bi, Weisheng Li 0001, Jianfeng Ma 0001
Inf. Sci.4
2020 A generative adversarial network with structural enhancement and spectral supplement for pan-sharpening
Liping Zhang 0012, Weisheng Li 0001, Dajiang Lei
Neural Comput. Appl.2
2020 BPGAN: Bidirectional CT-to-MRI prediction using multi-generative multi-adversarial nets with spectral normalization and localization
Liming Xu, Weisheng Li 0001, Jianbo Lei
Neural Networks4
2020 Fast Depth and Mode Decision in Intra Prediction for Quality SHVC
abstract
Scalable High Efficiency Video Coding (SHVC) is the extension of High Efficiency Video Coding (HEVC). In intra prediction for quality SHVC, a Coding Unit (CU) is recursively divided into a quadtree-based structure from the largest 64×64 CU to the smallest 8×8 CU, in which 35 intra prediction modes and Inter-Layer Reference (ILR) mode are checked to determine the best possible mode. This leads to very high coding efficiency but also results in an extremely high coding complexity. To improve coding speed while maintaining coding efficiency, in this paper, we propose a new efficient algorithm for fast intra prediction for enhancement layer in SHVC. First, temporal and spatial correlations, as well as their correlation degrees, are combined in a Naive Bayes classifier to predict depth probabilities and skip depths with low likelihood. Second, for a given depth candidate, we combine ILR mode probability with Partial Zero Blocks (PZBs) based on the Sum of Squared Differences (SSD) to determine whether the ILR mode is the best one. In that case, we can skip intra prediction, which requires very high complexity. Third, initial Intra Modes (IMs) are obtained through Sobel operator, and are combined with the relationship between IMs and their corresponding Hadamard Cost (HC) values to predict candidate IMs in Rough Mode Decision (RMD). Then, an analytical criterion of early termination is developed based on the HC values of two neighboring IMs in the Rate-Distortion Optimization (RDO) process. Finally, we combine depth probabilities and the distribution of residual coefficients at the current depth to early terminate depth selection. The proposed scheme can significantly decrease the complexity of depth determination while reducing the complexity of mode decision for a depth candidate. Our experimental results demonstrate that the proposed scheme can achieve a speed up gain of more than 80% in average, while maintaining coding efficiency.
Yu Sun 0003, Ce Zhu, Weisheng Li 0001, Frédéric Dufaux, Jiangtao Luo
IEEE Trans. Image Process.4
2020 Three-Layer Image Representation by an Enhanced Illumination-Based Image Fusion Method
abstract
The recently developed multiscale-based fusion methods can be improved with two approaches: an advanced image decomposition scheme and an advanced fusion rule. In this paper, three-layer image decomposition, enhanced illumination fusion rule-based method is proposed. The proposed method includes three steps. First, each input image is decomposed into its corresponding smooth, texture, and edge layers using defined local extrema and low-pass filters in the spatial domain. Second, three different strategies are applied as fusion rules for the three-layer representation. To preserve the illumination closely related to tumors, the illumination is corrected by applying a higher contrast to the decomposed image details, including the texture and edge inputs, such as those found in grayscale CT and MRI images. The final fused image is created by the addition of the normalized smooth, texture, and edge image layers. The experiments demonstrate that the proposed method performs better than the existing state-of-the-art fusion methods.
Jiao Du, Weisheng Li 0001, Hengliang Tan
IEEE J. Biomed. Health Informatics2
2020 Fast Depth and Inter Mode Prediction for Quality Scalable High Efficiency Video Coding
abstract
The scalable high efficiency video coding (SHVC) is an extension of high efficiency video coding (HEVC). It introduces multiple layers and inter-layer prediction, thus significantly increases the coding complexity on top of the already complicated HEVC encoder. In inter prediction for quality SHVC, in order to determine the best possible mode at each depth level, a coding tree unit can be recursively split into four depth levels, including merge mode, inter2N×2N, inter2N×N, interN×2N, interN×N, inter2N×nU, inter2N×nD, internL×2N and internRx×2N, intra modes and inter-layer reference (ILR) mode. This can obtain the highest coding efficiency, but also result in very high coding complexity. Therefore, it is crucial to improve coding speed while maintaining coding efficiency. In this research, we have proposed a new depth level and inter mode prediction algorithm for quality SHVC. First, the depth level candidates are predicted based on inter-layer correlation, spatial correlation and its correlation degree. Second, for a given depth candidate, we divide mode prediction into square and non-square mode predictions respectively. Third, in the square mode prediction, ILR and merge modes are predicted according to depth correlation, and early terminated whether residual distribution follows a Gaussian distribution. Moreover, ILR mode, merge mode and inter2N×2N are early terminated based on significant differences in Rate Distortion (RD) costs. Fourth, if the early termination condition cannot be satisfied, non-square modes are further predicted based on significant differences in expected values of residual coefficients. Finally, inter-layer and spatial correlations are combined with residual distribution to examine whether to early terminate depth selection. Experimental results have demonstrated that, on average, the proposed algorithm can achieve a time saving of 71.14%, with a bit rate increase of 1.27%.
Yu Sun 0003, Ce Zhu, Weisheng Li 0001, Frédéric Dufaux
IEEE Trans. Multim.4
2020 Multi-Focus Image Fusion by Hessian Matrix Based Decomposition
abstract
In this paper, a Hessian matrix based multi-focus image fusion method is proposed. First, the integral map is introduced for fast compute the Hessian matrix of source images at different scales, and the multi-scale Hessian matrix of source image is obtained. Second, the multi-scale Hessian matrix is used to decompose each source image into two kinds of regions: the feature and background regions. In order to improve the fusion performance, two new focus measures based on the multi-scale Hessian matrix and two different fusion strategies for both feature and background regions are utilized to obtain the initial decision maps, respectively. Finally, the final decision map for image fusion is achieved by post-processing on the results of the previous step. The proposed method is a primary attempt to introduce image feature and background regions decomposition strategies in the field of multi-focus image fusion. The experimental results also show that our method outperforms the existing image fusion methods in both visual perception and objective evaluations.
Bin Xiao 0002, Ge Ou, Xiuli Bi, Weisheng Li 0001
IEEE Trans. Multim.5
2019 Fast Inter Mode Predictions for SHVC
abstract
The Scalable High Efficiency Video Coding (SHVC) has very high coding efficiency, but its computational complexity is also very high. This definitely limits its wide applications, particularly for real-time video applications. Therefore, it is crucial to improve the coding speed. In this research, we have proposed a new inter mode prediction algorithm for quality SHVC, in order to improve the coding speed while maintaining coding efficiency. First, we divide mode prediction into square mode prediction and non-square mode prediction. Second, in the square mode prediction, Inter-Layer Reference (ILR) and merge modes are predicted based on depth correlation. Moreover, ILR mode, merge mode and inter 2N×2N are early terminated based on Rate Distortion (RD) cost. Third, if the early termination condition cannot be satisfied, nonsquare modes are further predicted based on the distribution of residual coef-ficients. Experimental results have demonstrated that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss.
Yu Sun 0003, Weisheng Li 0001, Ce Zhu, Frédéric Dufaux
ICME3
2019 LVC-Net: Medical Image Segmentation with Noisy Label Based on Local Visual Cues
Yucheng Shu, Weisheng Li 0001
MICCAI (6)3
2019 Automatic Segmentation of Liver from CT Scans with CCP-TSPM Algorithm
abstract
With the increase in the morbidity of liver cancer and its high mortality rate, liver segmentation in abdominal computed tomography (CT) scan images has received extensive attention. Segmentation results play an important role in computer-assisted diagnosis and therapy. However, it remains a challenging task because of the complexity of the liver’s anatomy, low contrast between the liver and its adjacent organs, and presence of lesions. This study presents an automatic method for liver segmentation from CT scan images based on the convex–concave point for tree structured part model (CCP–TSPM). First, TSPM is utilized as a coarse segmentation tool for capturing the topological shape variation. Then, the proposed CCP is implemented to adjust the position between adjacent points dynamically. As a result, the CCP–TSPM can locate the liver boundary adaptively. Furthermore, color space data provide abundant feature information, which can further improve the method’s effectiveness and efficiency. Finally, the curve is evolved by an iteration level set function to obtain the fine segmentation results. The experimental results show that the proposed method can extract the liver boundary successfully. Furthermore, a comparison of the results with those of the state-of-the-art methods demonstrates the superior performance of the proposed method.
Lifang Zhou, Weisheng Li 0001, Shan Liang 0004
Int. J. Pattern Recognit. Artif. Intell.3
2019 Adaptive illumination-invariant face recognition via local nonlinear multi-layer contrast feature
Li Zhou 0002, Weisheng Li 0001, Yue-Wei Du, Bang Jun Lei, Shan Liang 0004
J. Vis. Commun. Image Represent.2
2019 Principal Component Analysis based on Nuclear norm Minimization
Jian-Xun Mi, Zhihui Lai 0001, Weisheng Li 0001, Lifang Zhou, Fujin Zhong
Neural Networks4
2019 2D-LBP: An Enhanced Local Binary Feature for Texture Image Classification
abstract
The local binary pattern (LBP) and its variants have shown the effectiveness in texture images classification, face recognition, and other applications. However, most of these LBP methods only focus on the histogram of LBP patterns and ignore the spatial contextual information between LBP patterns. In this paper, we propose a 2D-LBP method which uses a sliding window to count the weighted occurrence number of the rotation invariant uniform LBP pattern pairs to obtain the spatial contextual information. The multi-resolution 2D-LBP features can also be obtained when the radius of 2D-LBP is changed. At last, a two-stage classifier which acts as an ensemble learning step is followed to achieve an accurate classification by combining the predictions on each 2D-LBP with single resolution. Theoretical proof shows that the proposed 2D-LBP is a general framework and can be integrated on other LBP variants to derive new feature extraction methods. Experimental results show that, the proposed method achieves 99.71%, 97.09%, 98.48%, and 49.00% classification accuracy on the public “Brodatz,” “CUReT,” “UIUC,” and “FMD” texture image databases, respectively. Compared with the original LBP and its variants, the proposed method obtains higher classification accuracy under different cases, and simultaneously owns shorter time complexity.
Bin Xiao 0002, Xiuli Bi, Weisheng Li 0001, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 A Geographically and Temporally Weighted Regression Model for Spatial Downscaling of MODIS Land Surface Temperatures Over Urban Heterogeneous Regions
abstract
The fine spatial resolution (~100 m) land surface temperature (LST) is a key variable of great concern in various environmental studies over urban heterogeneous regions. An improvement in the spatial resolution of the coarse spatial resolution LST is an effective way to extend its potential uses in applications that have strict requests on both the spatial and temporal resolutions. However, previous statistical downscaling algorithms were proposed mainly by addressing the spatial variability in the LST while neglecting the temporal variability. In this paper, we propose a new algorithm based on a geographically and temporally weighted regression (GTWR) model for spatial downscaling of the Moderate Resolution Imaging Spectroradiometer LST data from 1000 to 100 m. The GTWR-based algorithm with temporally and geographically varying regression coefficients can capture both the spatial and temporal variabilities in the LST from the time series data at a coarse spatial resolution for effectively reconstructing the subpixel variability in the LST at fine spatial resolution. In addition, because of a better ability to explain the LST variability over urban heterogeneous regions, a normalized difference built-up index and a digital elevation model were selected as auxiliary variables. Taking Beijing and Lanzhou as examples, the performance of the GTWR-based algorithm was assessed by comparing the results with the TsHARP and GWR-based algorithms and the Landsat-8 LST. The results indicate that the GTWR-based algorithm outperforms the above-mentioned algorithms with lower mean root mean square error (1.62 °C) and mean absolute error (1.28 °C) and better agreement between the GTWR downscaled LST and the Landsat-8 LST.
Yidong Peng, Weisheng Li 0001, Xiaobo Luo, Hua Li 0005
IEEE Trans. Geosci. Remote. Sens.2
2018 Brightness and contrast controllable image enhancement based on histogram specification
Bin Xiao 0002, Yanjun Jiang, Weisheng Li 0001, Guoyin Wang 0001
Neurocomputing4
2018 Fusion of anatomical and functional images using parallel saliency features
Jiao Du, Weisheng Li 0001, Bin Xiao 0002
Inf. Sci.2
2018 Pixel convolutional neural network for multi-focus image fusion
Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001
Inf. Sci.3
2018 Pose-robust face recognition with Huffman-LBP enhanced by Divide-and-Rule strategy
Lifang Zhou, Yue-Wei Du, Weisheng Li 0001, Jian-Xun Mi, Xiao Luan
Pattern Recognit.3
2017 Adaptive Class Preserving Representation for Image Classification
abstract
In linear representation-based image classification, an unlabeled sample is represented by the entire training set. To obtain a stable and discriminative solution, regularization on the vector of representation coefficients is necessary. For example, the representation in sparse representation-based classification (SRC) uses L1 norm penalty as regularization, which is equal to lasso. However, lasso overemphasizes the role of sparseness while ignoring the inherent structure among samples belonging to a same class. Many recent developed representation classifications have adopted lasso-type regressions to improve the performance. In this paper, we propose the adaptive class preserving representation for classification (ACPRC). Our method is related to group lasso based classification but different in two key points: When training samples in a class are uncorrelated, ACPRC turns into SRC, when samples in a class are highly correlated, it obtains similar result as group lasso. The superiority of ACPRC over other state-of-the-art regularization techniques including lasso, group lasso, sparse group lasso, etc. are evaluated by extensive experiments.
Jian-Xun Mi, Qiankun Fu, Weisheng Li 0001
CVPR3
2017 Synergistic integration of graph-cut and cloud model strategies for image segmentation
Weisheng Li 0001, Jiao Du, Jun Lai
Neurocomputing1
2017 Image analysis by fractional-order orthogonal moments
Bin Xiao 0002, Linping Li, Yu Li 0018, Weisheng Li 0001, Guoyin Wang 0001
Inf. Sci.4
2017 Color perception of diffusion tensor images using hierarchical manifold learning
Shanshan He, Weisheng Li 0001
Pattern Recognit.3
2017 Anatomical-Functional Image Fusion by Information of Interest in Local Laplacian Filtering Domain
abstract
A novel method for performing anatomical (MRI)-functional (PET or SPECT) image fusion is presented. The method merges specific feature information from input image signals of a single or multiple medical imaging modalities into a single fused image while preserving more information and generating less distortion. The proposed method uses a local Laplacian filtering based technique realized through a novel multi-scale system architecture. Firstly, the input images are generated in a multi-scale image representation and are processed using local Laplacian filtering. Secondly, at each scale, the decomposed images are combined to produce fused approximate images using a local energy maximum scheme and produce the fused residual images using an information of interest-based scheme. Finally, a fused image is obtained using a reconstruction process that is analogous to that of conventional Laplacian pyramid transform. Experimental results computed using individual multi-scale analysis-based decomposition schemes or fusion rules clearly demonstrate the superiority of the proposed method through subjective observation as well as objective metrics. Furthermore, the proposed method can obtain better performance, compared to the state-of-the-art fusion methods.
Jiao Du, Weisheng Li 0001, Bin Xiao 0002
IEEE Trans. Image Process.2
2017 The Recognition of the Point Symbols in the Scanned Topographic Maps
abstract
It is difficult to separate the point symbols from the scanned topographic maps accurately, which brings challenges for the recognition of the point symbols. In this paper, based on the framework of generalized Hough transform (GHT), we propose a new algorithm, which is named shear line segment GHT (SLS-GHT), to recognize the point symbols directly in the scanned topographic maps. SLS-GHT combines the line segment GHT (LS-GHT) and the shear transformation. On the one hand, LS-GHT is proposed to represent the features of the point symbols more completely. Its R-table has double level indices, the first one is the color information of the point symbols, and the other is the slope of the line segment connected a pair of the skeleton points. On the other hand, the shear transformation is introduced to increase the directional features of the point symbols; it can make up for the directional limitation of LS-GHT indirectly. In this way, the point symbols are detected in a series of the sheared maps by LS-GHT, and the final optimal coordinates of the setpoints are gotten from a series of the recognition results. SLS-GHT detects the point symbols directly in the scanned topographic maps, totally different from the traditional pattern of extraction before recognition. Moreover, several experiments demonstrate that the proposed method allows improved recognition in complex scenes than the existing methods.
Qiguang Miao, Pengfei Xu 0003, Xuelong Li 0001, Jianfeng Song, Weisheng Li 0001
IEEE Trans. Image Process.5
2016 An overview of multi-modal medical image fusion
Jiao Du, Weisheng Li 0001, Ke Lu 0002, Bin Xiao 0002
Neurocomputing2
2016 Union Laplacian pyramid with multiple features for medical image fusion
Jiao Du, Weisheng Li 0001, Bin Xiao 0002, Qamar Nawaz
Neurocomputing2
2016 Object recognition based on the Region of Interest and optimal Bag of Words model
Weisheng Li 0001, Bin Xiao 0002, Lifang Zhou
Neurocomputing1
2016 Single image haze removal based on haze physical characteristics and adaptive sky region detection
Yunan Li 0001, Qiguang Miao, Jianfeng Song, Yi-Ning Quan, Weisheng Li 0001
Neurocomputing5
2016 Lossless image compression based on integer Discrete Tchebichef Transform
Bin Xiao 0002, Yanhong Zhang, Weisheng Li 0001, Guoyin Wang 0001
Neurocomputing4
2016 Medical image fusion by combining parallel features on multi-scale local extrema scheme
Jiao Du, Weisheng Li 0001, Bin Xiao 0002, Qamar Nawaz
Knowl. Based Syst.2
2015 Moments and moment invariants in the Radon space
Bin Xiao 0002, Jiangtao Cui, Hongxing Qin, Weisheng Li 0001, Guoyin Wang 0001
Pattern Recognit.4
2015 Errata and comments on "Orthogonal moments based on exponent functions: Exponent-Fourier moments"
Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001
Pattern Recognit.2
2015 Facilitating Image Search With a Scalable and Compact Semantic Mapping
abstract
This paper introduces a novel approach to facilitating image search based on a compact semantic embedding. A novel method is developed to explicitly map concepts and image contents into a unified latent semantic space for the representation of semantic concept prototypes. Then, a linear embedding matrix is learned that maps images into the semantic space, such that each image is closer to its relevant concept prototype than other prototypes. In our approach, the semantic concepts equated with query keywords and the images mapped into the vicinity of the prototype are retrieved by our scheme. In addition, a computationally efficient method is introduced to incorporate new semantic concept prototypes into the semantic space by updating the embedding matrix. This novelty improves the scalability of the method and allows it to be applied to dynamic image repositories. Therefore, the proposed approach not only narrows semantic gap but also supports an efficient image search process. We have carried out extensive experiments on various cross-modality image search tasks over three widely-used benchmark image datasets. Results demonstrate the superior effectiveness, efficiency, and scalability of our proposed approach.
Meng Wang 0001, Weisheng Li 0001, Dong Liu 0001, Bingbing Ni, Jialie Shen 0001, Shuicheng Yan
IEEE Trans. Cybern.2
2014 Radial shifted Legendre moments for image analysis and invariant image recognition
Bin Xiao 0002, Guoyin Wang 0001, Weisheng Li 0001
Image Vis. Comput.3
2013 A new finer resolution land-use mapping method using time series of NDVI from HJ-1/CCD data
abstract
HJ-1/CCD data has both high spatial resolution and high temporal frequency. By using NDVI time series and SVM classifier with HJ-1/CCD data source, this paper proposed a new resolution for land use classification of Dahuofang Reservoir, Liaoning Province, China. The validation result demonstrates that this resolution could obtain finer result than others.
Peng Ma, Weisheng Li 0001, Qinhuo Liu
IGARSS3
2013 A new cloud detection method over Tibetan plateau and its surrounding area
abstract
To extract information about the Earth's surface from Earth Observation data, a key processing step is the separation of pixels representing clear-sky observations of land surface from observation influenced by cloud. This paper presents a new method used for MDOIS data over Tibetan plateau and desert. The method for cloud detecting based on the difference on band 6 (apparent reflectance of blue band) and band 31 (bright temperature in thermal infrared band) between cloud and land surface. Experimental results show that the method is feasible.
Shanlong Wu, Weisheng Li 0001, Qinhuo Liu
IGARSS3
2013 An Adaptive Fuzzy Fusion Framework for Face Recognition under Illumination variation Based on Local Multiple Patterns
abstract
Local binary pattern (LBP) operator offers an efficient way to recognize face under varying illumination, while it has the drawback of abandoning some important texture features. Local multiple patterns (LMP) has alleviated the problem by a hierarchical model. However, the LMP method can bring out the rapid expansion of feature dimension, so a special feature encoding method is adopted by this paper. Meanwhile, we find that the LMP features of different layers can be used to recognize face independently so that it would preserve more abundant recognition information. Most importantly, the contribution of the LMP features from different layers is blurry under varying illumination. We propose a fuzzy framework to fuse the recognition result of different layers and use adaptive weights to calculate contribution rates of different layers under varying illumination. Experimental results demonstrate that the proposed method outperforms other state-of-the-art methods on four databases such as Yale B, Extended Yale B, CMU PIE and Outdoor.
Lifang Zhou, Bin Fang 0001, Weisheng Li 0001, Hengxin Chen, Lidou Wang
Int. J. Pattern Recognit. Artif. Intell.3
2013 An algorithm on fairness verification of mobile sink routing in wireless sensor network
Guangquan Xu, Weisheng Li 0001, Yingyuan Xiao, Honghao Gao, Xiaohong Li 0001, Zhiyong Feng 0002, Jia Mei
Pers. Ubiquitous Comput.2
2013 Linear Feature Separation From Topographic Maps Using Energy Density and the Shear Transform
abstract
Linear features are difficult to be separated from complicated background in color scanned topographic maps, especially when the color of linear features approximate to that of background in some particular images. This paper presents a method, which is based on energy density and the shear transform, for the separation of lines from background. First, the shear transform, which could add the directional characteristics of the lines, is introduced to overcome the disadvantage that linear information loss would happen if the separation method is used in an image, which is in only one direction. Then templates in the horizontal and vertical directions are built to separate lines from background on account of the fact that the energy concentration of the lines usually reaches a higher level than that of the background in the negtive image. Furthermore, the remaining grid background can be wiped off by grid templates matching. The isolated patches, which include only one pixel or less than ten pixels, are removed according to the connected region area measurement. Finally, using the union operation, the linear features obtained in different sheared images could supplement each other, thus the lines of the final result are more complete. The basic property of this method is introducing the energy density instead of color information commonly used in traditional methods. The experiment results indicate that the proposed method could distinguish the linear features from the background more effectively, and obtain good results for its ability in changing the directions of the lines with the shear transform.
Qiguang Miao, Pengfei Xu 0003, Tiange Liu, Weisheng Li 0001
IEEE Trans. Image Process.6
2012 An edge detection algorithm based on the multi-direction shear transform
Pengfei Xu 0003, Qiguang Miao, Cheng Shi 0002, Weisheng Li 0001
J. Vis. Commun. Image Represent.5