Hu Chen 0002

dblp:04/1286-2 · DBLP profile ↗
← Back
37ranked-venue papers
2as first author
31since 2021 · last 2026
0000-0001-9300-6572ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prompt-level contrastive learning for context-aware multi-modal image representation in medical diagnosis
Guowei Dai 0001, Zhimin Tian, Duwei Dai, Chaoyu Wang 0001, Yi Zhang 0098, Hu Chen 0002
Pattern Recognit.7
2026 Multi-Task Learning Network for Medical Image Analysis Guided by Lesion Regions and Spatial Relationships of Tissues
abstract
Medical image analysis plays key role in computer-aided diagnosis, where segmentation and classification are essential and interconnected tasks. While multi-task learning (MTL) has been widely explored to leverage inter-task synergies, effectively guiding knowledge transfer to prevent task conflict and negative transfer remains a key challenge, particularly in anatomically complex diagnostic scenarios. This paper presents LTRMTL-Net, a novel multi-task learning framework for medical image analysis that simultaneously addresses segmentation and classification tasks guided by lesion regions and spatial relationships of tissues. The proposed architecture integrates an Enhanced Lesion Region Fusion (ELRF) module that leverages GradCAM-guided attention mechanisms to precisely locate and enhance lesion regions, providing critical prior knowledge for both tasks. Tissue Space Structure Prediction (TSSP) component captures local-global spatial dependencies through contrastive learning, establishing effective anatomical context modeling. The core encoder employs Hybrid Wavelet-State Attention blocks that combine modulated wavelet transform convolutions with structured state space models to extract multi-scale features while maintaining computational efficiency. Dual-stream inputs with symmetric architecture accommodate single-source scenarios across diverse medical imaging applications. Experimental results on mammography and breast ultrasound datasets demonstrate that the proposed method captures fine-grained lesion boundary details while providing accurate malignancy classification. Harnessing cooperative knowledge transfer between segmentation and classification, guided by anatomical priors, boosts diagnostic performance and provides comprehensive, interpretable clinical insights.
Guowei Dai 0001, Duwei Dai, Chaoyu Wang 0001, Qingfeng Tang 0001, Hu Chen 0002, Yi Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.6
2026 MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAM
abstract
The Medical Segment Anything Model (MedSAM) has demonstrated strong performance in medical image segmentation, attracting increasing attention in the medical imaging domain. However, as with many prompt-based segmentation models, its performance is highly sensitive to the type and location of input prompts. This sensitivity often leads to suboptimal segmentation outcomes and necessitates labor-intensive manual prompt tuning, which hampers both efficiency and robustness. To address this challenge, this paper proposes MedSAM-U, an uncertainty-guided framework designed to automatically refine prompt inputs and enhance segmentation reliability. Specifically, a Multi-Prompt Adapter is integrated into MedSAM, resulting in MPA-MedSAM, which enables the model to effectively accommodate diverse multi-prompt inputs. An uncertainty estimation module is then introduced to evaluate the reliability of the prompts and their initial segmentation results. Based on this, a novel uncertainty-guided prompt adaptation strategy is applied to automatically generate refined prompts and more accurate segmentation outputs. The proposed MedSAM-U framework is evaluated across multiple medical imaging modalities. Experimental results on five diverse datasets demonstrate that MedSAM-U achieves consistent performance improvements ranging from 1.7% to 20.5% over the baseline MedSAM, confirming its effectiveness and practicality for robust and efficient medical image segmentation.
Ke Zou, Mengting Luo, Linchao He, Meng Wang 0038, Yi Zhang 0018, Hu Chen 0002, Huazhu Fu
IEEE Trans. Circuits Syst. Video Technol.9
2025 Incorporating Improved Sinusoidal Threshold-based Semi-supervised Method and Diffusion Models for Osteoporosis Diagnosis
abstract
Osteoporosis is a common skeletal disease that seriously affects patients’ quality of life. Traditional osteoporosis diagnosis methods are expensive and complex. The semi-supervised model based on diffusion model and class threshold sinusoidal decay proposed in this paper can automatically diagnose osteoporosis based on patient’s imaging data, which has the advantages of convenience, accuracy, and low cost. Unlike previous semi-supervised models, all the unlabeled data used in this paper are generated by the diffusion model. Compared with real unlabeled data, synthetic data generated by the diffusion model show better performance. In addition, this paper proposes a novel pseudo-label threshold adjustment mechanism, Sinusoidal Threshold Decay, which can make the semi-supervised model converge more quickly and improve its performance. Specifically, the method is tested on a dataset including 749 dental panoramic images, and its achieved leading detect performance and produces a 80.10% accuracy.
Wenchi Ke, Hu Chen 0002, Xiong Deng
ICASSP2
2025 Learning Geometry-Aware Representation for Gaze Estimation
abstract
Appearance-based gaze estimation has achieved remarkable progress in recent years. However, the inherent geometry characteristics of eye and facial areas are not fully explored in existing methods, which limits the generalization and robustness of the model. In this paper, we propose a novel end-to-end framework for cross-domain gaze estimation by integrating latent geometric representation into appearance-based gaze framework. More specifically, we first exploit the 3DMM method to fit unconstrained faces and eyes in the wild, which would generate adaptive normal information with explicit 3D geometry prior. Then we joint the normal map and the corresponding RGB appearance information to infer the 3D gaze direction with carefully-designed spatial-frequency attention and local-global feature interaction modules. The key to our method is to integrate explicit 3D geometry representation into a 2D learning architecture, which leads to a better trade-off between performance and efficiency. Experiments on both MPIIGaze and EyeDiap datasets demonstrate that the proposed method achieves the state-of-the-art accuracy of 3.56° and 5.10° separately, and also presents superior generalization ability on cross-domain dataset evaluations.
Qida Tan, Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
ICIP4
2025 A Unified Multimodal Multi-Granularity Pre-Training Framework for Fine-Grained Medical Image Analysis
abstract
Medical image diagnosis relies heavily on subtle, fine-grained details, making it crucial to integrate global imaging representations with localized pathological information. Existing vision-language pretraining models have shown promise in utilizing free-text radiology reports for deeper semantic insights; however, they face significant challenges in precisely capturing and aligning region-specific, fine-grained information, which often results in suboptimal feature extraction for localized pathologies and reduced model interpretability. In this paper, we propose a unified multimodal multi-granularity pre-training framework (UMMPF) that explicitly models the local anatomical regions in chest X-ray images through a Region Querying Module (RQM), thereby establishing stronger correspondences between image sub-regions and textual descriptions. We further incorporate cross-modal bidirectional attention and unified multiple training objectives, facilitating the interaction of both global and local features across modalities. Our method enhances interpretability by explicitly mapping pathological findings to anatomical regions. Experimental results on benchmark datasets (RSNA, SIIM, and ChestX-ray14) demonstrate a significant improvement in diagnostic accuracy.
Zenan Gong, Linchao He, Peixi Liao, Hongjie Yang, Hu Chen 0002, Yi Zhang 0018
IJCNN7
2025 Lightweight Vision Transformer With Lite-AVPSO Hyperparameter Optimization for Agricultural Disease Recognition
abstract
Plant disease identification and management are vital for crop protection, productivity, and sustainable agriculture. This paper presents RepAgrViT, a lightweight vision transformer architecture for agricultural disease recognition in Internet of Things edge environments. The proposed model integrates CNN with vision transformer principles through a novel dual-stream design comprising token mixer and channel mixer arranged in series. A Bilinear Attention Transformation module performs extensive local-global attention via direct feature position connections. This enables comprehensive capture of long-range dependencies between healthy and diseased leaf regions. We introduce an efficient parameter fusion methodology that optimizes depthwise separable convolutions while integrating normalized weights with biases, reducing model complexity. Furthermore, we propose Lite-AVPSO, a hyperparameter optimization algorithm for precise model configuration. It incorporates adaptive weighted delayed velocity optimization and neighborhood-based local search strategies. Visualization techniques using multi-channel color heatmaps and three-dimensional feature spaces provide interpretable diagnostic results. Experiments across diverse plant disease datasets demonstrate RepAgrViT efficacy in capturing discriminative disease features while maintaining computational efficiency suitable for resource-constrained agricultural monitoring systems, establishing a foundation for efficient vision-based monitoring in real-time agricultural applications.
Guowei Dai 0001, Zhimin Tian, Chaoyu Wang 0001, Qingfeng Tang 0001, Hu Chen 0002, Yi Zhang 0018
IEEE Internet Things J.5
2025 Geometry-Aware Appearance Learning for Generalized Gaze Estimation
Qida Tan, Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
IEEE Signal Process. Lett.4
2025 Solving Zero-Shot Sparse-View CT Reconstruction With Variational Score Solver
abstract
Computed tomography (CT) stands as a ubiquitous medical diagnostic tool. Nonetheless, the radiation-related concerns associated with CT scans have raised public apprehensions. Mitigating radiation dosage in CT imaging poses an inherent challenge as it inevitably compromises the fidelity of CT reconstructions, impacting diagnostic accuracy. While previous deep learning techniques have exhibited promise in enhancing CT reconstruction quality, they remain hindered by the reliance on paired data, which is arduous to procure. In this study, we present a novel approach named Variational Score Solver (VSS) for sparse-view reconstruction without paired data. Our approach entails the acquisition of a probability distribution from densely sampled CT reconstructions, employing a latent diffusion model. High-quality reconstruction outcomes are achieved through an iterative process, wherein the diffusion model serves as the prior term, subsequently integrated with the data consistency term. Notably, rather than directly employing the prior diffusion model, we distill prior knowledge by finding the fixed point of the diffusion model. This framework empowers us to exercise precise control over the process. Moreover, we depart from modeling the reconstruction outcomes as deterministic values, opting instead for a distribution-based approach. This enables us to achieve more accurate reconstructions utilizing a trainable model. Our approach introduces a fresh perspective to the realm of zero-shot CT reconstruction, circumventing the constraints of supervised learning. Extensive qualitative and quantitative experiments unequivocally demonstrate that VSS surpasses other contemporary unsupervised and achieves comparable results compared to the most advanced supervised methods in sparse-view reconstruction tasks. Codes are available in https://github.com/fpsandnoob/vss.
Linchao He, Wenchao Du, Peixi Liao, Fenglei Fan, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging5
2025 Bi-Constraints Diffusion: A Conditional Diffusion Model With Degradation Guidance for Metal Artifact Reduction
abstract
In recent years, score-based diffusion models have emerged as effective tools for estimating score functions from empirical data distributions, particularly in integrating implicit priors with inverse problems like CT reconstruction. However, score-based diffusion models are rarely explored in challenging tasks such as metal artifact reduction (MAR). In this paper, we introduce a Bi-Constraints Diffusion Model for Metal Artifact Reduction (BCDMAR), an innovative approach that enhances iterative reconstruction with a conditional diffusion model for MAR. This method employs a metal artifact degradation operator in place of the traditional metal-excluded projection operator in the data-fidelity term, thereby preserving structure details around metal regions. However, score-based diffusion models tend to be susceptible to grayscale shifts and unreliable structures, making it challenging to reach an optimal solution. To address this, we utilize a pre-corrected image as a prior constraint, guiding the generation of the score-based diffusion model. By iteratively applying the score-based diffusion model and the data-fidelity step in each sampling iteration, BCDMAR effectively maintains reliable tissue representation around metal regions and produces highly consistent structures in non-metal regions. Through extensive experiments focused on metal artifact reduction tasks, BCDMAR demonstrates superior performance over other state-of-the-art unsupervised and supervised methods, both quantitatively and qualitatively.
Mengting Luo, Tao Wang 0167, Linchao He, Wang Wang, Hu Chen 0002, Peixi Liao, Yi Zhang 0018
IEEE Trans. Medical Imaging6
2024 A Progressively Prompt-guided Model for Sparse-View CT Reconstruction
abstract
While sparse-view Computed Tomography (CT) has a remarkable impact on reducing ionizing radiation dose while accelerating data acquisition, the reconstructed images have been compromised by streak-like artifacts, affecting clinical diagnostics. By integrating powerful regularization with deep learning technologies into iterative reconstruction algorithms, the deep-unrolling-based methods have achieved promising results in terms of reconstruction quality and theoretical interpretability. However, leading methods always focus on learning powerful content priors with diverse technologies and ignoring the latent noise distribution prior in the image domain, thereby limiting the ability of structure-preserving and detail reconstructing of the model. To alleviate this problem, we propose a Progressively Prompt-guided Model (shorted by PPM) for sparse-view CT reconstruction. Specifically, we inject the idea of prompt learning into an iterative unrolled neural network, in which a learnable prompt module is inserted into each unrolled block to perceive image content and noise distribution in a self-adaptive manner, which leads to the more powerful priors to guide high-quality CT image reconstruction. Furthermore, we construct a progressively guiding strategy to facilitate high-quality prompt generation while speeding model convergence. Extensive experiments demonstrate that our PPM achieves state-of-the-art performance in artifact suppression, structure fidelity, and visual perception similarity. The code is available at https://github.com/Wenchao-Du/PPM/.
Qiao Mu, Hu Chen 0002, Wenchao Du, Hongyu Yang 0002
BIBM4
2024 Dtpose: Learning Disentangled Token Representation For Effective Human Pose Estimation
abstract
Exploring rich visual clues and spatial geometric constraints to locate keypoints is essential for human pose estimation. Existing Transformer-based methods have presented unique advantages via token representation, where each keypoint is explicitly embedded as a token to learn visual appearance clues and geometric relationships simultaneously from images. However, it is difficult to learn powerful pose representation via self-attention mechanism due to latent interference, e.g., blurring and self-occlusion. To alleviate this challenge, we present a novel framework that Disentangles hybrid Token representation to explore more effective visual and keypoint information for Pose estimation (termed by DTPose). In detail, DTPose contains two key modules. First, the Disentangled Token Representation module is used to explore visual clues and geometry constraints sequentially, which alleviates the noise interference and enables the geometry and appearance clues to be exploited more sufficiently. Furthermore, the Hierarchical Spatial Decoding head is exploited to preserve the 2 D geometric structure information of keypoints as much as possible. Extensive experiments on COCO dataset demonstrate significant performance gains of our DTPose, which achieves 76.5 (+0.7) AP and 75.7 (+0.6) AP than the TokenPose-L on the COCO validation and test-dev sets separately.
Shiyang Ye, Hu Chen 0002, Wenchao Du, Hongyu Yang 0002
ICIP4
2024 Memory Coordinated Cross Perception for Few-Shot Object Detection
abstract
Significant advances have been made in the field of few-shot object detection. Few-shot object detection involves using limited examples to detect new categories. However, as the mainstream approach in FSOD, meta-learning methods face two main challenges: the prototypes they create lack sufficient representativeness and insufficient distinction between different prototypes. These issues stem from two factors. First, the limited quantity and diversity of support instances make it hard for the model to accurately capture the prototype’s essence. Second, similar categories complicate matching query features to the correct prototype. To address these issues, we propose an adaptive memory coordinated generation method. This method distinguishes each similar prototype while generating more representative prototypes using memory prototypes. Specifically, it enhances the representational information of support features through their interaction with memory prototypes. This approach also achieves information perception and adaptive fusion among support features, leading to the generation of easily distinguishable and more identifiable prototypes. Furthermore, we developed a prototype cross-perception module. This module consistently refines prototypes and query features, achieving deep integration of features while preserving the essential information of query features. It enhances prototypes’ directive function and improves query feature robustness. Our model demonstrates leading performance across most shot settings and evaluation metrics on multiple few-shot object detection benchmarks.
Yunfeng Kou, Hu Chen 0002, Wenchao Du
IJCNN4
2024 FSAD:Few-Shot Object Detection via Aggregation and Disentanglement
abstract
Few-shot Object Detection (FSOD) aims to leverage knowledge gained from general object detectors to enhance future detection tasks for novel object categories. In response to the poor performance observed in commonly used attention-based feature fusion methods, particularly in 1-shot, this paper proposes an enhanced Cross-Attention-Like Aggregation (CAL) Module and a GAN-Disentangled-Like Feature Enhancement (GDL) Module.CAL module utilizes an asymmetric mechanism and neutralized features resulting from concatenation for cross-attention, which improves the model’s generalization ability, enabling it to address issues where the target is entirely unrecognizable in 1-shot. The GDL module captures latent independent information from the support set. It employs a discriminator to stabilize the information extraction process, enabling the attention module to discern relevant information more effectively and accurately guide the query features.Extensive experiments conducted on the PASCAL VOC and COCO datasets demonstrate the superior performance of our approach over strong baselines, showcasing substantial advancements in one-shot and two-shot performance.our code is available at https://github.com/yun1232/FSAD
Yunfeng Kou, Kunming Wu, Hu Chen 0002, Wenchao Du
IJCNN4
2024 Diffusion Posterior Proximal Sampling for Image Restoration
abstract
Diffusion models have demonstrated remarkable efficacy in generating high-quality samples. Existing diffusion-based image restoration algorithms exploit pre-trained diffusion models to leverage data priors, yet they still preserve elements inherited from the unconditional generation paradigm. These strategies initiate the denoising process with pure white noise and incorporate random noise at each generative step, leading to over-smoothed results. In this paper, we present a refined paradigm for diffusion-based image restoration. Specifically, we opt for a sample consistent with the measurement identity at each generative step, exploiting the sampling selection as an avenue for output stability and enhancement. The number of candidate samples used for selection is adaptively determined based on the signal-to-noise ratio of the timestep. Additionally, we start the restoration process with an initialization combined with the measurement signal, providing supplementary information to better align the generative process. Extensive experimental results and analyses validate that our proposed method significantly enhances image restoration performance while consuming negligible additional computational resources.
Hongjie Wu, Linchao He, Mingqin Zhang, Dongdong Chen 0004, Kunming Luo, Mengting Luo, Jizhe Zhou 0001, Hu Chen 0002, Jiancheng Lv 0001
ACM Multimedia8
2024 Hierarchical disentangled representation for image denoising and beyond
Wenchao Du, Hu Chen 0002, Yi Zhang 0018, Hongyu Yang 0002
Image Vis. Comput.2
2024 SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) is an emerging technology, that generates a multienergy attenuation map for the interior of an object and extends the traditional image volume into a 4-D form. Compared with traditional CT based on energy-integrating detectors, spectral CT can make full use of spectral information, resulting in high resolution and providing accurate material quantification. Numerous model-based iterative reconstruction methods have been proposed for spectral CT reconstruction. However, these methods usually suffer from difficulties such as laborious parameter selection and expensive computational costs. In addition, due to the image similarity of different energy bins, spectral CT usually implies a strong low-rank prior, which has been widely adopted in current iterative reconstruction models. Singular value thresholding (SVT) is an effective algorithm to solve the low-rank constrained model. However, the SVT method requires a manual selection of thresholds, which may lead to suboptimal results. To relieve these problems, in this article, we propose a sparse and low-rank unrolling network (SOUL-Net) for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner. Furthermore, a Taylor expansion-based neural network backpropagation method is introduced to improve the numerical stability. The qualitative and quantitative results demonstrate that the proposed method outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction.
Xiang Chen 0015, Wenjun Xia, Ziyuan Yang 0001, Hu Chen 0002, Yan Liu 0052, Jiliu Zhou, Yang Chen 0008, Bihan Wen, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.4
2023 Differential evolution with variable leader-adjoint populations
Hongyu Yang 0002, Hu Chen 0002
Appl. Intell.4
2023 Enhancing differential evolution algorithm using leader-adjoint populations
Hongyu Yang 0002, Hu Chen 0002, Bo Yang 0063
Inf. Sci.4
2023 SemiMAR: Semi-Supervised Learning for CT Metal Artifact Reduction
abstract
Metal artifacts lead to CT imaging quality degradation. With the success of deep learning (DL) in medical imaging, a number of DL-based supervised methods have been developed for metal artifact reduction (MAR). Nonetheless, fully-supervised MAR methods based on simulated data do not perform well on clinical data due to the domain gap. Although this problem can be avoided in an unsupervised way to a certain degree, severe artifacts cannot be well suppressed in clinical practice. Recently, semi-supervised metal artifact reduction (MAR) methods have gained wide attention due to their ability in narrowing the domain gap and improving MAR performance in clinical data. However, these methods typically require large model sizes, posing challenges for optimization. To address this issue, we propose a novel semi-supervised MAR framework. In our framework, only the artifact-free parts are learned, and the artifacts are inferred by subtracting these clean parts from the metal-corrupted CT images. Our approach leverages a single generator to execute all complex transformations, thereby reducing the model's scale and preventing overlap between clean part and artifacts. To recover more tissue details, we distill the knowledge from the advanced dual-domain MAR network into our model in both image domain and latent feature space. The latent space constraint is achieved via contrastive learning. We also evaluate the impact of different generator architectures by investigating several mainstream deep learning-based MAR backbones. Our experiments demonstrate that the proposed method competes favorably with several state-of-the-art semi-supervised MAR techniques in both qualitative and quantitative aspects.
Tao Wang 0167, Zhiwen Wang 0002, Hu Chen 0002, Yan Liu 0052, Jingfeng Lu, Yi Zhang 0018
IEEE J. Biomed. Health Informatics4
2023 M3NAS: Multi-Scale and Multi-Level Memory-Efficient Neural Architecture Search for Low-Dose CT Denoising
abstract
Lowering the radiation dose in computed tomography (CT) can greatly reduce the potential risk to public health. However, the reconstructed images from dose-reduced CT or low-dose CT (LDCT) suffer from severe noise which compromises the subsequent diagnosis and analysis. Recently, convolutional neural networks have achieved promising results in removing noise from LDCT images. The network architectures that are used are either handcrafted or built on top of conventional networks such as ResNet and U-Net. Recent advances in neural network architecture search (NAS) have shown that the network architecture has a dramatic effect on the model performance. This indicates that current network architectures for LDCT may be suboptimal. Therefore, in this paper, we make the first attempt to apply NAS to LDCT and propose a multi-scale and multi-level memory-efficient NAS for LDCT denoising, termed M3NAS. On the one hand, the proposed M3NAS fuses features extracted by different scale cells to capture multi-scale image structural details. On the other hand, the proposed M3NAS can search a hybrid cell- and network-level structure for better performance. In addition, M3NAS can effectively reduce the number of model parameters and increase the speed of inference. Extensive experimental results on two different datasets demonstrate that the proposed M3NAS can achieve better performance and fewer parameters than several state-of-the-art methods. In addition, we also validate the effectiveness of the multi-scale and multi-level architecture for LDCT denoising, and present further analysis for different configurations of super-net.
Wenjun Xia, Yongqiang Huang 0003, Mingzheng Hou, Hu Chen 0002, Jiliu Zhou, Hongming Shan, Yi Zhang 0018
IEEE Trans. Medical Imaging5
2023 MLF-IOSC: Multi-Level Fusion Network With Independent Operation Search Cell for Low-Dose CT Denoising
abstract
Computed tomography (CT) is widely used in clinical medicine, and low-dose CT (LDCT) has become popular to reduce potential patient harm during CT acquisition. However, LDCT aggravates the problem of noise and artifacts in CT images, increasing diagnosis difficulty. Through deep learning, denoising CT images by artificial neural network has aroused great interest for medical imaging and has been hugely successful. We propose a framework to achieve excellent LDCT noise reduction using independent operation search cells, inspired by neural architecture search, and introduce the Laplacian to further improve image quality. Employing patch-based training, the proposed method can effectively eliminate CT image noise while retaining the original structures and details, hence significantly improving diagnosis efficiency and promoting LDCT clinical applications.
Jinbo Shen, Mengting Luo, Peixi Liao, Hu Chen 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging5
2023 DHI-GAN: Improving Dental-Based Human Identification Using Generative Adversarial Networks
abstract
In this work, a novel semisupervised framework is proposed to tackle the small-sample problem of dental-based human identification (DHI), achieving enhanced performance via a "classifying while generating" paradigm. A generative adversarial network (GAN), called the DHI-GAN, is presented to implement this idea, in which an extra classifier is also dedicatedly proposed to achieve an efficient training procedure. Considering the complex specificities of this problem, except for the noise input of the generator, an identity embedding-guided architecture is proposed to retain informative features for each individual. A parallel spatial and channel fusion attention block is innovatively designed to encourage the model to learn discriminative and informative features by focusing on different regional details and abstract concepts. The attention block is also widely applied to the overall classifier to learn identity-dependent information. A loss combination of the ArcFace and focal loss is utilized to address the small-sample problem. Two parameters are proposed to control the generated samples that are fed into the classifier during the optimization procedure. The proposed DHI-GAN framework is finally validated on a real-world dataset, and the experimental results demonstrate that it outperforms other baselines, achieving a 92.5% top-one accuracy rate. Most importantly, the proposed GAN-based semisupervised training strategy is able to reduce the required number of training samples (individuals) and can also be incorporated into other classification models. Our code will be available at https://github.com/sculyi/MedicalImages/.
Yi Lin 0006, Jianwei Zhang 0013, Jizhe Zhou 0001, Peixi Liao, Hu Chen 0002, Zhenhua Deng, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.6
2022 Learning Interval-Aware Embedding for Macro and Micro-expression Spotting
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
ACCV (4)4
2022 Depth Completion Using Geometry-Aware Embedding
abstract
Exploiting internal spatial geometric constraints of sparse LiDARs is beneficial to depth completion, however, has been not explored well. This paper proposes an efficient method to learn geometry-aware embedding, which encodes the local and global geometric structure information from 3D points, e.g., scene layout, object's sizes and shapes, to guide dense depth estimation. Specifically, we utilize the dynamic graph representation to model generalized geometric relationship from irregular point clouds in a flexible and efficient manner. Further, we joint this embedding and corresponded RGB appearance information to infer missing depths of the scene with well structure-preserved details. The key to our method is to integrate implicit 3D geometric representation into a 2D learning architecture, which leads to a better trade-off between the performance and efficiency. Extensive experiments demonstrate that the proposed method outperforms previous works and could reconstruct fine depths with crisp boundaries in regions that are over-smoothed by them. The ablation study gives more insights into our method that could achieve significant gains with a simple design, while having better generalization capability and stability. The code is available at https://github.com/Wenchao-Du/GAENet.
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018
ICRA2
2022 FONT-SIR: Fourth-Order Nonlocal Tensor Decomposition Model for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) reconstructs images from different spectral data through photon counting detectors (PCDs). However, due to the limited number of photons and the counting rate in the corresponding spectral segment, the reconstructed spectral images are usually affected by severe noise. In this paper, we propose a fourth-order nonlocal tensor decomposition model for spectral CT image reconstruction (FONT-SIR). To maintain the original spatial relationships among similar patches and improve the imaging quality, similar patches without vectorization are grouped in both spectral and spatial domains simultaneously to form the fourth-order processing tensor unit. The similarity of different patches is measured with the cosine similarity of latent features extracted using principal component analysis (PCA). By imposing the constraints of the weighted nuclear and total variation (TV) norms, each fourth-order tensor unit is decomposed into a low-rank component and a sparse component, which can efficiently remove noise and artifacts while preserving the structural details. Moreover, the alternating direction method of multipliers (ADMM) is employed to solve the decomposition model. Extensive experimental results on both simulated and real data sets demonstrate that the proposed FONT-SIR achieves superior qualitative and quantitative performance compared with several state-of-the-art methods.
Xiang Chen 0015, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Zhiyuan Zha, Bihan Wen, Yi Zhang 0018
IEEE Trans. Medical Imaging4
2021 Dual-Domain Adaptive-Scaling Non-local Network for CT Metal Artifact Reduction
Tao Wang 0167, Wenjun Xia, Yongqiang Huang 0003, Huaiqiang Sun, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Yi Zhang 0018
MICCAI (6)6
2021 Noise-Powered Disentangled Representation for Unsupervised Speckle Reduction of Optical Coherence Tomography Images
abstract
Due to its noninvasive character, optical coherence tomography (OCT) has become a popular diagnostic method in clinical settings. However, the low-coherence interferometric imaging procedure is inevitably contaminated by heavy speckle noise, which impairs both visual quality and diagnosis of various ocular diseases. Although deep learning has been applied for image denoising and achieved promising results, the lack of well-registered clean and noisy image pairs makes it impractical for supervised learning-based approaches to achieve satisfactory OCT image denoising results. In this paper, we propose an unsupervised OCT image speckle reduction algorithm that does not rely on well-registered image pairs. Specifically, by employing the ideas of disentangled representation and generative adversarial network, the proposed method first disentangles the noisy image into content and noise spaces by corresponding encoders. Then, the generator is used to predict the denoised OCT image with the extracted content features. In addition, the noise patches cropped from the noisy image are utilized to facilitate more accurate disentanglement. Extensive experiments have been conducted, and the results suggest that our proposed method is superior to the classic methods and demonstrates competitive performance to several recently proposed learning-based approaches in both quantitative and qualitative aspects. Code is available at: https://github.com/tsmotlp/DRGAN-OCT.
Yongqiang Huang 0003, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Leyuan Fang, Yi Zhang 0018
IEEE Trans. Medical Imaging5
2021 LCANet: Learnable Connected Attention Network for Human Identification Using Dental Images
abstract
Forensic odontology is regarded as an important branch of forensics dealing with human identification based on dental identification. This paper proposes a novel method that uses deep convolution neural networks to assist in human identification by automatically and accurately matching 2-D panoramic dental X-ray images. Designed as a top-down architecture, the network incorporates an improved channel attention module and a learnable connected module to better extract features for matching. By integrating associated features among all channel maps, the channel attention module can selectively emphasize interdependent channel information, which contributes to more precise recognition results. The learnable connected module not only connects different layers in a feed-forward fashion but also searches the optimal connections for each connected layer, resulting in automatically and adaptively learning the connections among layers. Extensive experiments demonstrate that our method can achieve new state-of-the-art performance in human identification using dental images. Specifically, the method is tested on a dataset including 1,168 dental panoramic images of 503 different subjects, and its dental image recognition accuracy for human identification reaches 87.21% rank-1 accuracy and 95.34% rank-5 accuracy. Code has been released on Github. (https://github.com/cclaiyc/TIdentify).
Yancun Lai, Qingsong Wu, Wenchi Ke, Peixi Liao, Zhenhua Deng, Hu Chen 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging7
2021 CT Reconstruction With PDF: Parameter-Dependent Framework for Data From Multiple Geometries and Dose Levels
abstract
The current mainstream computed tomography (CT) reconstruction methods based on deep learning usually need to fix the scanning geometry and dose level, which significantly aggravates the training costs and requires more training data for real clinical applications. In this paper, we propose a parameter-dependent framework (PDF) that trains a reconstruction network with data originating from multiple alternative geometries and dose levels simultaneously. In the proposed PDF, the geometry and dose level are parameterized and fed into two multilayer perceptrons (MLPs). The outputs of the MLPs are used to modulate the feature maps of the CT reconstruction network, which condition the network outputs on different geometries and dose levels. The experiments show that our proposed method can obtain competitive performance compared to the original network trained with either specific or mixed geometry and dose level, which can efficiently save extra training costs for multiple geometries and dose levels.
Wenjun Xia, Yongqiang Huang 0003, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Yi Zhang 0018
IEEE Trans. Medical Imaging5
2021 MAGIC: Manifold and Graph Integrative Convolutional Network for Low-Dose CT Reconstruction
abstract
Low-dose computed tomography (LDCT) scans, which can effectively alleviate the radiation problem, will degrade the imaging quality. In this paper, we propose a novel LDCT reconstruction network that unrolls the iterative scheme and performs in both image and manifold spaces. Because patch manifolds of medical images have low-dimensional structures, we can build graphs from the manifolds. Then, we simultaneously leverage the spatial convolution to extract the local pixel-level features from the images and incorporate the graph convolution to analyze the nonlocal topological features in manifold space. The experiments show that our proposed method outperforms both the quantitative and qualitative aspects of state-of-the-art methods. In addition, aided by a projection loss component, our proposed method also demonstrates superior performance for semi-supervised learning. The network can remove most noise while maintaining the details of only 10% (40 slices) of the training data labeled.
Wenjun Xia, Yongqiang Huang 0003, Zuoqiang Shi, Yan Liu 0052, Hu Chen 0002, Yang Chen 0008, Jiliu Zhou, Yi Zhang 0018
IEEE Trans. Medical Imaging6
2020 Learning Invariant Representation for Unsupervised Image Restoration
abstract
Recently, cross domain transfer has been applied for unsupervised image restoration tasks. However, directly applying existing frameworks would lead to domain-shift problems in translated images due to lack of effective supervision. Instead, we propose an unsupervised learning method that explicitly learns invariant presentation from noisy data and reconstructs clear observations. To do so, we introduce discrete disentangling representation and adversarial domain adaption into general domain transfer framework, aided by extra self-supervised modules including background and semantic consistency constraints, learning robust representation under dual domain constraints, such as feature and image domains. Experiments on synthetic and real noise removal tasks show the proposed method achieves comparable performance with other stateof-the-art supervised and unsupervised methods, while having faster and stable convergence than other domain adaption methods.
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
CVPR2
2019 Denoising of 3D magnetic resonance images using a residual encoder-decoder Wasserstein generative adversarial network
Maosong Ran, Jinrong Hu, Yang Chen 0008, Hu Chen 0002, Huaiqiang Sun, Jiliu Zhou, Yi Zhang 0018
Medical Image Anal.4
2019 Visual Attention Network for Low-Dose CT
abstract
Noise and artifacts are intrinsic to low-dose computed tomography (LDCT) data acquisition, and will significantly affect the imaging performance. Perfect noise removal and image restoration is intractable in the context of LDCT due to the statistical and the technical uncertainties. In this letter, we apply the generative adversarial network (GAN) framework with a visual attention mechanism to deal with this problem in a data-driven/machine learning fashion. Our main idea is to inject visual attention knowledge into the learning process of GAN to provide a powerful prior of the noise distribution. By doing this, both the generator and discriminator networks are empowered with visual attention information so that they will not only pay special attention to noisy regions and surrounding structures but also explicitly assess the local consistency of the recovered regions. Our experiments qualitatively and quantitatively demonstrate the effectiveness of the proposed method with clinic CT images.
Wenchao Du, Hu Chen 0002, Peixi Liao, Hongyu Yang 0002, Ge Wang 0001, Yi Zhang 0018
IEEE Signal Process. Lett.2
2018 LEARN: Learned Experts' Assessment-Based Reconstruction Network for Sparse-Data CT
abstract
Compressive sensing (CS) has proved effective for tomographic reconstruction from sparsely collected data or under-sampled measurements, which are practically important for few-view computed tomography (CT), tomosynthesis, interior tomography, and so on. To perform sparse-data CT, the iterative reconstruction commonly uses regularizers in the CS framework. Currently, how to choose the parameters adaptively for regularization is a major open problem. In this paper, inspired by the idea of machine learning especially deep learning, we unfold the state-of-the-art "fields of experts"-based iterative reconstruction scheme up to a number of iterations for data-driven training, construct a learned experts' assessment-based reconstruction network (LEARN) for sparse-data CT, and demonstrate the feasibility and merits of our LEARN network. The experimental results with our proposed LEARN network produces a superior performance with the well-known Mayo Clinic low-dose challenge data set relative to the several state-of-the-art methods, in terms of artifact reduction, feature preservation, and computational speed. This is consistent to our insight that because all the regularization terms and parameters used in the iterative reconstruction are now learned from the training data, our LEARN network utilizes application-oriented knowledge more effectively and recovers underlying images more favorably than competing algorithms. Also, the number of layers in the LEARN network is only 50, reducing the computational complexity of typical iterative algorithms by orders of magnitude.
Hu Chen 0002, Yi Zhang 0018, Yunjin Chen, Huaiqiang Sun, Yang Lu 0011, Peixi Liao, Jiliu Zhou, Ge Wang 0001
IEEE Trans. Medical Imaging1
2017 Low-Dose CT With a Residual Encoder-Decoder Convolutional Neural Network
abstract
Given the potential risk of X-ray radiation to the patient, low-dose CT has attracted a considerable interest in the medical imaging field. Currently, the main stream low-dose CT methods include vendor-specific sinogram domain filtration and iterative reconstruction algorithms, but they need to access raw data, whose formats are not transparent to most users. Due to the difficulty of modeling the statistical characteristics in the image domain, the existing methods for directly processing reconstructed images cannot eliminate image noise very well while keeping structural details. Inspired by the idea of deep learning, here we combine the autoencoder, deconvolution network, and shortcut connections into the residual encoder-decoder convolutional neural network (RED-CNN) for low-dose CT imaging. After patch-based training, the proposed RED-CNN achieves a competitive performance relative to the-state-of-art methods in both simulated and clinical cases. Especially, our method has been favorably evaluated in terms of noise suppression, structural preservation, and lesion detection.
Hu Chen 0002, Yi Zhang 0018, Mannudeep K. Kalra, Feng Lin 0010, Yang Chen 0008, Peixi Liao, Jiliu Zhou, Ge Wang 0001
IEEE Trans. Medical Imaging1
2016 Unseen head pose prediction using dense multivariate label distribution
abstract
Accurate head poses are useful for many face-related tasks such as face recognition, gaze estimation, and emotion analysis. Most existing methods estimate head poses that are included in the training data (i.e., previously seen head poses). To predict head poses that are not seen in the training data, some regression-based methods have been proposed. However, they focus on estimating continuous head pose angles, and thus do not systematically evaluate the performance on predicting unseen head poses. In this paper, we use a dense multivariate label distribution (MLD) to represent the pose angle of a face image. By incorporating both seen and unseen pose angles into MLD, the head pose predictor can estimate unseen head poses with an accuracy comparable to that of estimating seen head poses. On the Pointing’04 database, the mean absolute errors of results for yaw and pitch are 4.01° and 2.13°, respectively. In addition, experiments on the CAS-PEAL and CMU Multi-PIE databases show that the proposed dense MLD-based head pose estimation method can obtain the state-of-the-art performance when compared to some existing methods.
Gaoli Sang, Hu Chen 0002, Ge Huang, Qijun Zhao
Frontiers Inf. Technol. Electron. Eng.2