VLDB 2026 Research / reviewers in the wild / expert
Hongming Shan
dblp:184/8229
· DBLP profile ↗
86ranked-venue papers
7as first author
71since 2021 · last 2026
0000-0002-0604-3197ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 1 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 3 first-author · 27 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 26 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E 2 AD: Enhanced and explainable Alzheimer's disease detection framework via anatomy- and relation-aware cross-modal knowledge distillation
Sirong Piao, Tao Chen 0055, Zhaoyang Li 0016, Tongrui Zhang, Xing-Ming Zhao, Hongming Shan |
Medical Image Anal. | 9 |
| 2026 | Prompt-in-prompt learning for all-in-one image restoration
Zilong Li 0001, Chenglong Ma 0002, Junping Zhang, Hongming Shan |
Pattern Recognit. | 5 |
| 2026 | Universal pre-training for generalizable incomplete-view CT reconstruction
Chenglong Ma 0002, Zilong Li 0001, Junjun He, Junping Zhang, Yi Zhang 0018, Hongming Shan |
Pattern Recognit. | 6 |
| 2026 | Adaptive multi-view consistency clustering via structure-enhanced contrastive learning
Xuqian Xue, Zhanwei Zhang, Hongming Shan, Junping Zhang |
Pattern Recognit. | 5 |
| 2026 | Enhancing federated learning through exploring filter-aware relationships and personalizing local structures
Ziyuan Yang 0001, Zerui Shao, Huijie Huangfu, Andrew Beng Jin Teoh, Hongming Shan, Yi Zhang 0018 |
Pattern Recognit. | 7 |
| 2026 | Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation LearningabstractSelf-supervised learning (SSL) has emerged as a crucial technique in image processing, encoding, and understanding, especially for developing today's vision foundation models that utilize large-scale datasets without annotations to enhance various downstream tasks. This study introduces a novel SSL approach, Information-Maximized Soft Variable Discretization (IMSVD), for image representation learning. Specifically, IMSVD softly discretizes each variable in the latent space, enabling the estimation of their probability distributions over training batches and allowing the learning process to be directly guided by information measures. Motivated by the MultiView assumption, we propose an information-theoretic objective function to learn transform-invariant, non-trivial, and redundancy-minimized representation features. We then derive a cross-joint entropy loss function for self-supervised image representation learning, which theoretically enjoys superiority over the existing methods in reducing feature redundancy. Notably, our non-contrastive IMSVD method statistically performs contrastive learning. Extensive experimental results demonstrate the effectiveness of IMSVD on various downstream tasks in terms of both accuracy and efficiency. Thanks to our variable discretization, the embedding features optimized by IMSVD offer unique explainability at the variable level. IMSVD has the potential to be adapted to other learning paradigms. Our code is publicly available at https://github.com/niuchuangnn/IMSVD. Chuang Niu, Wenjun Xia, Hongming Shan, Ge Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT DenoisingabstractLow-dose computed tomography (CT) denoising is crucial for reduced radiation exposure while ensuring diagnostically acceptable image quality. Despite significant advancements driven by deep learning (DL) in recent years, existing DL-based methods, typically trained on a specific dose level and anatomical region, struggle to handle diverse noise characteristics and anatomical heterogeneity during varied scanning conditions, limiting their generalizability and robustness in clinical scenarios. In this paper, we propose FoundDiff, a foundational diffusion model for unified and generalizable LDCT denoising across various dose levels and anatomical regions. FoundDiff employs a two-stage strategy: (i) dose-anatomy perception and (ii) adaptive denoising. First, we develop a dose- and anatomy-aware contrastive language-image pre-training model (DA-CLIP) to achieve robust dose and anatomy perception by leveraging specialized contrastive learning strategies to learn continuous representations that quantify ordinal dose variations and identify salient anatomical regions. Second, we design a dose- and anatomy-aware diffusion model (DA-Diff) to perform adaptive and generalizable denoising by synergistically integrating the learned dose and anatomy embeddings from DA-CLIP into diffusion process via a novel dose and anatomy conditional block (DACB) based on Mamba. Extensive experiments on a large simulated multi-dose CT dataset spanning three anatomical regions, together with cross-dataset evaluations on Mayo-2016, CQ500, and piglet datasets, demonstrate superior denoising performance and strong generalization to unseen dose levels and anatomical regions. The codes and models are available at https://github.com/hao1635/FoundDiff. Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Jun Zhao 0010, Hongming Shan |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Editorial AI Reviewer (AIR) Trial for Responsible, Secure, and Efficient Peer ReviewabstractPeer review is central to the integrity of scientific publishing. At IEEE Transactions on Medical Imaging (TMI), thousands of reviewers and editors work each year to ensure that accepted papers meet our high standards of significance, innovation, evaluation, and reproducibility (SIER) [1]. Yet the rapid growth in submissions, the increasing complexity of papers, and the decreasing availability of reviewers place mounting pressure on the TMI peer review system. Ge Wang 0001, Tolga Çukur, Uwe Krüger 0001, Jennifer Ferina, Hongming Shan |
IEEE Trans. Medical Imaging | 5 |
| 2026 | SMART: Self-Supervised Learning for Metal Artifact Reduction in Computed Tomography Using Range Null Space DecompositionabstractMetal artifacts in computed tomography (CT) imaging significantly hinder diagnostic accuracy and clinical decision-making. While deep learning-based metal artifact reduction (MAR) methods have demonstrated promising progress, their clinical application is still constrained by three major challenges: 1) balancing metal artifact reduction with the preservation of critical anatomical structures, 2) effectively capturing the clinical priors of metal artifacts, and 3) dynamically adapting to polychromatic spectral variations. To address these limitations, in this paper, we propose a Self-supervised MAR method for computed Tomography (SMART) that leverages range-null space decomposition (RND) to model metal and tissue LACs separately, and employs implicit neural representation (INR) to learn their respective clinical characteristics without explicit supervision. Specifically, RND decouples metal and tissue LACs into a residual range component for metal LAC modeling, which captures metal artifacts, thus facilitating metal artifact reduction, and a null component for tissue LAC modeling, which focuses on preserving tissue details. To deal with the lack of paired data in clinical settings, we utilize INR to learn the clinical characteristics of these components in a self-supervised manner. Furthermore, SMART incorporates polychromatic spectra into the implicit representation, allowing dynamic adaptation to spectral variations across different imaging conditions. Extensive experiments on one synthetic and two clinical datasets demonstrate the strong potential of SMART in real-world scenarios. By flexibly adapting to spectral variations, it achieves superior generalizability to out-of-distribution clinical data. Yanxin Cao, Yongqiang Huang 0003, Jingfeng Lu, Fenglei Fan, Hongming Shan, Yi Zhang 0018 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Re-Thinking the Nature of Planning for Safe and Personalized Treatment Management Planning Using Large Language ModelsabstractWhile large language models have advanced di-agnostic reasoning in clinical domains, but to fully support the patient care, accurate and personalized treatment (illness) management plans are also needed. Unlike conventional planning tasks with defined goals and constraints, illness management planning navigates through uncertainty, incomplete data, and nonlinear, cross-disease effects. As a result, existing methods (e.g., Chain-of- Thought and Reflexion), which rely on assumptions of defined goals and linear reasoning, fall short in the complexities and ambiguities inherent in illness management planning. Another major challenge is the lack of a high-quality dataset that pairs real-world patient narratives with actionable illness management plans, which are crucial for evaluating the capabilities of LLMs in illness management planning. To address these challenges, we propose a novel planning method that reconceives treatment planning as the modulation of a patient's current illness state toward a healthy state without any defined goal. The approach introduces Attractive Tendencies, a latent, directional vectors that define desirable shifts toward healthier states, then uses field mapping to identify modifiable life domains that define personalized illness management goals, and applies field sculpting to generate safe, individualized, and actionable interventions. To enable standardized evaluation, we release an evaluation dataset comprising 1,015 patient cases, each paired with a real-world, complex narrative and a personalized treatment plan. Our method outperforms baseline approaches, achieving improvements of up to +12 BLEU-4 and +11 METEOR, while maintaining strong clinical relevance and computational efficiency. Muhammad Ayoub, Hai Zhao 0001, Dongjie Yang, Hongming Shan, Lifeng Li |
BIBM | 4 |
| 2025 | Patient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language ModelabstractReducing radiation doses benefits patients, but the resultant low-dose computed tomography (LDCT) images often suffer from clinically unacceptable noise and artifacts. While deep learning (DL) has shown promise in LDCT reconstruction, it requires large-scale data collection from multiple clients, raising privacy concerns. Federated learning (FL) has been introduced to mitigate these privacy concerns; however, current methods are typically tailored to specific scanning protocols, which limits their generalizability and makes them less effective for unseen protocols. To address these issues, we propose SCANPhysFed, a novel SCanning- and ANatomy-level personalized Physics-Driven Federated learning paradigm for LDCT reconstruction. Since the noise distribution in LDCT data is closely tied to scanning protocols and anatomical structures, we propose a dual-level physics-informed way to address these challenges. Specifically, we incorporate physical and anatomical prompts into our physics-informed hypernetworks to capture scanning- and anatomy-specific information, enabling dual-level physics-driven personalization of imaging features. These prompts are derived from the scanning protocol and the radiology report generated by a medical large language model (MLLM). Subsequently, client-specific decoders project these dual-level personalized imaging features back into the image domain. Besides, to tackle the challenge of unseen data, we introduce a novel protocol vector-quantization strategy (PVQS), which ensures consistent performance across new clients by quantifying unseen scanning codes to the closest match in the scanning codebook. Extensive experimental results demonstrate the superior performance of SCAN-PhysFed on public datasets1. Ziyuan Yang 0001, Zhiwen Wang 0002, Hongming Shan, Yang Chen 0008, Yi Zhang 0018 |
CVPR | 4 |
| 2025 | DepMamba: Progressive Fusion Mamba for Multimodal Depression DetectionabstractDepression is a common mental disorder that affects millions of people worldwide. Although promising, current multimodal methods hinge on aligned or aggregated multi-modal fusion, suffering two significant limitations: (i) inefficient long-range temporal modeling, and (ii) sub-optimal multimodal fusion between intermodal fusion and intramodal processing. In this paper, we propose an audio-visual progressive fusion Mamba for multimodal depression detection, termed DepMamba. DepMamba features two core designs: hierarchical contextual modeling and progressive multimodal fusion. On the one hand, hierarchical modeling introduces convolution neural networks and Mamba to extract the local-to-global features within long-range sequences. On the other hand, the progressive fusion first presents a multimodal collaborative State Space Model (SSM) extracting intermodal and intramodal information for each modality, and then utilizes a multimodal enhanced SSM for modality cohesion. Extensive experimental results on two large-scale depression datasets demonstrate the superior performance of our DepMamba over existing state-of-the-art methods. Code is available at https://github.com/Jiaxin-Ye/DepMamba. Jiaxin Ye, Junping Zhang, Hongming Shan |
ICASSP | 3 |
| 2025 | DreamRelation: Relation-Centric Video Customization
Yujie Wei 0001, Shiwei Zhang 0001, Hangjie Yuan, Biao Gong, Longxiang Tang, Xiang Wang 0012, Haonan Qiu, Hengjia Li, Yingya Zhang, Hongming Shan |
ICCV | 11 |
| 2025 | PROTOCOL: Partial Optimal Transport-enhanced Contrastive Learning for Imbalanced Multi-view ClusteringabstractWhile contrastive multi-view clustering has achieved remarkable success, it implicitly assumes balanced class distribution. However, real-world multi-view data primarily exhibits class imbalance distribution. Consequently, existing methods suffer performance degradation due to their inability to perceive and model such imbalance. To address this challenge, we present the first systematic study of imbalanced multi-view clustering, focusing on two fundamental problems: i. perceiving class imbalance distribution, and ii. mitigating representation degradation of minority samples. We propose PROTOCOL, a novel PaRtial Optimal TranspOrt-enhanced COntrastive Learning framework for imbalanced multi-view clustering. First, for class imbalance perception, we map multi-view features into a consensus space and reformulate the imbalanced clustering as a partial optimal transport (POT) problem, augmented with progressive mass constraints and weighted KL divergence for class distributions. Second, we develop a POT-enhanced class-rebalanced contrastive learning at both feature and class levels, incorporating logit adjustment and class-sensitive learning to enhance minority sample representations. Extensive experiments demonstrate that PROTOCOL significantly improves clustering performance on imbalanced multi-view data, filling a critical research gap in this field. Xuqian Xue, Hongming Shan, Junping Zhang |
ICML | 4 |
| 2025 | Emotional Face-to-SpeechabstractHow much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders. Existing face-to-speech methods offer great promise in capturing identity characteristics but struggle to generate diverse vocal styles with emotional expression. In this paper, we explore a new task, termed emotional face-to-speech, aiming to synthesize emotional speech directly from expressive facial cues. To that end, we introduce DEmoFace, a novel generative framework that leverages a discrete diffusion transformer (DiT) with curriculum learning, built upon a multi-level neural audio codec. Specifically, we propose multimodal DiT blocks to dynamically align text and speech while tailoring vocal styles based on facial emotion and identity. To enhance training efficiency and generation quality, we further introduce a coarse-to-fine curriculum learning algorithm for multi-level token processing. In addition, we develop an enhanced predictor-free guidance to handle diverse conditioning scenarios, enabling multi-conditional generation and disentangling complex attributes effectively. Extensive experimental results demonstrate that DEmoFace generates more natural and consistent speech compared to baselines, even surpassing speech-driven methods. Demos of DEmoFace are shown at our project https://demoface.github.io. Jiaxin Ye, Boyuan Cao, Hongming Shan |
ICML | 3 |
| 2025 | Autoregressive Medical Image Segmentation via Next-Scale Mask Prediction
Tao Chen 0055, Hongming Shan |
MICCAI (2) | 4 |
| 2025 | Towards Interpretable Counterfactual Generation via Multimodal Autoregression
Chenglong Ma 0002, Yuanfeng Ji, Jin Ye 0002, Lu Zhang 0060, Tianbin Li, Mingjie Li 0006, Junjun He, Hongming Shan |
MICCAI (2) | 9 |
| 2025 | RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image GenerationabstractWhile latent diffusion models (LDMs), such as Stable Diffusion, are designed for high-resolution image generation, they often struggle with significant structural distortions when generating images at resolutions higher than their training one.
Instead of relying on extensive retraining, a more resource-efficient approach is to reprogram the pretrained model for high-resolution (HR) image generation; however, existing methods often result in poor image quality and long inference time.
We introduce RepLDM, a novel reprogramming framework for pretrained LDMs that enables high-quality, high-efficiency, high-resolution image generation; see Fig. 1. RepLDM consists of two stages: (i) an attention guidance stage, which generates a latent representation of a higher-quality training-resolution image using a novel parameter-free self-attention mechanism to enhance the structural consistency; and (ii) a progressive upsampling stage, which progressively performs upsampling in pixel space to mitigate the severe artifacts caused by latent space upsampling. The effective initialization from the first stage allows for denoising at higher resolutions with significantly fewer steps, improving the efficiency.
Extensive experimental results demonstrate that RepLDM significantly outperforms state-of-the-art methods in both quality and efficiency for HR image generation, underscoring its advantages for real-world applications.
Codes: https://github.com/kmittle/RepLDM. Boyuan Cao, Jiaxin Ye, Yujie Wei 0001, Hongming Shan |
NeurIPS | 4 |
| 2025 | Noise-inspired diffusion model for generalizable low-dose CT reconstruction
Dong Zeng, Junping Zhang, Hongming Shan |
Medical Image Anal. | 6 |
| 2025 | Radiologist-in-the-Loop Self-Training for Generalizable CT Metal Artifact ReductionabstractMetal artifacts in computed tomography (CT) images can significantly degrade image quality and impede accurate diagnosis. Supervised metal artifact reduction (MAR) methods, trained using simulated datasets, often struggle to perform well on real clinical CT images due to a substantial domain gap. Although state-of-the-art semi-supervised methods use pseudo ground-truths generated by a prior network to mitigate this issue, their reliance on a fixed prior limits both the quality and quantity of these pseudo ground-truths, introducing confirmation bias and reducing clinical applicability. To address these limitations, we propose a novel radiologist-in-the-loop self-training framework for MAR, termed RISE-MAR, which can integrate radiologists' feedback into the semi-supervised learning process, progressively improving the quality and quantity of pseudo ground-truths for enhanced generalization on real clinical CT images. For quality assurance, we introduce a clinical quality assessor model that emulates radiologist evaluations, effectively selecting high-quality pseudo ground-truths for semi-supervised training. For quantity assurance, our self-training framework iteratively generates additional high-quality pseudo ground-truths, expanding the clinical dataset and further improving model generalization. Extensive experimental results on multiple clinical datasets demonstrate the superior generalization performance of our RISE-MAR over state-of-the-art methods, advancing the development of MAR models for practical application. The source code is available at https://github.com/Masaaki-75/rise-mar. Chenglong Ma 0002, Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Jiannan Liu, Hongming Shan |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Editorial Criteria for TMI Papers - Significance, Innovation, Evaluation, and ReproducibilityabstractIEEE Transactions on Medical Imaging (TMI) publishes high-quality work that innovates imaging methods and advances medicine, science, and engineering. While artificial intelligence (AI) is currently prominent, the journal's scope extends well beyond AI-based imaging to encompass a full spectrum of imaging methods involving CT, MRI, PET, SPECT, ultrasound, optical, and hybrid systems, image reconstruction and processing (ranging from analytical and iterative algorithms to emerging deep imaging approaches), quantitative imaging and analysis (radiomics, biomarkers, and health analytics), image-guided interventions and therapy, as well as multimodal and multiscale imaging with integration of imaging and nonimaging data. Hongming Shan, Uwe Krüger 0001, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | UniAda: Domain Unifying and Adapting Network for Generalizable Medical Image SegmentationabstractLearning a generalizable medical image segmentation model is an important but challenging task since the unseen (testing) domains may have significant discrepancies from seen (training) domains due to different vendors and scanning protocols. Existing segmentation methods, typically built upon domain generalization (DG), aim to learn multi-source domain-invariant features through data or feature augmentation techniques, but the resulting models either fail to characterize global domains during training or cannot sense unseen domain information during testing. To tackle these challenges, we propose a domain Unifying and Adapting network (UniAda) for generalizable medical image segmentation, a novel "unifying while training, adapting while testing" paradigm that can learn a domain-aware base model during training and dynamically adapt it to unseen target domains during testing. First, we propose to unify the multi-source domains into a global inter-source domain via a novel feature statistics update mechanism, which can sample new features for the unseen domains, facilitating the training of a domain base model. Second, we leverage the uncertainty map to guide the adaptation of the trained model for each testing sample, considering the specific target domain may be outside the global inter-source domain. Extensive experimental results on two public cross-domain medical datasets and one in-house cross-domain dataset demonstrate the strong generalization capacity of the proposed UniAda over state-of-the-art DG methods. The source code of our UniAda is available at https://github.com/ZhouZhang233/UniAda. Zhongzhou Zhang, Zhiwen Wang 0002, Shanshan Wang 0008, Fenglei Fan, Hongming Shan, Yi Zhang 0018 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Deep Rank-Consistent Pyramid Model for Enhanced Crowd CountingabstractMost conventional crowd counting methods utilize a fully-supervised learning framework to establish a mapping between scene images and crowd density maps. They usually rely on a large quantity of costly and time-intensive pixel-level annotations for training supervision. One way to mitigate the intensive labeling effort and improve counting accuracy is to leverage large amounts of unlabeled images. This is attributed to the inherent self-structural information and rank consistency within a single image, offering additional qualitative relation supervision during training. Contrary to earlier methods that utilized the rank relations at the original image level, we explore such rank-consistency relation within the latent feature spaces. This approach enables the incorporation of numerous pyramid partial orders, strengthening the model representation capability. A notable advantage is that it can also increase the utilization ratio of unlabeled samples. Specifically, we propose a Deep Rank-consist Ent pyrAmid Model (DREAM), which makes full use of rank consistency across coarse-to-fine pyramid features in latent spaces for enhanced crowd counting with massive unlabeled images. In addition, we have collected a new unlabeled crowd counting dataset, FUDAN-UCC, comprising 4000 images for training purposes. Extensive experiments on four benchmark datasets, namely UCF-QNRF, ShanghaiTech PartA and PartB, and UCF-CC-50, show the effectiveness of our method compared with previous semi-supervised methods. The codes are available at https://github.com/bridgeqiqi/DREAM. Zhizhong Huang, Hongming Shan, James Z. Wang 0001, Fei-Yue Wang 0001, Junping Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Low-dose CT Denoising with Language-engaged Dual-space AlignmentabstractWhile various deep learning methods were proposed for low-dose computed tomography (CT) denoising, they often suffer from over-smoothing, blurring, and lack of explainability. To alleviate these issues, we propose a plug-and-play Language-Engaged Dual-space Alignment loss (LEDA) to optimize low-dose CT denoising models. Our idea is to leverage large language models (LLMs) to align denoised CT and normal-dose CT images in both the continuous perceptual space and discrete semantic space, which is the first LLM-based scheme for low-dose CT denoising. LEDA involves two steps: the first is to pretrain an LLM-guided CT autoencoder, which can encode a CT image into continuous high-level features and quantize them into a token space to produce semantic tokens derived from the LLM’s vocabulary; and the second is to minimize the discrepancy between the denoised CT images and normal-dose CT in terms of both encoded high-level features and quantized token embeddings derived by the LLM-guided CT autoencoder. Extensive experimental results demonstrate that our LEDA can enhance existing denoising models in terms of quantitative metrics and qualitative evaluation, and also provide explainability through language-level image understanding. The code is publicly available at https://github.com/hao1635/LEDA. Tao Chen 0055, Chuang Niu, Ge Wang 0001, Hongming Shan |
BIBM | 7 |
| 2024 | Dream Video: Composing Your Dream Videos with Customized Subject and MotionabstractCustomized generation using diffusion models has made impressive progress in image generation, but remains unsatisfactory in the challenging video generation task, as it requires the controllability of both subjects and motions. To that end, we present DreamVideo, a novel approach to generating personalized videos from a few static images of the desired subject and a few videos of target motion. DreamVideo decouples this task into two stages, subject learning and motion learning, by leveraging a pre-trained video diffusion model. The subject learning aims to accurately capture the fine appearance of the subject from provided images, which is achieved by combining textual inversion and fine-tuning of our carefully designed identity adapter. In motion learning, we architect a motion adapter and fine-tune it on the given videos to effectively model the target motion pattern. Combining these two lightweight and efficient adapters allows for flexible customization of any subject with any motion. Extensive experimental results demonstrate the superior performance of our DreamVideo over the state-of-the-art methods for customized video generation. Our project page is at https://dreamvideo-t2v.github.io. Yujie Wei 0001, Shiwei Zhang 0001, Zhiwu Qing, Hangjie Yuan, Yu Liu 0063, Yingya Zhang, Jingren Zhou 0001, Hongming Shan |
CVPR | 9 |
| 2024 | Point, Segment and Count: A Generalized Framework for Object CountingabstractClass-agnostic object counting aims to count all objects in an image with respect to example boxes or class names, a.k.a few-shot and zero-shot counting. In this paper, we propose a generalized framework for both few-shot and zero-shot object counting based on detection. Our framework combines the superior advantages of two foundation models without compromising their zero-shot capability: (i) SAM to segment all possible objects as mask proposals, and (ii) CLIP to classify proposals to obtain accurate object counts. However, this strategy meets the obstacles of efficiency over-head and the small crowded objects that cannot be localized and distinguished. To address these issues, our framework, termed PseCo, follows three steps: point, segment, and count. Specifically, we first propose a class-agnostic object localization to provide accurate but least point prompts for SAM, which consequently not only reduces computation costs but also avoids missing small objects. Furthermore, we propose a generalized object classification that leverages CLIP image/text embeddings as the classifier, following a hierarhical knowledge distillation to obtain discriminative classifications among hierarchical mask proposals. Extensive experimental results on FSC-147, COCO, and LVIS demonstrate that PseCo achieves state-of-the-art performance in both few-shot/zero-shot object counting/detection. Zhizhong Huang, Mingliang Dai, Yi Zhang 0018, Junping Zhang, Hongming Shan |
CVPR | 5 |
| 2024 | Semantic Latent Decomposition with Normalizing Flows for Face EditingabstractNavigating in the latent space of StyleGAN has shown effectiveness for face editing. However, the resulting methods usually encounter challenges in complicated navigation due to the entanglement among different attributes in the latent space. To address this issue, this paper proposes a novel framework, termed SDFlow, with a semantic decomposition in original latent space using continuous conditional normalizing flows. Specifically, SDFlow decomposes the original latent code into different irrelevant variables by jointly optimizing two components: (i) a semantic encoder to estimate semantic variables from input faces and (ii) a flow-based transformation module to map the latent code into a semantic-irrelevant variable in Gaussian distribution, conditioned on the learned semantic variables. To eliminate the entanglement between variables, we employ a disentangled learning strategy under a mutual information framework, thereby providing precise manipulation controls. Experimental results demonstrate that SDFlow outperforms existing state-of-the-art face editing methods both qualitatively and quantitatively. The source code is available at https://github.com/phil329/SDFlow. Binglei Li, Zhizhong Huang, Hongming Shan, Junping Zhang |
ICASSP | 3 |
| 2024 | SIAM: A Simple Alternating Mixer for Video PredictionabstractVideo prediction, predicting future frames from the previous ones, has broad applications such as autonomous driving and weather forecasting. Existing state-of-the-art methods typically focus on extracting either spatial, temporal, or spatiotemporal features from videos. Different feature focuses, resulting from different network architectures, may make the resultant models excel at some video prediction tasks but perform poorly on others. Towards a more generic video prediction solution, we explicitly model these features in a unified encoder-decoder framework and propose a simple alternating Mixer (SIAM). The novelty of SIAM lies in the design of dimension alternating mixing (DaMi) blocks, which can model spatial, temporal, and spatiotemporal features through alternating the dimensions of the feature maps. Extensive experimental results demonstrate the superior performance of the proposed SIAM on four benchmark video datasets covering both synthetic and real-world scenarios. Ziang Peng, Hongming Shan, Junping Zhang |
ICME | 4 |
| 2024 | FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on
Tao Chen 0055, Zhizhong Huang, Taoran Jiang, Hongming Shan |
IJCAI | 7 |
| 2024 | Denoising Diffusion Path: Attribution Noise Reduction with An Auxiliary Diffusion ModelabstractThe explainability of deep neural networks (DNNs) is critical for trust and reliability in AI systems. Path-based attribution methods, such as integrated gradients (IG), aim to explain predictions by accumulating gradients along a path from a baseline to the target image. However, noise accumulated during this process can significantly distort the explanation. While existing methods primarily concentrate on finding alternative paths to circumvent noise, they overlook a critical issue: intermediate-step images frequently diverge from the distribution of training data, further intensifying the impact of noise. This work presents a novel Denoising Diffusion Path (DDPath) to tackle this challenge by harnessing the power of diffusionmodels for denoising. By exploiting the inherent ability of diffusion models to progressively remove noise from an image, DDPath constructs a piece-wise linear path. Each segment of this path ensures that samples drawn from a Gaussian distribution are centered around the target image. This approach facilitates a gradual reduction of noise along the path. We further demonstrate that DDPath adheres to essential axiomatic properties for attribution methods and can be seamlessly integrated with existing methods such as IG. Extensive experimental results demonstrate that DDPath can significantly reduce noise in the attributions—resulting in clearer explanations—and achieves better quantitative results than traditional path-based methods. Zilong Li 0001, Junping Zhang, Hongming Shan |
NeurIPS | 4 |
| 2024 | Prompt learning in computer vision: a surveyabstractPrompt learning has attracted broad attention in computer vision since the large pre-trained vision-language models (VLMs) exploded. Based on the close relationship between vision and language information built by VLM, prompt learning becomes a crucial technique in many important applications such as artificial intelligence generated content (AIGC). In this survey, we provide a progressive and comprehensive review of visual prompt learning as related to AIGC. We begin by introducing VLM, the foundation of visual prompt learning. Then, we review the vision prompt learning methods and prompt-guided generative models, and discuss how to improve the efficiency of adapting AIGC models to specific downstream tasks. Finally, we provide some promising research directions concerning prompt learning. Zilong Li 0001, Hongming Shan |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2024 | Joint learning framework of cross-modal synthesis and diagnosis for Alzheimer's disease by mining underlying shared modality informationabstractAlzheimer's disease (AD) is one of the most common neurodegenerative disorders presenting irreversible progression of cognitive impairment. How to identify AD as early as possible is critical for intervention with potential preventive measures. Among various neuroimaging modalities used to diagnose AD, functional positron emission tomography (PET) has higher sensitivity than structural magnetic resonance imaging (MRI), but it is also costlier and often not available in many hospitals. How to leverage massive unpaired unlabeled PET to improve the diagnosis performance of AD from MRI becomes rather important. To address this challenge, this paper proposes a novel joint learning framework of unsupervised cross-modal synthesis and AD diagnosis by mining underlying shared modality information, improving the AD diagnosis from MRI while synthesizing more discriminative PET images. We mine underlying shared modality information in two aspects: diversifying modality information through the cross-modal synthesis network and locating critical diagnosis-related patterns through the AD diagnosis network. First, to diversify the modality information, we propose a novel unsupervised cross-modal synthesis network, which implements the inter-conversion between 3D PET and MRI in a single model modulated by the AdaIN module. Second, to locate shared critical diagnosis-related patterns, we propose an interpretable diagnosis network based on fully 2D convolutions, which takes either 3D synthesized PET or original MRI as input. Extensive experimental results on the ADNI dataset show that our framework can synthesize more realistic images, outperform the state-of-the-art AD diagnosis methods, and have better generalization on external AIBL and NACC datasets. Sirong Piao, Zhizhong Huang, Junping Zhang, Hongming Shan |
Medical Image Anal. | 7 |
| 2024 | CORE: Learning consistent ordinal representations with convex optimization for image ordinal estimation
Zilong Li 0001, Junping Zhang, Hongming Shan |
Pattern Recognit. | 5 |
| 2024 | HOPE: Hybrid-Granularity Ordinal Prototype Learning for Progression Prediction of Mild Cognitive ImpairmentabstractMild cognitive impairment (MCI) is often at high risk of progression to Alzheimer's disease (AD). Existing works to identify the progressive MCI (pMCI) typically require MCI subtype labels, pMCI vs. stable MCI (sMCI), determined by whether or not an MCI patient will progress to AD after a long follow-up. However, prospectively acquiring MCI subtype data is time-consuming and resource-intensive; the resultant small datasets could lead to severe overfitting and difficulty in extracting discriminative information. Inspired by that various longitudinal biomarkers and cognitive measurements present an ordinal pathway on AD progression, we propose a novel Hybrid-granularity Ordinal PrototypE learning (HOPE) method to characterize AD ordinal progression for MCI progression prediction. First, HOPE learns an ordinal metric space that enables progression prediction by prototype comparison. Second, HOPE leverages a novel hybrid-granularity ordinal loss to learn the ordinal nature of AD via effectively integrating instance-to-instance ordinality, instance-to-class compactness, and class-to-class separation. Third, to make the prototype learning more stable, HOPE employs an exponential moving average strategy to learn the global prototypes of NC and AD dynamically. Experimental results on the internal ADNI and the external NACC datasets demonstrate the superiority of the proposed HOPE over existing state-of-the-art methods as well as its interpretability. Tao Chen 0055, Junping Zhang, Hongming Shan |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | LIT-Former: Linking In-Plane and Through-Plane Transformers for Simultaneous CT Image Denoising and DeblurringabstractThis paper studies 3D low-dose computed tomography (CT) imaging. Although various deep learning methods were developed in this context, typically they focus on 2D images and perform denoising due to low-dose and deblurring for super-resolution separately. Up to date, little work was done for simultaneous in-plane denoising and through-plane deblurring, which is important to obtain high-quality 3D CT images with lower radiation and faster imaging speed. For this task, a straightforward method is to directly train an end-to-end 3D network. However, it demands much more training data and expensive computational costs. Here, we propose to link in-plane and through-plane transformers for simultaneous in-plane denoising and through-plane deblurring, termed as LIT-Former, which can efficiently synergize in-plane and through-plane sub-tasks for 3D CT imaging and enjoy the advantages of both convolution and transformer networks. LIT-Former has two novel designs: efficient multi-head self-attention modules (eMSM) and efficient convolutional feed-forward networks (eCFN). First, eMSM integrates in-plane 2D self-attention and through-plane 1D self-attention to efficiently capture global interactions of 3D self-attention, the core unit of transformer networks. Second, eCFN integrates 2D convolution and 1D convolution to extract local information of 3D convolution in the same fashion. As a result, the proposed LIT-Former synergizes these two sub-tasks, significantly reducing the computational complexity as compared to 3D counterparts and enabling rapid convergence. Extensive experimental results on simulated and clinical datasets demonstrate superior performance over state-of-the-art models. The source code is made available at https://github.com/hao1635/LIT-Former. Chuang Niu, Ge Wang 0001, Hongming Shan |
IEEE Trans. Medical Imaging | 5 |
| 2024 | HiDiff: Hybrid Diffusion Framework for Medical Image SegmentationabstractMedical image segmentation has been significantly advanced with the rapid development of deep learning (DL) techniques. Existing DL-based segmentation models are typically discriminative; i.e., they aim to learn a mapping from the input image to segmentation masks. However, these discriminative methods neglect the underlying data distribution and intrinsic class characteristics, suffering from unstable feature space. In this work, we propose to complement discriminative segmentation methods with the knowledge of underlying data distribution from generative models. To that end, we propose a novel hybrid diffusion framework for medical image segmentation, termed HiDiff, which can synergize the strengths of existing discriminative segmentation models and new generative diffusion models. HiDiff comprises two key components: discriminative segmentor and diffusion refiner. First, we utilize any conventional trained segmentation models as discriminative segmentor, which can provide a segmentation mask prior for diffusion refiner. Second, we propose a novel binary Bernoulli diffusion model (BBDM) as the diffusion refiner, which can effectively, efficiently, and interactively refine the segmentation mask by modeling the underlying data distribution. Third, we train the segmentor and BBDM in an alternate-collaborative manner to mutually boost each other. Extensive experimental results on abdomen organ, brain tumor, polyps, and retinal vessels segmentation datasets, covering four widely-used modalities, demonstrate the superior performance of HiDiff over existing medical segmentation algorithms, including the state-of-the-art transformer- and diffusion-based ones. In addition, HiDiff excels at segmenting small objects and generalizing to new datasets. Source codes are made available at https://github.com/takimailto/HiDiff. Tao Chen 0055, Hongming Shan |
IEEE Trans. Medical Imaging | 5 |
| 2024 | CoreDiff: Contextual Error-Modulated Generalized Diffusion Model for Low-Dose CT Denoising and GeneralizationabstractLow-dose computed tomography (CT) images suffer from noise and artifacts due to photon starvation and electronic noise. Recently, some works have attempted to use diffusion models to address the over-smoothness and training instability encountered by previous deep-learning-based denoising models. However, diffusion models suffer from long inference time due to a large number of sampling steps involved. Very recently, cold diffusion model generalizes classical diffusion models and has greater flexibility. Inspired by cold diffusion, this paper presents a novel COntextual eRror-modulated gEneralized Diffusion model for low-dose CT (LDCT) denoising, termed CoreDiff. First, CoreDiff utilizes LDCT images to displace the random Gaussian noise and employs a novel mean-preserving degradation operator to mimic the physical process of CT degradation, significantly reducing sampling steps thanks to the informative LDCT images as the starting point of the sampling process. Second, to alleviate the error accumulation problem caused by the imperfect restoration operator in the sampling process, we propose a novel ContextuaL Error-modulAted Restoration Network (CLEAR-Net), which can leverage contextual information to constrain the sampling process from structural distortion and modulate time step embedding features for better alignment with the input at the next time step. Third, to rapidly generalize the trained model to a new, unseen dose level with as few resources as possible, we devise a one-shot learning framework to make CoreDiff generalize faster and better using only one single LDCT image (un)paired with normal-dose CT (NDCT). Extensive experimental results on four datasets demonstrate that our CoreDiff outperforms competing methods in denoising and generalization performance, with clinically acceptable inference time. Source code is made available at https://github.com/qgao21/CoreDiff. Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Hongming Shan |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Quad-Net: Quad-Domain Network for CT Metal Artifact ReductionabstractMetal implants and other high-density objects in patients introduce severe streaking artifacts in CT images, compromising image quality and diagnostic performance. Although various methods were developed for CT metal artifact reduction over the past decades, including the latest dual-domain deep networks, remaining metal artifacts are still clinically challenging in many cases. Here we extend the state-of-the-art dual-domain deep network approach into a quad-domain counterpart so that all the features in the sinogram, image, and their corresponding Fourier domains are synergized to eliminate metal artifacts optimally without compromising structural subtleties. Our proposed quad-domain network for MAR, referred to as Quad-Net, takes little additional computational cost since the Fourier transform is highly efficient, and works across the four receptive fields to learn both global and local features as well as their relations. Specifically, we first design a Sinogram-Fourier Restoration Network (SFR-Net) in the sinogram domain and its Fourier space to faithfully inpaint metal-corrupted traces. Then, we couple SFR-Net with an Image-Fourier Refinement Network (IFR-Net) which takes both an image and its Fourier spectrum to improve a CT image reconstructed from the SFR-Net output using cross-domain contextual information. Quad-Net is trained on clinical datasets to minimize a composite loss function. Quad-Net does not require precise metal masks, which is of great importance in clinical practice. Our experimental results demonstrate the superiority of Quad-Net over the state-of-the-art MAR methods quantitatively, visually, and statistically. The Quad-Net code is publicly available at https://github.com/longzilicart/Quad-Net. Zilong Li 0001, Yaping Wu, Chuang Niu, Junping Zhang, Ge Wang 0001, Hongming Shan |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Twin Contrastive Learning with Noisy LabelsabstractLearning from noisy data is a challenging task that sig-nificantly degenerates the model performance. In this paper, we present TCL, a novel twin contrastive learning model to learn robust representations and handle noisy labels for classification. Specifically, we construct a Gaussian mixture model (GMM) over the representations by injecting the supervised model predictions into GMM to link label- free latent variables in GMM with label-noisy annotations. Then, TCL detects the examples with wrong labels as the out- of-distribution examples by another two-component GMM, taking into account the data distribution. We further propose a cross-supervision with an entropy regularization loss that bootstraps the true targets from model predictions to handle the noisy labels. As a result, TCL can learn discriminative representations aligned with estimated labels through mixup and contrastive learning. Extensive experimental results on several standard benchmarks and real-world datasets demonstrate the superior performance of TCL. In particular, TCL achieves 7.5% improvements on CIFAR-10 with 90% noisy label-an extremely noisy scenario. The source code is available at https://github.com/Hzzone/TCL. Zhizhong Huang, Junping Zhang, Hongming Shan |
CVPR | 3 |
| 2023 | Mutual Information Based Reweighting for Precipitation NowcastingabstractPrecipitation nowcasting uses previous rainfall observations to forecast future rainfall intensities in a local area. In rainfall data, the rain-less samples usually well exceed the heavy rainfall samples, and it causes the data imbalance problem in precipitation nowcasting tasks. In this paper, we find that if the imbalance ratio is fixed, tasks with higher mutual information make the nowcasting model more robust to the data imbalance problem. Based on this observation, we propose a mutual information-based reweighting strategy. The reweighting strategy allows the neural network models to achieve better performance on minorities without compromising the performance of majorities and overall nowcasting image quality. Extensive experimental results demonstrate that this proposed approach is effective and compatible with state-of-the-art models. Danchen Zhang, Hongming Shan, Junping Zhang |
ICASSP | 4 |
| 2023 | Cross-Head Supervision for Crowd Counting with Noisy AnnotationsabstractNoisy annotations such as missing annotations and location shifts often exist in crowd counting datasets due to multi-scale head sizes, high occlusion, etc. These noisy annotations severely affect the model training, especially for density map-based methods. To alleviate the negative impact of noisy annotations, we propose a novel crowd counting model with one convolution head and one transformer head, in which these two heads can supervise each other in noisy areas, called Cross-Head Supervision. The resultant model, CHS-Net, can synergize different types of inductive biases for better counting. In addition, we develop a progressive cross-head supervision learning strategy to stabilize the training process and provide more reliable supervision. Extensive experimental results on Shang-haiTech and QNRF datasets demonstrate superior performance over state-of-the-art methods. Code is available at https://github.com/RaccoonDML/CHSNet. Mingliang Dai, Zhizhong Huang, Hongming Shan, Junping Zhang |
ICASSP | 4 |
| 2023 | Motion Matters: A Novel Motion Modeling for Cross-View Gait Feature LearningabstractAs a unique biometric that can be perceived at a distance, gait has broad applications in person authentication, social security and so on. Existing gait recognition methods suffer from changes in viewpoint and clothing and barely consider extracting diverse motion features, a fundamental characteristic in gaits, from gait sequences. This paper proposes a novel motion modeling method to extract the discriminative and robust representation. Specifically, we first extract the motion features from the encoded motion sequences in the shallow layer. Then we continuously enhance the motion feature in deep layers. This motion modeling approach is independent of mainstream work in building network architectures. As a result, one can apply this motion modeling method to any backbone to improve gait recognition performance. In this paper, we combine motion modeling with one commonly used backbone (GaitGL) as GaitGL-M to illustrate motion modeling. Extensive experimental results on two commonly-used crossview gait datasets demonstrate the superior performance of GaitGL-M over existing state-of-the-art methods. Hongming Shan, Junping Zhang |
ICASSP | 4 |
| 2023 | Gaitcotr: Improved Spatial-Temporal Representation for Gait Recognition with a Hybrid Convolution-Transformer FrameworkabstractThis work presents a novel hybrid convolution-transformer framework for gait recognition, termed GaitCoTr. The developed framework captures the appearance and short-term temporal features by convolution and extracts the long-term temporal features by transformer architecture, achieving a comprehensive spatial-temporal representation of gait. To unleash the potential of this hybrid framework for extracting richness and generalized temporal features, we propose a new variant of transformer tailored for gait, including temporally shifted tokenization, length-flexible position embedding, and inter-frame encoder. In addition, we introduce an auxiliary task—view label prediction—aiming to disentangle view from ID information. Extensive experimental results on two well-known gait benchmark datasets, CASIA-B and GREW, demonstrate the superior performance of the proposed Gait-CoTr. Hongming Shan, Junping Zhang |
ICASSP | 3 |
| 2023 | Temporal Modeling Matters: A Novel Temporal Emotional Modeling Approach for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) plays a vital role in improving the interactions between humans and machines by inferring human emotion and affective states from speech signals. Whereas recent works primarily focus on mining spatiotemporal information from hand-crafted features, we explore how to model the temporal patterns of speech emotions from dynamic temporal scales. Towards that goal, we introduce a novel temporal emotional modeling approach for SER, termed Temporal-aware bI-direction Multi-scale Network (TIM-Net), which learns multi-scale contextual affective representations from various time scales. Specifically, TIM-Net first employs temporal-aware blocks to learn temporal affective representation, then integrates complementary information from the past and the future to enrich contextual representations, and finally fuses multiple time scale features for better adaptation to the emotional variation. Extensive experimental results on six benchmark SER datasets demonstrate the superior performance of TIM-Net, gaining 2.34% and 2.61% improvements of the average UAR and WAR over the second-best on each corpus. The source code is available at https://github.com/Jiaxin-Ye/TIM-Net_SER. Jiaxin Ye, Xin-Cheng Wen, Yujie Wei 0001, Yong Xu 0009, Kunhong Liu 0001, Hongming Shan |
ICASSP | 6 |
| 2023 | Fan-Net: Fourier-Based Adaptive Normalization for Cross-Domain Stroke Lesion SegmentationabstractSince stroke is the main cause of various cerebrovascular diseases, deep learning-based stroke lesion segmentation on magnetic resonance (MR) images has attracted considerable attention. However, the existing methods often neglect the domain shift among MR images collected from different sites, which has limited performance improvement. To address this problem, we intend to change style information without affecting high-level semantics via adaptively changing the low-frequency amplitude components of the Fourier transform so as to enhance model robustness to varying domains. Thus, we propose a novel FAN-Net, a U-Net–based segmentation network incorporated with a Fourier-based adaptive normalization (FAN) and a domain classifier with a gradient reversal layer. The FAN module is tailored for learning adaptive affine parameters for the amplitude components of different domains, which can dynamically normalize the style information of source images. Then, the domain classifier provides domain-agnostic knowledge to endow FAN with strong domain generalizability. The experimental results on the ATLAS dataset, which consists of MR images from 9 sites, show the superior performance of the proposed FAN-Net compared with baseline methods. Weiyi Yu, Hongming Shan |
ICASSP | 3 |
| 2023 | DO-FAM: Disentangled Non-Linear Latent Navigation For Facial Attribute ManipulationabstractFacial attribute manipulation (FAM) aims to edit the semantic attributes of facial images according to the user’s requirements. Unfortunately, the majority of existing FAM methods struggle in meeting at least one of the two requirements: high reconstruction quality and high irrelevance preservation. To alleviate these two limitations, we propose a novel Disentangled nOn-linear latent navigation framework for FAM, termed DO-FAM. To promote the reconstruction quality, we leverage hypernetworks to fine-tune a pre-trained StyleGAN2 generator. To decouple entangled attributes, we propose a novel Disentangled nOn-Linear Latent transformation module, named DOLL, which consists of three components: (1) a decomposer to factorize input latent codes into two parts: attribute-related and attribute-unrelated; (2) a non-linear Latent Transformation Network (LTNet) to navigate the attribute-related latent codes to the target one with respect to the designed attribute(s); and (3) a latent classifier tasked with predicting latent codes’ attributes to guide the latent code navigation. Extensive experimental results on a widely-used benchmark facial editing dataset, CelebA-HQ, demonstrate the superiority of our method over state-of-the-art methods. Yifan Yuan 0001, Siteng Ma, Hongming Shan, Junping Zhang |
ICASSP | 3 |
| 2023 | Online Prototype Learning for Online Continual LearningabstractOnline continual learning (CL) studies the problem of learning continuously from a single-pass data stream while adapting to new data and mitigating catastrophic forgetting. Recently, by storing a small subset of old data, replay-based methods have shown promising performance. Unlike previous methods that focus on sample storage or knowledge distillation against catastrophic forgetting, this paper aims to understand why the online learning models fail to generalize well from a new perspective of shortcut learning. We identify shortcut learning as the key limiting factor for online CL, where the learned features may be biased, not generalizable to new tasks, and may have an adverse impact on knowledge distillation. To tackle this issue, we present the online prototype learning (OnPro) framework for online CL. First, we propose online prototype equilibrium to learn representative features against shortcut learning and discriminative features to avoid class confusion, ultimately achieving an equilibrium status that separates all seen classes well while learning new classes. Second, with the feedback of online prototypes, we devise a novel adaptive prototypical feedback mechanism to sense the classes that are easily misclassified and then enhance their boundaries. Extensive experimental results on widely-used benchmark datasets demonstrate the superior performance of OnPro over the state-of-the-art baseline methods. Source code is available at https://github.com/weilllllls/OnPro. Yujie Wei 0001, Jiaxin Ye, Zhizhong Huang, Junping Zhang, Hongming Shan |
ICCV | 5 |
| 2023 | Adaptive Nonlinear Latent Transformation for Conditional Face EditingabstractRecent works for face editing usually manipulate the latent space of StyleGAN via the linear semantic directions. However, they usually suffer from the entanglement of facial attributes, need to tune the optimal editing strength, and are limited to binary attributes with strong supervision signals. This paper proposes a novel adaptive nonlinear latent transformation for disentangled and conditional face editing, termed AdaTrans. Specifically, our AdaTrans divides the manipulation process into several finer steps; i.e., the direction and size at each step are conditioned on both the facial attributes and the latent codes. In this way, AdaTrans describes an adaptive nonlinear transformation trajectory to manipulate the faces into target attributes while keeping other attributes unchanged. Then, AdaTrans leverages a predefined density model to constrain the learned trajectory in the distribution of latent codes by maximizing the likelihood of transformed latent code. Moreover, we also propose a disentangled learning strategy under a mutual information framework to eliminate the entanglement among attributes, which can further relax the need for labeled data. Consequently, AdaTrans enables a controllable face editing with the advantages of disentanglement, flexibility with non-binary attributes, and high fidelity. Extensive experimental results on various facial attributes demonstrate the qualitative and quantitative effectiveness of the proposed AdaTrans over existing state-of-the-art methods, especially in the most challenging scenarios with a large age gap and few labeled examples. The source code is available at https://github.com/Hzzone/AdaTrans. Zhizhong Huang, Siteng Ma, Junping Zhang, Hongming Shan |
ICCV | 4 |
| 2023 | Learning to Distill Global Representation for Sparse-View CTabstractSparse-view computed tomography (CT)—using a small number of projections for tomographic reconstruction—enables much lower radiation dose to patients and accelerated data acquisition. The reconstructed images, however, suffer from strong artifacts, greatly limiting their diagnostic value. Current trends for sparse-view CT turn to the raw data for better information recovery. The resultant dual-domain methods, nonetheless, suffer from secondary artifacts, especially in ultra-sparse view scenarios, and their generalization to other scanners/protocols is greatly limited. A crucial question arises: have the image post-processing methods reached the limit? Our answer is not yet. In this paper, we stick to image post-processing methods due to great flexibility and propose global representation(GloRe) distillation framework for sparse-view CT, termed GloReDi. First, we propose to learn GloRe with Fourier convolution, so each element in GloRe has an image-wide receptive field. Second, unlike methods that only use the full-view images for supervision, we propose to distill GloRe from intermediate-view reconstructed images that are readily available but not explored in previous literature. The success of GloRe distillation is attributed to two key components: representation directional distillation to align the GloRe directions, and band-pass-specific contrastive distillation to gain clinically important details. Extensive experiments demonstrate the superiority of the proposed GloReDi over the state-of-the-art methods, including dual-domain ones. The source code is available at https://github.com/longzilicart/GloReDi. Zilong Li 0001, Chenglong Ma 0002, Jie Chen 0001, Junping Zhang, Hongming Shan |
ICCV | 5 |
| 2023 | ASCON: Anatomy-Aware Supervised Contrastive Learning Framework for Low-Dose CT Denoising
Yi Zhang 0018, Hongming Shan |
MICCAI (10) | 4 |
| 2023 | BerDiff: Conditional Bernoulli Diffusion Model for Medical Image Segmentation
Tao Chen 0055, Hongming Shan |
MICCAI (4) | 3 |
| 2023 | CLIP-Lung: Textual Knowledge-Guided Lung Nodule Malignancy Prediction
Zilong Li 0001, Junping Zhang, Hongming Shan |
MICCAI (7) | 5 |
| 2023 | FreeSeed: Frequency-Band-Aware and Self-guided Network for Sparse-View CT Reconstruction
Chenglong Ma 0002, Zilong Li 0001, Junping Zhang, Yi Zhang 0018, Hongming Shan |
MICCAI (10) | 5 |
| 2023 | Emo-DNA: Emotion Decoupling and Alignment Learning for Cross-Corpus Speech Emotion RecognitionabstractCross-corpus speech emotion recognition (SER) seeks to generalize the ability of inferring speech emotion from a well-labeled corpus to an unlabeled one, which is a rather challenging task due to the significant discrepancy between two corpora. Existing methods, typically based on unsupervised domain adaptation (UDA), struggle to learn corpus-invariant features by global distribution alignment, but unfortunately, the resulting features are mixed with corpus-specific features or not class-discriminative. To tackle these challenges, we propose a novel Emotion Decoupling aNd Alignment learning framework (EMO-DNA) for cross-corpus SER, a novel UDA method to learn emotion-relevant corpus-invariant features. The novelties of EMO-DNA are two-fold: contrastive emotion decoupling and dual-level emotion alignment. On one hand, our contrastive emotion decoupling achieves decoupling learning via a contrastive decoupling loss to strengthen the separability of emotion-relevant features from corpus-specific ones. On the other hand, our dual-level emotion alignment introduces an adaptive threshold pseudo-labeling to select confident target samples for class-level alignment, and performs corpus-level alignment to jointly guide model for learning class-discriminative corpus-invariant features across corpora. Extensive experimental results demonstrate the superior performance of EMO-DNA over the state-of-the-art methods in several cross-corpus scenarios. Source code is available at https://github.com/Jiaxin-Ye/Emo-DNA. Jiaxin Ye, Yujie Wei 0001, Xin-Cheng Wen, Chenglong Ma 0002, Zhizhong Huang, Kunhong Liu 0001, Hongming Shan |
ACM Multimedia | 7 |
| 2023 | LICO: Explainable Models with Language-Image COnsistencyabstractInterpreting the decisions of deep learning models has been actively studied since the explosion of deep neural networks. One of the most convincing interpretation approaches is salience-based visual interpretation, such as Grad-CAM, where the generation of attention maps depends merely on categorical labels. Although existing interpretation methods can provide explainable decision clues, they often yield partial correspondence between image and saliency maps due to the limited discriminative information from one-hot labels. This paper develops a Language-Image COnsistency model for explainable image classification, termed LICO, by correlating learnable linguistic prompts with corresponding visual features in a coarse-to-fine manner. Specifically, we first establish a coarse global manifold structure alignment by minimizing the distance between the distributions of image and language features. We then achieve fine-grained saliency maps by applying optimal transport (OT) theory to assign local feature maps with class-specific prompts. Extensive experimental results on eight benchmark datasets demonstrate that the proposed LICO achieves a significant improvement in generating more explainable attention maps in conjunction with existing interpretation methods such as Grad-CAM. Remarkably, LICO improves the classification performance of existing models without introducing any computational overhead during inference. Zilong Li 0001, Junping Zhang, Hongming Shan |
NeurIPS | 5 |
| 2023 | Impact of loss functions on the performance of a deep neural network designed to restore low-dose digital mammography
Hongming Shan, Rodrigo de Barros Vimieiro, Lucas R. Borges, Marcelo A. C. Vieira, Ge Wang 0001 |
Artif. Intell. Medicine | 1 |
| 2023 | Forget less, count better: a domain-incremental self-distillation learning benchmark for lifelong crowd countingabstractCrowd counting has important applications in public safety and pandemic control. A robust and practical crowd counting system has to be capable of continuously learning with the newly incoming domain data in real-world scenarios instead of fitting one domain only. Off-the-shelf methods have some drawbacks when handling multiple domains: (1) the models will achieve limited performance (even drop dramatically) among old domains after training images from new domains due to the discrepancies in intrinsic data distributions from various domains, which is called catastrophic forgetting; (2) the well-trained model in a specific domain achieves imperfect performance among other unseen domains because of domain shift; (3) it leads to linearly increasing storage overhead, either mixing all the data for training or simply training dozens of separate models for different domains when new ones are available. To overcome these issues, we investigate a new crowd counting task in incremental domain training setting called lifelong crowd counting. Its goal is to alleviate catastrophic forgetting and improve the generalization ability using a single model updated by the incremental domains. Specifically, we propose a self-distillation learning framework as a benchmark (forget less, count better, or FLCB) for lifelong crowd counting, which helps the model leverage previous meaningful knowledge in a sustainable manner for better crowd counting to mitigate the forgetting when new data arrive. A new quantitative metric, normalized Backward Transfer (nBwT), is developed to evaluate the forgetting degree of the model in the lifelong learning process. Extensive experimental results demonstrate the superiority of our proposed benchmark in achieving a low catastrophic forgetting degree and strong generalization ability. Hongming Shan, Yanyun Qu, James Z. Wang 0001, Fei-Yue Wang 0001, Junping Zhang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | Learning Representation for Clustering Via Prototype Scattering and Positive SamplingabstractExisting deep clustering methods rely on either contrastive or non-contrastive representation learning for downstream clustering task. Contrastive-based methods thanks to negative pairs learn uniform representations for clustering, in which negative pairs, however, may inevitably lead to the class collision issue and consequently compromise the clustering performance. Non-contrastive-based methods, on the other hand, avoid class collision issue, but the resulting non-uniform representations may cause the collapse of clustering. To enjoy the strengths of both worlds, this paper presents a novel end-to-end deep clustering method with prototype scattering and positive sampling, termed ProPos. Specifically, we first maximize the distance between prototypical representations, named prototype scattering loss, which improves the uniformity of representations. Second, we align one augmented view of instance with the sampled neighbors of another view-assumed to be truly positive pair in the embedding space-to improve the within-cluster compactness, termed positive sampling alignment. The strengths of ProPos are avoidable class collision issue, uniform representations, well-separated clusters, and within-cluster compactness. By optimizing ProPos in an end-to-end expectation-maximization framework, extensive experimental results demonstrate that ProPos achieves competing performance on moderate-scale clustering benchmark datasets and establishes new state-of-the-art performance on large-scale datasets. Source code is available at https://github.com/Hzzone/ProPos. Zhizhong Huang, Jie Chen 0001, Junping Zhang, Hongming Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | When Age-Invariant Face Recognition Meets Face Age Synthesis: A Multi-Task Learning Framework and a New BenchmarkabstractTo minimize the impact of age variation on face recognition, age-invariant face recognition (AIFR) extracts identity-related discriminative features by minimizing the correlation between identity- and age-related features while face age synthesis (FAS) eliminates age variation by converting the faces in different age groups to the same group. However, AIFR lacks visual results for model interpretation and FAS compromises downstream recognition due to artifacts. Therefore, we propose a unified, multi-task framework to jointly handle these two tasks, termed MTLFace, which can learn the age-invariant identity-related representation for face recognition while achieving pleasing face synthesis for model interpretation. Specifically, we propose an attention-based feature decomposition to decompose the mixed face features into two uncorrelated components-identity- and age-related features-in a spatially constrained way. Unlike the conventional one-hot encoding that achieves group-level FAS, we propose a novel identity conditional module to achieve identity-level FAS, which can improve the age smoothness of synthesized faces through a weight-sharing strategy. Benefiting from the proposed multi-task framework, we then leverage those high-quality synthesized faces from FAS to further boost AIFR via a novel selective fine-tuning strategy. Furthermore, to advance both AIFR and FAS, we collect and release a large cross-age face dataset with age and gender annotations, and a new benchmark specifically designed for tracing long-missing children. Extensive experimental results on five benchmark cross-age datasets demonstrate that MTLFace yields superior performance than state-of-the-art methods for both AIFR and FAS. We further validate MTLFace on two popular general face recognition datasets, obtaining competitive performance on face recognition in the wild. The source code and datasets are available at http://hzzone.github.io/MTLFace. Zhizhong Huang, Junping Zhang, Hongming Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | M3NAS: Multi-Scale and Multi-Level Memory-Efficient Neural Architecture Search for Low-Dose CT DenoisingabstractLowering the radiation dose in computed tomography (CT) can greatly reduce the potential risk to public health. However, the reconstructed images from dose-reduced CT or low-dose CT (LDCT) suffer from severe noise which compromises the subsequent diagnosis and analysis. Recently, convolutional neural networks have achieved promising results in removing noise from LDCT images. The network architectures that are used are either handcrafted or built on top of conventional networks such as ResNet and U-Net. Recent advances in neural network architecture search (NAS) have shown that the network architecture has a dramatic effect on the model performance. This indicates that current network architectures for LDCT may be suboptimal. Therefore, in this paper, we make the first attempt to apply NAS to LDCT and propose a multi-scale and multi-level memory-efficient NAS for LDCT denoising, termed M3NAS. On the one hand, the proposed M3NAS fuses features extracted by different scale cells to capture multi-scale image structural details. On the other hand, the proposed M3NAS can search a hybrid cell- and network-level structure for better performance. In addition, M3NAS can effectively reduce the number of model parameters and increase the speed of inference. Extensive experimental results on two different datasets demonstrate that the proposed M3NAS can achieve better performance and fewer parameters than several state-of-the-art methods. In addition, we also validate the effectiveness of the multi-scale and multi-level architecture for LDCT denoising, and present further analysis for different configurations of super-net. Wenjun Xia, Yongqiang Huang 0003, Mingzheng Hou, Hu Chen 0002, Jiliu Zhou, Hongming Shan, Yi Zhang 0018 |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Hybrid Weighting Loss for Precipitation Nowcasting from Radar ImagesabstractPrecipitation nowcasting is gaining increasing attention in the signal processing community. Existing deep learning-based studies focus on designing an effective model architecture, neglecting the influence of the severe imbalanced distribution of rainfall data that can compromise the predictive accuracy on heavy rainfall intensities. To address the uneven distribution of precipitation nowcasting data, we propose a novel data reweighting strategy, termed Hybrid Weighting, which hybrids reweighting and non-weighting strategies together, boosting the precipitation nowcasting performance. Experimental results on two natural radar echo benchmark datasets demonstrate the superior performance of our proposed approach for precipitation nowcasting over existing loss functions on high rainfall intensities, without degenerating on low rainfall intensities compared with state-of-art reweighting methods. Danchen Zhang, Leiming Ma, Hongming Shan |
ICASSP | 5 |
| 2022 | SPICE: Semantic Pseudo-Labeling for Image ClusteringabstractThe similarity among samples and the discrepancy among clusters are two crucial aspects of image clustering. However, current deep clustering methods suffer from inaccurate estimation of either feature similarity or semantic discrepancy. In this paper, we present a Semantic Pseudo-labeling-based Image ClustEring (SPICE) framework, which divides the clustering network into a feature model for measuring the instance-level similarity and a clustering head for identifying the cluster-level discrepancy. We design two semantics-aware pseudo-labeling algorithms, prototype pseudo-labeling and reliable pseudo-labeling, which enable accurate and reliable self-supervision over clustering. Without using any ground-truth label, we optimize the clustering network in three stages: 1) train the feature model through contrastive learning to measure the instance similarity; 2) train the clustering head with the prototype pseudo-labeling algorithm to identify cluster semantics; and 3) jointly train the feature model and clustering head with the reliable pseudo-labeling algorithm to improve the clustering performance. Extensive experimental results demonstrate that SPICE achieves significant improvements (~10%) over existing methods and establishes the new state-of-the-art clustering results on six balanced benchmark datasets in terms of three popular metrics. Importantly, SPICE significantly reduces the gap between unsupervised and fully-supervised classification; e.g. there is only 2% (91.8% vs 93.8%) accuracy difference on CIFAR-10. Our code is made publicly available at https://github.com/niuchuangnn/SPICE. Chuang Niu, Hongming Shan, Ge Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Content-Noise Complementary Learning for Medical Image DenoisingabstractMedical imaging denoising faces great challenges, yet is in great demand. With its distinctive characteristics, medical imaging denoising in the image domain requires innovative deep learning strategies. In this study, we propose a simple yet effective strategy, the content-noise complementary learning (CNCL) strategy, in which two deep learning predictors are used to learn the respective content and noise of the image dataset complementarily. A medical image denoising pipeline based on the CNCL strategy is presented, and is implemented as a generative adversarial network, where various representative networks (including U-Net, DnCNN, and SRDenseNet) are investigated as the predictors. The performance of these implemented models has been validated on medical imaging datasets including CT, MR, and PET. The results show that this strategy outperforms state-of-the-art denoising algorithms in terms of visual quality and quantitative metrics, and the strategy demonstrates a robust generalization capability. These findings validate that this simple yet effective strategy demonstrates promising potential for medical image denoising tasks, which could exert a clinical impact in the future. Code is available at: https://github.com/gengmufeng/CNCL-denoising. Mufeng Geng, Xiangxi Meng 0001, Jiangyuan Yu, Lei Zhu 0012, Lujia Jin, Bin Qiu, Hanjing Kong, Jianmin Yuan, Hongming Shan, Hongbin Han, Qiushi Ren, Yanye Lu |
IEEE Trans. Medical Imaging | 12 |
| 2022 | Convolutional Ordinal Regression Forest for Image Ordinal EstimationabstractImage ordinal estimation is to predict the ordinal label of a given image, which can be categorized as an ordinal regression (OR) problem. Recent methods formulate an OR problem as a series of binary classification problems. Such methods cannot ensure that the global ordinal relationship is preserved since the relationships among different binary classifiers are neglected. We propose a novel OR approach, termed convolutional OR forest (CORF), for image ordinal estimation, which can integrate OR and differentiable decision trees with a convolutional neural network for obtaining precise and stable global ordinal relationships. The advantages of the proposed CORF are twofold. First, instead of learning a series of binary classifiers independently, the proposed method aims at learning an ordinal distribution for OR by optimizing those binary classifiers simultaneously. Second, the differentiable decision trees in the proposed CORF can be trained together with the ordinal distribution in an end-to-end manner. The effectiveness of the proposed CORF is verified on two image ordinal estimation tasks, i.e., facial age estimation and image esthetic assessment, showing significant improvements and better stability over the state-of-the-art OR methods. Hongming Shan, Lingfu Che, Junping Zhang, Jianbo Shi, Fei-Yue Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | When Age-Invariant Face Recognition Meets Face Age Synthesis: A Multi-Task Learning FrameworkabstractTo minimize the effects of age variation in face recognition, previous work either extracts identity-related discriminative features by minimizing the correlation between identity- and age-related features, called age-invariant face recognition (AIFR), or removes age variation by transforming the faces of different age groups into the same age group, called face age synthesis (FAS); however, the former lacks visual results for model interpretation while the latter suffers from artifacts compromising downstream recognition. Therefore, this paper proposes a unified, multi-task framework to jointly handle these two tasks, termed MTL-Face, which can learn age-invariant identity-related representation while achieving pleasing face synthesis. Specifically, we first decompose the mixed face features into two uncorrelated components—identity- and age-related features—through an attention mechanism, and then decorrelate these two components using multi-task training and continuous domain adaption. In contrast to the conventional one-hot encoding that achieves group-level FAS, we propose a novel identity conditional module to achieve identity-level FAS, with a weight-sharing strategy to improve the age smoothness of synthesized faces. In addition, we collect and release a large cross-age face dataset with age and gender annotations to advance AIFR and FAS. Extensive experiments on five benchmark cross-age datasets demonstrate the superior performance of our proposed MTLFace over state-of-the-art methods for AIFR and FAS. We further validate MTLFace on two popular general face recognition datasets, showing competitive performance for face recognition in the wild. The source code and dataset are available at https://github.com/Hzzone/MTLFace. Zhizhong Huang, Junping Zhang, Hongming Shan |
CVPR | 3 |
| 2021 | Selfgait: A Spatiotemporal Representation Learning Method for Self-Supervised Gait RecognitionabstractGait recognition plays a vital role in human identification since gait is a unique biometric feature that can be perceived at a distance. Although existing gait recognition methods can learn gait features from gait sequences in different ways, the performance of gait recognition suffers from insufficient labeled data, especially in some practical scenarios associated with short gait sequences or various clothing styles. It is unpractical to label the numerous gait data. In this work, we propose a self-supervised gait recognition method, termed SelfGait, which takes advantage of the massive, diverse, unlabeled gait data as a pre-training process to improve the representation abilities of spatiotemporal backbones. Specifically, we employ the horizontal pyramid mapping (HPM) and micro-motion template builder (MTB) as our spatiotemporal backbones to capture the multi-scale spatiotemporal representations. Experiments on CASIA-B and OU-MVLP benchmark gait datasets demonstrate the effectiveness of the proposed SelfGait compared with four state-of-the-art gait recognition methods. The source code has been released at https://github.com/EchoItLiu/SelfGait. Yiqun Liu 0009, Jian Pu, Hongming Shan, Peiyang He, Junping Zhang |
ICASSP | 4 |
| 2021 | Routinggan: Routing Age Progression and Regression with Disentangled LearningabstractAlthough impressive results have been achieved for age progression and regression, there remain two major issues in generative adversarial networks (GANs)-based methods: 1) conditional GANs (cGANs)-based methods can learn various effects between any two age groups in a single model, but are insufficient to characterize some specific patterns due to completely shared convolutions filters; and 2) GANs-based methods can, by utilizing several models to learn effects independently, learn some specific patterns, however, they are cumbersome and require age label in advance. To address these deficiencies and have the best of both worlds, this paper introduces a dropout-like method based on GAN (RoutingGAN) to route different effects in a high-level semantic feature space. Specifically, we first disentangle the age-invariant features from the input face, and then gradually add the effects to the features by residual routers that assign the convolution filters to different age groups by dropping out the outputs of others. As a result, the proposed RoutingGAN can simultaneously learn various effects in a single model, with convolution filters being shared in part to learn some specific effects. Experimental results on two benchmarked datasets demonstrate superior performance over existing methods both qualitatively and quantitatively. Zhizhong Huang, Junping Zhang, Hongming Shan |
ICASSP | 3 |
| 2021 | Meta Ordinal Weighting Net For Improving Lung Nodule ClassificationabstractThe progression of lung cancer implies the intrinsic ordinal relationship of lung nodules at different stages—from benign to unsure then to malignant. This problem can be solved by ordinal regression methods, which is between classification and regression due to its ordinal label. However, existing convolutional neural network-based ordinal regression methods only focus on modifying classification head based on a randomly sampled mini-batch of data, ignoring the ordinal relationship resided in the data itself. In this paper, we propose a Meta Ordinal Weighting Network (MOW-Net) to explicitly align each training sample with a meta ordinal set (MOS) containing a few samples from all classes. During the training process, the MOW-Net learns a mapping from samples in MOS to corresponding class-specific weight. We further propose a meta cross-entropy loss to optimize the network in a meta-learning scheme. Experimental results demonstrate that the MOW-Net achieves better accuracy than the state-of-the-art ordinal regression methods, especially for the unsure class. Hongming Shan, Junping Zhang |
ICASSP | 2 |
| 2021 | AgeFlow: Conditional Age Progression and Regression with Normalizing FlowsabstractAge progression and regression aim to synthesize photorealistic appearance of a given face image with aging and rejuvenation effects, respectively. Existing generative adversarial networks (GANs) based methods suffer from the following three major issues: 1) unstable training introducing strong ghost artifacts in the generated faces, 2) unpaired training leading to unexpected changes in facial attributes such as genders and races, and 3) non-bijective age mappings increasing the uncertainty in the face transformation. To overcome these issues, this paper proposes a novel framework, termed AgeFlow, to integrate the advantages of both flow-based models and GANs. The proposed AgeFlow contains three parts: an encoder that maps a given face to a latent space through an invertible neural network, a novel invertible conditional translation module (ICTM) that translates the source latent vector to target one, and a decoder that reconstructs the generated face from the target latent vector using the same encoder network; all parts are invertible achieving bijective age mappings. The novelties of ICTM are two-fold. First, we propose an attribute-aware knowledge distillation to learn the manipulation direction of age progression while keeping other unrelated attributes unchanged, alleviating unexpected changes in facial attributes. Second, we propose to use GANs in the latent space to ensure the learned latent vector indistinguishable from the real ones, which is much easier than traditional use of GANs in the image domain. Experimental results demonstrate superior performance over existing GANs-based methods on two benchmarked datasets. The source code is available at https://github.com/Hzzone/AgeFlow. Zhizhong Huang, Shouzhen Chen, Junping Zhang, Hongming Shan |
IJCAI | 4 |
| 2021 | PFA-GAN: Progressive Face Aging With Generative Adversarial NetworkabstractFace aging is to render a given face to predict its future appearance, which plays an important role in the information forensics and security field as the appearance of the face typically varies with age. Although impressive results have been achieved with conditional generative adversarial networks (cGANs), the existing cGANs-based methods typically use a single network to learn various aging effects between any two different age groups. However, they cannot simultaneously meet three essential requirements of face aging-including image quality, aging accuracy, and identity preservation-and usually generate aged faces with strong ghost artifacts when the age gap becomes large. Inspired by the fact that faces gradually age over time, this paper proposes a novel progressive face aging framework based on generative adversarial network (PFA-GAN) to mitigate these issues. Unlike the existing cGANs-based methods, the proposed framework contains several sub-networks to mimic the face aging process from young to old, each of which only learns some specific aging effects between two adjacent age groups. The proposed framework can be trained in an end-to-end manner to eliminate accumulative artifacts and blurriness. Moreover, this paper introduces an age estimation loss to take into account the age distribution for an improved aging accuracy, and proposes to use the Pearson correlation coefficient as an evaluation metric measuring the aging smoothness for face aging methods. Extensively experimental results demonstrate superior performance over existing (c)GANs-based methods, including the state-of-the-art one; e.g., PFA-GAN reduces the aging estimation errors by 0.23 and 0.35 and increases the identity preservation rates by 0.49 and 0.63 on two benchmarked datasets compared to the second best method for the challenging face aging from 30- to 51+. The source code is available at https://github.com/Hzzone/PFA-GAN. Zhizhong Huang, Shouzhen Chen, Junping Zhang, Hongming Shan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Cine Cardiac MRI Motion Artifact Reduction Using a Recurrent Neural NetworkabstractCine cardiac magnetic resonance imaging (MRI) is widely used for the diagnosis of cardiac diseases thanks to its ability to present cardiovascular features in excellent contrast. As compared to computed tomography (CT), MRI, however, requires a long scan time, which inevitably induces motion artifacts and causes patients' discomfort. Thus, there has been a strong clinical motivation to develop techniques to reduce both the scan time and motion artifacts. Given its successful applications in other medical imaging tasks such as MRI super-resolution and CT metal artifact reduction, deep learning is a promising approach for cardiac MRI motion artifact reduction. In this paper, we propose a novel recurrent generative adversarial network model for cardiac MRI motion artifact reduction. This model utilizes bi-directional convolutional long short-term memory (ConvLSTM) and multi-scale convolutions to improve the performance of the proposed network, in which bi-directional ConvLSTMs handle long-range temporal features while multi-scale convolutions gather both local and global features. We demonstrate a decent generalizability of the proposed method thanks to the novel architecture of our deep network that captures the essential relationship of cardiovascular dynamics. Indeed, our extensive experiments show that our method achieves better image quality for cine cardiac MRI images than existing state-of-the-art methods. In addition, our method can generate reliable missing intermediate frames based on their adjacent frames, improving the temporal resolution of cine cardiac MRI sequences. Qing Lyu 0003, Hongming Shan, Yibin Xie, Alan C. Kwan, Yuka Otaki, Keiichiro Kuronuma, Debiao Li, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Meta Ordinal Regression Forest For Learning with Unsure Lung NodulesabstractDeep learning-based methods have achieved promising performance in early detection and classification of lung nodules, most of which discard unsure nodules and simply deal with a binary classification-malignant vs benign. Recently, an unsure data model (UDM) was proposed to incorporate those unsure nodules by formulating this problem as an ordinal regression, showing better performance over traditional binary classification. To further explore the ordinal relationship for lung nodule classification, this paper proposes a meta ordinal regression forest (MORF), which improves upon the state-of the-art ordinal regression method, deep ordinal regression forest (DORF), in three major ways. First, MORF can alleviate the biases of the predictions by making full use of deep features while DORF needs to fix the composition of decision trees before training. Second, MORF has a novel grouped feature selection (GFS) module to re-sample the split nodes of decision trees. Last, combined with GFS, MORF is equipped with a meta learning based weighting scheme to map the features selected by GFS to tree-wise weights while DORF assigns equal weights for all trees. Experimental results on LIDC-IDRI dataset demonstrate superior performance over existing methods, including the state of-the-art DORF. Junping Zhang, Hongming Shan |
BIBM | 4 |
| 2020 | Look Globally, Age Locally: Face Aging With an Attention MechanismabstractFace aging is of great importance for cross-age recognition and entertainment-related applications. Recently, conditional generative adversarial networks (cGANs) have achieved impressive results for face aging. Existing cGANs-based methods usually require a pixel-wise loss to keep the identity and background consistent. However, minimizing the pixel-wise loss between the input and synthesized images likely resulting in a ghosted or blurry face. To address this deficiency, this paper introduces an Attention Conditional GANs (AcGANs) approach for face aging, which utilizes attention mechanism to only alert the regions relevant to face aging. In doing so, the synthesized face can well preserve the background information and personal identity without using the pixel-wise loss, and the ghost artifacts and blurriness can be significantly reduced. Based on the benchmarked dataset Morph, both qualitative and quantitative experiment results demonstrate superior performance over existing algorithms in terms of image quality, personal identity, and age accuracy. Codes are available on https://github.com/JensonZhu14/AcGAN. Zhizhong Huang, Hongming Shan, Junping Zhang |
ICASSP | 3 |
| 2020 | Ordinal distribution regression for gait-based age estimation
Guohao Li 0005, Junping Zhang, Hongming Shan |
Sci. China Inf. Sci. | 5 |
| 2020 | Shape and margin-aware lung nodule classification in low-dose CT images via soft activation mapping
Yukun Tian, Hongming Shan, Junping Zhang, Ge Wang 0001, Mannudeep K. Kalra |
Medical Image Anal. | 3 |
| 2020 | Quadratic Autoencoder (Q-AE) for Low-Dose CT DenoisingabstractInspired by complexity and diversity of biological neurons, our group proposed quadratic neurons by replacing the inner product in current artificial neurons with a quadratic operation on input data, thereby enhancing the capability of an individual neuron. Along this direction, we are motivated to evaluate the power of quadratic neurons in popular network architectures, simulating human-like learning in the form of "quadratic-neuron-based deep learning". Our prior theoretical studies have shown important merits of quadratic neurons and networks in representation, efficiency, and interpretability. In this paper, we use quadratic neurons to construct an encoder-decoder structure, referred as the quadratic autoencoder, and apply it to low-dose CT denoising. The experimental results on the Mayo low-dose CT dataset demonstrate the utility and robustness of quadratic autoencoder in terms of image denoising and model efficiency. To our best knowledge, this is the first time that the deep learning approach is implemented with a new type of neurons and demonstrates a significant potential in the medical imaging field. Fenglei Fan, Hongming Shan, Mannudeep K. Kalra, Guhan Qian, Matthew Getzin, Yueyang Teng, Juergen Hahn, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Multi-Contrast Super-Resolution MRI Through a Progressive NetworkabstractMagnetic resonance imaging (MRI) is widely used for screening, diagnosis, image-guided therapy, and scientific research. A significant advantage of MRI over other imaging modalities such as computed tomography (CT) and nuclear imaging is that it clearly shows soft tissues in multi-contrasts. Compared with other medical image super-resolution methods that are in a single contrast, multi-contrast super-resolution studies can synergize multiple contrast images to achieve better super-resolution results. In this paper, we propose a one-level non-progressive neural network for low up-sampling multi-contrast super-resolution and a two-level progressive network for high up-sampling multi-contrast super-resolution. The proposed networks integrate multi-contrast information in a high-level feature space and optimize the imaging performance by minimizing a composite loss function, which includes mean-squared-error, adversarial loss, perceptual loss, and textural loss. Our experimental results demonstrate that 1) the proposed networks can produce MRI super-resolution images with good image quality and outperform other multi-contrast super-resolution methods in terms of structural similarity and peak signal-to-noise ratio; 2) combining multi-contrast information in a high-level feature space leads to a significantly improved result than a combination in the low-level pixel space; and 3) the progressive network produces a better super-resolution image quality than the non-progressive network, even if the original low-resolution images were highly down-sampled. Qing Lyu 0003, Hongming Shan, Cole Steber, Corbin Helis, Christopher T. Whitlow, Michael D. Chan, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | CT Super-Resolution GAN Constrained by the Identical, Residual, and Cycle Learning Ensemble (GAN-CIRCLE)abstractIn this paper, we present a semi-supervised deep learning approach to accurately recover high-resolution (HR) CT images from low-resolution (LR) counterparts. Specifically, with the generative adversarial network (GAN) as the building block, we enforce the cycle-consistency in terms of the Wasserstein distance to establish a nonlinear end-to-end mapping from noisy LR input images to denoised and deblurred HR outputs. We also include the joint constraints in the loss function to facilitate structural preservation. In this process, we incorporate deep convolutional neural network (CNN), residual learning, and network in network techniques for feature extraction and restoration. In contrast to the current trend of increasing network depth and complexity to boost the imaging performance, we apply a parallel 1×1 CNN to compress the output of the hidden layer and optimize the number of layers and the number of filters for each convolutional layer. The quantitative and qualitative evaluative results demonstrate that our proposed model is accurate, efficient and robust for super-resolution (SR) image restoration from noisy LR input images. In particular, we validate our composite SR networks on three large-scale CT datasets, and obtain promising results as compared to the other state-of-the-art methods. Chenyu You, Wenxiang Cong, Michael W. Vannier, Punam K. Saha, Eric A. Hoffman, Ge Wang 0001, Guang Li 0011, Yi Zhang 0018, Xiaoliu Zhang, Hongming Shan, Mengzhou Li, Shenghong Ju, Zhen Zhao 0003, Zhuiyang Zhang |
IEEE Trans. Medical Imaging | 10 |
| 2019 | Framework of Randomized Distribution Features for Visual Representation and CategorizationabstractThis paper introduces a framework to deal with the distribution of descriptive features, which preserves the advantages of the vectorial representation and computational efficiency of histogram-based techniques, and inherits the rigorous theoretical guarantee and competitive performance of metric-based ones. The methods developed under this framework describe the underlying distribution of a set of features as a vectorial feature by utilizing random features. Moreover, the proposed methods asymptotically converge to metric-based methods in terms of the similarity and distance and, depending on a specific kernel function, reduce to histogram-based methods. The experimental results show the benefits of a comparable performance on categorization tasks compared to conventional metric-based methods at a significantly reduced computational cost. Hongming Shan, Junping Zhang, Uwe Krüger 0001 |
IEEE Trans. Cybern. | 1 |
| 2019 | Multi-Task GANs for View-Specific Feature Learning in Gait RecognitionabstractGait recognition is of great importance in the fields of surveillance and forensics to identify human beings since gait is the unique biometric feature that can be perceived efficiently at a distance. However, the accuracy of gait recognition to some extent suffers from both the variation of view angles and the deficient gait templates. On one hand, the existing cross-view methods focus on transforming gait templates among different views, which may accumulate the transformation error in a large variation of view angles. On the other hand, a commonly used gait energy image template loses temporal information of a gait sequence. To address these problems, this paper proposes multi-task generative adversarial networks (MGANs) for learning view-specific feature representations. In order to preserve more temporal information, we also propose a new multi-channel gait template, called period energy image (PEI). Based on the assumption of view angle manifold, the MGANs can leverage adversarial training to extract more discriminative features from gait sequences. Experiments on OU-ISIR, CASIA-B, and USF benchmark data sets indicate that compared with several recently published approaches, PEI + MGANs achieves competitive performance and is more interpretable to cross-view gait recognition. Yiwei He, Junping Zhang, Hongming Shan, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Crowd Counting With Limited Labeling Through Submodular Frame SelectionabstractAutomated crowd counting is valuable for intelligent transportation systems, as it can help to improve the emergency planning and prevent congestion in transit hubs such as train stations and airports. Semi-supervised crowd counting aims to estimate the number of pedestrians in an ongoing scene using a combination of a small number of labeled frames and a large number of unlabeled ones. However, existing methods do not incorporate ways to effectively select informative frames as labeled training samples, resulting in low accuracy on unseen crowd scenes. We propose a submodular method to select the most informative frames from the image sequences of crowds. Specifically, the method selects the most representative images to guarantee the information coverage, by maximizing the similarities between the group of selected images and the image sequence. In addition, these frames are chosen to avoid redundancies and preserve diversity. Finally, our semi-supervised method incorporates graph Laplacian regularization and spatiotemporal constraints. Extensive experiments on three benchmark data sets demonstrate that our proposed approach achieves higher accuracy compared with the state-of-the-art regression methods and competitive performance with deep convolutional models, especially when the number of labeled data is exceptionally small. Junping Zhang, Lingfu Che, Hongming Shan, James Z. Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2018 | 3-D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2-D Trained NetworkabstractLow-dose computed tomography (LDCT) has attracted major attention in the medical imaging field, since CT-associated X-ray radiation carries health risks for patients. The reduction of the CT radiation dose, however, compromises the signal-to-noise ratio, which affects image quality and diagnostic performance. Recently, deep-learning-based algorithms have achieved promising results in LDCT denoising, especially convolutional neural network (CNN) and generative adversarial network (GAN) architectures. This paper introduces a conveying path-based convolutional encoder-decoder (CPCE) network in 2-D and 3-D configurations within the GAN framework for LDCT denoising. A novel feature of this approach is that an initial 3-D CPCE denoising model can be directly obtained by extending a trained 2-D CNN, which is then fine-tuned to incorporate 3-D spatial information from adjacent slices. Based on the transfer learning from 2-D to 3-D, the 3-D network converges faster and achieves a better denoising performance when compared with a training from scratch. By comparing the CPCE network with recently published work based on the simulated Mayo data set and the real MGH data set, we demonstrate that the 3-D CPCE denoising model has a better performance in that it suppresses image noise and preserves subtle structures. Hongming Shan, Yi Zhang 0018, Qingsong Yang, Uwe Krüger 0001, Mannudeep K. Kalra, Ling Sun 0006, Wenxiang Cong, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Correction for "3D Convolutional Encoder-Decoder Network for Low-Dose CT via Transfer Learning From a 2D Trained Network"abstractIn[1], please note the updated figure captions for Figures 5, 6, 7, and 8 as follows: Hongming Shan, Yi Zhang 0018, Qingsong Yang, Uwe Krüger 0001, Mannudeep K. Kalra, Ling Sun 0006, Wenxiang Cong, Ge Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Randomized Distribution Feature for Image ClassificationabstractLocal image features can be assumed to be drawn from an unknown distribution. For image classification, such features are compared through the histogram-based model or the metric-based model. By quantizing these local features into a set of histograms, the histogram-based model is convenient and has vectorial representation of image but information could be lost in vector quantization. Unlike the histogram-based model, the metric-based model estimates the metrics over the underlying distribution of local features immediately, achieving better predictive performance. However, the model requires higher computational cost and loses the benefit of vectorial representation of image. Hongming Shan, Junping Zhang |
ECAI | 1 |
| 2016 | Group Information-Based Dimensionality Reduction via Canonical Correlation Analysis
Hongming Shan, Yiwei He, Junping Zhang |
ICONIP (2) | 2 |
| 2016 | Learning Linear Representation of Space Partitioning Trees Based on Unsupervised Kernel Dimension ReductionabstractSpace partitioning trees, which sequentially divide and subdivide a space into disjoint subsets using splitting hyperplanes, play a key role in accelerating the query of samples in the cybernetics and computer vision domains. Associated methods, however, suffer from the curse of dimensionality or stringent assumptions on the data distribution. This paper presents a new concept, termed kernel dimension reduction-tree (KDR-tree), that relies on linear projections computed based on an unsupervised kernel dimension reduction approach. The proposed concept does not rely on any assumption on the data distribution and can capture higher-order statistical information encapsulated within the data. This paper then develops two variants of the KDR-tree concept: 1) to handle residual data [i.e., the residual-based KDR-tree (rKDR-tree) algorithm] and 2) to cope with larger datasets, [i.e., the sampling-based KDR-tree (sKDR-tree) algorithm]. By directly comparing the KDR-tree concept to competitive techniques, involving several benchmark datasets, this paper shows that the sKDR-tree yields a better performance for non-Gaussian distributed datasets. Based on the analysis of three datasets, this paper highlights, experimentally, that the rKDR-tree has the potential to discover the intrinsic dimension. This paper also provides a theoretical analysis about the KDR-tree concept to outline why it outperforms existing techniques if the data distribution is non-Gaussian. Hongming Shan, Junping Zhang, Uwe Krüger 0001 |
IEEE Trans. Cybern. | 1 |