Nianchang Huang

dblp:258/2761 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0001-9530-3490ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 'Knowledge and experience' for visible-infrared person re-identification
Nianchang Huang, Qiang Zhang 0020, Jungong Han, Jin Huang 0004
Pattern Recognit.1
2026 Mitigating fusion bias for RGB-D salient object detection
Yang Yang 0009, Nianchang Huang, Qiang Zhang 0020, Jungong Han
Pattern Recognit.2
2026 Modality Adaptive Network for Arbitrary Modality Salient Object Detection
abstract
This paper delves into the task of arbitrary modality salient object detection (AM SOD), aiming to detect salient objects from the images with arbitrary modality types or arbitrary modality numbers by using a single model trained once. Specifically, we develop a novel model, termed modality adaptive network (MAN), for AM SOD, which addresses two fundamental challenges in AM SOD: the diverse modality discrepancies arising from varying modality types and the dynamic fusion dilemma resulting from an unfixed number of modalities in the input data. Technically, MAN first introduces a novel Modality-Adaptive Feature Extractor (MAFE) to adaptively extract features from different input modalities based on their characteristics by utilizing a set of learnable modality prompts. Concurrently, a new modality translation contractive (MTC) loss is devised to facilitate the training of MAFE as well as modality prompts, thereby effectively addressing the inherent modality discrepancies and extracting more discriminative features from each modality image. Subsequently, MAN presents a hybrid dynamic fusion (HDF) strategy to effectively resolve the challenge of dynamic inputs in multi-modal feature fusion as well as enhance the exploitation of complementary information across different modalities. This is specially achieved by a Channel- wise Dynamic Fusion Module (CDFM) and a Spatial- wise Dynamic Fusion Module (SDFM). Experimental results show that by virtue of MAFE, MTC loss and HDF strategy, our proposed method achieves significant increasements over existing models on benchmark datasets.
Yang Yang 0132, Nianchang Huang, Qiang Zhang 0020, Jungong Han, Jin Huang 0004
IEEE Trans. Multim.2
2025 CSANet: Cross-Modality Self-Paced Association Network for Unsupervised Visible-Infrared Person Re-Identification
abstract
For preeminent unsupervised visible-infrared person re-identification (US-VI-ReID), existing studies typically adhere to a two-step paradigm,i.e., intra-modality clustering and inter-modality matching. Nevertheless, high intra-modality variations may result in suboptimal clusters containing intricate pedestrians, while significant inter-modality discrepancies further complicate their cross-modality associations. Most existing methods fail to adopt a differentiated approach for samples of varying difficulty, especially intricate ones. To address this, we propose enabling the model to gradually establish cross-modality associations from easy to hard, mimicking human learning patterns to avoid error accumulation caused by intricate pedestrians. To this end, we propose a Cross-modality Self-paced Association Network, termed CSANet, embracing Twain Bipartite Graph Matching (TBGM), Cross-curriculum Association Prompter (CAP) and Instance-Prototype Consistency Constraint (IPCC) modules. TBGM conceives a graph-driven metric to tailor athree-levelcurriculum (plain,moderateandintricate) for self-paced cross-modality learning. CAP transfers high-confidence associations deduced from the plain subsets to intricate ones, prompting exploring more complex cross-modality relationships. Alongside CAP, IPCC further enforces the intricate instances to mimic their prototype characteristics, facilitating their discriminative feature learning. Extensive experiments demonstrate CSANet’s superiority over state-of-the-art methods, highlighting the potential of self-paced learning for US-VI-ReID.
Ruida Xi, Zhenyang Fu, Nianchang Huang, Xiaowei Zhao 0002, Qiang Zhang 0020, Jungong Han
IEEE Trans. Inf. Forensics Secur.3
2025 FMCNet+: Feature-Level Modality Compensation for Visible-Infrared Person Re-Identification
abstract
For visible-infrared person re-identification (VI-ReID), current models that compensate modality-specific information strive to generate missing modality images from existing ones to bridge the cross-modality discrepancies. Despite that, those generated images often suffer from low qualities due to the significant modality gap and include interfering information, e.g., inconsistent colors, thus severely degrading the subsequent VI-ReID performance. Alternatively, we propose a feature-level modality compensation network, i.e., FMCNet+, for VI-ReID in this article as an improved version of our previous work (FMCNet). The core of FMCNet+ is to compensate for the missing modality-specific information at the feature level, rather than at the image level, enabling our model to generate more person-related and discriminative modality-specific features for VI-ReID. Concretely, FMCNet+ aims to progressively generate missing modality-specific features by fully exploring the relationships among single-modality features, modality-shared features, and modality-specific features, instead of directly generating them through a generative adversarial way as in the previous FMCNet. To this end, three modules, i.e., single-modality feature decomposition (SFD), modality characteristic dictionary learning (MCDL), and missing modality-specific feature compensation (MMFC), are incorporated in FMCNet+. Experimental results demonstrate the superiority of our proposed FMCNet+ over existing ones, especially for those that compensate for modality-specific information at the image level. Our intriguing findings highlight the necessity of feature-level modality compensation in VI-ReID. Our code and pre-trained models will be released on https://github.com/jssyzsfzy/FMCNet_series.
Ruida Xi, Nianchang Huang, Changzhou Lai, Qiang Zhang 0020, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.2
2024 Lightweight cross-modal transformer for RGB-D salient object detection
Nianchang Huang, Yang Yang 0009, Qiang Zhang 0020, Jungong Han, Jin Huang 0004
Comput. Vis. Image Underst.1
2024 Salient Object Detection From Arbitrary Modalities
abstract
Toward desirable saliency prediction, the types and numbers of inputs for a salient object detection (SOD) algorithm may dynamically change in many real-life applications. However, existing SOD algorithms are mainly designed or trained for one particular type of inputs, failing to be generalized to other types of inputs. Consequentially, more types of SOD algorithms need to be prepared in advance for handling different types of inputs, raising huge hardware and research costs. Differently, in this paper, we propose a new type of SOD task, termed Arbitrary Modality SOD (AM SOD). The most prominent characteristics of AM SOD are that the modality types and modality numbers will be arbitrary or dynamically changed. The former means that the inputs to the AM SOD algorithm may be arbitrary modalities such as RGB, depths, or even any combination of them. While, the latter indicates that the inputs may have arbitrary modality numbers as the input type is changed, e.g. single-modality RGB image, dual-modality RGB-Depth (RGB-D) images or triple-modality RGB-Depth-Thermal (RGB-D-T) images. Accordingly, a preliminary solution to the above challenges, i.e. a modality switch network (MSN), is proposed in this paper. In particular, a modality switch feature extractor (MSFE) is first designed to extract discriminative features from each modality effectively by introducing some modality indicators, which will generate some weights for modality switching. Subsequently, a dynamic fusion module (DFM) is proposed to adaptively fuse features from a variable number of modalities based on a novel Transformer structure. Finally, a new dataset, named AM-XD, is constructed to facilitate research on AM SOD. Extensive experiments demonstrate that our AM SOD method can effectively cope with changes in the type and number of input modalities for robust salient object detection. Our code and AM-XD dataset will be released on https://github.com/nexiakele/AMSODFirst.
Nianchang Huang, Yang Yang 0132, Ruida Xi, Qiang Zhang 0020, Jungong Han, Jin Huang 0004
IEEE Trans. Image Process.1
2023 On exploring pose estimation as an auxiliary learning task for Visible-Infrared Person Re-identification
Yunqi Miao, Nianchang Huang, Xiao Ma 0013, Qiang Zhang 0020, Jungong Han
Neurocomputing2
2023 Exploring modality-shared appearance features and modality-invariant relation features for cross-modality person Re-IDentification
Nianchang Huang, Yongjiang Luo, Qiang Zhang 0020, Jungong Han
Pattern Recognit.1
2022 FMCNet: Feature-Level Modality Compensation for Visible-Infrared Person Re-Identification
abstract
For Visible-Infrared person ReIDentification (VI-ReID), existing modality-specific information compensation based models try to generate the images of missing modality from existing ones for reducing cross-modality discrepancy. However, because of the large modality discrepancy between visible and infrared images, the generated images usually have low qualities and introduce much more interfering information (e.g., color inconsistency). This greatly degrades the subsequent VI-ReID performance. Alternatively, we present a novel Feature-level Modality Compensation Network (FMCNet) for VI-ReID in this paper, which aims to compensate the missing modality-specific information in the feature level rather than in the image level, i.e., directly generating those missing modality-specific features of one modality from existing modality-shared features of the other modality. This will enable our model to mainly generate some discriminative person related modality-specific features and discard those non-discriminative ones for benefiting VI-ReID. For that, a single-modality feature decomposition module is first designed to decompose single-modality features into modality-specific ones and modality-shared ones. Then, a feature-level modality compensation module is present to generate those missing modality-specific features from existing modality-shared ones. Finally, a shared-specific feature fusion module is proposed to combine the existing and generated features for VI-ReID. The effectiveness of our proposed model is verified on two benchmark datasets.
Qiang Zhang 0020, Changzhou Lai, Nianchang Huang, Jungong Han
CVPR4
2022 Enabling modality interactions for RGB-T salient object detection
Qiang Zhang 0020, Ruida Xi, Tonglin Xiao, Nianchang Huang, Yongjiang Luo
Comput. Vis. Image Underst.4
2022 Cross-modality person re-identification via multi-task learning
Nianchang Huang, Kunlong Liu, Yang Liu 0069, Qiang Zhang 0020, Jungong Han
Pattern Recognit.1
2022 Discriminative unimodal feature selection and fusion for RGB-D salient object detection
Nianchang Huang, Yongjiang Luo, Qiang Zhang 0020, Jungong Han
Pattern Recognit.1
2022 Revisiting Modality-Specific Feature Compensation for Visible-Infrared Person Re-Identification
abstract
Although modality-specific feature compensation becomes a prevailing paradigm for Visible-Infrared Person Re-Identification (VI-ReID) to learn features, it, performance-wise, is not promising, especially when compared to modality-shared feature learning. In this paper, by revisiting the modality-specific feature compensation based models, we reveal that the reasons for being under-performed are: (1) generated images of one modality from another modality may be poor in quality; (2) such existing models usually achieve the modality-specific feature compensation just via simple pixel-level fusion strategies; (3) generated images cannot fully replace corresponding missing ones, which brings in extra modality discrepancy. To address these issues, we propose a new Two-Stage Modality Enhancement Network (TSME) for VI-ReID. Concretely, it first considers the modality discrepancy for cross-modality style translation and optimizes the structures of image generators by involving a new Deeper Skip-connection Generative Adversarial Networks (DSGAN) to generate high-quality images. Then, it presents an attention mechanism based feature-level fusion module, i.e., Pair-wise Image Fusion (PwIF) module, and an auxiliary learning module, i.e., Invoking All-Images (IAI) module, to better exploit the generated and original images for reducing modality discrepancy from the perspectives of feature fusion and feature constraints, respectively. Comprehensive experiments are carried out to demonstrate the success of TSME in tackling the modality discrepancy issue exposed in VI-ReID.
Nianchang Huang, Qiang Zhang 0020, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.3
2022 Middle-Level Feature Fusion for Lightweight RGB-D Salient Object Detection
abstract
Most existing RGB-D salient object detection (SOD) models adopt a two-stream structure to extract the information from the input RGB and depth images. Since they use two subnetworks for unimodal feature extraction and multiple multi-modal feature fusion modules for extracting cross-modal complementary information, these models require a huge number of parameters, thus hindering their real-life applications. To remedy this situation, we propose a novel middle-level feature fusion structure that allows to design a lightweight RGB-D SOD model. Specifically, the proposed structure first employs two shallow subnetworks to extract low- and middle-level unimodal RGB and depth features, respectively. Afterward, instead of integrating middle-level unimodal features multiple times at different layers, we just fuse them once via a specially designed fusion module. On top of that, high-level multi-modal semantic features are further extracted for final salient object detection via an additional subnetwork. This will greatly reduce the network's parameters. Moreover, to compensate for the performance loss due to parameter deduction, a relation-aware multi-modal feature fusion module is specially designed to effectively capture the cross-modal complementary information during the fusion of middle-level multi-modal features. By enabling the feature-level and decision-level information to interact, we maximize the usage of the fused cross-modal middle-level features and the extracted cross-modal high-level features for saliency prediction. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed method over some state-of-the-art methods. Remarkably, our proposed model has only 3.9M parameters and runs at 33 FPS.
Nianchang Huang, Qiang Jiao, Qiang Zhang 0020, Jungong Han
IEEE Trans. Image Process.1
2022 Employing Bilinear Fusion and Saliency Prior Information for RGB-D Salient Object Detection
abstract
Multi-modal feature fusion and saliency reasoning are two core sub-tasks of RGB-D salient object detection. However, most existing models employ linear fusion strategies (e.g., concatenation) for multi-modal feature fusion and use a simple coarse-to-fine structure for saliency reasoning. Despite their simpleness, they can neither fully capture the cross-modal complementary information nor exploit the multi-level complementary information among the cross-modal features at different levels. To address these issues, a novel RGB-D salient object detection model is presented, where we pay special attention to the aforementioned two sub-tasks. Concretely, a multi-modal feature interaction module is first presented to explore more interactions between the unimodal RGB and depth features. It helps to capture their cross-modal complementary information by jointly using some simple linear fusion strategies and bilinear fusion ones. Then, a saliency prior information guided fusion module is presented to exploit the multi-level complementary information among the fused cross-modal features at different levels. Instead of employing a simple convolutional layer for the final saliency prediction, a saliency refinement and prediction module is designed to better exploit those extracted multi-level cross-modal information for RGB-D saliency detection. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed framework over some state-of-the-art methods.
Nianchang Huang, Yang Yang 0132, Dingwen Zhang, Qiang Zhang 0020, Jungong Han
IEEE Trans. Multim.1
2021 ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic Segmentation
abstract
Semantic segmentation models gain robustness against poor lighting conditions by virtue of complementary information from visible (RGB) and thermal images. Despite its importance, most existing RGB-T semantic segmentation models perform primitive fusion strategies, such as concatenation, element-wise summation and weighted summation, to fuse features from different modalities. These strategies, unfortunately, overlook the modality differences due to different imaging mechanisms, so that they suffer from the reduced discriminability of the fused features. To address such an issue, we propose, for the first time, the strategy of bridging-then-fusing, where the innovation lies in a novel Adaptive-weighted Bi-directional Modality Difference Reduction Network (ABMDRNet). Concretely, a Modality Difference Reduction and Fusion (MDRF) subnetwork is designed, which first employs a bi-directional image-to-image translation based method to reduce the modality differences between RGB features and thermal features, and then adaptively selects those discriminative multi-modality features for RGB-T semantic segmentation in a channel-wise weighted fusion way. Furthermore, considering the importance of contextual information in semantic segmentation, a Multi-Scale Spatial Context (MSC) module and a Multi-Scale Channel Context (MCC) module are proposed to exploit the interactions among multi-scale contextual information of cross-modality features together with their long-range dependencies along spatial and channel dimensions, respectively. Comprehensive experiments on MFNet dataset demonstrate that our method achieves new state-of-the-art results.
Qiang Zhang 0020, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang, Nianchang Huang, Jungong Han
CVPR5
2021 Revisiting Feature Fusion for RGB-T Salient Object Detection
abstract
While many RGB-based saliency detection algorithms have recently shown the capability of segmenting salient objects from an image, they still suffer from unsatisfactory performance when dealing with complex scenarios, insufficient illumination or occluded appearances. To overcome this problem, this article studies RGB-T saliency detection, where we take advantage of thermal modality's robustness against illumination and occlusion. To achieve this goal, we revisit feature fusion for mining intrinsic RGB-T saliency patterns and propose a novel deep feature fusion network, which consists of the multi-scale, multi-modality, and multi-level feature fusion modules. Specifically, the multi-scale feature fusion module captures rich contexture features from each modality feature, while the multi-modality and multi-level feature fusion modules integrate complementary features from different modality features and different level of features, respectively. To demonstrate the effectiveness of the proposed approach, we conduct comprehensive experiments on the RGB-T saliency detection benchmark. The experimental results demonstrate that our approach outperforms other state-of-the-art methods and the conventional feature fusion modules by a large margin.
Qiang Zhang 0020, Tonglin Xiao, Nianchang Huang, Dingwen Zhang, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.3
2021 Joint Cross-Modal and Unimodal Features for RGB-D Salient Object Detection
abstract
RGB-D salient object detection is one of the basic tasks in computer vision. Most existing models focus on investigating efficient ways of fusing the complementary information from RGB and depth images for better saliency detection. However, for many real-life cases, where one of the input images has poor visual quality or contains affluent saliency cues, fusing cross-modal features does not help to improve the detection accuracy, when compared to using unimodal features only. In view of this, a novel RGB-D salient object detection model is proposed by simultaneously exploiting the cross-modal features from the RGB-D images and the unimodal features from the input RGB and depth images for saliency detection. To this end, a Multi-branch Feature Fusion Module is presented to effectively capture the cross-level and cross-modal complementary information between RGB-D images, as well as the cross-level unimodal features from the RGB images and the depth images separately. On top of that, a Feature Selection Module is designed to adaptively select those highly discriminative features for the final saliency prediction from the fused cross-modal features and the unimodal features. Extensive evaluations on four benchmark datasets demonstrate that the proposed model outperforms the state-of-the-art approaches by a large margin.
Nianchang Huang, Yi Liu 0038, Qiang Zhang 0020, Jungong Han
IEEE Trans. Multim.1
2020 RGB-T Salient Object Detection via Fusing Multi-Level CNN Features
abstract
RGB-induced salient object detection has recently witnessed substantial progress, which is attributed to the superior feature learning capability of deep convolutional neural networks (CNNs). However, such detections suffer from challenging scenarios characterized by cluttered backgrounds, low-light conditions and variations in illumination. Instead of improving RGB based saliency detection, this paper takes advantage of the complementary benefits of RGB and thermal infrared images. Specifically, we propose a novel end-to-end network for multi-modal salient object detection, which turns the challenge of RGB-T saliency detection to a CNN feature fusion problem. To this end, a backbone network (e.g., VGG-16) is first adopted to extract the coarse features from each RGB or thermal infrared image individually, and then several adjacent-depth feature combination (ADFC) modules are designed to extract multi-level refined features for each single-modal input image, considering that features captured at different depths differ in semantic information and visual details. Subsequently, a multi-branch group fusion (MGF) module is employed to capture the cross-modal features by fusing those features from ADFC modules for a RGB-T image pair at each level. Finally, a joint attention guided bi-directional message passing (JABMP) module undertakes the task of saliency prediction via integrating the multi-level fused features from MGF modules. Experimental results on several public RGB-T salient object detection datasets demonstrate the superiorities of our proposed algorithm over the state-of-the-art approaches, especially under challenging conditions, such as poor illumination, complex background and low contrast.
Qiang Zhang 0020, Nianchang Huang, Dingwen Zhang, Caifeng Shan, Jungong Han
IEEE Trans. Image Process.2