EDBT 2026 Demo / reviewers in the wild / expert
Qiang Zhang 0020
dblp:72/3527-20
· DBLP profile ↗
76ranked-venue papers
24as first author
56since 2021 · last 2026
0000-0002-2828-9905ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 14 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 12 first-author · 26 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracking and Segmenting Anything in Any ModalityabstractTracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite their shared objective, existing approaches often tackle these tasks using specialized architectures or modality-specific parameters, limiting their generalization and scalability. Recent efforts have attempted to unify multiple tracking and segmentation sub-tasks from the perspectives of any modality input or multi-task inference. However, these approaches tend to overlook two critical challenges: the distributional gap across different modalities and the feature representation gap across tasks. These issues hinder effective cross-task and cross-modal knowledge sharing, ultimately constraining the development of a true generalist model. To address these limitations, we propose a universal tracking and segmentation framework named SATA, which unifies a broad spectrum of tracking and segmentation subtasks with any modality input. Specifically, a Decoupled Mixture-of-Expert (DeMoE) mechanism is presented to decouple the unified representation learning task into the modeling process of cross-modal shared knowledge and specific information, thus enabling the model to maintain flexibility while enhancing generalization. Additionally, we introduce a Task-aware Multi-object Tracking (TaMOT) pipeline to unify all the task outputs as a unified set of instances with calibrated ID information, thereby alleviating the degradation of task-specific knowledge during multi-task training. SATA demonstrates superior performance on 18 challenging tracking and segmentation benchmarks, offering a novel perspective for more generalizable video understanding. Tianlu Zhang, Qiang Zhang 0020, Guiguang Ding, Jungong Han |
AAAI | 2 |
| 2026 | FC$^{2}$2: Fast Co-Clustering With Small-Scale Similarity Graph and Bipartite Graph LearningabstractBipartite graph-based co-clustering is efficient in modeling cluster manifold structures. However, existing methods decouple bipartite graph construction from the learning of pseudo-labels for samples and anchors, often leading to suboptimal clustering performance. Moreover, neglecting local manifold relationships among anchors yields inferior anchor pseudo-labels, which further degrades the quality of sample pseudo-labels. To overcome these limitations, we propose a novel model termed Fast Co-Clustering (FC$^{2}$2), which jointly captures both local and global correlations between samples and anchors. Specifically, to model the coupling between the one-hot pseudo-labels of samples and anchors, we construct a bipartite graph with adaptively updated weights during the clustering process. To prevent severely imbalanced cluster assignments, we prove the equivalence between maximizing pseudo-label covariance and balancing cluster proportions, and incorporate a balanced regularization term to enhance the rationality of the resulting clusters. Furthermore, the local smoothness of anchor pseudo-labels is preserved via a low-rank decomposition of a compact anchor similarity graph. These two components jointly ensure that spatially adjacent anchors tend to share similar cluster identities, and that samples and anchors in close proximity are also assigned to similar clusters. We develop an efficient iterative optimization algorithm to update all model variables. Extensive experiments on benchmark and synthetic datasets validate the superior performance and efficiency of the proposed method compared with state-of-the-art approaches. Xiaowei Zhao 0002, Linrui Xie, Xiaojun Chang, Feiping Nie 0001, Qiang Zhang 0020 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | 'Knowledge and experience' for visible-infrared person re-identification
Nianchang Huang, Qiang Zhang 0020, Jungong Han, Jin Huang 0004 |
Pattern Recognit. | 3 |
| 2026 | Mitigating fusion bias for RGB-D salient object detection
Yang Yang 0009, Nianchang Huang, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 3 |
| 2026 | QANet: Query-Aware Multi-Modal Prior Refinement for Few-Shot SegmentationabstractRecent Few-Shot Segmentation (FSS) approaches incorporate vision-language models to improve segmentation by constructing multi-modal priors, including visual, CAM and textual priors. However, these priors are often suboptimal, suffering from background interference, incomplete activation and semantic misalignment. To address these limitations, we propose a query-aware multi-modal prior refinement framework (QANet) for FSS, in which the multi-modal priors are jointly refined for more accurate segmentation by fully exploiting target-relevant cues from query features. Specifically, QANet comprises two modules, i.e., a Query-aware Multi-modal Prior Refinement (QMPR) module for visual and CAM prior refinement, and a Query-Aware Textual Embedding and Prior Refinement (QTEPR) module for textual prior refinement. More specifically, in QMPR, a query self-attention bilateral-guided enhancement strategy and a query self-similarity prior refinement strategy are carefully designed for suppressing background interferences and recovering foreground completeness, respectively. And in QTEPR, such refined visual and CAM priors are first leveraged to extract target-relevant cues for visual-textual domain gap mitigation and text embedding enhancement, producing query-aware textual priors. These textual priors are then enhanced under the guidance of those refined multi-modal priors. Extensive experiments on PASCAL-5iand COCO-20idatasets demonstrate that QANet achieves new state-of-the-art performance. Moreover, QANet remains competitive on cross-domain FSS and weak-label FSS tasks. Qiang Jiao, Mengrui Shi, Qiang Zhang 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Efficient Co-Clustering via Bipartite Graph Factorization
Xiaowei Zhao 0002, Liuyun Guo, Xiaojun Chang, Jun Guo 0020, Feiping Nie 0001, Qiang Zhang 0020 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2026 | Modality Adaptive Network for Arbitrary Modality Salient Object DetectionabstractThis paper delves into the task of arbitrary modality salient object detection (AM SOD), aiming to detect salient objects from the images with arbitrary modality types or arbitrary modality numbers by using a single model trained once. Specifically, we develop a novel model, termed modality adaptive network (MAN), for AM SOD, which addresses two fundamental challenges in AM SOD: the diverse modality discrepancies arising from varying modality types and the dynamic fusion dilemma resulting from an unfixed number of modalities in the input data. Technically, MAN first introduces a novel Modality-Adaptive Feature Extractor (MAFE) to adaptively extract features from different input modalities based on their characteristics by utilizing a set of learnable modality prompts. Concurrently, a new modality translation contractive (MTC) loss is devised to facilitate the training of MAFE as well as modality prompts, thereby effectively addressing the inherent modality discrepancies and extracting more discriminative features from each modality image. Subsequently, MAN presents a hybrid dynamic fusion (HDF) strategy to effectively resolve the challenge of dynamic inputs in multi-modal feature fusion as well as enhance the exploitation of complementary information across different modalities. This is specially achieved by a Channel- wise Dynamic Fusion Module (CDFM) and a Spatial- wise Dynamic Fusion Module (SDFM). Experimental results show that by virtue of MAFE, MTC loss and HDF strategy, our proposed method achieves significant increasements over existing models on benchmark datasets. Yang Yang 0132, Nianchang Huang, Qiang Zhang 0020, Jungong Han, Jin Huang 0004 |
IEEE Trans. Multim. | 3 |
| 2025 | Cross-Modality Distillation for Multi-Modal TrackingabstractContemporary multi-modal trackers achieve strong performance by leveraging complex backbones and fusion strategies, but this comes at the cost of computational efficiency, limiting their deployment in resource-constrained settings. On the other hand, compact multi-modal trackers are more efficient but often suffer from reduced performance due to limited feature representation. To mitigate the performance gap between compact and more complex trackers, we introduce a cross-modality distillation framework. This framework includes a complementarity-aware mask autoencoder designed to enhance cross-modal interactions by selectively masking patches within a modality, thereby forcing the model to learn more robust multi-modal representations. Additionally, we present a specific-common feature distillation module that transfers both modality-specific and shared information from a more powerful model's backbone to the compact model. Moreover, we develop a multi-path selection distillation module to guide a simple fusion module in learning more accurate multi-modal information from a sophisticated fusion mechanism using multiple paths. Extensive experiments on six multi-modal tracking benchmarks demonstrate that the proposed tracker, despite being lightweight, outperforms most state-of-the-art methods, highlighting its effectiveness. Notably, our tiny variant achieves a PR score of 67.5% on LasHeR, a PR score of 58.5% on DepthTrack, and a PR score of 73.1% on VisEvent with only 6.5 M parameters, while operating at 126 FPS on an NVIDIA 2080Ti GPU. Tianlu Zhang, Qiang Zhang 0020, Kurt Debattista, Jungong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Resolving semantic conflicts in RGB-T semantic segmentation
Shenlu Zhao, Ziniu Jin, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 4 |
| 2025 | Modality Adaptive Representation Network for Efficient RGB-T Semantic Segmentation
Shenlu Zhao, Ziniu Jin, Qiang Zhang 0020 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Scalable Multi-View Regression Clustering for Large-Scale DataabstractIn recent years, unsupervised linear regression has attracted attention for its ability to directly capture the mapping relationship between samples and targets. However, existing algorithms can only utilize limited information from a single view, which often leads to unsatisfactory results. To address this problem, we propose a regression clustering model based on multi-view information fusion, called Scalable Multi-view Regression Clustering. This model consists of two parts: intra-view information fusion and inter-view information fusion. In the first part, to capture the local correlations among samples, we propose constructing view-specific bipartite graphs. Unlike traditional single-view and multi-view clustering algorithms, we treat the weights of the bipartite graph as additional features of the samples, thereby directly incorporating the local manifold structure of the samples at the feature level. Furthermore, since the original features of the samples also contain valuable information, we perform unsupervised linear regression separately on the samples represented by the original features and those represented by the bipartite graph weights in each view. The results are then integrated in a weighted manner. In the second part, we propose adaptively weighting the clustering results from each view to capture complementary information across views, thereby enhancing clustering performance. This strategy not only avoids the bipartite graph alignment issue in multi-view clustering but also enables clustering with linear time complexity, making it effective for handling large-scale data. An iterative optimization algorithm is developed to update all variables alternately. Experiments conducted on benchmark datasets demonstrate the superiority of our proposed model. Xiaowei Zhao 0002, Xiaojun Chang, Feiping Nie 0001, Qiang Zhang 0020, Jun Guo 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | C⁴Net: Excavating Cross-Modal Context- and Content-Complementarity for RGB-T Semantic SegmentationabstractThe complementary properties exhibited upon RGB-T data involve context complementarity as well as content complementarity. During cross-modal feature fusion, most existing RGB-T semantic segmentation methods are dedicated to highlighting the exploitation of content-complementary information. Unfortunately, these methods usually overlook the excavation of cross-modal context-complementary information (i.e., the contextual dependencies among different regions that only exist in one certain modality data) or try to exploit such cross-modal context-complementary information in an implicit way, yielding fragmentary semantic segmentation results. To remedy this problem, in this paper, a novel Cross-modal Context- and Content-Complementarity Network (${\mathbf { C}}^{4}$Net) is presented for RGB-T semantic segmentation, in which both the cross-modal context-complementary information and the cross-modal content-complementary information are fully excavated and exploited during cross-modal feature fusion. Specifically, a Context-Complementary Information Aggregation (CxCIA) module is carefully designed, in which the cross-modal context-complementary information is explicitly excavated by measuring the discrepancies between contextual dependencies from different modality data. Then, such cross-modal context-complementary information is further exploited to enhance the original RGB and thermal contextual dependencies for boosting the integrity of objects in the fused features. In the meantime, a Content-Complementary Information Aggregation (CnCIA) module is presented, which highlights the utilization of cross-modal content-complementary information from a multi-scale perspective. Furthermore, an MLP-based Multi-level Feature Interaction (MFI) decoder is presented, in which the semantic gaps among different levels of fused features are mitigated by establishing the interactions of multi-level fused features along spatial and channel dimensions. Comprehensive experimental results on several public datasets demonstrate that our proposed${\mathbf { C}}^{4}$Net surpasses other state-of-the-art models. Shenlu Zhao, Qiang Zhang 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | CSANet: Cross-Modality Self-Paced Association Network for Unsupervised Visible-Infrared Person Re-IdentificationabstractFor preeminent unsupervised visible-infrared person re-identification (US-VI-ReID), existing studies typically adhere to a two-step paradigm,i.e., intra-modality clustering and inter-modality matching. Nevertheless, high intra-modality variations may result in suboptimal clusters containing intricate pedestrians, while significant inter-modality discrepancies further complicate their cross-modality associations. Most existing methods fail to adopt a differentiated approach for samples of varying difficulty, especially intricate ones. To address this, we propose enabling the model to gradually establish cross-modality associations from easy to hard, mimicking human learning patterns to avoid error accumulation caused by intricate pedestrians. To this end, we propose a Cross-modality Self-paced Association Network, termed CSANet, embracing Twain Bipartite Graph Matching (TBGM), Cross-curriculum Association Prompter (CAP) and Instance-Prototype Consistency Constraint (IPCC) modules. TBGM conceives a graph-driven metric to tailor athree-levelcurriculum (plain,moderateandintricate) for self-paced cross-modality learning. CAP transfers high-confidence associations deduced from the plain subsets to intricate ones, prompting exploring more complex cross-modality relationships. Alongside CAP, IPCC further enforces the intricate instances to mimic their prototype characteristics, facilitating their discriminative feature learning. Extensive experiments demonstrate CSANet’s superiority over state-of-the-art methods, highlighting the potential of self-paced learning for US-VI-ReID. Ruida Xi, Zhenyang Fu, Nianchang Huang, Xiaowei Zhao 0002, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Cross-Modality Prompts: Few-Shot Multi-Label Recognition With Single-Label TrainingabstractFew-shot multi-label recognition (FS-MLR) presents a significant challenge due to the need to assign multiple labels to images with limited examples. Existing methods often struggle to balance the learning of novel classes and the retention of knowledge from base classes. To address this issue, we propose a novel Cross-Modality Prompts (CMP) approach. Unlike conventional methods that rely on additional semantic information to mitigate the impact of limited samples, our approach leverages multimodal prompts to adaptively tune the feature extraction network. A new FS-MLR benchmark is also proposed, which includes single-label training and multi-label testing, accompanied by benchmark datasets constructed from MS-COCO and NUS-WIDE. Extensive experiments on these datasets demonstrate the superior performance of our CMP approach, highlighting its effectiveness and adaptability. Our results show that CMP outperforms CoOp on the MS-COCO dataset with a maximal improvement of 19.47% and 23.94% in mAPharmonicfor 5-way 1-shot and 5-way 5-shot settings, respectively. Zixuan Ding, Hui Chen 0013, Tianxiang Hao 0001, Yizhe Xiong, Sicheng Zhao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Multim. | 7 |
| 2025 | FMCNet+: Feature-Level Modality Compensation for Visible-Infrared Person Re-IdentificationabstractFor visible-infrared person re-identification (VI-ReID), current models that compensate modality-specific information strive to generate missing modality images from existing ones to bridge the cross-modality discrepancies. Despite that, those generated images often suffer from low qualities due to the significant modality gap and include interfering information, e.g., inconsistent colors, thus severely degrading the subsequent VI-ReID performance. Alternatively, we propose a feature-level modality compensation network, i.e., FMCNet+, for VI-ReID in this article as an improved version of our previous work (FMCNet). The core of FMCNet+ is to compensate for the missing modality-specific information at the feature level, rather than at the image level, enabling our model to generate more person-related and discriminative modality-specific features for VI-ReID. Concretely, FMCNet+ aims to progressively generate missing modality-specific features by fully exploring the relationships among single-modality features, modality-shared features, and modality-specific features, instead of directly generating them through a generative adversarial way as in the previous FMCNet. To this end, three modules, i.e., single-modality feature decomposition (SFD), modality characteristic dictionary learning (MCDL), and missing modality-specific feature compensation (MMFC), are incorporated in FMCNet+. Experimental results demonstrate the superiority of our proposed FMCNet+ over existing ones, especially for those that compensate for modality-specific information at the image level. Our intriguing findings highlight the necessity of feature-level modality compensation in VI-ReID. Our code and pre-trained models will be released on https://github.com/jssyzsfzy/FMCNet_series. Ruida Xi, Nianchang Huang, Changzhou Lai, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Revisiting motion information for RGB-Event tracking with MOT philosophyabstractRGB-Event single object tracking (SOT) aims to leverage the merits of RGB and event data to achieve higher performance. However, existing frameworks focus on exploring complementary appearance information within multi-modal data, and struggle to address the association problem of targets and distractors in the temporal domain using motion information from the event stream. In this paper, we introduce the Multi-Object Tracking (MOT) philosophy into RGB-E SOT to keep track of targets as well as distractors by using both RGB and event data, thereby improving the robustness of the tracker. Specifically, an appearance model is employed to predict the initial candidates. Subsequently, the initially predicted tracking results, in combination with the RGB-E features, are encoded into appearance and motion embeddings, respectively. Furthermore, a Spatial-Temporal Transformer Encoder is proposed to model the spatial-temporal relationships and learn discriminative features for each candidate through guidance of the appearance-motion embeddings. Simultaneously, a Dual-Branch Transformer Decoder is designed to adopt such motion and appearance information for candidate matching, thus distinguishing between targets and distractors. The proposed method is evaluated on multiple benchmark datasets and achieves state-of-the-art performance on all the datasets tested. Tianlu Zhang, Kurt Debattista, Qiang Zhang 0020, Guiguang Ding, Jungong Han |
NeurIPS | 3 |
| 2024 | Lightweight cross-modal transformer for RGB-D salient object detection
Nianchang Huang, Yang Yang 0009, Qiang Zhang 0020, Jungong Han, Jin Huang 0004 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Deep unsupervised shadow detection with curriculum learning and self-training
Qiang Zhang 0020, Hongyuan Guo, Guanghe Li, Tianlu Zhang, Qiang Jiao |
Comput. Vis. Image Underst. | 1 |
| 2024 | Multi-level modality-specific and modality-common features fusion network for RGB-IR person re-identification
Qiang Zhang 0020 |
Neurocomputing | 2 |
| 2024 | Exploring target-related information with reliable global pixel relationships for robust RGB-T tracking
Tianlu Zhang, Xiaoyi He, Yongjiang Luo, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 4 |
| 2024 | Finding Camouflaged Objects Along the Camouflage MechanismsabstractCommon mechanisms for achieving object camouflage include reducing differences and increasing distractions. Such camouflage mechanisms hinder the object detectors to accurately distinguish the camouflaged objects from their surroundings. Considering that, we reexamine the camouflaged object detection (COD) task from the perspective of camouflage mechanisms and make the first attempt to discover the target objects in a de-camouflaging manner. We argue that this process can not only lead to a better understanding of camouflage, but also provide a new perspective for detecting camouflaged objects. For that, we first analyze some existing camouflage mechanisms together with their induced problems. Afterwards, considering the inner relationships between SOD and COD, we resort to the SOD task to synergistically achieve de-camouflaging for COD. Specifically, we incorporate the SOD task into the COD model and present a multi-task learning framework for COD, which models the intrinsic relationships between the two tasks from different perspectives, i.e., task-conflicting attribute and task-consistent attribute, to destroy the camouflage conditions for highlighting those inconspicuous yet valuable cues of camouflaged objects. In more detail, modeling the task-conflicting attribute is to well identify camouflaged objects by alleviating such interfering information from salient ones, and is achieved by a Gate Classification (GC) strategy and a Region Distraction Module (RDM). While, modeling the task-consistent attribute, which is achieved by an adversarial learning (AL) scheme and a Boundary Injection Module (BIM), is intended to enhance the boundary differences between the camouflaged objects and their backgrounds for fully segmenting the camouflaged objects. Extensive results demonstrate the superiorities of our proposed model over existing ones in camouflaged object detection. Yang Yang 0132, Qiang Zhang 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | AMNet: Learning to Align Multi-Modality for RGB-T TrackingabstractRGB-T tracking has attracted increasing attention recently due to the all-weather and all-day working capability. However, most current RGB-T trackers usually assume that RGB data and thermal infrared (TIR) data are well spatially aligned, which is difficult to be achieved in practice. Such spatial misalignment between RGB data and TIR data may lead to the ineffective cross-modal information propagation during multi-modal feature fusion, thus reducing the tracking performance. In addition, due to the discrepancy in imaging characteristics of RGB images and TIR images, there also exist great differences between the information captured by the two modality data. The differences in characteristics of RGB and TIR modalities in different local areas will cause a single fusion strategy to be unable to fully explore the complementary information within multi-modal data. For that, we propose an RGB-T tracker, referred to as AMNet, to specifically solve such two problems with two dedicated modules, i.e., a Mutual-interacted Spatial Alignment (MSA) module and an Information Matching Fusion (IMF) module. The former spatially aligns the two modality data through three essential parts, including interactions of multi-modal features, prediction of cross-modal offset map, and enhancement of the aligned features. While the latter first discriminates different types of local regions by employing several intra-modal attention modules and then uses a divide-and-conquer fusion strategy to exploit such discriminative information within RGB and TIR features of different cases for tracking. We validate the effectiveness of our AMNet with extensive experiments on three RGB-T benchmarks, which achieves new state-of-the-art performance. Tianlu Zhang, Xiaoyi He, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Feature Calibrating and Fusing Network for RGB-D Salient Object DetectionabstractDue to their imaging mechanisms and techniques, some depth images inevitably have low visual qualities or have some inconsistent foregrounds with their corresponding RGB images. Directly using such depth images will deteriorate the performance of RGB-D SOD. In view of this, a novel RGB-D salient object detection model is presented, which follows the principle of calibration-then-fusion to effectively suppress the influence of such two types of depth images on final saliency prediction. Specifically, the proposed model is composed of two stages, i.e., an image generation stage and a saliency reasoning stage. The former generates high-quality and foreground-consistent pseudo depth images via an image generation network. While the latter first calibrates the original depth information with the aid of those newly generated pseudo depth images and then performs cross-modal feature fusion for the final saliency reasoning. Especially, in the first stage, a Two-steps Sample Selection (TSS) strategy is employed to select such reliable depth images from the original RGB-D image pairs as supervision information to optimize the image generation network. Afterwards, in the second stage, a Feature Calibrating and Fusing Network (FCFNet) is proposed to achieve the calibration-then-fusion of cross-modal information for the final saliency prediction, which is achieved by a Depth Feature Calibration (DFC) module, a Shallow-level Feature Injection (SFI) module and a Multi-modal Multi-scale Fusion (MMF) module. Moreover, a loss function, i.e., Region Consistency Aware (RCA) loss, is presented as an auxiliary loss for FCFNet to facilitate the completeness of salient objects together with the reduction of background interference by considering the local regional consistency in the saliency maps. Experiments on six benchmark datasets demonstrate the superiorities of our proposed RGB-D SOD model over some state-of-the-arts. Qiang Zhang 0020, Yang Yang 0132, Qiang Jiao, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Salient Object Detection From Arbitrary ModalitiesabstractToward desirable saliency prediction, the types and numbers of inputs for a salient object detection (SOD) algorithm may dynamically change in many real-life applications. However, existing SOD algorithms are mainly designed or trained for one particular type of inputs, failing to be generalized to other types of inputs. Consequentially, more types of SOD algorithms need to be prepared in advance for handling different types of inputs, raising huge hardware and research costs. Differently, in this paper, we propose a new type of SOD task, termed Arbitrary Modality SOD (AM SOD). The most prominent characteristics of AM SOD are that the modality types and modality numbers will be arbitrary or dynamically changed. The former means that the inputs to the AM SOD algorithm may be arbitrary modalities such as RGB, depths, or even any combination of them. While, the latter indicates that the inputs may have arbitrary modality numbers as the input type is changed, e.g. single-modality RGB image, dual-modality RGB-Depth (RGB-D) images or triple-modality RGB-Depth-Thermal (RGB-D-T) images. Accordingly, a preliminary solution to the above challenges, i.e. a modality switch network (MSN), is proposed in this paper. In particular, a modality switch feature extractor (MSFE) is first designed to extract discriminative features from each modality effectively by introducing some modality indicators, which will generate some weights for modality switching. Subsequently, a dynamic fusion module (DFM) is proposed to adaptively fuse features from a variable number of modalities based on a novel Transformer structure. Finally, a new dataset, named AM-XD, is constructed to facilitate research on AM SOD. Extensive experiments demonstrate that our AM SOD method can effectively cope with changes in the type and number of input modalities for robust salient object detection. Our code and AM-XD dataset will be released on https://github.com/nexiakele/AMSODFirst. Nianchang Huang, Yang Yang 0132, Ruida Xi, Qiang Zhang 0020, Jungong Han, Jin Huang 0004 |
IEEE Trans. Image Process. | 4 |
| 2024 | Exploring Multi-Modal Spatial-Temporal Contexts for High-Performance RGB-T TrackingabstractIn RGB-T tracking, there exist rich spatial relationships between the target and backgrounds within multi-modal data as well as sound consistencies of spatial relationships among successive frames, which are crucial for boosting the tracking performance. However, most existing RGB-T trackers overlook such multi-modal spatial relationships and temporal consistencies within RGB-T videos, hindering them from robust tracking and practical applications in complex scenarios. In this paper, we propose a novel Multi-modal Spatial-Temporal Context (MMSTC) network for RGB-T tracking, which employs a Transformer architecture for the construction of reliable multi-modal spatial context information and the effective propagation of temporal context information. Specifically, a Multi-modal Transformer Encoder (MMTE) is designed to achieve the encoding of reliable multi-modal spatial contexts as well as the fusion of multi-modal features. Furthermore, a Quality-aware Transformer Decoder (QATD) is proposed to effectively propagate the tracking cues from historical frames to the current frame, which facilitates the object searching process. Moreover, the proposed MMSTC network can be easily extended to various tracking frameworks. New state-of-the-art results on five prevalent RGB-T tracking benchmarks demonstrate the superiorities of our proposed trackers over existing ones. Tianlu Zhang, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Image Process. | 3 |
| 2024 | Mitigating Modality Discrepancies for RGB-T Semantic SegmentationabstractSemantic segmentation models gain robustness against adverse illumination conditions by taking advantage of complementary information from visible and thermal infrared (RGB-T) images. Despite its importance, most existing RGB-T semantic segmentation models directly adopt primitive fusion strategies, such as elementwise summation, to integrate multimodal features. Such strategies, unfortunately, overlook the modality discrepancies caused by inconsistent unimodal features obtained by two independent feature extractors, thus hindering the exploitation of cross-modal complementary information within the multimodal data. For that, we propose a novel network for RGB-T semantic segmentation, i.e. MDRNet+, which is an improved version of our previous work ABMDRNet. The core of MDRNet+ is a brand new idea, termed the strategy of bridging-then-fusing, which mitigates modality discrepancies before cross-modal feature fusion. Concretely, an improved Modality Discrepancy Reduction (MDR+) subnetwork is designed, which first extracts unimodal features and reduces their modality discrepancies. Afterward, discriminative multimodal features for RGB-T semantic segmentation are adaptively selected and integrated via several channel-weighted fusion (CWF) modules. Furthermore, a multiscale spatial context (MSC) module and a multiscale channel context (MCC) module are presented to effectively capture the contextual information. Finally, we elaborately assemble a challenging RGB-T semantic segmentation dataset, i.e., RTSS, for urban scene understanding to mitigate the lack of well-annotated training data. Comprehensive experiments demonstrate that our proposed model surpasses other state-of-the-art models on the MFNet, PST900, and RTSS datasets remarkably. Shenlu Zhao, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Exploring Structured Semantic Prior for Multi Label Recognition with Incomplete LabelsabstractMulti-label recognition (MLR) with incomplete labels is very challenging. Recent works strive to explore the image-to-label correspondence in the vision-language model, i.e., CLIP [22], to compensate for insufficient annotations. In spite of promising performance, they generally overlook the valuable prior about the label-to-label correspondence. In this paper, we advocate remedying the deficiency of label supervision for the MLR with incomplete labels by deriving a structured semantic prior about the label-to-label corre-spondence via a semantic prior prompter. We then present a novel Semantic Correspondence Prompt Network (SCP-Net), which can thoroughly explore the structured semantic prior. A Prior-Enhanced Self-Supervised Learning method is further introduced to enhance the use of the prior. Comprehensive experiments and analyses on several widely used benchmark datasets show that our method significantly out-performs existing methods on all datasets, well demonstrating the effectiveness and the superiority of our method. Our code will be available at https://github.com/jameslahm/SCPNet. Zixuan Ding, Hui Chen 0013, Qiang Zhang 0020, Pengzhang Liu, Yongjun Bao, Weipeng Yan, Jungong Han |
CVPR | 4 |
| 2023 | Efficient RGB-T Tracking via Cross-Modality DistillationabstractMost current RGB-T trackers adopt a two-stream structure to extract unimodal RGB and thermal features and complex fusion strategies to achieve multi-modal feature fusion, which require a huge number of parameters, thus hindering their real-life applications. On the other hand, a compact RGB-T tracker may be computationally efficient but encounter non-negligible performance degradation, due to the weakening of feature representation ability. To remedy this situation, a cross-modality distillation framework is presented to bridge the performance gap between a compact tracker and a powerful tracker. Specifically, a specific-common feature distillation module is proposed to transform the modality-common information as well as the modality-specific information from a deeper two-stream network to a shallower single-stream network. In addition, a multi-path selection distillation module is proposed to instruct a simple fusion module to learn more accurate multi-modal information from a well-designed fusion mechanism by using multiple paths. We validate the effectiveness of our method with extensive experiments on three RGB-T benchmarks, which achieves state-of-the-art performance but consumes much less computational resources. Tianlu Zhang, Hongyuan Guo, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
CVPR | 4 |
| 2023 | M2FINet: Modality-specific and Modality-shared Features Interaction Network for RGB-IR Person Re-Identification
Qiang Zhang 0020 |
Comput. Vis. Image Underst. | 3 |
| 2023 | On exploring pose estimation as an auxiliary learning task for Visible-Infrared Person Re-identification
Yunqi Miao, Nianchang Huang, Xiao Ma 0013, Qiang Zhang 0020, Jungong Han |
Neurocomputing | 4 |
| 2023 | Hybrid routing transformer for zero-shot learning
De Cheng, Gerong Wang, Bo Wang 0011, Qiang Zhang 0020, Jungong Han, Dingwen Zhang |
Pattern Recognit. | 4 |
| 2023 | Exploring modality-shared appearance features and modality-invariant relation features for cross-modality person Re-IDentification
Nianchang Huang, Yongjiang Luo, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 4 |
| 2023 | Discriminative and Robust Attribute Alignment for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to learn models that can recognize images of semantically related unseen categories, through transferring attribute-based knowledge learned from training data of seen classes to unseen testing data. As visual attributes play a vital role in ZSL, recent embedding-based methods usually focus on learning a compatibility function between the visual representation and the class semantic attributes. While in this work, in addition to simply learning the region embedding of different semantic attributes to maintain the generalization capability of the learned model, we further consider to improve the discrimination power of the learned visual features themselves by contrastive embedding. It exploits both the class-wise and instance-wise supervision for GZSL, under the attribute guided weakly supervised representation learning framework. To further improve the robustness of the ZSL model, we also propose to train the model under the consistency regularization constraint, through taking full advantages of self-supervised signals of the image under various perturbed augmentation situations, which could make the model robust to some occluded or un-related attribute regions. Extensive experimental results demonstrate the effectiveness of the proposed ZSL method, achieving superior performances to state-of-the-art methods on three widely-used benchmark datasets, namely CUB, SUN, and AWA2. Our source code is released athttps://github.com/KORIYN/CC-ZSL. De Cheng, Gerong Wang, Nannan Wang 0001, Dingwen Zhang, Qiang Zhang 0020, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | A Feature Divide-and-Conquer Network for RGB-T Semantic SegmentationabstractSimilar to other multi-modal pixel-level prediction tasks, existing RGB-T semantic segmentation methods usually employ a two-stream structure to extract RGB and thermal infrared (TIR) features, respectively, and adopt the same fusion strategies to integrate different levels of unimodal features. This will result in inadequate extraction of unimodal features and exploitation of cross-modal information from the paired RGB and TIR images. Alternatively, in this paper, we present a novel RGB-T semantic segmentation model, i.e., FDCNet, where a feature divide-and-conquer strategy performs unimodal feature extraction and cross-modal feature fusion in one go. Concretely, we first employ a two-stream structure to extract unimodal low-level features, followed by a Siamese structure to extract unimodal high-level features from the paired RGB and TIR images. This concise but efficient structure enables to take into account both the modality discrepancies of low-level features and the underlying semantic consistency of high-level features across the paired RGB and TIR images. Furthermore, considering the characteristics of different layers of features, a Cross-modal Spatial Activation (CSA) module and a Cross-modal Channel Activation (CCA) module are presented for the fusion of low-level RGB and TIR features and for the fusion of high-level RGB and TIR features, respectively, thus facilitating the capture of cross-modal information. On top of that, with an embedded Cross-scale Interaction Context (CIC) module for mining multi-scale contextual information, our proposed model (i.e., FDCNet) for RGB-T semantic segmentation achieves new state-of-the-art experimental results on MFNet dataset and PST900 dataset. Shenlu Zhao, Qiang Zhang 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | FMCNet: Feature-Level Modality Compensation for Visible-Infrared Person Re-IdentificationabstractFor Visible-Infrared person ReIDentification (VI-ReID), existing modality-specific information compensation based models try to generate the images of missing modality from existing ones for reducing cross-modality discrepancy. However, because of the large modality discrepancy between visible and infrared images, the generated images usually have low qualities and introduce much more interfering information (e.g., color inconsistency). This greatly degrades the subsequent VI-ReID performance. Alternatively, we present a novel Feature-level Modality Compensation Network (FMCNet) for VI-ReID in this paper, which aims to compensate the missing modality-specific information in the feature level rather than in the image level, i.e., directly generating those missing modality-specific features of one modality from existing modality-shared features of the other modality. This will enable our model to mainly generate some discriminative person related modality-specific features and discard those non-discriminative ones for benefiting VI-ReID. For that, a single-modality feature decomposition module is first designed to decompose single-modality features into modality-specific ones and modality-shared ones. Then, a feature-level modality compensation module is present to generate those missing modality-specific features from existing modality-shared ones. Finally, a shared-specific feature fusion module is proposed to combine the existing and generated features for VI-ReID. The effectiveness of our proposed model is verified on two benchmark datasets. Qiang Zhang 0020, Changzhou Lai, Nianchang Huang, Jungong Han |
CVPR | 1 |
| 2022 | Densely nested top-down flows for salient object detection
Chaowei Fang, Haibin Tian, Dingwen Zhang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001 |
Sci. China Inf. Sci. | 4 |
| 2022 | Onfocus detection: identifying individual-camera eye contact from unconstrained imagesabstractAbstract Onfocus detection aims at identifying whether the focus of the individual captured by a camera is on the camera or not. Based on the behavioral research, the focus of an individual during face-to-camera communication leads to a special type of eye contact, i.e., the individual-camera eye contact, which is a powerful signal in social communication and plays a crucial role in recognizing irregular individual status (e.g., lying or suffering mental disease) and special purposes (e.g., seeking help or attracting fans). Thus, developing effective onfocus detection algorithms is of significance for assisting the criminal investigation, disease discovery, and social behavior analysis. However, the review of the literature shows that very few efforts have been made toward the development of onfocus detector owing to the lack of large-scale public available datasets as well as the challenging nature of this task. To this end, this paper engages in the onfocus detection research by addressing the above two issues. Firstly, we build a large-scale onfocus detection dataset, named as the onfocus detection in the wild (OFDIW). It consists of 20623 images in unconstrained capture conditions (thus called “in the wild”) and contains individuals with diverse emotions, ages, facial characteristics, and rich interactions with surrounding objects and background scenes. On top of that, we propose a novel end-to-end deep model, i.e., the eye-context interaction inferring network (ECIIN), for onfocus detection, which explores eye-context interaction via dynamic capsule routing. Finally, comprehensive experiments are conducted on the proposed OFDIW dataset to benchmark the existing learning models and demonstrate the effectiveness of the proposed ECIIN. Dingwen Zhang, Bo Wang 0011, Gerong Wang, Qiang Zhang 0020, Jungong Han, Zheng You |
Sci. China Inf. Sci. | 4 |
| 2022 | Enabling modality interactions for RGB-T salient object detection
Qiang Zhang 0020, Ruida Xi, Tonglin Xiao, Nianchang Huang, Yongjiang Luo |
Comput. Vis. Image Underst. | 1 |
| 2022 | RGB-T tracking by modality difference reduction and feature re-selection
Qiang Zhang 0020, Xueru Liu, Tianlu Zhang |
Image Vis. Comput. | 1 |
| 2022 | Part-Object Relational Visual SaliencyabstractRecent years have witnessed a big leap in automatic visual saliency detection attributed to advances in deep learning, especially Convolutional Neural Networks (CNNs). However, inferring the saliency of each image part separately, as was adopted by most CNNs methods, inevitably leads to an incomplete segmentation of the salient object. In this paper, we describe how to use the property of part-object relations endowed by the Capsule Network (CapsNet) to solve the problems that fundamentally hinge on relational inference for visual saliency detection. Concretely, we put in place a two-stream strategy, termed Two-Stream Part-Object RelaTional Network (TSPORTNet), to implement CapsNet, aiming to reduce both the network complexity and the possible redundancy during capsule routing. Additionally, taking into account the correlations of capsule types from the preceding training images, a correlation-aware capsule routing algorithm is developed for more accurate capsule assignments at the training stage, which also speeds up the training dramatically. By exploring part-object relationships, TSPORTNet produces a capsule wholeness map, which in turn aids multi-level features in generating the final saliency map. Experimental results on five widely-used benchmarks show that our framework consistently achieves state-of-the-art performance. The code can be found on https://github.com/liuyi1989/TSPORTNet. Yi Liu 0038, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Cross-modality person re-identification via multi-task learning
Nianchang Huang, Kunlong Liu, Yang Liu 0069, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 4 |
| 2022 | Discriminative unimodal feature selection and fusion for RGB-D salient object detection
Nianchang Huang, Yongjiang Luo, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 3 |
| 2022 | Revisiting Modality-Specific Feature Compensation for Visible-Infrared Person Re-IdentificationabstractAlthough modality-specific feature compensation becomes a prevailing paradigm for Visible-Infrared Person Re-Identification (VI-ReID) to learn features, it, performance-wise, is not promising, especially when compared to modality-shared feature learning. In this paper, by revisiting the modality-specific feature compensation based models, we reveal that the reasons for being under-performed are: (1) generated images of one modality from another modality may be poor in quality; (2) such existing models usually achieve the modality-specific feature compensation just via simple pixel-level fusion strategies; (3) generated images cannot fully replace corresponding missing ones, which brings in extra modality discrepancy. To address these issues, we propose a new Two-Stage Modality Enhancement Network (TSME) for VI-ReID. Concretely, it first considers the modality discrepancy for cross-modality style translation and optimizes the structures of image generators by involving a new Deeper Skip-connection Generative Adversarial Networks (DSGAN) to generate high-quality images. Then, it presents an attention mechanism based feature-level fusion module, i.e., Pair-wise Image Fusion (PwIF) module, and an auxiliary learning module, i.e., Invoking All-Images (IAI) module, to better exploit the generated and original images for reducing modality discrepancy from the perspectives of feature fusion and feature constraints, respectively. Comprehensive experiments are carried out to demonstrate the success of TSME in tackling the modality discrepancy issue exposed in VI-ReID. Nianchang Huang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Bi-Directional Progressive Guidance Network for RGB-D Salient Object DetectionabstractMost existing RGB-D salient detection models pay more attention to the quality of the depth images, while in some special cases, the quality of RGB images may even have greater impacts on saliency detection, which has long been ignored and underestimated. To address this problem, in this paper, we present a Bi-directional Progressive Guidance Network (BPGNet) for RGB-D salient object detection, where the qualities of both RGB and depth images are involved. Since it is usually difficult to determine which modality data have low quality in advance, a bi-directional framework based on progressive guidance (PG) strategy is employed to extract and enhance the unimodal features with the aid of another modality data via the alternative interactions between the saliency prediction results and the extracted features from the multi-modality input data. Specifically, the proposed PG strategy is achieved by using the proposed Global Context Awareness (GCA), Auxiliary Feature Extraction (AFE) and Cross-modality Feature Enhancement (CFE) modules. Benefiting from the proposed PG strategy, the disturbing information within the input RGB and depth images can be well suppressed, while the discriminative information within the input images gets enhanced. On top of that, a Fusion Prediction Module (FPM) is further designed to adaptively select those features with higher discriminability as well as enhancing the common information for the final saliency prediction. Experimental results demonstrate that our proposed model is comparable to those of state-of-the-art RGB-D SOD models. Yang Yang 0132, Yongjiang Luo, Yi Liu 0038, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Engaging Part-Whole Hierarchies and Contrast Cues for Salient Object DetectionabstractReal-world scenes always exhibit objects with clutter backgrounds, posing great challenges for deep salient object detection models. In this paper, we propose salient object detection by engaging two saliency cues,i.e., the part-whole hierarchies and contrast cues, resulting in a PWHCNet. Specifically, two branches, which consists of a Dynamic Grouping Capsules (DGC) branch and a DenseHRNet branch, are put in place to learn the part-whole hierarchies and contrast cues, respectively. Moreover, to help highlight the whole salient object in complex scenes, a Background Suppression (BS) module is proposed to guide the shallow features of DenseHRNet with the aid of the part-whole relational cues captured by DGC. Subsequently, these two saliency cues are integrated via a Self-Channel and Mutual-Spatial (SCMS) attention mechanism. Experimental results on five benchmarks demonstrate that the proposed PWHCNet achieves state-of-the-art performance while obtaining the whole salient objects with fine details. Qiang Zhang 0020, Mingxing Duanmu, Yongjiang Luo, Yi Liu 0038, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | SiamCDA: Complementarity- and Distractor-Aware RGB-T Tracking Based on Siamese NetworkabstractRecent years have witnessed the prevalence of using the Siamese network for RGB-T tracking because of its remarkable success in RGB object tracking. Despite their faster than real-time speeds, existing RGB-T Siamese trackers suffer from low accuracy and poor robustness, compared to other state-of-the-art RGB-T trackers. To address such issues, a new complementarity- and distractor-aware RGB-T tracker based on Siamese network (referred to as SiamCDA) is developed in this paper. To this end, several modules are presented, where the feature pyramid network (FPN) is incorporated into the Siamese network to capture the cross-level information within unimodal features extracted from the RGB or the thermal images. Next, a complementarity-aware multi-modal feature fusion module (CA-MF) is specially designed to capture the cross-modal information between RGB features and thermal features. In the final bounding box selection phase, a distractor-aware region proposal selection module (DAS) further enhances the robustness of our tracker. On top of the technical modules, we also build a large-scale, diverse synthetic RGB-T tracking dataset, containing more than 4831 pairs of synthetic RGB-T videos and 12K synthetic RGB-T images. Extensive experiments on three RGB-T tracking benchmark datasets demonstrate the outstanding performance of our proposed tracker with a tracking speed over 37 frames per second (FPS). Tianlu Zhang, Xueru Liu, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Middle-Level Feature Fusion for Lightweight RGB-D Salient Object DetectionabstractMost existing RGB-D salient object detection (SOD) models adopt a two-stream structure to extract the information from the input RGB and depth images. Since they use two subnetworks for unimodal feature extraction and multiple multi-modal feature fusion modules for extracting cross-modal complementary information, these models require a huge number of parameters, thus hindering their real-life applications. To remedy this situation, we propose a novel middle-level feature fusion structure that allows to design a lightweight RGB-D SOD model. Specifically, the proposed structure first employs two shallow subnetworks to extract low- and middle-level unimodal RGB and depth features, respectively. Afterward, instead of integrating middle-level unimodal features multiple times at different layers, we just fuse them once via a specially designed fusion module. On top of that, high-level multi-modal semantic features are further extracted for final salient object detection via an additional subnetwork. This will greatly reduce the network's parameters. Moreover, to compensate for the performance loss due to parameter deduction, a relation-aware multi-modal feature fusion module is specially designed to effectively capture the cross-modal complementary information during the fusion of middle-level multi-modal features. By enabling the feature-level and decision-level information to interact, we maximize the usage of the fused cross-modal middle-level features and the extracted cross-modal high-level features for saliency prediction. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed method over some state-of-the-art methods. Remarkably, our proposed model has only 3.9M parameters and runs at 33 FPS. Nianchang Huang, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Image Process. | 3 |
| 2022 | Employing Bilinear Fusion and Saliency Prior Information for RGB-D Salient Object DetectionabstractMulti-modal feature fusion and saliency reasoning are two core sub-tasks of RGB-D salient object detection. However, most existing models employ linear fusion strategies (e.g., concatenation) for multi-modal feature fusion and use a simple coarse-to-fine structure for saliency reasoning. Despite their simpleness, they can neither fully capture the cross-modal complementary information nor exploit the multi-level complementary information among the cross-modal features at different levels. To address these issues, a novel RGB-D salient object detection model is presented, where we pay special attention to the aforementioned two sub-tasks. Concretely, a multi-modal feature interaction module is first presented to explore more interactions between the unimodal RGB and depth features. It helps to capture their cross-modal complementary information by jointly using some simple linear fusion strategies and bilinear fusion ones. Then, a saliency prior information guided fusion module is presented to exploit the multi-level complementary information among the fused cross-modal features at different levels. Instead of employing a simple convolutional layer for the final saliency prediction, a saliency refinement and prediction module is designed to better exploit those extracted multi-level cross-modal information for RGB-D saliency detection. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed framework over some state-of-the-art methods. Nianchang Huang, Yang Yang 0132, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Multim. | 4 |
| 2021 | ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationabstractSemantic segmentation models gain robustness against poor lighting conditions by virtue of complementary information from visible (RGB) and thermal images. Despite its importance, most existing RGB-T semantic segmentation models perform primitive fusion strategies, such as concatenation, element-wise summation and weighted summation, to fuse features from different modalities. These strategies, unfortunately, overlook the modality differences due to different imaging mechanisms, so that they suffer from the reduced discriminability of the fused features. To address such an issue, we propose, for the first time, the strategy of bridging-then-fusing, where the innovation lies in a novel Adaptive-weighted Bi-directional Modality Difference Reduction Network (ABMDRNet). Concretely, a Modality Difference Reduction and Fusion (MDRF) subnetwork is designed, which first employs a bi-directional image-to-image translation based method to reduce the modality differences between RGB features and thermal features, and then adaptively selects those discriminative multi-modality features for RGB-T semantic segmentation in a channel-wise weighted fusion way. Furthermore, considering the importance of contextual information in semantic segmentation, a Multi-Scale Spatial Context (MSC) module and a Multi-Scale Channel Context (MCC) module are proposed to exploit the interactions among multi-scale contextual information of cross-modality features together with their long-range dependencies along spatial and channel dimensions, respectively. Comprehensive experiments on MFNet dataset demonstrate that our method achieves new state-of-the-art results. Qiang Zhang 0020, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang, Nianchang Huang, Jungong Han |
CVPR | 1 |
| 2021 | Exploring multi-scale deformable context and channel-wise attention for salient object detection
Yi Liu 0038, Mingxing Duanmu, Zhen Huo, Zuntian Chen, Qiang Zhang 0020 |
Neurocomputing | 7 |
| 2021 | Cross-modality deep feature learning for brain tumor segmentation
Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Yu |
Pattern Recognit. | 3 |
| 2021 | Exploring a unified low rank representation for multi-focus image fusion
Qiang Zhang 0020, Yongjiang Luo, Jungong Han |
Pattern Recognit. | 1 |
| 2021 | Automatic pancreas segmentation based on lightweight DCNN modules and spatial prior propagation
Dingwen Zhang, Qiang Zhang 0020, Jungong Han, Shu Zhang 0001, Junwei Han 0001 |
Pattern Recognit. | 3 |
| 2021 | Revisiting Feature Fusion for RGB-T Salient Object DetectionabstractWhile many RGB-based saliency detection algorithms have recently shown the capability of segmenting salient objects from an image, they still suffer from unsatisfactory performance when dealing with complex scenarios, insufficient illumination or occluded appearances. To overcome this problem, this article studies RGB-T saliency detection, where we take advantage of thermal modality's robustness against illumination and occlusion. To achieve this goal, we revisit feature fusion for mining intrinsic RGB-T saliency patterns and propose a novel deep feature fusion network, which consists of the multi-scale, multi-modality, and multi-level feature fusion modules. Specifically, the multi-scale feature fusion module captures rich contexture features from each modality feature, while the multi-modality and multi-level feature fusion modules integrate complementary features from different modality features and different level of features, respectively. To demonstrate the effectiveness of the proposed approach, we conduct comprehensive experiments on the RGB-T saliency detection benchmark. The experimental results demonstrate that our approach outperforms other state-of-the-art methods and the conventional feature fusion modules by a large margin. Qiang Zhang 0020, Tonglin Xiao, Nianchang Huang, Dingwen Zhang, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Integrating Part-Object Relationship and Contrast for Camouflaged Object DetectionabstractObject detectors that solely rely on image contrast are struggling to detect camouflaged objects in images because of the high similarity between camouflaged objects and their surroundings. To address this issue, in this paper, we investigate the role of the part-object relationship for camouflaged object detection. Specifically, we propose a Part-Object relationship and Contrast Integrated Network (POCINet) covering both search and identification stages, where each stage adopts an appropriate scheme to engage the contrast information and part-object relational knowledge for camouflaged pattern decoding. Besides, we bridge these two stages via a Search-to-Identification Guidance (SIG) module, in which the search result, as well as decoded semantic knowledge, jointly enhances the features encoding ability of the identification stage. Experimental results demonstrate the superiority of our algorithm on three datasets. Notably, our algorithm raises Fβ of the best existing method by approximately 17 points on the CPD1K dataset. The source code will be released soon. Yi Liu 0038, Dingwen Zhang, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Joint Cross-Modal and Unimodal Features for RGB-D Salient Object DetectionabstractRGB-D salient object detection is one of the basic tasks in computer vision. Most existing models focus on investigating efficient ways of fusing the complementary information from RGB and depth images for better saliency detection. However, for many real-life cases, where one of the input images has poor visual quality or contains affluent saliency cues, fusing cross-modal features does not help to improve the detection accuracy, when compared to using unimodal features only. In view of this, a novel RGB-D salient object detection model is proposed by simultaneously exploiting the cross-modal features from the RGB-D images and the unimodal features from the input RGB and depth images for saliency detection. To this end, a Multi-branch Feature Fusion Module is presented to effectively capture the cross-level and cross-modal complementary information between RGB-D images, as well as the cross-level unimodal features from the RGB images and the depth images separately. On top of that, a Feature Selection Module is designed to adaptively select those highly discriminative features for the final saliency prediction from the fused cross-modal features and the unimodal features. Extensive evaluations on four benchmark datasets demonstrate that the proposed model outperforms the state-of-the-art approaches by a large margin. Nianchang Huang, Yi Liu 0038, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Multim. | 3 |
| 2020 | Multi-focus image fusion based on non-negative sparse representation and patch-level consistency rectification
Qiang Zhang 0020, Guanghe Li, Jungong Han |
Pattern Recognit. | 1 |
| 2020 | A structure-aware splitting framework for separating cell clumps in biomedical images
Qiang Zhang 0020, Zaihao Liu, Dingwen Zhang |
Signal Process. | 1 |
| 2020 | Deep Salient Object Detection With Contextual Information GuidanceabstractIntegration of multi-level contextual information, such as feature maps and side outputs, is crucial for Convolutional Neural Networks (CNNs) based salient object detection. However, most existing methods either simply concatenate multi-level feature maps or calculate element-wise addition of multi-level side outputs, thus failing to take full advantages of them. In this work, we propose a new strategy for guiding multi-level contextual information integration, where feature maps and side outputs across layers are fully engaged. Specifically, shallower-level feature maps are guided by the deeper-level side outputs to learn more accurate properties of the salient object. In turn, the deeper-level side outputs can be propagated to high-resolution versions with spatial details complemented by means of shallower-level feature maps. Moreover, a group convolution module is proposed with the aim to achieve high-discriminative feature maps, in which the backbone feature maps are divided into a number of groups and then the convolution is applied to the channels of backbone feature maps within each group. Eventually, the group convolution module is incorporated in the guidance module to further promote the guidance role. Experiments on three public benchmark datasets verify the effectiveness and superiority of the proposed method over the state-of-the-art methods. Yi Liu 0038, Jungong Han, Qiang Zhang 0020, Caifeng Shan |
IEEE Trans. Image Process. | 3 |
| 2020 | RGB-T Salient Object Detection via Fusing Multi-Level CNN FeaturesabstractRGB-induced salient object detection has recently witnessed substantial progress, which is attributed to the superior feature learning capability of deep convolutional neural networks (CNNs). However, such detections suffer from challenging scenarios characterized by cluttered backgrounds, low-light conditions and variations in illumination. Instead of improving RGB based saliency detection, this paper takes advantage of the complementary benefits of RGB and thermal infrared images. Specifically, we propose a novel end-to-end network for multi-modal salient object detection, which turns the challenge of RGB-T saliency detection to a CNN feature fusion problem. To this end, a backbone network (e.g., VGG-16) is first adopted to extract the coarse features from each RGB or thermal infrared image individually, and then several adjacent-depth feature combination (ADFC) modules are designed to extract multi-level refined features for each single-modal input image, considering that features captured at different depths differ in semantic information and visual details. Subsequently, a multi-branch group fusion (MGF) module is employed to capture the cross-modal features by fusing those features from ADFC modules for a RGB-T image pair at each level. Finally, a joint attention guided bi-directional message passing (JABMP) module undertakes the task of saliency prediction via integrating the multi-level fused features from MGF modules. Experimental results on several public RGB-T salient object detection datasets demonstrate the superiorities of our proposed algorithm over the state-of-the-art approaches, especially under challenging conditions, such as poor illumination, complex background and low contrast. Qiang Zhang 0020, Nianchang Huang, Dingwen Zhang, Caifeng Shan, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2020 | Exploring Task Structure for Brain Tumor Segmentation From Multi-Modality MR ImagesabstractBrain tumor segmentation, which aims at segmenting the whole tumor area, enhancing tumor core area, and tumor core area from each input multi-modality bioimaging data, has received considerable attention from both academia and industry. However, the existing approaches usually treat this problem as a common semantic segmentation task without taking into account the underlying rules in clinical practice. In reality, physicians tend to discover different tumor areas by weighing different modality volume data. Also, they initially segment the most distinct tumor area, and then gradually search around to find the other two. We refer to the first property as the task-modality structure while the second property as the task-task structure, based on which we propose a novel task-structured brain tumor segmentation network (TSBTS net). Specifically, to explore the task-modality structure, we design a modality-aware feature embedding mechanism to infer the important weights of the modality data during network learning. To explore the tasktask structure, we formulate the prediction of the different tumor areas as conditional dependency sub-tasks and encode such dependency in the network stream. Experiments on BraTS benchmarks show that the proposed method achieves superior performance in segmenting the desired brain tumor areas while requiring relatively lower computational costs, compared to other state-of-the-art methods and baseline models. Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Wang 0001, Yizhou Yu |
IEEE Trans. Image Process. | 3 |
| 2019 | Employing Deep Part-Object Relationships for Salient Object DetectionabstractDespite Convolutional Neural Networks (CNNs) based methods have been successful in detecting salient objects, their underlying mechanism that decides the salient intensity of each image part separately cannot avoid inconsistency of parts within the same salient object. This would ultimately result in an incomplete shape of the detected salient object. To solve this problem, we dig into part-object relationships and take the unprecedented attempt to employ these relationships endowed by the Capsule Network (CapsNet) for salient object detection. The entire salient object detection system is built directly on a Two-Stream Part-Object Assignment Network (TSPOANet) consisting of three algorithmic steps. In the first step, the learned deep feature maps of the input image are transformed to a group of primary capsules. In the second step, we feed the primary capsules into two identical streams, within each of which low-level capsules (parts) will be assigned to their familiar high-level capsules (object) via a locally connected routing. In the final step, the two streams are integrated in the form of a fully connected layer, where the relevant parts can be clustered together to form a complete salient object. Experimental results demonstrate the superiority of the proposed salient object detection network over the state-of-the-art methods. Yi Liu 0038, Qiang Zhang 0020, Dingwen Zhang, Jungong Han |
ICCV | 2 |
| 2019 | Video Synchronization Based on Projective-Invariant DescriptorabstractIn this paper, we present a novel trajectory-based method to synchronize two videos shooting the same dynamic scene, which are recorded by stationary un-calibrated cameras from different viewpoints. The core algorithm is carried out in two steps: projective-invariant descriptor construction and trajectory points matching. In the first step, a new five-coplanar-points structure is proposed to compute the cross ratio during the construction of the projective-invariant descriptor. The five points include one trajectory point and four fixed points induced from the background scene, which are co - planar in the 3D coordinate. In the second step, the matched trajectory points are initially estimated by the primitive nearest neighbor method, and are further refined by using epipolar geometric constraints and post processing. Experimental results demonstrate that the proposed method significantly outperforms the existing state-of-the-arts. More importantly, the proposed method is more generic in the sense that it works well for those videos captured under different conditions, including different frame rates, wide baseline, multiple moving objects, planar or non-planar motion trajectories. Qiang Zhang 0020, Jungong Han |
Neural Process. Lett. | 1 |
| 2019 | Salient object detection employing a local tree-structured low-rank representation and foreground consistency
Qiang Zhang 0020, Zhen Huo, Yi Liu 0038, Yunhui Pan, Caifeng Shan, Jungong Han |
Pattern Recognit. | 1 |
| 2019 | Salient Object Detection via Two-Stage GraphsabstractDespite recent advances made in salient object detection using graph theory, the approach still suffers from accuracy problems when the image is characterized by a complex structure, either in the foreground or background, causing erroneous saliency segmentation. This fundamental challenge is mainly attributed to the fact that most existing graph-based methods take only the adjacently spatial consistency among graph nodes into consideration. In this paper, we tackle this issue from a coarse-to-fine perspective and propose a two-stage-graphs approach for salient object detection, in which two graphs having the same nodes but different edges are employed. Specifically, a weighted joint robust sparse representation model, rather than the commonly used manifold ranking model, helps to compute the saliency value of each node in the first-stage graph, thereby providing a saliency map at the coarse level. In the second-stage graph, along with the adjacently spatial consistency, a new regionally spatial consistency among graph nodes is considered in order to refine the coarse saliency map, assuring uniform saliency assignment even in complex scenes. Particularly, the second stage is generic enough to be integrated in existing salient object detectors, enabling improved performance. Experimental results on benchmark data sets validate the effectiveness and superiority of the proposed scheme over related state-of-the-art methods. Yi Liu 0038, Jungong Han, Qiang Zhang 0020, Long Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Salient object detection employing robust sparse representation and local consistency
Liu Yi, Qiang Zhang 0020, Jungong Han, Long Wang 0001 |
Image Vis. Comput. | 2 |
| 2018 | Robust sparse representation based multi-focus image fusion with dictionary construction and local spatial consistency
Qiang Zhang 0020, Rick S. Blum, Jungong Han |
Pattern Recognit. | 1 |
| 2017 | Salient object detection based on super-pixel clustering and unified low-rank representation
Qiang Zhang 0020, Yi Liu 0038, Siyang Zhu, Jungong Han |
Comput. Vis. Image Underst. | 1 |
| 2016 | Matching of images with projective distortion using transform invariant low-rank textures
Qiang Zhang 0020, Rick S. Blum |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Robust Multi-Focus Image Fusion Using Multi-Task Sparse Representation and Spatial ContextabstractWe present a novel fusion method based on a multi-task robust sparse representation (MRSR) model and spatial context information to address the fusion of multi-focus gray-level images with misregistration. First, we present a robust sparse representation (RSR) model by replacing the conventional least-squared reconstruction error by a sparse reconstruction error. We then propose a multi-task version of the RSR model, viz., the MRSR model. The latter is then applied to multi-focus image fusion by employing the detailed information regarding each image patch and its spatial neighbors to collaboratively determine both the focused and defocused regions in the input images. To achieve this, we formulate the problem of extracting details from multiple image patches as a joint multi-task sparsity pursuit based on the MRSR model. Experimental results demonstrate that the suggested algorithm is competitive with the current state-of-the-art and superior to some approaches that use traditional sparse representation methods when input images are misregistered. Qiang Zhang 0020, Martin D. Levine |
IEEE Trans. Image Process. | 1 |
| 2015 | Registration of images with affine geometric distortion based on Maximally Stable Extremal Regions and phase congruency
Qiang Zhang 0020, Long Wang 0001 |
Image Vis. Comput. | 1 |
| 2014 | Video fusion performance assessment based on spatial-temporal phase congruency
Qiang Zhang 0020, Sheng Hua, Rick S. Blum, Minli Chen |
Signal Process. | 1 |
| 2013 | Multimodality image fusion by using both phase and magnitude information
Qiang Zhang 0020, Zhaokun Ma, Long Wang 0001 |
Pattern Recognit. Lett. | 1 |
| 2013 | Multisensor video fusion based on spatial-temporal salience detection
Qiang Zhang 0020, Yueling Chen, Long Wang 0001 |
Signal Process. | 1 |
| 2012 | Video fusion performance evaluation based on structural similarity and human visual perception
Qiang Zhang 0020, Long Wang 0001, Zhaokun Ma |
Signal Process. | 1 |
| 2011 | Similarity-based multimodality image fusion with shiftable complex directional pyramid
Qiang Zhang 0020, Long Wang 0001, Zhaokun Ma |
Pattern Recognit. Lett. | 1 |