VLDB 2026 Research / reviewers in the wild / expert
Zhengzheng Tu
dblp:138/5016
· DBLP profile ↗
36ranked-venue papers
15as first author
29since 2021 · last 2026
0000-0002-9689-8657ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 10 first-author · 20 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ImageBind Guided Progressive Transformation Network for Alignment-free RGBT Video Object Detection
Zhengzheng Tu, Chuanwang Guo, Qishun Wang, Chenglong Li 0002, Jin Tang 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation NetworkabstractAlignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecting and annotating image pairs limits the scale of existing benchmarks, hindering the advancement of alignment-free RGB-T SOD. In this paper, we construct a large-scale and high-diversity unaligned RGB-T SOD dataset named UVT20K, comprising 20,000 image pairs, 407 scenes, and 1256 object categories. All samples are collected from real-world scenarios with various challenges, such as low illumination, image clutter, complex salient objects, and so on. To support the exploration for further research, each sample in UVT20K is annotated with a comprehensive set of ground truths, including saliency masks, scribbles, boundaries, and challenge attributes. In addition, we propose a Progressive Correlation Network (PCNet), which models inter- and intra-modal correlations on the basis of explicit alignment to achieve accurate predictions in unaligned image pairs. Extensive experiments conducted on two unaligned three weakly aligned three aligned datasets demonstrate the effectiveness of our method. Kunpeng Wang 0005, Keke Chen, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001 |
AAAI | 4 |
| 2025 | Spatial-Temporal Memory Filtering SAM for Lesion Segmentation in Breast Ultrasound Videos
Zhengzheng Tu, Liang Zong, Bo Jiang 0002, Chaoxue Zhang |
MICCAI (2) | 1 |
| 2025 | Erasure-based interaction network for red-green-blue and thermal object detection and a unified benchmark
Qishun Wang, Zhengzheng Tu, Chenglong Li 0002, Hongshun Wang, Kunpeng Wang 0005 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Challenge-aware U-net for breast lesion segmentation in ultrasound images
Dengdi Sun, Changxu Dong, Bo Jiang 0002, Yayang Duan, Zhengzheng Tu, Chaoxue Zhang |
Pattern Recognit. | 6 |
| 2025 | Unified-Modal Salient Object Detection via Adaptive Prompt LearningabstractExisting single-modal and multi-modal salient object detection (SOD) methods focus on designing specific architectures tailored for their respective tasks. However, developing completely different models for different tasks leads to labor and time consumption, as well as high computational and practical deployment costs. In this paper, we attempt to address both single-modal and multi-modal SOD in a unified framework called UniSOD, which fully exploits the overlapping prior knowledge between different tasks. Nevertheless, assigning appropriate strategies to modality variable inputs is challenging. To this end, UniSOD learns modality-aware prompts with task-specific hints through adaptive prompt learning, which are seamlessly plugged into the proposed pre-trained baseline SOD model to handle corresponding tasks, while only requiring few learnable parameters compared to training the entire model from scratch. In particular, each modality-aware prompt is solely generated from a homogeneous switchable prompt generation (SPG) block, which adaptively performs structural switching based on single-modal and multi-modal inputs without manual intervention, ensuring that the framework can effectively handle diverse input cases (e.g., RGB-only, RGB-D, RGB-T) with a unified approach. Through end-to-end joint training, UniSOD achieves ovrall competitive performance on 14 benchmark datasets, demonstrating its ability to efficiently unify single-modal and multi-modal SOD tasks. Code has been available athttps://github.com/Angknpng/UniSOD Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Zhengyi Liu, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DiffRSD: Diffusion-Based and Integrity-Aware RGB-D Rail Surface Defect Inspection
Zhengyi Liu, Junnan Zhou, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Unity Is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-IdentificationabstractPerson Re-identification (ReID) aims to retrieve the specific person across non-overlapping cameras, which greatly helps intelligent transportation systems. As we all know, Convolutional Neural Networks (CNNs) and Transformers have the unique strengths to extract local and global features, respectively. Considering this fact, we focus on the mutual fusion between them to learn more comprehensive representations for persons. In particular, we utilize the complementary integration of deep features from different model structures. We propose a novel fusion framework called FusionReID to unify the strengths of CNNs and Transformers for image-based person ReID. More specifically, we first deploy a Dual-branch Feature Extraction (DFE) to extract features through CNNs and Transformers from a single image. Moreover, we design a novel Dual-attention Mutual Fusion (DMF) to achieve sufficient feature fusions. The DMF comprises Local Refinement Units (LRU) and Heterogenous Transmission Modules (HTM). LRU utilizes depth-separable convolutions to align deep features in channel dimensions and spatial sizes. HTM consists of a Shared Encoding Unit (SEU) and two Mutual Fusion Units (MFU). Through the continuous stacking of HTM, deep features after LRU are repeatedly utilized to generate more discriminative features. Extensive experiments on three public ReID benchmarks demonstrate that our method can attain superior performances than most state-of-the-arts. The source code is available athttps://github.com/924973292/FusionReID. Xuehu Liu, Zhengzheng Tu, Huchuan Lu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | A Spatial-Temporal Progressive Fusion Network for Breast Lesion Segmentation in Ultrasound VideosabstractUltrasound video-based breast lesion segmentation provides valuable assistance in early breast lesion detection and discrimination. However, this field faces two key challenges: the first is how to simultaneously utilize both intra-frame and inter-frame lesion cues to accurately segment breast lesions, and the second is that the availability of breast ultrasound video datasets is quite limited. In this paper, we propose a novel Spatial-Temporal Progressive Fusion Network (STPFNet) for video-based breast lesion segmentation problem. The proposed STPFNet comprises three main components. First, we propose to adopt a unified network architecture to capture spatial dependencies within each ultrasound frame and temporal correlations between different frames together for feature representation of ultrasound video. Second, we propose a new fusion module called Multi-Granularity Feature Fusion (MGFF) to fuse the extracted information with different granularities for lesion segmentation. MGFF can help improve the issue of lesion boundary blurring. Third, we propose to take the segmentation result of the previous frame as prior knowledge to suppress the noisy background and learn a more robust representation. To further promote the research in this field, we construct a new ultrasound video breast lesion segmentation dataset, called UVBLS200, comprising 200 videos (80 benign and 120 malignant lesions). Experiments on the proposed dataset demonstrate that the proposed STPFNet achieves a better breast lesion detection performance than state-of-the-art methods. Zhengzheng Tu, Zigang Zhu, Yayang Duan, Bo Jiang 0002, Qishun Wang, Chaoxue Zhang |
IEEE Trans. Multim. | 1 |
| 2024 | TOP-ReID: Multi-Spectral Object Re-identification with Token PermutationabstractMulti-spectral object Re-identification (ReID) aims to retrieve specific objects by leveraging complementary information from different image spectra. It delivers great advantages over traditional single-spectral ReID in complex visual environment. However, the significant distribution gap among different image spectra poses great challenges for effective multi-spectral feature representations. In addition, most of current Transformer-based ReID methods only utilize the global feature of class tokens to achieve the holistic retrieval, ignoring the local discriminative ones. To address the above issues, we step further to utilize all the tokens of Transformers and propose a cyclic token permutation framework for multi-spectral object ReID, dubbled TOP-ReID. More specifically, we first deploy a multi-stream deep network based on vision Transformers to preserve distinct information from different image spectra. Then, we propose a Token Permutation Module (TPM) for cyclic multi-spectral feature aggregation. It not only facilitates the spatial feature alignment across different image spectra, but also allows the class token of each spectrum to perceive the local details of other spectra. Meanwhile, we propose a Complementary Reconstruction Module (CRM), which introduces dense token-level reconstruction constraints to reduce the distribution gap across different image spectra. With the above modules, our proposed framework can generate more discriminative multi-spectral features for robust object ReID. Extensive experiments on three ReID benchmarks (i.e., RGBNT201, RGBNT100 and MSVR310) verify the effectiveness of our methods. The code is available at https://github.com/924973292/TOP-ReID. Xuehu Liu, Hu Lu, Zhengzheng Tu, Huchuan Lu |
AAAI | 5 |
| 2024 | Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationabstractSingle-modal object re-identification (ReID) faces great challenges in maintaining robustness within complex visual scenarios. In contrast, multi-modal object ReID utilizes complementary information from diverse modalities, showing great potentials for practical applications. How-ever, previous methods may be easily affected by irrele-vant backgrounds and usually ignore the modality gaps. To address above issues, we propose a novel learning frame-work named EDITOR to select diverse tokens from vision Transformers for multi-modal object ReID. We be-gin with a shared vision Transformer to extract tokenized features from different input modalities. Then, we intro-duce a Spatial-Frequency Token Selection (SFTS) module to adaptively select object-centric tokens with both spa-tial and frequency information. Afterwards, we employ a Hierarchical Masked Aggregation (HMA) module to fa-cilitate feature interactions within and across modalities. Finally, to further reduce the effect of backgrounds, we propose a Background Consistency Constraint (BCC) and an Object-Centric Feature Refinement (OCFR). They are formulated as two new loss functions, which improve the feature discrimination with background suppression. As a result, our framework can generate more discriminative features for multi-modal object ReID. Extensive ex-periments on three multi-modal ReID benchmarks verify the effectiveness of our methods. The code is available at https://github.com/924973292/EDITOR. Zhengzheng Tu, Huchuan Lu |
CVPR | 4 |
| 2024 | MGDR: Multi-modal Graph Disentangled Representation for Brain Disease Prediction
Bo Jiang 0002, Xixi Wan, Yuan Chen 0012, Zhengzheng Tu, Yumiao Zhao, Jin Tang 0001 |
MICCAI (2) | 5 |
| 2024 | ACENet: Adaptive Context Enhancement Network for RGB-T Video Object Detection
Zhengzheng Tu, Le Gu, Danying Lin |
PRCV (8) | 1 |
| 2024 | Bidirectional Alternating Fusion Network for RGB-T Salient Object Detection
Zhengzheng Tu, Danying Lin, Le Gu, Sulan Zhai |
PRCV (8) | 1 |
| 2024 | LFSamba: Marry SAM With Mamba for Light Field Salient Object DetectionabstractA light field camera can reconstruct 3D scenes using captured multi-focus images that contain rich spatial geometric information, enhancing applications in stereoscopic photography, virtual reality, and robotic vision. In this work, a state-of-the-art salient object detection model for multi-focus light field images, called LFSamba, is introduced to emphasize four main insights: (a) Efficient feature extraction, where SAM is used to extract modality-aware discriminative features; (b) Inter-slice relation modeling, leveraging Mamba to capture long-range dependencies across multiple focal slices, thus extracting implicit depth cues; (c) Inter-modal relation modeling, utilizing Mamba to integrate all-focus and multi-focus images, enabling mutual enhancement; (d) Weakly supervised learning capability, developing a scribble annotation dataset from an existing pixel-level mask dataset, establishing the first scribble-supervised baseline for light field salient object detection. Zhengyi Liu, Longzhen Wang, Xianyong Fang, Zhengzheng Tu, Linbo Wang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | AST-GCN: Augmented Spatial Temporal Graph Convolutional Neural Network for Gait Emotion RecognitionabstractSkeleton-based methods have recently achieved good performance in deep learning-based gait emotion recognition (DL-GER). However, the current methods have two drawbacks that limit the ability to learn discriminative emotional features from gait. First, these methods do not exclude the effect of the subject’s walking orientation on emotion classification. Second, they do not sufficiently learn the implicit connections between the joints during human walking. In this paper, an augmented spatial-temporal graph convolutional neural network (AST-GCN) is introduced to solve these two problems. The interframe shift encoding (ISE) module acquires interframe shifts of joints to make the network sensitive to changes in emotion-related joint movements regardless of the subject’s walking orientation. A multichannel implicit connection inference method learns more implicit connection relations related to emotions. Notably, we unify current skeleton-based methods into a common framework that validates the most powerful feature representation capability of our AST-GCN from a theoretical perspective. In addition, we extend the skeleton-based gait dataset using posture estimation software. Experiments demonstrate that our AST-GCN outperforms state-of-the-art methods on three datasets on two tasks. Xiao Sun 0003, Zhengzheng Tu, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Learning Adaptive Fusion Bank for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods generally design different fusion schemes to handle certain issues or challenges. Although these fusion schemes are effective at addressing specific issues or challenges, they may struggle to handle multiple complex challenges simultaneously. To solve this problem, we propose a novel adaptive fusion bank that makes full use of the complementary benefits from a set of basic fusion schemes to handle different challenges simultaneously for robust MSOD. We focus on handling five major challenges in MSOD, namely center bias, scale variation, image clutter, low illumination, and thermal crossover or depth ambiguity. The fusion bank proposed consists of five representative fusion schemes, which are specifically designed based on the characteristics of each challenge, respectively. The bank is scalable, and more fusion schemes could be incorporated into the bank for more challenges. To adaptively select the appropriate fusion scheme for multi-modal input, we introduce an adaptive ensemble module that forms the adaptive fusion bank, which is embedded into hierarchical layers for sufficient fusion of different source data. Moreover, we design an indirect interactive guidance module to accurately detect salient hollow objects via the skip integration of high-level semantic information and low-level spatial details. Extensive experiments on three RGBT datasets and seven RGBD datasets demonstrate that the proposed method achieves the outstanding performance compared to the state-of-the-art methods. Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Cheng Zhang 0010, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Alignment-Free RGBT Salient Object Detection: Semantics-Guided Asymmetric Correlation Network and a Unified BenchmarkabstractRGB and Thermal (RGBT) Salient Object Detection (SOD) aims to achieve high-quality saliency prediction by exploiting the complementary information of visible and thermal image pairs, which are initially captured in an unaligned manner. However, existing methods are tailored for manually aligned image pairs, which are labor-intensive, and directly applying these methods to original unaligned image pairs could significantly degrade their performance. In this paper, we make the first attempt to address RGBT SOD for initially captured RGB and thermal image pairs without manual alignment. Specifically, we propose a Semantics-guided Asymmetric Correlation Network (SACNet) that consists of two novel components: 1) an asymmetric correlation module utilizing semantics-guided attention to model cross-modal correlations specific to unaligned salient regions; 2) an associated feature sampling module to sample relevant thermal features according to the corresponding RGB features for multi-modal feature integration. In addition, we construct a unified benchmark dataset called UVT2000, containing 2000 RGB and thermal image pairs directly captured from various real-world scenes without any alignment, to facilitate research on alignment-free RGBT SOD. Extensive experiments on both aligned and unaligned datasets demonstrate the effectiveness and superior performance of our method. Kunpeng Wang 0005, Danying Lin, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Multimodal salient object detection via adversarial learning with collaborative generator
Zhengzheng Tu, Wenfang Yang, Kunpeng Wang 0005, Amir Hussain 0001, Bin Luo 0001, Chenglong Li 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Centralized sub-critic based hierarchical-structured reinforcement learning for temporal sentence grounding
Yingyuan Zhao, Zhiyi Tan 0002, Bing-Kun Bao, Zhengzheng Tu |
Multim. Syst. | 4 |
| 2023 | RGBT Salient Object Detection: A Large-Scale Dataset and BenchmarkabstractSalient object detection in complex scenes and environments is a challenging research topic. Most works focus on RGB-based salient object detection, which limits its performance of real-life applications when confronted with adverse conditions such as dark environments and complex backgrounds. Taking advantage of RGB and thermal infrared(RGBT) images becomes a new research direction for detecting salient objects in complex scenes, since the thermal infrared spectrum provides the complementary information and has been used in many computer vision tasks. However, current research for RGBT salient object detection is limited by the lack of a large-scale dataset and comprehensive benchmark. This work contributes such a RGBT image dataset named VT5000, including 5000 spatially aligned RGBT image pairs with ground truth annotations. VT5000 has 11 challenges collected in different scenes and environments for exploring the robustness of algorithms. With this dataset, we propose a powerful baseline approach, which extracts multilevel features of each modality and aggregates these features of all modalities with the attention mechanism for accurate RGBT salient object detection. To further solve the problem of blur boundaries of salient objects, we also use an edge loss to refine the boundaries. Extensive experiments show that the proposed baseline approach outperforms the state-of-the-art methods on VT5000 dataset and other two public datasets. In addition, we carry out a comprehensive analysis of different algorithms of RGBT salient object detection on VT5000 dataset, and then make several valuable conclusions and provide some potential research directions for RGBT salient object detection. Our new VT5000 dataset is made publicly available at https://github.com/lz118/RGBT-Salient-Object-Detection. Zhengzheng Tu, Chenglong Li 0002, Jieming Xu |
IEEE Trans. Multim. | 1 |
| 2022 | RGBT tracking via reliable feature configuration
Zhengzheng Tu, Wenli Pan, Yunsheng Duan, Jin Tang 0001, Chenglong Li 0002 |
Sci. China Inf. Sci. | 1 |
| 2022 | ORSI Salient Object Detection via Multiscale Joint Region and Boundary ModelabstractSalient object detection (SOD) in optical remote sense images (ORSIs) is a valuable and challenging task. The factors in ORSI, such as background clutter, lighting shadows, imaging blur, and low resolution, significantly degrade the completeness and accuracy of salient objects. To handle this problem, we propose a novel model to learn robust multiscale region features of salient objects by simultaneously optimizing their boundaries. First, we extract multiscale region features of salient objects through a hierarchical attention module. Second, we generate the boundary features by combining the local cues and the global information generated by pyramid pooling. Finally, we embed the boundary features into region features at multiple scales. In particular, we design a joint learning scheme based on a bidirectional feature transformation to optimize boundary and region features simultaneously for accurate ORSI SOD. To provide a comprehensive evaluation platform, we construct a new dataset called ORSI-4199 for ORSI SOD. It contains 4199 finely annotated image pairs with diverse scenes, in which nine attributes (i.e., challenge types) are annotated to facilitate analyzing the strengths and weaknesses of SOD models from different perspectives. Extensive experiments on the public dataset ORSSD, EORRSD, and the newly created dataset ORSI-4199 show that the proposed approach achieves promising results against state-of-the-art methods.https://github.com/wchao1213/ORSI-SOD. Zhengzheng Tu, Chenglong Li 0002, Minghao Fan, Haifeng Zhao 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Weakly Alignment-Free RGBT Salient Object Detection With Deep Correlation NetworkabstractRGBT Salient Object Detection (SOD) focuses on common salient regions of a pair of visible and thermal infrared images. Existing methods perform on the well-aligned RGBT image pairs, but the captured image pairs are always unaligned and aligning them requires much labor cost. To handle this problem, we propose a novel deep correlation network (DCNet), which explores the correlations across RGB and thermal modalities, for weakly alignment-free RGBT SOD. In particular, DCNet includes a modality alignment module based on the spatial affine transformation, the feature-wise affine transformation and the dynamic convolution to model the strong correlation of two modalities. Moreover, we propose a novel bi-directional decoder model, which combines the coarse-to-fine and fine-to-coarse processes for better feature enhancement. In particular, we design a modality correlation ConvLSTM by adding the first two components of modality alignment module and a global context reinforcement module into ConvLSTM, which is used to decode hierarchical features in both top-down and button-up manners. Extensive experiments on three public benchmark datasets show the remarkable performance of our method against state-of-the-art methods. Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | M5L: Multi-Modal Multi-Margin Metric Learning for RGBT TrackingabstractClassifying hard samples in the course of RGBT tracking is a quite challenging problem. Existing methods only focus on enlarging the boundary between positive and negative samples, but ignore the relations of multilevel hard samples, which are crucial for the robustness of hard sample classification. To handle this problem, we propose a novel Multi-Modal Multi-Margin Metric Learning framework named M5L for RGBT tracking. In particular, we divided all samples into four parts including normal positive, normal negative, hard positive and hard negative ones, and aim to leverage their relations to improve the robustness of feature embeddings, e.g., normal positive samples are closer to the ground truth than hard positive ones. To this end, we design a multi-modal multi-margin structural loss to preserve the relations of multilevel hard samples in the training stage. In addition, we introduce an attention-based fusion module to achieve quality-aware integration of different source data. Extensive experiments on large-scale datasets testify that our framework clearly improves the tracking performance and performs favorably the state-of-the-art RGBT trackers. Zhengzheng Tu, Chun Lin, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding NetworkabstractSalient object detection is the pixel-level dense prediction task which can highlight the prominent object in the scene. Recently U-Net framework is widely used, and continuous convolution and pooling operations generate multi-level features which are complementary with each other. In view of the more contribution of high-level features for the performance, we propose a triplet transformer embedding module to enhance them by learning long-range dependencies across layers. It is the first to use three transformer encoders with shared weights to enhance multi-level features. By further designing scale adjustment module to process the input, devising three-stream decoder to process the output and attaching depth features to color features for the multi-modal fusion, the proposed triplet transformer embedding network (TriTransNet) achieves the state-of-the-art performance in RGB-D salient object detection, and pushes the performance to a new level. Experimental results demonstrate the effectiveness of the proposed modules and the competition of TriTransNet. Zhengyi Liu, Zhengzheng Tu, Yun Xiao 0003, Bin Tang 0003 |
ACM Multimedia | 3 |
| 2021 | A novel domain activation mapping-guided network (DA-GNT) for visual tracking
Zhengzheng Tu, Ajian Zhou, Chuang Gan 0003, Bo Jiang 0002, Amir Hussain 0001, Bin Luo 0001 |
Neurocomputing | 1 |
| 2021 | Edge-Guided Non-Local Fully Convolutional Network for Salient Object DetectionabstractFully Convolutional Neural Network (FCN) has been widely applied to salient object detection recently by virtue of high-level semantic feature extraction, but existing FCN-based methods still suffer from continuous striding and pooling operations leading to loss of spatial structure and blurred edges. To maintain the clear edge structure of salient objects, we propose a novel Edge-guided Non-local FCN (ENFNet) to perform edge-guided feature learning for accurate salient object detection. In a specific, we extract hierarchical global and local information in FCN to incorporate non-local features for effective feature representations. To preserve good boundaries of salient objects, we propose a guidance block to embed edge prior knowledge into hierarchical feature maps. The guidance block not only performs feature-wise manipulation but also spatial-wise transformation for effective edge embeddings. Our model is trained on the MSRA-B dataset and tested on five popular benchmark datasets. Comparing with the state-of-the-art methods, the proposed method performance well on five datasets. Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Multi-Interactive Dual-Decoder for RGB-Thermal Salient Object DetectionabstractRGB-thermal salient object detection (SOD) aims to segment the common prominent regions of visible image and corresponding thermal infrared image that we call it RGBT SOD. Existing methods don't fully explore and exploit the potentials of complementarity of different modalities and multi-type cues of image contents, which play a vital role in achieving accurate results. In this paper, we propose a multi-interactive dual-decoder to mine and model the multi-type interactions for accurate RGBT SOD. In specific, we first encode two modalities into multi-level multi-modal feature representations. Then, we design a novel dual-decoder to conduct the interactions of multi-level features, two modalities and global contexts. With these interactions, our method works well in diversely challenging scenarios even in the presence of invalid modality. Finally, we carry out extensive experiments on public RGBT and RGBD SOD datasets, and the results show that the proposed method achieves the outstanding performance against state-of-the-art algorithms. The source code has been released at: https://github.com/lz118/Multi-interactive-Dual-decoder. Zhengzheng Tu, Chenglong Li 0002, Yang Lang, Jin Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Investigating the Visual Lombard Effect with Gabor Based Features
Waito Chiu, Andrew Abel, Chun Lin, Zhengzheng Tu |
INTERSPEECH | 5 |
| 2020 | RGBT Salient Object Detection: Benchmark and A Novel Cooperative Ranking ApproachabstractDespite significant progress, image saliency detection still remains a challenging task in complex scenes and environments. Integrating multiple different but complementary cues, like RGB and Thermal infrared (RGBT), may be an effective way for boosting saliency detection performance. This work contributes a RGBT image dataset, which includes 821 spatially aligned RGBT image pairs and their ground truth annotations for saliency detection purpose. Moreover, 11 challenges are annotated on these image pairs for performing the challenge-sensitive analysis and 3 kinds of baseline methods are implemented to provide a comprehensive comparison platform. With this benchmark, we propose a novel approach based on a cooperative ranking algorithm for RGBT saliency detection. In particular, we introduce a weight for each modality to describe the reliability and a ℓ1-based cross-modal consistency in a unified ranking model, and design an efficient solver to iteratively optimize several subproblems with closed-form solutions. Extensive experiments against baseline methods demonstrate the effectiveness of the proposed approach on both our introduced dataset and a public dataset. Jin Tang 0001, Dongzhe Fan, Xiaoxiao Wang 0003, Zhengzheng Tu, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | RGB-T Image Saliency Detection via Collaborative Graph LearningabstractImage saliency detection is an active research topic in the community of computer vision and multimedia. Fusing complementary RGB and thermal infrared data has been proven to be effective for image saliency detection. In this paper, we propose an effective approach for RGB-T image saliency detection. Our approach relies on a novel collaborative graph learning algorithm. In particular, we take superpixels as graph nodes, and collaboratively use hierarchical deep features to jointly learn graph affinity and node saliency in a unified optimization framework. Moreover, we contribute a more challenging dataset for the purpose of RGB-T image saliency detection, which contains 1000 spatially aligned RGB-T image pairs and their ground truth annotations. Extensive experiments on the public dataset and the newly created dataset suggest that the proposed approach performs favorably against the state-of-the-art RGB-T saliency detection methods. Zhengzheng Tu, Chenglong Li 0002, Xiaoxiao Wang 0003, Jin Tang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | A Novel Method for Thermal Image Based Electrical-Equipment Detection
Futian Wang, Songjian Hua, Xiao Wang 0014, Zhengzheng Tu, Cheng Zhang 0010, Jin Tang 0001 |
PRCV (1) | 4 |
| 2019 | Robust pixelwise saliency detection via progressive graph rankings
Bo Jiang 0002, Zhengzheng Tu, Amir Hussain 0001, Jin Tang 0001 |
Neurocomputing | 3 |
| 2018 | A prior regularized multi-layer graph ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jixin Ma 0001, Jin Tang 0001 |
Neurocomputing | 3 |
| 2017 | A global and local consistent ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jin Tang 0001 |
J. Vis. Commun. Image Represent. | 4 |