VLDB 2026 Research / reviewers in the wild / expert
Yuhui Zheng
dblp:155/0258
· DBLP profile ↗
138ranked-venue papers
12as first author
105since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 7 first-author · 66 since 2021Artificial intelligence and machine learning · 41 · 3 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAC-DMVC: Reliability-Aware Contrastive Deep Multi-View Clustering Under Multi-Source NoiseabstractMulti-view clustering (MVC), which aims to separate the multi-view data into distinct clusters in an unsupervised manner, is a fundamental yet challenging task. To enhance its applicability in real-world scenarios, this paper addresses a more challenging task: MVC under multi-source noises, including missing noise and observation noise. To this end, we propose a novel framework, Reliability-Aware Contrastive Deep Multi-View Clustering (RAC-DMVC), which constructs a reliability graph to guide robust representation learning under noisy environments. Specifically, to address observation noise, we introduce a cross-view reconstruction to enhances robustness at the data level, and a reliability-aware noise contrastive learning to mitigates bias in positive and negative pairs selection caused by noisy representations. To handle missing noise, we design a dual-attention imputation to capture shared information across views while preserving view-specific features. In addition, a self-supervised cluster distillation module further refines the learned representations and improves the clustering performance. Extensive experiments on five benchmark datasets demonstrate that RAC-DMVC outperforms SOTA methods on multiple evaluation metrics and maintains excellent performance under varying ratios of noise. Shihao Dong, Yue Liu 0008, Xiaotong Zhou, Yuhui Zheng, Xinzhong Zhu |
AAAI | 4 |
| 2026 | Semi-supervised medical image segmentation method via dual-view graph contrastive learning and latent space uncertainty rectification
Dongxu Cheng, Qiwei Dong, Ruian Zhu, Yuhui Zheng |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | 3D-MolGL: A multimodal framework for integrating 3D molecular graphs into language models
Huizhi Li, Dagang Li 0001, Jinglin Zhang 0001, Yuhui Zheng, Cong Bai |
Expert Syst. Appl. | 4 |
| 2026 | Entropy-Guided Condensing for Vision Transformer
Sihao Lin, Pumeng Lyu, Dongrui Liu, Zhihui Li 0001, Wenguan Wang, Xiaojun Chang, Yuhui Zheng |
Int. J. Comput. Vis. | 7 |
| 2026 | Advancing open-set object detection with SAM knowledge transfer and variational feature reconstruction
Yuhui Zheng |
Neurocomputing | 6 |
| 2026 | Network resilience prediction based on adaptive spatio-temporal feature perception
Yuzhi Xiao, Yuhui Zheng, Zhonglin Ye, Haixing Zhao |
Neurocomputing | 3 |
| 2026 | CSCA: Channel-specific information contrast and aggregation for weakly supervised semantic segmentation
Wenxin Sun, Yuhui Zheng, Zhonglin Ye |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | Implicit Alignment with Complementary Information for Text-based Person Re-identification
Guoqing Zhang 0002, Yadang Chen, Le Sun 0002, Yulin Cao, Yuhui Zheng |
Knowl. Based Syst. | 6 |
| 2026 | Target-agnostic common attributes learning for few-shot semantic segmentation
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu |
Pattern Recognit. | 3 |
| 2026 | SFF-CycleGAN: Spatial-Frequency Fusion CycleGAN by frequency-selective modeling of scattering and absorption for underwater image enhancement
Gangping Zhang, Tuxin Guan, Yuhui Zheng |
Pattern Recognit. | 6 |
| 2026 | Alpha-aware neural style transfer in RGBA space via soft alpha-guided feature propagation
Xiaotong Zhou, Yuhui Zheng, Shihao Dong |
Pattern Recognit. | 2 |
| 2026 | Boosting Video Object Segmentation With Discriminative Core Features and Adaptive Position Refinement
Yadang Chen, Guolong Li, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | IdentityGuard: Disrupting Both Identity Aggregation and Binding Against Diffusion-Based PersonalizationabstractDiffusion-based personalization brings convenience to users in text-to-image generation but it also poses risks of rights infringement and content misuse. To address this issue, researchers have proposed several proactive defense methods by adversarial attacks. However, most of these methods directly attack the noise prediction results during the fine-tuning process, overlooking the unique characteristics of diffusion-based personalization, which results in limited defense performance. Therefore, this paper summarizes the two core tasks of personalized fine-tuning as identity aggregation and identity binding, and proposes a defense method named IdentityGuard to specifically attack these two core tasks. The IdentityGuard designs a training sample decorrelation (TSD) attack and a text-image decoupling (TID) attack respectively for the two core tasks. The TSD attack disrupts the learning of common features by reducing the correlations among training samples. The TID attack targets all tokens by using the value-inverted attention map of each token as its adaptive target, aiming to suppress high-attention regions and strengthen low-attention regions. In addition, a token-level adaptive weighting strategy is designed to dynamically allocate attack weights across different tokens during fine-tuning. Experimental results demonstrate that the IdentityGuard effectively enhances proactive defense performance against diffusion-based personalization, achieving an average improvement of 22.26% in terms of Identity Score Matching (ISM) metric compared to the state-of-the-art (SOTA) methods. The source code is available at https://github.com/imagecbj/IdentityGuard. Beijing Chen, Ziqiang Li 0001, Yuhui Zheng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DC2MNet: Lightweight and Efficient Discrete Cosine Channel Modulation Network for Image RestorationabstractImage restoration aims to remove degradation factors (such as blur, snow e.g.) from the damaged image and reconstruct a clean image. Although some methods seek solutions from the frequency domain and are proven to be effective, they are still faced two challenges: (i) Degradation blurs cannot be removed well, and (ii) Inverse transform in frequency domain is computationally expensive. To this end, we propose a lightweight and efficient Discrete Cosine Channel Modulation Network (DC2MNet) for recovering images of multiple degraded conditions from the frequency and spatial perspectives. Specifically, we propose a Discrete Cosine Channel Modulation (DCCM) module to extract the most informative lowest-frequency components of features, and subsequently utilize the channel modulation to reconstruct the global structure of the corresponding feature, avoiding inverse transform in high-dimensional spaces. Furthermore, to effectively remove degradation, we propose a Spatial Mask Modulation (SMM) module to suppress degradation blurs in high-frequency features and emphasize local details that are beneficial to image restoration via pixel-level spatial attention. Finally, we embed the DCCM module and SMM module into the Channel Spatial Modulation Block (CSMB) to form the basic component of DC2MNet, which achieves SOTA performance on various restoration tasks through extensive experiments, including image dehazing, deraining, desnowing and multi-weather restoration. The code and pre-trained models will be open source in this repository. Guoqing Zhang 0002, Wenxuan Fang 0001, Yupeng Shang, Yuhui Zheng, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Decoupling Localization and Semantics for Open-Set Object DetectionabstractOpen-set object detection (OSOD) is an important research direction in computer vision, focusing on enhancing a model’s ability to detect unknown categories. Current methods are overly dependent on supervision from known categories, resulting in detection bias that substantially impairs the model’s capacity to recognize unknown classes. In this study, we propose a decoupled localization and semantic OSOD method (DLS-OSOD) that refines supervision granularity to improve unknown category perception. Specifically, to reduce the impact of inaccurate localization on classification, we propose a class-agnostic region proposal network (CA-RPN), which removes the binary classification module in the traditional RPN, allowing the model to focus on region positioning. Furthermore, to mitigate misclassification effects on localization, we design a prototype-based region filtering module (PBF), which constructs a compact prototype space using category semantics during training and pre-filters unknown regions before classification based on region-prototype distance during inference. Additionally, we propose the Unknown Feature Expansion (UFE) and Known Feature Preservation (KFP) modules. UFE enhances supervision for unknown categories by synthesizing unknown category features, improving the model’s ability to detect unknown regions. KFP constrains known-category features through textual anchors, preserving the detection performance of known categories. Experiments on benchmark datasets demonstrate the superior performance of our method. Guoqing Zhang 0002, Yuhui Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Learnable Object Queries for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment unseen-category objects given only a few annotated samples. Although significant progress has been made in the field of FSS, selecting an appropriate feature matching method remains a challenge. Traditional prototype-based methods can preserve high-level semantic features, but they tend to lose detailed information. On the other hand, pixel-level comparison methods retain fine-grained details but are vulnerable to distractors and noise, leading to poor robustness. To address these issues, this paper proposes a target-agnostic object-based method. Specifically, we propose a set of learnable "object queries" to extract object features, which preserve both high-level semantic information and fine-grained details. Additionally, during the training phase, we exploit the prior knowledge of foreground and background embedded in the samples to enhance the model's performance. In the inference phase, the model utilizes both the support set and the learned prior knowledge to perform segmentation tasks, mitigating the data distribution bias caused by limited samples. Extensive experiments on benchmark datasets demonstrate that our method outperforms state-of-the-art approaches in both accuracy and robustness. Code is available at https://github.com/wenbo456/OTBNet. Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu |
IEEE Trans. Image Process. | 3 |
| 2026 | Refinement on Both Foreground and Background Prototype for Few-Shot SegmentationabstractAlthough few-shot segmentation (FSS) methods have achieved remarkable results, there remain challenges associated with the limited number of support samples.i)The objects in support and query images may have substantially different appearances even though they belong to the same category, which is known as the prototype bias problem.ii)Most methods neglect the background information, especially the query background during the inference stage. To address these problems, we propose DPRNet, a novel network with dual branch of foreground and background prototype refinement modules. Specifically, we first present a Variational Feature Semantic Enhancement (VFSE) module, in which we refine the object prototype with a variational autoencoder and word-text labels. In this way, the biased class-wise prototype caused by the limited support samples can be aligned, achieving better performance. Second, we design a Background Prototype Refinement (BPR) module that effectively explores the potential information in the background for both the support and query images. More importantly, it is designed to generate online predictions of the query background during the training stage to fully mimic the inference stage. These advancements enhance the robustness and generalizability of our method, and the results of experiments demonstrate its effectiveness. In the 1-way 5-shot setting on PASCAL-$5^{i}$, our method achieves a mean-IoU improvement of 1.59% over the competing method. Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu |
IEEE Trans. Multim. | 3 |
| 2026 | Identity Clue Refinement and Enhancement for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) is a challenging cross-modal matching task due to significant modality discrepancies. While current methods mainly focus on learning modality-invariant features through unified embedding spaces, they often focus solely on the common discriminative semantics across modalities while disregarding the critical role of modality-specific identity-aware knowledge in discriminative feature learning. To bridge this gap, we propose a novel Identity Clue Refinement and Enhancement (ICRE) network to mine and utilize the implicit discriminative knowledge inherent in modality-specific attributes. Initially, we design a Multi-Perception Feature Refinement (MPFR) module that aggregates shallow features from shared branches, aiming to capture modality-specific attributes that are easily overlooked. Then, we propose a Semantic Distillation Cascade Enhancement (SDCE) module, which distills identity-aware knowledge from the aggregated shallow features and guide the learning of modality-invariant features. Finally, an Identity Clues Guided (ICG) Loss is proposed to alleviate the modality discrepancies within the enhanced features and promote the learning of a diverse representation space. Extensive experiments across multiple public datasets clearly show that our proposed ICRE outperforms existing SOTA methods. Guoqing Zhang 0002, Zhun Wang, Zhonglin Ye, Yuhui Zheng |
IEEE Trans. Multim. | 5 |
| 2025 | Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image CaptionsabstractWhile densely annotated image captions significantly facilitate the learning of robust visionlanguage alignment, methodologies for systematically optimizing human annotation efforts remain underexplored.We introduce CHAIN-OF-TALKERS (COTALK), an AI-in-the-loop methodology designed to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints (e.g., total human annotation time).The framework is built upon two key insights.First, sequential annotation reduces redundant workload compared to conventional parallel annotation, as subsequent annotators only need to annotate the "residual"-the missing visual information that previous annotations have not covered.Second, humans process textual input faster by reading while outputting annotations with much higher throughput via talking; thus a multimodal interface enables optimized efficiency.We evaluate our framework from two aspects: intrinsic evaluations that assess the comprehensiveness of semantic units, obtained by parsing detailed captions into object-attribute trees and analyzing their effective connections; extrinsic evaluation measures the practical usage of the annotated captions in facilitating vision-language alignment.Experiments with eight participants show our CHAIN-OF-TALKERS (CoTalk) improves annotation speed (0.42 vs. 0.30 units/sec) and retrieval performance (41.13% vs. 40.52%)over the parallel method.per minute? a review and meta-analysis of reading rate. Delong Chen, Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001, Yuhui Zheng |
EMNLP | 7 |
| 2025 | Center-Oriented Prototype Contrastive ClusteringabstractContrastive learning is widely used in clustering tasks due to its discriminative representation. However, the conflict problem between classes is difficult to solve effectively. Existing methods try to solve this problem through prototype contrast, but there is a deviation between the calculation of hard prototypes and the true cluster center. To address this problem, we propose a center-oriented prototype contrastive clustering framework, which consists of a soft prototype contrastive module and a dual consistency learning module. In short, the soft prototype contrastive module uses the probability that the sample belongs to the cluster center as a weight to calculate the prototype of each category, while avoiding inter-class conflicts and reducing prototype drift. The dual consistency learning module aligns different transformations of the same sample and the neighborhoods of different samples respectively, ensuring that the features have transformation-invariant semantic information and compact intra-cluster distribution, while providing reliable guarantees for the calculation of prototypes. Extensive experiments on five datasets show that the proposed method is effective compared to the SOTA. Our code is published on https://github.com/LouisDong95/CPCC. Shihao Dong, Xiaotong Zhou, Yuhui Zheng, Xinzhong Zhu |
ICME | 3 |
| 2025 | Multi-view Clustering via Bi-level Decoupling and Consistency LearningabstractMulti-view clustering has shown to be an effective method for analyzing underlying patterns in multi-view data. The performance of clustering can be improved by learning the consistency and complementarity between multi-view features, however, cluster-oriented representation learning is often over-looked. In this paper, we propose a novel Bi-level Decoupling and Consistency Learning framework (BDCL) to further explore the effective representation for multi-view data to enhance inter-cluster discriminability and intra-cluster compactness of features in multi-view clustering. Our framework comprises three modules: 1) The multi-view instance learning module aligns the consistent information while preserving the private features between views through reconstruction autoencoder and contrastive learning. 2) The bi-level decoupling of features and clusters enhances the discriminability of feature space and cluster space. 3) The consistency learning module treats the different views of the sample and their neighbors as positive pairs, learns the consistency of their clustering assignments, and further compresses the intra-cluster space. Experimental results on five benchmark datasets demonstrate the superiority of the proposed method compared with the SOTA methods. Our code is published on https://github.com/LouisDong95/BDCL. Shihao Dong, Yuhui Zheng, Xinzhong Zhu |
IJCNN | 2 |
| 2025 | RemoteSAM: Towards Segment Anything for Earth ObservationabstractWe aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM. Liang Yao 0001, Fan Liu 0003, Delong Chen, Chuanyi Zhang, Ziyun Chen 0004, Shimin Di, Yuhui Zheng |
ACM Multimedia | 9 |
| 2025 | Self-learning weight network based on label distribution training for facial expression recognitionabstractAbstract The recent widespread utilization of facial expression recognition (FER) has garnered significant attention in the affective computing field. To address the issue of dominant features being suppressed during feature fusion in FER, this study proposes a self‐learning weight network based on label distribution training (SLW‐LDT). First, based on the ShuffleNet‐V2 backbone model, SLW‐LDT introduced a local feature extraction branch that highlights specific expression‐related features by cropping the facial image into four local regions. Subsequently, the SLW algorithm is devised to allocate learnable weights to global and local features from different branches before their fusion. Moreover, considering the challenge associated with accessing emotional distribution in facial images directly, a label distribution training module (LDT) is introduced during the training phase to generate label distributions for effective training purposes. Experimental results demonstrate that the proposed method achieves accuracies of 89.77% and 64.21% on two in‐the‐wild datasets (RAF‐DB and AffectNet‐7), and 98.90% on the lab‐controlled CK+ dataset. Comparative analysis against state‐of‐the‐art methods reveals slight improvements in recognition accuracy along with robust performance exhibited by the model. Yangbo Chen, Chunyan Peng, Yuhui Zheng |
IET Image Process. | 4 |
| 2025 | The Attack and Defense Researches on the Dual-Layer Network of Multivariable Anomaly CausesabstractMultivariate anomaly causes interpretation provides insight into the root cause of information system anomalies, identifying the direct factors that trigger anomalies and revealing potential systemic flaws. However, current research generally focuses on two directions: on the one hand, anomaly diagnosis research for nodes with high anomaly degree; on the other hand, single‐layer anomaly causes interpretation graph construction based on explicit features capturing anomaly locations and their neighborhood structures. These approaches pay insufficient attention to the attack defense of anomaly causes interpretation graph, thereby weakening the credibility and reliability of anomaly causation interpretation. Therefore, we systematically explore the attack strategy and defense mechanism of the multivariate anomaly causes interpretation graph. Firstly, we propose an adaptive learning method for constructing a dual‐layer anomaly causes interpretation graph. The method reduces the dependence on artificial a priori assumptions by introducing an adaptive mechanism and realizes the dynamic decoupling of the spatiotemporal coupling relationships of multivariate data, thus providing a diversified perspective for the multivariate anomaly causes interpretation. Second, considering the vulnerability of the multivariate spatiotemporal correlation after decoupling and the structural characteristics of the dual‐layer anomaly causes interpretation graph, we further propose a structural protection mechanism based on dual‐layer complex networks to improve the structural robustness and resistance to the interference of anomaly causes interpretation graph. Finally, we verify the effectiveness of the proposed model by testing various attack defense scenarios such as noise attack, gradient attack, and structure attack. The experimental results show that the model in this paper can effectively defend against multiple attack methods and ensure the integrity and reliability of the anomaly causes interpretation graph. Jiaxin Han, Zhonglin Ye, Xuanrong Huo, Yuzhi Xiao, Yuhui Zheng |
Int. J. Intell. Syst. | 6 |
| 2025 | Bridging the metrics gap in image style transfer: A comprehensive survey of models and criteria
Xiaotong Zhou, Yuhui Zheng |
Neurocomputing | 2 |
| 2025 | Single stage weakly supervised semantic segmentation via enhanced patch affinity
Jingjie Jiang, Yuhui Zheng, Guoqing Zhang 0002 |
Image Vis. Comput. | 2 |
| 2025 | Structure-aware contrastive learning for glomerulus segmentation in renal pathology
Xiangbo Shu, Yuhui Zheng, Jin Ding, Xianghui Fu |
Image Vis. Comput. | 4 |
| 2025 | Expert-scoring guided global information interaction network for lightweight image super-resolution
Runtao Liu, Xiaotong Zhou, Yuhui Zheng |
Image Vis. Comput. | 4 |
| 2025 | DBFAM: A dual-branch network with efficient feature fusion and attention-enhanced gating for medical image segmentation
Benzhe Ren, Yuhui Zheng, Jin Ding |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Local-enhanced representation for text-based person search
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng, Gaven Martin, Ruili Wang 0001 |
Pattern Recognit. | 3 |
| 2025 | Adaptive transformer with Pyramid Fusion for cloth-changing Person Re-Identification
Guoqing Zhang 0002, Jieqiong Zhou, Yuhui Zheng, Gaven J. Martin, Ruili Wang 0001 |
Pattern Recognit. | 3 |
| 2025 | Synth-Tracker: Recoverable and Traceable Defense Watermark Against Face SynthesisabstractThe fast development of face synthesis technology has brought an increasing number of face synthesis services. However, when such services are misused in malicious activities, the face synthesis service providers (FSSPs) may face serious legal risks. Therefore, this letter proposes a recoverable and traceable defense watermarking method to protect FSSPs. The method designs a decoupled data hiding framework to separate two embedding tasks, where the source image and operator’s ID are inserted into the synthesized image respectively by invertible neural network and convolutional neural network. Furthermore, a facial masking strategy is employed to exclude background information of source image for enhancing the imperceptibility. The proposed method enables forensic traceability for the FSSPs to track back to the malicious users and recover the original source images, building a chain of evidence and inferring the forger’s intent after forgery. Experimental results show that compared to the existing methods, the proposed method has superior performance in source image recovery and ID extraction. In addition, the plug-and-play design of the proposed method allows for seamless integration into current face synthesis services. Beijing Chen, Yuhui Zheng |
IEEE Signal Process. Lett. | 3 |
| 2025 | Cross-Attention With Conditional Matching for Multi-Target Domain AdaptationabstractAs an emerging direction of machine learning, multi-target domain adaptation (MTDA) aims to address the challenges of adapting models to multiple target domains. However, existing studies often focus on single-target domain adaptation or fail to delve into the complexities associated with multiple target domains. So there is a notable lack of comprehensive research and exploration in MTDA. Consequently, we propose a cross-attention with conditional matching for MTDA that intends to overcome the challenges posed by domain discrepancy, multi-target domain heterogeneity, and scalability. Foremost, we design a novel multi-target conditional matching that aims to align the sample distribution by leveraging nearest neighbor principle. This strategy takes into account the unique characteristics of each target domain, facilitating adaptive adaptation across multiple domains. Furthermore, we use the transformer module and well-design a cross-attention mechanism to facilitate the alignment of distributions across the source and target domains, as well as among the target domains, thus mitigating discrepancies among multiple domains. Through integrating the cross-attention mechanism into the training phase, attaining effective alignment of cross-domain distributions, we improve the adaptability and performance of the method. By the end, our approach demonstrates effective and superior experimental results indicating the significance of our work. Qing Tian 0001, Yuhui Zheng, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | InfinitePerson: Innovating Synthetic Data Creation for Generalization Person Re-IdentificationabstractRecently, large-scale synthetic datasets have effectively alleviated the issue of insufficient person re-identification (Re-ID) datasets. However, synthetic datasets grapple with inherent challenges, including the subpar quality of synthetic pedestrians and single data collection. This paper presents InfinitePerson, a costless pipeline that fully utilizes the infinite generation capability of diffusion models to produce diverse UV texture images and effortlessly constructs high-quality synthetic datasets by simulating a real surveillance network. Specifically, we innovatively propose the utilization of diffusion models to generate high-quality, realistic, and diverse UV texture images to address the limitations of clothing textures. This ensures that our 3D character models have complete clothing texture information and look very similar to real-world pedestrians. Moreover, in response to the challenges in replicating synthetic data collection pipelines, we propose a sub-monitoring network data collection method, which can collect pedestrians data from different viewpoints, backgrounds, and lighting conditions through simple scene layout. Finally, a more scalable and realistic large synthetic dataset called InfinitePerson is created, containing 4,700 identities and 535,636 images. Experimental evidence demonstrates show that models trained on InfinitePerson exhibit superior generalization performance, surpassing those trained on both popular real-world and synthetic person Re-ID datasets. The InfinitePerson project is available athttps://github.com/zhguoqing/InfinitePerson. Guoqing Zhang 0002, Jin Li 0074, Yuhui Zheng, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Mask-Aware Hierarchical Aggregation Transformer for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID) is a challenging problem due to the absence of notable discriminative features resulting from incomplete body part images and interference from occluded regions. Recently, some transformer-based methods have demonstrated excellent capabilities in resolving this problem, however these methods are not able to precisely focus on the non-occluded body parts and cannot capture fine-grained local features. To achieve these we propose a Mask-Aware Hierarchical Aggregation TrAnsforMer (MAHATMA) method to enhance occluded person Re-ID. Specifically, we propose a Mask Information Embedding (MIE) module, which directs the model to focus on non-occluded body parts by incorporating the mask semantic information of a human body. Furthermore, to effectively capture fine-grained local features, we propose a Hierarchical Feature Aggregation (HFA) module that mines more exploitable high-quality detail information by aggregating hierarchical image patch representations. To further alleviate the feature loss problem, we propose a Diverse Feature Completion (DFC) module, which is able to complete global features through multi-path feature integration. Extensive experimental evaluations demonstrate that our method exhibits superior performance in dealing with occluded and holistic person datasets. Guoqing Zhang 0002, Yuhui Zheng, Gaven J. Martin, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Local-Global Information Perception Network for Salient Object Detection in Optical Remote Sensing ImagesabstractIn the field of salient object detection (SOD), optical remote sensing images (ORSI) differ significantly from natural sensing images (NSI). Existing research in ORSI-based SOD is constrained by the limitations of convolutional neural networks (CNNs) in feature extraction and by the underutilization of feature information in Transformer-based approaches. To address these challenges, this paper presents a Transformer-based Local-Global Information Perception Network (LGIPNet) for ORSI, which enhances encoder-generated features at multiple levels to highlight salient targets through three specialized feature enhancement modules. The Edge Adaptive Enhancement Module (EAEM) focuses on extracting local edge features to guide precise edge generation. The Multi-scale Grouped Weighted Attention Module (MGWAM) extracts local information from low-level features, scales features, and uses multi-scale channel and learnable weighted spatial attention to locate salient targets. For high-level features, the Dual-Domain Attention Module (DDAM) integrates a Channel Enhancement Attention Block (CEA) and an Adaptive Spatial Attention Block (ASA) to refine both local and global information. Specifically, the EAEM sharpens the edges of salient objects, thereby ensuring the precision of boundary detection. The MGWAM, on the other hand, enriches the feature representation across multiple scales, enhancing the network’s capability to encapsulate both fine-grained details and broader contextual information. The DDAM further strengthens the balance between local and global information, preserving feature integrity across levels. Finally, multi-scale features are cascaded to produce the final saliency map. This holistic strategy empowers LGIPNet to accurately identify and emphasize salient objects in ORSI. Experiments on three datasets demonstrate that LGIPNet outperforms existing state-of-the-art methods, establishing its effectiveness and robustness in ORSI-based SOD. The source code is available at https://github.com/sCauliflower/LGIPNet.git. Le Sun 0002, Hongxin Liu, Yuhui Zheng, Qiao Chen 0004, Zebin Wu 0001, Liyong Fu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Deformable Convolution-Enhanced Hierarchical Transformer With Spectral-Spatial Cluster Attention for Hyperspectral Image ClassificationabstractVision Transformer (ViT), known for capturing non-local features, is an effective tool for hyperspectral image classification (HSIC). However, ViT's multi-head self-attention (MHSA) mechanism often struggles to balance local details and long-range relationships for complex high-dimensional data, leading to a loss in spectral-spatial information representation. To address this issue, we propose a deformable convolution-enhanced hierarchical Transformer with spectral-spatial cluster attention (SClusterFormer) for HSIC. The model incorporates a unique cluster attention mechanism that utilizes spectral angle similarity and Euclidean distance metrics to enhance the representation of fine-grained homogenous local details and improve discrimination of non-local structures in 3-D HSI and 2-D morphological data, respectively. Additionally, a dual-branch multiscale deformable convolution framework augmented with frequency-based spectral attention is designed to capture both the discrepancy patterns in high-frequency and overall trend of the spectral profile in low-frequency. Finally, we utilize a cross-feature pixel-level fusion module for collaborative cross-learning and fusion of the results from the dual-branch framework. Comprehensive experiments conducted on multiple HSIC datasets validate the superiority of our proposed SClusterFormer model, which outperforms existing methods. The source code of SClusterFormer is available at https://github.com/Fang666666/HSIC SClusterFormer. Yu Fang 0012, Le Sun 0002, Yuhui Zheng, Zebin Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | CLIP-Based Multi-Modal Feature Learning for Cloth-Changing Person Re-IdentificationabstractContrastive Language-Image Pre-training (CLIP) has achieved remarkable results in the field of person re-identification (ReID) due to its excellent cross-modal understanding ability and high scalability. Since the text encoder of CLIP mainly focuses on easy-to-describe attributes such as clothing, and clothing is the main interference factor that reduces the recognition accuracy in cloth-changing person ReID (CC ReID). Consequently, directly applying CLIP to cloth-changing scenario may be difficult to adapt to such dynamic feature changes, thereby affecting the precision of identification. To solve this challenge, we propose a CLIP-based multi-modal feature learning framework (CMFF) for CC ReID. Specifically, we first design a pose-aware identity enhancement module (PIE) to enhance the model's perception of identity-intrinsic information. In this branch, to weaken the interference of clothing information, we apply a ranking loss to minimize the difference between appearance and pose in the feature space. Secondly, we propose a global-local hybrid attention module (GLHA), which fuses head and global features through a cross-attention mechanism, enhancing the global recognition ability of key head information. Finally, considering that existing CLIP-based methods often ignore the potential importance of shallow features, we propose a graph-based multi-layer interactive enhancement module (GMIE), which groups and integrates multi-layer features of the image encoder, aiming to enhance the contextual awareness of multi-scale features. Extensive experiments on multiple popular pedestrian datasets validate the outstanding performance of our proposed CMFF. Guoqing Zhang 0002, Jieqiong Zhou, Yuhui Zheng, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2025 | Graph Convolutional Multi-Label Hashing for Cross-Modal RetrievalabstractCross-modal hashing encodes different modalities of multimodal data into low-dimensional Hamming space for fast cross-modal retrieval. In multi-label cross-modal retrieval, multimodal data are often annotated with multiple labels, and some labels, e.g., "ocean" and "cloud," often co-occur. However, existing cross-modal hashing methods overlook label dependency that is crucial for improving performance. To fulfill this gap, this article proposes graph convolutional multi-label hashing (GCMLH) for effective multi-label cross-modal retrieval. Specifically, GCMLH first generates word embedding of each label and develops label encoder to learn highly correlated label embedding via graph convolutional network (GCN). In addition, GCMLH develops feature encoder for each modality, and feature fusion module to generate highly semantic feature via GCN. GCMLH uses teacher-student learning scheme to transfer knowledge from the teacher modules, i.e., label encoder and feature fusion module, to the student module, i.e., feature encoder, such that learned hash code can well exploit multi-label dependency and multimodal semantic structure. Extensive empirical results on several benchmarks demonstrate the superiority of the proposed method over existing state-of-the-arts. Xiaobo Shen 0001, Yinfan Chen, Weiwei Liu 0003, Yuhui Zheng, Quan-Sen Sun, Shirui Pan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Distributed Manifold Hashing for Image Set Classification and RetrievalabstractConventional image set methods typically learn from image sets stored in one location. However, in real-world applications, image sets are often distributed or collected across different positions. Learning from such distributed image sets presents a challenge that has not been studied thus far. Moreover, efficiency is seldom addressed in large-scale image set applications. To fulfill these gaps, this paper proposes Distributed Manifold Hashing (DMH), which models distributed image sets as a connected graph. DMH employs Riemannian manifold to effectively represent each image set and further suggests learning hash code for each image set to achieve efficient computation and storage. DMH is formally formulated as a distributed learning problem with local consistency constraint on global variables among neighbor nodes, and can be optimized in parallel. Extensive experiments on three benchmark datasets demonstrate that DMH achieves highly competitive accuracies in a distributed setting and provides faster classification and retrieval than state-of-the-arts. Xiaobo Shen 0001, Peizhuo Song, Yun-Hao Yuan 0001, Yuhui Zheng |
AAAI | 4 |
| 2024 | Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationabstractThe performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distribution (i.i.d.) assumption, leading to domain shift. In this paper, we propose a generalizable fourier augmentation method during training to improve the generalization ability of the model. To achieve this, we perform Fast Fourier Transform (FFT) over the intermediate spatial domain features in each layer to yield corresponding frequency representations, including amplitude components (encoding scene-aware styles such as texture, color, contrast of the scene) and phase components (encoding rich semantics). We produce a variety of style features via Gaussian sampling to augment the training data, thereby improving the generalization capability of the model. To further improve the cross-domain generalization performance of the model, we design a phase feature update strategy via exponential moving average using phase features from past frames in an online update manner, which could help the model to learn cross-domain-invariant features. Extensive experiments show that our proposed method achieves the state-of-the-art performance on popular benchmarks. Huihui Song 0003, Tiankang Su, Yuhui Zheng, Kaihua Zhang 0001, Bo Liu 0005, Dong Liu 0002 |
AAAI | 3 |
| 2024 | How Can We Design a Standardized and Efficient Health Data Management System for Large-Scale Heterogeneous TCM Data?abstractWith the rapid development of medical information technology, the standardization management and storage of large-scale, multi-source, and heterogeneous traditional Chinese medicine (TCM) data has become one of the critical challenges in the current TCM industry. The vast scale and diverse structures of multi-source, heterogeneous TCM data exhibit characteristics such as dispersed data patterns and fragmented data content. These features present numerous governance challenges for TCM data, including inconsistent storage modes, data loss, lack of structure, standardization, and non-intuitive display methods. These challenges severely restrict the effective utilization and retrieval of TCM data, hindering efficient analysis and exploration of patient information. In light of the aforementioned research background, this paper proposes a standardized processing method for large-scale, multi-source, and heterogeneous TCM data. This method employs data cleaning, unified file naming, and other processing techniques to filter and process data. Building upon this foundation, a unified approach based on structured data is designed to link various types of data through user’s basic information using file linkage. This approach can establish effective connections among various test data obtained by patients in the same batch without affecting the original distribution of various types of data, thereby providing a more convenient pathway for the storage, analysis, and standardization of patient data. Furthermore, to enhance the efficiency of TCM information management and the value of data utilization, this paper constructs a health data management system based on standardized data. This system encompasses functions such as data storage, information input and retrieval, and intelligent analysis, enabling TCM medical institutions to rationally allocate medical resources, improve resource utilization efficiency, and reduce medical costs. At the data storage stage, this paper designs a health data storage system based on blockchain technology, which can enhance the security and integrity of user data and promote data sharing among various institutions. The proposed standardized processing method for TCM data and the health data management system can promote the informatization construction of TCM, enhance the level of TCM services, and significantly contribute to improving people’s health status and enhancing the overall benefits of the TCM healthcare system. Yuhui Zheng, Yankun Zhang, Weiqiang Lin |
BIBM | 1 |
| 2024 | Glance, Focus and Refinement Network for Remote Sensing Change DetectionabstractExisting change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively capture the change targets due to the imbalance ratio between the change regions and the whole scene. To this end, this paper presents a glance, focus, and refinement network (GFRNet), which formulates CD as a continuous, step-by-step focusing process to mimic the human visual system. Specifically, the GFRNet first employs a transformer encoder to extract the global features from the bi-temporal images, where each feature takes a glance at the whole scene. Then, the GFRNet gradually pays attention to a cascade of salient regions, and ultimately progressively refines its focus on the desired areas of change. Comprehensive evaluations on two extensively utilized benchmark datasets, including LEVIR-CD and WHU-CD, demonstrate the superiority of our GFR-Net to a variety of state-of-the-art methods. Zixuan Sun, Yuhui Zheng, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao |
ICASSP | 3 |
| 2024 | Contrastive Transformer Masked Image Hashing for Degraded Image Retrieval
Xiaobo Shen 0001, Haoyu Cai, Xiuwen Gong, Yuhui Zheng |
IJCAI | 4 |
| 2024 | Contrastive Transformer Cross-Modal Hashing for Video-Text Retrieval
Xiaobo Shen 0001, Qianxin Huang, Long Lan, Yuhui Zheng |
IJCAI | 4 |
| 2024 | A Transformer-Based Adaptive Prototype Matching Network for Few-Shot Semantic Segmentation
Yadang Chen, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu |
IJCAI | 3 |
| 2024 | Graph Convolutional Semi-Supervised Cross-Modal HashingabstractCross-modal hashing encodes different modalities of multi-modal data into a low-dimensional Hamming space for fast cross-modal retrieval. Most existing cross-modal hashing methods heavily rely on label semantics to boost retrieval performance; however, semantics are expensive to collect in real applications. To mitigate the heavy reliance on semantics, this work proposes a new semi-supervised deep cross-modal hashing method, namely, Graph Convolutional Semi-Supervised Cross-Modal Hashing (GCSCH), which is trained with limited label supervision. The proposed GCSCH first generates pseudo-multi-labels of the unlabeled samples using the simple yet effective idea of consistency regularization and pseudo-labeling. GCSCH designs a fusion network that merges the two modalities and employs Graph Convolutional Network (GCN) to capture semantic information among ground-truth-labeled and pseudo-labeled multi-modal data. Using the idea of knowledge distillation, GCSCH employs a teacher-student learning scheme that can successfully transfer knowledge from the fusion module to the image and text hashing networks. Empirical studies on three multi-modal benchmark datasets demonstrate the superiority of the proposed GCSCH over state-of-the-art cross-modal hashing methods with limited label supervision. Xiaobo Shen 0001, Gaoyao Yu, Yinfan Chen, Xichen Yang, Yuhui Zheng |
ACM Multimedia | 5 |
| 2024 | Local Feature-Emphasizing Transformer for Cloth-Changing Person Re-identification
Jieqiong Zhou, Guoqing Zhang 0002, Yuhui Zheng, Fuguo Zhang |
MMAsia | 3 |
| 2024 | Multi-task image restoration network based on spatial aggregation attention and multi-feature fusionabstractAbstract The main purpose of image restoration is to recover high‐quality image content from degraded versions. However, current mainstream models tend to focus solely on spatial details or contextual semantics, resulting in poor repair effects. To address this issue, a multi‐task image repair network based on spatial aggregation attention and multi‐feature fusion (SAAM) is proposed. It utilizes the global semantic information from the low‐resolution subnetwork to guide the local feature extraction of the high‐resolution subnetwork, thereby preserving the overall image structure while enhancing local details. Additionally, to enhance the model's understanding and representation capabilities of images, the feature fusion mechanism (FFM) is designed to merge feature information from different levels. Finally, the spatial aggregation attention mechanism SAAM enhances the accuracy and quality of image restoration by weighting the importance of different regions in the image at multiple scales. The experimental results demonstrate that the proposed SAAM method outperforms similar approaches in image denoising, deraining and decracking tasks in peak signal‐to‐noise ratio, structural similarity and learned perceptual image patch similarity metrics. The model also exhibits promising performance in restoring real old photos and murals which demonstrates its generalizability. Chunyan Peng, Xueya Zhao, Yangbo Chen, Wanqing Zhang, Yuhui Zheng |
IET Image Process. | 5 |
| 2024 | Progressive discrepancy elimination for visible-infrared person re-identification
Guoqing Zhang 0002, Zhun Wang, Jieqiong Zhou, Yuhui Zheng |
Neurocomputing | 5 |
| 2024 | Learning dual attention enhancement feature for visible-infrared person re-identification
Guoqing Zhang 0002, Yinyin Zhang, Yuhao Chen 0002, Yuhui Zheng |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Boosting Video Object Segmentation via Robust and Efficient Memory NetworkabstractRecently, memory-based methods have exhibited remarkable performance in Video Object Segmentation (VOS) by employing non-local pixel-wise matching between the query and memory. Nevertheless, these methods suffer from two limitations: 1) Non-local pixel-wise matching can result in the incorrect segmentation of background distractor objects, and 2) memory features with substantial temporal redundancy consume significant computing resources and reduce the inference speed. To address the limitations, we first propose a local attention mechanism to suppress background features, and we introduce a novel training framework based on contrast learning to ensure the network learns reliable and robust pixel-wise correspondence between query and memory. We adaptively determine whether to update the memory based on the variation of foreground objects. Next, we propose a dynamic memory bank, which utilizes a lightweight and differentiable soft modulation gate to determine the number of memory features to remove along the temporal dimension. This allows efficient and flexible management of memory features. Our network achieves competitive results (e.g., 92.1% on DAVIS 2016 val, 87.6%/81.3% on DAVIS 2017 val/test, 87.0% on YouTube-VOS 2018 val) compared with the state-of-the-art methods while maintaining a faster inference speed of 25+FPS. Moreover, our network demonstrates a favorable balance between performance and speed when dealing with the long-time video dataset. Yadang Chen, Dingwei Zhang, Yuhui Zheng, Zhi-Xin Yang 0001, Enhua Wu, Haixing Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DCL: Dipolar Confidence Learning for Source-Free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation (SFUDA) aims to conduct prediction on the target domain by leveraging knowledge from the well-trained source model. Due to the absence of source data in the SFUDA setting, the existing methods mainly build the target classifier by fine-tuning the source model incorporated with empirical adaptation losses. Although these methods have achieved somewhat promising results, nearly all of them typically suffer from the closed-fitting dilemma that their models are dominantly affected by these easy-to-distinguish instances than those hard-to-distinguish ones, resulting from the absence of the labeled source data. To address aforementioned issues, we propose the Dipolar Confidence Learning (DCL) for SFUDA. Specifically, we conduct positive confidence learning on the samples with standard outputs to avoid overfitting of the model to these samples. In contrast, we perform negative confidence learning for the samples with abnormal outputs to optimize the complementary label, which forces the network to pay more attention to these confusing samples. Furthermore, to achieve more generalized domain alignment, both the confidence-based fuzzy mixup and rotation-based self-supervised learning are respectively constructed to boost the representation ability of the target model. Finally, extensive experiments are conducted to demonstrate the effectiveness and performance superiority of the proposed method. Qing Tian 0001, Heyang Sun, Shun Peng, Yuhui Zheng, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | SDBAD-Net: A Spatial Dual-Branch Attention Dehazing Network Based on Meta-Former ParadigmabstractImage dehazing is an emblematical low-level vision task that aims at restoring haze-free images from haze images. Recently, some methods adopts deep learning techniques to rebuild haze-free images. However, in real-world scenarios, complex degradation of captured images and non-uniform spatial distributions of haze will significantly weaken the generalization ability of these models. Accordingly, we propose a novel Spatial Dual-Branch Attention Dehazing network (SDBAD-Net) based on the Meta-Former paradigm for end-to-end dehazing. Specifically, we firstly design a robust Spatial Dual-Branch Attention (SDBA) module to filter the haze distribution features from different densities, which is suitable for both uniform and non-uniform situations. Secondly, we introduce a Structural Features Supplementary (SFS) module to dynamically fuse the contextual structural features in a nonlinear manner, so as to correct the image distortion caused by the lack of structural details. Finally, the quantitative and qualitative experiments are carried out on two challenging datasets, and the results show that our method outperforms most of state-of-the-art algorithms with fewer parameters and faster speed, especially surpassing FFA-Net with only 50% parameters and 7% computational costs. In addition, we ulteriorly explore its performance on object detection in foggy weather with our model on the challenging Real-world Task-driven Testing Set (RTTS), and the surprising results further prove the robustness and wide-applicability of our method. Guoqing Zhang 0002, Wenxuan Fang 0001, Yuhui Zheng, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Multiscale 3-D-2-D Mixed CNN and Lightweight Attention-Free Transformer for Hyperspectral and LiDAR ClassificationabstractThe effective combination of hyperspectral image (HSI) and light detection and ranging (LiDAR) data can be utilized for land cover classification. Recently, deep learning-based classification methods, especially those utilizing Transformer networks, have achieved remarkable success. However, deep learning classification methods for multi-source data still encounter various technical challenges, such as the comprehensive utilization of multi-scale information, the lightweight network design, and the efficient fusion strategies for heterogeneous data. To address these challenges, we propose a novel and efficient deep neural network, namely multi-scale 3D-2D mixed CNN feature extraction and multi-source data lightweight attention-free fusion network (M2FNet) based on CNN and Transformer. Through end-to-end training, this network effectively combines heterogeneous information from multiple sources, leading to improved performance in joint classification. Specifically, M2FNet employs a multi-scale 3D-2D mixed CNN design to extract both the spatial-spectral features of HSI and the depth-based elevation features of LiDAR data. Subsequently, the extracted features are fed into a novel encoder comprising a feature enhancement module, designed with mathematical morphology and a dilated convolutional module derived from the self-attention of the conventional Transformer encoder (DConvformer), which plays a crucial role in integrating multi-source information within the network. The well-designed architecture enables the network to acquire multi-scale depth and high-order features, significantly reducing the number of training parameters. Comparative experimental results and ablation studies demonstrate that M2FNet outperforms other advanced methods. The source code is publicly available at https://github.com/cupid6868/M2FNet.git. Le Sun 0002, Yuhui Zheng, Zebin Wu 0001, Liyong Fu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MASSFormer: Memory-Augmented Spectral-Spatial Transformer for Hyperspectral Image ClassificationabstractIn recent years, convolutional neural networks (CNNs) have achieved remarkable success in hyperspectral image (HSI) classification tasks, primarily due to their outstanding spatial feature extraction capabilities. However, CNNs struggle to capture the diagnostic spectral information inherent in HSI. In contrast, vision transformers exhibit formidable prowess in handling spectral sequence information and excelling at capturing long-range correlations between pixels and bands. Nevertheless, due to the information loss during propagation, some existing transformer-based classification methods struggle to form sufficient spectral-spatial information mixing. To mitigate these limitations, we propose a memory-augmented spectral-spatial transformer (MASSFormer) for HSI classification. Specifically, MASSFormer incorporates two efficacious modules, the memory tokenizer (MT) and the memory-augmented transformer encoder (MATE). The former serves to transform spectral-spatial features into memory tokens for storing prior knowledge. The latter aims to extend traditional multi-head self-attention (MHSA) operations by incorporating these memory tokens, enabling ample information blending while alleviating the potential depth decay in the model, and consequently improving the model’s classification performance. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms state-of-the-art methods. The source code is available at https://github.com/hz63/MASSFormer for the sake of reproducibility. Le Sun 0002, Yuhui Zheng, Zebin Wu 0001, Zhonglin Ye, Haixing Zhao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Dual Branch Multi-Level Semantic Learning for Few-Shot SegmentationabstractFew-shot semantic segmentation aims to segment novel-class objects in a query image with only a few annotated examples in support images. Although progress has been made recently by combining prototype-based metric learning, existing methods still face two main challenges. First, various intra-class objects between the support and query images or semantically similar inter-class objects can seriously harm the segmentation performance due to their poor feature representations. Second, the latent novel classes are treated as the background in most methods, leading to a learning bias, whereby these novel classes are difficult to correctly segment as foreground. To solve these problems, we propose a dual-branch learning method. The class-specific branch encourages representations of objects to be more distinguishable by increasing the inter-class distance while decreasing the intra-class distance. In parallel, the class-agnostic branch focuses on minimizing the foreground class feature distribution and maximizing the features between the foreground and background, thus increasing the generalizability to novel classes in the test stage. Furthermore, to obtain more representative features, pixel-level and prototype-level semantic learning are both involved in the two branches. The method is evaluated on PASCAL-5i1-shot, PASCAL-5i5-shot, COCO-20i1-shot, and COCO-20i5-shot, and extensive experiments show that our approach is effective for few-shot semantic segmentation despite its simplicity. Yadang Chen, Ren Jiang, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu |
IEEE Trans. Image Process. | 3 |
| 2024 | Dual-Stream Complex-Valued Convolutional Network for Authentic Dehazed Image Quality AssessmentabstractEffectively evaluating the perceptual quality of dehazed images remains an under-explored research issue. In this paper, we propose a no-reference complex-valued convolutional neural network (CV-CNN) model to conduct automatic dehazed image quality evaluation. Specifically, a novel CV-CNN is employed that exploits the advantages of complex-valued representations, achieving better generalization capability on perceptual feature learning than real-valued ones. To learn more discriminative features to analyze the perceptual quality of dehazed images, we design a dual-stream CV-CNN architecture. The dual-stream model comprises a distortion-sensitive stream that operates on the dehazed RGB image, and a haze-aware stream on a novel dark channel difference image. The distortion-sensitive stream accounts for perceptual distortion artifacts, while the haze-aware stream addresses the possible presence of residual haze. Experimental results on three publicly available dehazed image quality assessment (DQA) databases demonstrate the effectiveness and generalization of our proposed CV-CNN DQA model as compared to state-of-the-art no-reference image quality assessment algorithms. Tuxin Guan, Yuhui Zheng, Xiaojun Wu 0001, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2024 | Multiple Riemannian Kernel Hashing for Large-Scale Image Set Classification and RetrievalabstractConventional image set methods typically learn from small to medium-sized image set datasets. However, when applied to large-scale image set applications such as classification and retrieval, they face two primary challenges: 1) effectively modeling complex image sets; and 2) efficiently performing tasks. To address the above issues, we propose a novel Multiple Riemannian Kernel Hashing (MRKH) method that leverages the powerful capabilities of Riemannian manifold and Hashing on effective and efficient image set representation. MRKH considers multiple heterogeneous Riemannian manifolds to represent each image set. It introduces a multiple kernel learning framework designed to effectively combine statistics from multiple manifolds, and constructs kernels by selecting a small set of anchor points, enabling efficient scalability for large-scale applications. In addition, MRKH further exploits inter- and intra-modal semantic structure to enhance discrimination. Instead of employing continuous feature to represent each image set, MRKH suggests learning hash code for each image set, thereby achieving efficient computation and storage. We present an iterative algorithm with theoretical convergence guarantee to optimize MRKH, and the computational complexity is linear with the size of dataset. Extensive experiments on five image set benchmark datasets including three large-scale ones demonstrate the proposed method outperforms state-of-the-arts in accuracy and efficiency particularly in large-scale image set classification and retrieval. Xiaobo Shen 0001, Xiaxin Wang, Yuhui Zheng |
IEEE Trans. Image Process. | 4 |
| 2024 | One-Click-Based Perception for Interactive Image SegmentationabstractExisting deep learning-based interactive image segmentation methods have significantly reduced the user's interaction burden with simple click interactions. However, they still require excessive numbers of clicks to continuously correct the segmentation for satisfactory results. This article explores how to harvest accurate segmentation of interested targets while minimizing the user interaction cost. To achieve the above goal, we propose a one-click-based interactive segmentation approach in this work. For this particularly challenging problem in the interactive segmentation task, we build a top-down framework dividing the original problem into a one-click-based coarse localization followed by a fine segmentation. A two-stage interactive object localization network is first designed, which aims to completely enclose the target of interest based on the supervision of object integrity (OI). Click centrality (CC) is also utilized to overcome the overlapping problem between objects. This coarse localization helps to reduce the search space and increase the focus of the click at a higher resolution. A principled multilayer segmentation network is then designed by a progressive layer-by-layer structure, which aims to accurately perceive the target with extremely limited prior guidance. A diffusion module is also designed to enhance the information flow between layers. Besides, the proposed model can be naturally extended to multiobject segmentation task. Our method achieves the state-of-the-art performance under one-click interaction on several benchmarks. Tao Wang 0020, Yuhui Zheng, Quan-Sen Sun |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-level Part-aware Feature Disentangling for Text-based Person SearchabstractText-based person search is an important sub-task in cross-modality image retrieval, aiming to capture interested person images by giving textual descriptions. The huge information differences between image and text modalities make this task challenging. Recent methods take local-aligned feature learning strategy into consideration, but lack sufficient mining of more local information. Accordingly, we explore a Multi-level Part-aware Feature Disentangling (MPFD) framework to more fully extract visual and textual representations from multiple angles. Specifically, we introduce a Textual Part-aware Matching (TPM) module into the existing baseline, to disentangle local features for detailed information from both visual and textual part-aware aspects. Besides, in order to fuse multiple local features and improve discrimination of global features, we propose a Multi-level Feature Integration (MFI) module which is capable to perceive the relations between features. We carry out adequate experiments on CUHK-PEDES and ICFG-PEDES datasets to verify our proposed framework, and the results demonstrate that MPFD framework performs favorably against the state-of-the-art methods. Yuhao Chen 0002, Guoqing Zhang 0002, Yuhui Zheng, Weisi Lin |
ICME | 4 |
| 2023 | Graph Convolutional Incomplete Multi-modal HashingabstractMulti-modal hashing (MMH) encodes multi-modal data into latent hash code, and has been widely applied for efficient large-scale multi-modal retrieval. In practice it is common that multi-modal data is often corrupted with missing modalities, e.g., social image often lacks its tags in image-text retrieval. Conventional MMHs can only learn on complete modalities, which however wastes a considerable amount of collected data. To fulfill this gap, this paper proposes Graph Convolutional Incomplete Multi-modal Hashing (GCIMH) to learn hash code on incomplete multi-modal data. GCIMH develops Graph Convolutional Autoencoder to reconstruct incomplete multi-modal data with effective exploit of its semantic structure. GCIMH further develops multi-modal and label networks to encode multiple modalities and label respectively. GCIMH can successfully transfer knowledge of autoencoder and label network to multi-modal hashing network using teacher-student learning framework. GCIMH can handle missing modalities in both offline training and online query stages. Extensive empirical studies on three benchmark datasets demonstrate the superiority of the proposed GCIMH over the state-of-the-arts on both complete and incomplete multi-modal retrieval. Xiaobo Shen 0001, Yinfan Chen, Shirui Pan, Weiwei Liu 0003, Yuhui Zheng |
ACM Multimedia | 5 |
| 2023 | A novel complex-valued convolutional network for real-world single image dehazing
Xinxiu Xie, Tuxin Guan, Yuhui Zheng, Xiaojun Wu 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Transformer-based global-local feature learning model for occluded person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | Compact network embedding for fast node classification
Xiaobo Shen 0001, Yew-Soon Ong, Zheng Mao, Shirui Pan, Weiwei Liu 0003, Yuhui Zheng |
Pattern Recognit. | 6 |
| 2023 | A black-box reversible adversarial example for authorizable recognition to shared images
Lizhi Xiong, Peipeng Yu, Yuhui Zheng |
Pattern Recognit. | 4 |
| 2023 | Motion Stimulation for Compositional Action RecognitionabstractRecognizing the unseen combinations of action and different objects, namely (zero-shot) compositional action recognition, is extremely challenging for conventional action recognition algorithms in real-world applications. Previous methods focus on enhancing the dynamic clues of objects that appear in the scene by building region features or tracklet embedding from ground-truths or detected bounding boxes. These methods rely heavily on manual annotation or the quality of detectors, which are inflexible for practical applications. In this work, we aim to mining the temporal clues from moving objects or hands without explicit supervision. Thus, we propose a novel Motion Stimulation (MS) block, which is specifically designed to mine dynamic clues of the local regions autonomously from adjacent frames. Furthermore, MS consists of the following three steps: motion feature extraction, motion feature recalibration, and action-centric excitation. The proposed MS block can be directly and conveniently integrated into existing video backbones to enhance the ability of compositional generalization for action recognition algorithms. Extensive experimental results on three action recognition datasets, the Something-Else, IKEA-Assembly and EPIC-KITCHENS datasets, indicate the effectiveness and interpretability of our MS block. Yuhui Zheng, Zhao Zhang 0001, Yazhou Yao, Xijian Fan, Qiaolin Ye |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Target-Aware Transformer TrackingabstractObject tracking is aimed at locating a specific object in the image sequence, such as pedestrians, vehicles, and so on. The existing algorithms based on siamese neural network predict the target through similarity matching. Although these algorithms have achieved satisfactory performance, in the process of similarity calculation between template image and search image, only local information is often concerned, which makes the algorithms difficult to obtain the optimal solution. To deal with the abovementioned problems, we propose a model based on Transformer, named TaTrack. Specifically, we first use the encoders to enhance the features. Then, the dependency between template features and search features is established through the target-aware module. Finally, we utilize the classification regression network to locate the target, and use the classification score to adapt to update the template image. Experiments show that our model can achieve great performance on GOT-10k, LaSOT, and TrackingNet datasets. Yuhui Zheng, Yan Zhang 0108, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Mixed Noise Removal for Hyperspectral Images Based on Global Tensor Low-Rankness and Nonlocal SVD-Aided Group SparsityabstractIn hyperspectral images (HSIs), mixed noise (e.g., Gaussian noise, impulse noise, stripe noise, and deadlines) contamination is a common phenomenon that greatly reduces the visual quality of the image. In recent years, methods combining global and non-local low-rankness have been widely used in the field of HSI denoising. However, most methods apply original space-based denoising strategies (low-rank tensor decomposition, total variation, and tensor sparse representation, etc.) directly to the modeling of non-local low-rank tensors in subspace, without fully exploiting the intrinsic and latent properties of the non-local similar tensors. In this paper, we propose a hybrid prior denoising method based on global tensor low-rankness and non-local SVD-aided group sparsity (GTL_NSGS). This method introduces a novel plug-and-play NSGS denoiser that uses singular value decomposition as assistance to successively explore self-similarity of spatial dimension, low-rankness of spectral dimension, and group sparsity of difference domain in subspace non-local similar tensors. Globally, we utilize the existing three-way log-based tensor nuclear norm (3DLogTNN) to approximate the HSI tensor fibered rank and introduce a difference continuity regularization to obtain a continuous smooth spectral basis. Finally, we combine the Alternating Direction Method of Multipliers (ADMM) with the Augmented Lagrangian Multiplication (ALM) algorithm to solve the proposed model effectively. Extensive experiments on simulated and real data sets demonstrate that the proposed method has superior performance in removing mixed noise compared to state-of-the-art denoisers. Le Sun 0002, Qiujie Cao, Yuwen Chen 0001, Yuhui Zheng, Zebin Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multiattention Joint Convolution Feature Representation With Lightweight Transformer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is currently a hot topic in the field of remote sensing. The goal is to utilize the spectral and spatial information from HSI to accurately identify land covers. Convolution neural network (CNN) is a powerful approach for HSI classification. However, CNN has limited ability to capture non-local information to represent complex features. Recently, vision transformers (ViTs) have gained attention due to their ability to process non-local information. Yet, under the HSI classification scenario with ultra-small sample rates, the spectral-spatial information given to ViTs for global modeling is insufficient, resulting in limited classification capability. Therefore, in this article, Multi-Attention Joint Convolution Feature Representation with Lightweight Transformer (MAR-LWFormer) is proposed, which effectively combines the spectral and spatial features of HSI to achieve efficient classification performance at ultra-small sample rates. Specifically, we use a three-branch network architecture to extract multi-scale convolved 3D-CNN, EMAP, and LBP features of HSI, respectively, by taking full exploitation of ultra-small training samples. Second, we design a series of multi-attention modules to enhance spectral-spatial representation for the three types of features and to improve the coupling and fusion of multiple features. Third, we propose an explicit feature attention tokenizer to transform the feature information, which maximizes the effective spectral-spatial information retained in the flat tokens. Finally, the generated tokens are input to the designed lightweight transformer for encoding and classification. Experimental results on three datasets validate that MAR-LWFormer has an excellent performance in HSI classification at ultra-small sample rates when compared to several state-of-the-art classifiers. Yu Fang 0012, Qiaolin Ye, Le Sun 0002, Yuhui Zheng, Zebin Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | CRNet: Channel-Enhanced Remodeling-Based Network for Salient Object Detection in Optical Remote Sensing ImagesabstractDespite the remarkable progress made by the salient object detection of natural sensing images (NSI-SOD), the complex background and scale diversity issues of remote sensing images (RSIs) still pose a substantial obstacle. In this study, we build an end-to-end channel-enhanced remodeling-based network (CRNet) for optical RSIs (ORSIs) to highlight salient objects through feature augmentation. First, the backbone convolutional block is used to suggest the fundamental characteristics. Then, we use the channel enhance module (CEM) to enhance the shallow features. CEM primarily relies on the channel attention mechanism and employs a no-downscaling strategy to produce local cross-channel interaction, which lowers model complexity while enhancing extraction performance. Meanwhile, we use the redefined feature module (RFM) to reconstruct the deep features and generate global attention features by dimensional transformation and feature relationship aggregation to achieve the role of locating salient targets. Finally, the cascade combines the multi-scale features to provide the final saliency map. To further enhance the representational power of the network, we use a hybrid loss function to improve performance. The proposed approach outperforms current state-of-the-art methods, as shown by several experiments on three available datasets. The source code of the proposed CRNet is available publicly at https://github.com/hilitteq/CRNet.git. Le Sun 0002, Yuwen Chen 0001, Yuhui Zheng, Zebin Wu 0001, Liyong Fu, Byeungwoo Jeon |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Contrastive Transformer Hashing for Compact Video RepresentationabstractVideo hashing learns compact representation by mapping video into low-dimensional Hamming space and has achieved promising performance in large-scale video retrieval. It is challenging to effectively exploit temporal and spatial structure in an unsupervised setting. To fulfill this gap, this paper proposes Contrastive Transformer Hashing (CTH) for effective video retrieval. Specifically, CTH develops a bidirectional transformer autoencoder, based on which visual reconstruction loss is proposed. CTH is more powerful to capture bidirectional correlations among frames than conventional unidirectional models. In addition, CTH devises multi-modality contrastive loss to reveal intrinsic structure among videos. CTH constructs inter-modality and intra-modality triplet sets and proposes multi-modality contrastive loss to exploit inter-modality and intra-modality similarities simultaneously. We perform video retrieval tasks on four benchmark datasets, i.e., UCF101, HMDB51, SVW30, FCVID using the learned compact hash representation, and extensive empirical results demonstrate the proposed CTH outperforms several state-of-the-art video hashing methods. Xiaobo Shen 0001, Yun-Hao Yuan 0001, Xichen Yang, Long Lan, Yuhui Zheng |
IEEE Trans. Image Process. | 6 |
| 2023 | Tensor Cascaded-Rank Minimization in Subspace: A Unified Regime for Hyperspectral Image Low-Level VisionabstractLow-rank tensor representation philosophy has enjoyed a reputation in many hyperspectral image (HSI) low-level vision applications, but previous studies often failed to comprehensively exploit the low-rank nature of HSI along different modes in low-dimensional subspace, and unsurprisingly handled only one specific task. To address these challenges, in this paper, we figured out that in addition to the spatial correlation, the spectral dependency of HSI also implicitly exists in the coefficient tensor of its subspace, this crucial dependency that was not fully utilized by previous studies yet can be effectively exploited in a cascaded manner. This led us to propose a unified subspace low-rank learning regime with a new tensor cascaded rank minimization, named STCR, to fully couple the low-rankness of HSI in different domains for various low-level vision tasks. Technically, the high-dimensional HSI was first projected into a low-dimensional tensor subspace, then a novel tensor low-cascaded-rank decomposition was designed to collapse the constructed tensor into three core tensors in succession to more thoroughly exploit the correlations in spatial, nonlocal, and spectral modes of the coefficient tensor. Next, difference continuity-regularization was introduced to learn a basis that more closely approximates the HSI's endmembers. The proposed regime realizes a comprehensive delineation of the self-portrait of HSI tensor. Extensive evaluations conducted with dozens of state-of-the-art (SOTA) baselines on eight datasets verified that the proposed regime is highly effective and robust to typical HSI low-level vision tasks, including denoising, compressive sensing reconstruction, inpainting, and destriping. The source code of our method is released at https://github.com/CX-He/STCR.git. Le Sun 0002, Chengxun He, Yuhui Zheng, Zebin Wu 0001, Byeungwoo Jeon |
IEEE Trans. Image Process. | 3 |
| 2023 | Multi-Biometric Unified Network for Cloth-Changing Person Re-IdentificationabstractPerson re-identification (re-ID) aims to match the same person across different cameras. However, most existing re-ID methods assume that people wear the same clothes in different views, which limit their performance in identifying target pedestrians who change clothes. Cloth-changing re-ID is a quite challenging problem as clothes occupying a large number of pixels in an image becomes invalid or even misleads information. To tackle this problem, we propose a novel Multi-biometric Unified Network (MBUNet) for learning the robustness of cloth-changing re-ID model by exploiting clothing-independent cues. Specifically, we first introduce a multi-biological feature branch to extract a variety of biological features, such as the head, neck, and shoulders to resist cloth-changing. Then, a differential feature attention module (DFAM) is embedded in this branch, which can extract discriminative fine-grained biological features. Besides, we design a differential recombination on max pooling (DRMP) strategy and simultaneously apply a direction-adaptive graph convolutional layer to mine more robust global and pose features. Finally, we propose a Lightweight Domain Adaptation Module (LDAM) that combines the attention mechanism before and after the waveblock to capture and enhance transferable features across scenarios. To further improve the performance of the model, we also integrate mAP optimization into the objective function of our model for joint training to solve the discrete optimization problem of mAP. Extensive experiments on five cloth-changing re-ID datasets demonstrate the advantages of our proposed MBUNet. The code is available at https://github.com/liyeabc/MBUNet. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
IEEE Trans. Image Process. | 4 |
| 2023 | Visibility and Distortion Measurement for No-Reference Dehazed Image Quality Assessment via Complex Contourlet TransformabstractRecently, most dehazed image quality assessment (DQA) methods have focused on estimating remaining haze and omitting distortion impact from the side effect of dehazing algorithms, which leads to their limited performance. Addressing this problem, we propose a method for learning both visibility and distortion-aware features no-reference (NR) dehazed image quality assessment (VDA-DQA). Visibility-aware features are exploited to characterize clarity optimization after dehazing, including the brightness-, contrast-, and sharpness-aware features extracted by the complex contourlet transform (CCT). Then, distortion-aware features are employed to measure the distortion artifacts of images, including the normalized histogram of the local binary pattern (LBP) from the reconstructed dehazed image and the statistics of the CCT subbands corresponding to the chroma and saturation map. Finally, all the above features are mapped into quality scores by support vector regression (SVR). Extensive experimental results on six public DQA datasets verify the superiority of the proposed VDA-DQA method in terms of consistency with subjective visual perception and outperform state-of-the-art methods. Tuxin Guan, Ke Gu 0001, Hantao Liu, Yuhui Zheng, Xiaojun Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Domain Adaptive Transformer Tracking Under OcclusionsabstractDue to their excellent performance on aggregating global features, Transformer structures are being widely employed in deep learning-based visual object tracking algorithms, recently. Nevertheless, existing Transformer-based trackers still fail to handle occlusion problems due to drift in feature distributions. To address this issue, we introduce domain adaptation techniques into a novel object tracking framework, DATransT, including feature extraction, domain adaptive Transformer module and prediction head. The domain adaptive Transformer module consists of three weight-sharing branches with self and cross attention mechanisms: the source, the target and the source-target branches. Specifically, the source-target branch employs cross-attention to effectively align the feature distributions of the source and target branches. Meanwhile, we present a pseudo-labeling strategy to generate high-quality training samples. Extensive experiments show that DATransT obtains promising results on several popular datasets, containing LaSOT, TrackingNet, GOT-10k, NfS, OTB2015 and UAV123. Moreover, our method outperforms existing state-of-the-art trackers under full occlusions and partial occlusions. Qianqian Yu 0001, Keqi Fan, Yuhui Zheng |
IEEE Trans. Multim. | 3 |
| 2022 | Multi-Biometric Unified Network for Cloth-Changing Person Re-IdentificationabstractPerson re-identification (re-ID) aims at matching the same person across different cameras. Most of the existing meth-ods for re- ID assume that people wear the same clothes on different cameras. However, Cloth-Changing re- ID is a quite challenging problem since people are likely to change clothes as the time span increases. To tackle this problem, a Multi-Biometric Unified Network (MBUNet) is proposed to ex-ploit clothing-unrelated cues. We firstly introduce a multi-biological feature branch that aims at extracting a variety of biological features, such as the head, neck, and shoulders to resist clothing changes. To extract discriminative fine-grained biological features, we embed a differential feature attention module (DFAM) for it. Besides, we adopt differ-ential recombination on max pooling (DRMP) and apply a direction-adaptive graph convolutional layer to extract more robust global features and pose features. Extensive experi-ments on three Cloth-Changing re-ID datasets show the ad-vantages of our proposed MBUNet. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
ICME | 4 |
| 2022 | TIPCB: A simple but effective part-based convolutional baseline for text-based person search
Yuhao Chen 0002, Guoqing Zhang 0002, Yujiang Lu, Yuhui Zheng |
Neurocomputing | 5 |
| 2022 | Close-set camera style distribution alignment for single camera person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
Neurocomputing | 4 |
| 2022 | DCA-CycleGAN: Unsupervised single image dehazing using Dark Channel Attention optimized CycleGAN
Yaozong Mo, Yuhui Zheng, Xiaojun Wu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Fine-grained-based multi-feature fusion for occluded person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | Dynamic Temporal-Spatial Regularization-Based Channel Weight Correlation Filter for Aerial Object TrackingabstractCorrelation filter (CF) has drawn extensive interest in aerial object tracking due to its remarkable performance. Recently, the popular CF methods based on temporal–spatial regularization have been proved to be able to effectively improve the tracking results. However, the boundary effect and filter template degradation still influence the speed and accuracy of the trackers. To handle the two problems, a novel dynamic temporal–spatial regularization-based channel weighted tracking (DTSCT) method was proposed in this work. First, we attempted to employ the saliency detection technique to describe object variation for weakening the boundary effect. Then, the filter template was introduced to the temporal regularization to alleviate the template degradation. In addition, an adaptive weighting strategy was utilized to remove data redundancy in the feature channels. Experiments on three benchmark datasets showed the competitive performance of our DTSCT approach compared to the state-of-the-art methods. Licheng Jiang, Yuhui Zheng, Xu Cheng 0003, Byeungwoo Jeon |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | No-reference stereoscopic image quality assessment on both complex contourlet and spatial domain via Kernel ELM
Tuxin Guan, Yuhui Zheng, Shenghu Zhao, Xiaojun Wu 0001 |
Signal Process. Image Commun. | 3 |
| 2022 | A Robust GAN-Generated Face Detection Method Based on Dual-Color Spaces and an Improved XceptionabstractIn recent years, generative adversarial networks (GANs) have been widely used to generate realistic fake face images, which can easily deceive human beings. To detect these images, some methods have been proposed. However, their detection performance will be degraded greatly when the testing samples are post-processed. In this paper, some experimental studies on detecting post-processed GAN-generated face images find that (a) both the luminance component and chrominance components play an important role, and (b) the RGB and YCbCr color spaces achieve better performance than the HSV and Lab color spaces. Therefore, to enhance the robustness, both the luminance component and chrominance components of dual-color spaces (RGB and YCbCr) are considered to utilize color information effectively. In addition, the convolutional block attention module and multilayer feature aggregation module are introduced into the Xception model to enhance its feature representation power and aggregate multilayer features, respectively. Finally, a robust dual-stream network is designed by integrating dual-color spaces RGB and YCbCr and using an improved Xception model. Experimental results demonstrate that our method outperforms some existing methods, especially in its robustness against different types of post-processing operations, such as JPEG compression, Gaussian blurring, gamma correction, and median filtering. Beijing Chen, Xin Liu 0012, Yuhui Zheng, Guoying Zhao 0001, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Unsupervised Multiview Distributed Hashing for Large-Scale RetrievalabstractMulti-view hashing (MvH) learns compact hash code by efficiently integrating multi-view data, and has achieved promising performance in large-scale retrieval task. In real-world applications, multi-view data is often stored or collected in different locations, and learning hash code in such case is more challenging yet less studied. In addition, unsupervised MvHs hardly achieve impressive retrieval performance due to absence of supervision. To fulfill this gap, this paper introduces a novel unsupervised multi-view distributed hashing (UMvDisH) to learn hash code from multi-view data, which is distributed in different nodes of a network. UMvDisH jointly performs latent factor model and spectral clustering to generate latent hash code and pseudo label respectively in each node. The consistency between hash code and pseudo label improves discrimination of hash code. The proposed distributed learning problem is divided into a set of decentralized subproblems by imposing local consistency among neighbor nodes. As such, the subproblems can be solved in parallel, and training time can be reduced. The communication cost is low due to no exchange of training data. Experimental results on four benchmark image datasets including a very large-scale image dataset show that UMvDisH achieves comparable retrieval performance and trains faster than state-of-the-art unsupervised MvHs in the distributed setting. Xiaobo Shen 0001, Yunpeng Tang, Yuhui Zheng, Yun-Hao Yuan 0001, Quan-Sen Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Detecting Aligned Double JPEG Compressed Color Image With Same Quantization Matrix Based on the Stability of ImageabstractJoint photographic experts group (JPEG) compression is widely used in image processing and computer vision. Detecting double compressed JPEG images is a common problem in forensics and detecting compressed images with the same quantization matrix remains a challenging task. However, most existing methods were designed for detection in grayscale images and cannot fully use the unique characteristics of color images (such as the relationship between channels and color information). In addition, the performance of existing methods is unsatisfactory for low JPEG quality factors and in cross detection experiments. To solve these problems, we analyze the stability of a color image to obtain the convergence error and transposition error. According to the convergence characteristics of color JPEG images, the continuous compression by the same quantization matrix can make the JPEG image tend to be stable. The final stable state and the convergence process are determined by the number of compressions of the original image. Thus, continuously compressed JPEG images can be regarded as a continuous frame to obtain the convergence error. As the color image converges, its ability to resist interference decreases. To reflect the changes in anti-interference ability, the transposition operation is used to disturb the color JPEG image to obtain the transposition error. In addition, quaternion mapping is used to retain the relationship between continuously compressed JPEG images and enlarge the influence caused by transposition operation. In our experiments on several image databases, the proposed method outperforms existing methods in different settings. Hao Wang 0060, Xiangyang Luo 0001, Yuhui Zheng, Bin Ma 0003, Jinsheng Sun, Sunil Kr. Jha |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Illumination Unification for Person Re-IdentificationabstractThe performance of person re-identification (re-ID) is easily affected by illumination variations caused by different shooting times, places and cameras. Existing illumination-adaptive methods usually require annotating cross-camera pedestrians on each illumination scale, which is unaffordable for a long-term person retrieval system. The cross-illumination person retrieval problem presents a great challenge for accurate person matching. In this paper, we propose a novel method to tackle this task, which only needs to annotate pedestrians on one illumination scale. Specifically, (i) we propose a novel Illumination Estimation and Restoring framework (IER) to estimate the illumination scale of testing images taken at different illumination conditions and restore them to the illumination scale of training images, such that the disparities between training images with uniform illumination and testing images with varying illuminations are reduced. IER achieves promising results on illumination-adaptive dataset and proving itself a proper baseline for cross-illumination person re-ID. (ii) we propose a Mixed Training strategy using both Original and Reconstructed images (MTOR) to further improve model performance. We generate reconstructed images that are consistent with the original training images in content but more similar to the restored images in style. The reconstructed images are combined with the original training images for supervised training to further reduce the domain gap between original training images and restored testing images. To verify the effectiveness of our method, some simulated illumination-adaptive datasets are constructed with various illumination conditions. Extensive experimental results on the simulated datasets validate the effectiveness of the proposed method. The source code is available athttps://github.com/FadeOrigin/IUReId. Guoqing Zhang 0002, Zhiyuan Luo 0003, Yuhao Chen 0002, Yuhui Zheng, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Global Relation-Aware Contrast Learning for Unsupervised Person Re-IdentificationabstractThe goal of unsupervised person re-identification (Re-ID) is to use unlabeled person images to learn discriminative features. In recent years, many approaches have adopted clustered pseudo labels to construct proxies for contrastive learning, and have thereby achieved great success. However, existing methods of this kind only utilize local structures within IDs to design their proxies while ignoring the relations between samples of different IDs, which limits the improvement for inter-ID discriminative ability. To resolve this issue, we propose a Global Relation-Aware Contrast Learning (GRACL) method for the task of unsupervised Re-ID. Our method first sets up two proxies for each cluster to capture the inter- and intra-ID relations respectively, which enables us to both effectively increase inter-ID variances and reduce the intra-ID discrepancies. Specifically, the samples that are most different from those in different clusters are selected as inter-ID relation-aware proxies, while those that are least similar to samples from the same clusters are employed as intra-ID relation-aware proxies. With the aid of these proxies, we design both inter- and intra-ID relation-aware contrastive learning modules to facilitate model learning. By pulling each sample close to the positive proxy, we can obtain identity-invariant discriminative features. Experiments on five widely-used Re-ID datasets prove that our GRACL model outperforms current state-of-the-art approaches to a remarkable extent. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Multi-Task Convolution Operators With Object Detection for Visual TrackingabstractRecently, multi-task correlation filters has drawn much attention in the object tracking field, which utilizes the multi-task learning (MTL) approach to explore the interdependencies among deep features for object tracking. However, the existing multi-task correlation filters based method fails to consider the relations between the correlation filters. To address this problem, a novel correlation filters based visual tracking method is proposed in this paper, with the integration of multi-task convolution operators and object detection. In our method, convolution and correction filters are jointly learnt through using the MTL technique, with the purpose of exploring not only the interdependencies of deep features but also the internal relevance of the convolution filters. In addition, object detection is introduced into our algorithm to handle the problem of object missing to ensure a better performance of our tracking method. Experiments on five benchmark datasets demonstrate that the proposed visual tracking method outperforms existing state-of-the-art approaches. Yuhui Zheng, Xinyan Liu 0002, Bin Xiao 0002, Xu Cheng 0003, Yi Wu 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multiview Learning With Robust Double-Sided Twin SVMabstractMultiview learning (MVL), which enhances the learners' performance by coordinating complementarity and consistency among different views, has attracted much attention. The multiview generalized eigenvalue proximal support vector machine (MvGSVM) is a recently proposed effective binary classification method, which introduces the concept of MVL into the classical generalized eigenvalue proximal support vector machine (GEPSVM). However, this approach cannot guarantee good classification performance and robustness yet. In this article, we develop multiview robust double-sided twin SVM (MvRDTSVM) with SVM-type problems, which introduces a set of double-sided constraints into the proposed model to promote classification performance. To improve the robustness of MvRDTSVM against outliers, we take L1-norm as the distance metric. Also, a fast version of MvRDTSVM (called MvFRDTSVM) is further presented. The reformulated problems are complex, and solving them are very challenging. As one of the main contributions of this article, we design two effective iterative algorithms to optimize the proposed nonconvex problems and then conduct theoretical analysis on the algorithms. The experimental results verify the effectiveness of our proposed methods. Qiaolin Ye, Zhao Zhang 0001, Yuhui Zheng, Liyong Fu, Wankou Yang |
IEEE Trans. Cybern. | 4 |
| 2022 | Spectral-Spatial Feature Tokenization Transformer for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification, each pixel sample is assigned to a land-cover category. In the recent past, convolutional neural network (CNN)-based HSI classification methods have greatly improved performance due to their superior ability to represent features. However, these methods have limited ability to obtain deep semantic features, and as the layer’s number increases, computational costs rise significantly. The transformer framework can represent high-level semantic features well. In this article, a spectral–spatial feature tokenization transformer (SSFTT) method is proposed to capture spectral–spatial features and high-level semantic features. First, a spectral–spatial feature extraction module is built to extract low-level features. This module is composed of a 3-D convolution layer and a 2-D convolution layer, which are used to extract the shallow spectral and spatial features. Second, a Gaussian weighted feature tokenizer is introduced for features transformation. Third, the transformed features are input into the transformer encoder module for feature representation and learning. Finally, a linear layer is used to identify the first learnable token to obtain the sample label. Using three standard datasets, experimental analysis confirms that the computation time is less than other deep learning methods and the performance of the classification outperforms several current state-of-the-art methods. The code of this work is available athttps://github.com/zgr6010/HSI_SSFTTfor the sake of reproducibility. Le Sun 0002, Guangrui Zhao, Yuhui Zheng, Zebin Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Deep Co-Image-Label Hashing for Multi-Label Image RetrievalabstractDeep supervised hashing has greatly improved retrieval performance with the powerful learning capability of deep neural network. In multi-label image retrieval, existing deep hashing simply indicates whether two images are similar by constructing a similarity matrix. However, it ignores the dependency among multiple labels that has been shown important in multi-label application. To fulfill this gap, this paper proposes Deep Co-Image-Label Hashing (DCILH) to discover label dependency. Specifically, DCILH regards image and label as two views, and maps the two views into a common deep Hamming space. DCILH proposes to learn prototype for each label, and preserve similarity among images, labels, and prototypes. To exploit label dependency, DCILH further employs the label-correlation aware loss on the predicted labels, such that predicted output on positive label is enforced to be larger than that on negative label. Extensive experiments on several multi-label benchmarks demonstrate the proposed DCILH outperforms state-of-the-art deep supervised hashing on large-scale multi-label image retrieval. Xiaobo Shen 0001, Guohua Dong, Yuhui Zheng, Long Lan, Ivor W. Tsang, Quan-Sen Sun |
IEEE Trans. Multim. | 3 |
| 2022 | SmsNet: A New Deep Convolutional Neural Network Model for Adversarial Example DetectionabstractThe emergence of adversarial examples has had a significant impact on the development and application of deep learning. In this paper, a novel convolutional neural network model, the stochastic multifilter statistical network (SmsNet), is proposed for the detection of adversarial examples. A feature statistical layer is constructed to collect statistical data of feature map output from each convolutional layer in SmsNet by combining manual features with a neural network. The entire model is an end-to-end detection model, so the feature statistical layer is not independent of the network, and its output is directly transmitted to the fully connected layer by a short-cut connection called the SmsConnection. Additionally, a dynamic pruning strategy is introduced to simplify the model structure for better performance. The experiments demonstrate the effectiveness of the network structure and pruning strategy, and the proposed model achieves high detection rates against state-of-the-art adversarial attacks. Qilin Yin, Xiangyang Luo 0001, Yuhui Zheng, Yun Q. Shi 0001, Sunil Kr. Jha |
IEEE Trans. Multim. | 5 |
| 2021 | Reference-Aided Part-Aligned Feature Disentangling for Video Person Re-IdentificationabstractRecently, video-based person re-identification (re-ID) has drawn increasing attention in compute vision community because of its practical application prospects. Due to the inaccurate person detections and pose changes, pedestrian misalignment significantly increases the difficulty of feature extraction and matching. To address this problem, in this paper, we propose a Reference-Aided Part-Aligned (RAPA) framework to disentangle robust features of different parts. Firstly, in order to obtain better references between different videos, a pose-based reference feature learning module is introduced. Secondly, an effective relation-based part feature disentangling module is explored to align frames within each video. By means of using both modules, the informative parts of pedestrian in videos are well aligned and more discriminative feature representation is generated. Comprehensive experiments on three widely-used benchmarks, i.e. iLIDS-VID, PRID-2011 and MARS datasets verify the effectiveness of the proposed framework. Our code will be made publicly available. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng, Yi Wu 0001 |
ICME | 4 |
| 2021 | Completely blind image quality assessment via contourlet energy statisticsabstractAbstract An aim of completely blind image quality assessment (BIQA) is to develop algorithms which can grade image quality without any prior knowledge of the images. Here, a new contourlet energy statistics based completely on blind opinion‐unaware BIQA (OU‐BIQA) method is proposed, which can predict the perceptual severity of a range of image distortion types without requiring any prior knowledge. According to the energy distribution of the contourlet sub‐bands of natural images in log‐domain, the lower‐scale sub‐band energy can be predicted by the corresponding higher‐scale sub‐band energies of distorted images. A quality model is then constructed by quantifying the difference between predicted energy and realistic energy. Meanwhile, an effective method for adjusting and compensating an undesired distortion is integrated into the quality model. Experimental results show that the proposed new method outperforms state‐of‐the‐art OU‐BIQA models on relevant portions of TID2013 database, and is competitive on the LIVE IQA database. Moreover, the proposed model is very fast, suggesting a real‐time solution to high‐performance BIQA. Tuxin Guan, Yuhui Zheng, Bo Jin 0001, Xiaojun Wu 0001, Alan C. Bovik |
IET Image Process. | 3 |
| 2021 | Locally GAN-generated face detection based on an improved Xception
Beijing Chen, Xingwang Ju, Bin Xiao 0002, Weiping Ding 0001, Yuhui Zheng, Victor Hugo C. de Albuquerque |
Inf. Sci. | 5 |
| 2021 | Optimal discriminative feature and dictionary learning for image set classification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Zhiyuan Luo 0003 |
Inf. Sci. | 3 |
| 2021 | Hybrid-attention guided network with multiple resolution features for person re-identification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Yi Wu 0001, Shengyong Chen |
Inf. Sci. | 3 |
| 2021 | Cross-view kernel collaborative representation classification for person re-identification
Guoqing Zhang 0002, Tong Jiang, Junchuan Yang, Yuhui Zheng |
Multim. Tools Appl. | 5 |
| 2021 | TSLRLN: Tensor subspace low-rank learning with non-local prior for hyperspectral image mixed denoising
Chengxun He, Le Sun 0002, Wei Huang 0013, Jianwei Zhang 0005, Yuhui Zheng, Byeungwoo Jeon |
Signal Process. | 5 |
| 2021 | Blind image quality assessment in the contourlet domain
Tuxin Guan, Yuhui Zheng, Xiaochun Zhong, Xiaojun Wu 0001, Alan C. Bovik |
Signal Process. Image Commun. | 3 |
| 2021 | Weighted LIC-Based Structure Tensor With Application to Image Content Perception and ProcessingabstractAs a famous visual content perception and processing tool, structure tensor has been widely studied in the past decades. Among them, the anisotropic nonlocal structure tensor (ANLST) has received much attention, recently. However, the existing ANLST calculation methods fail to fully utilize the anisotropic characteristic of the tensor field, thus resulting in limited performance. For this problem, in this article, we present a novel ANLST construction method, by means of combining tensor decomposition with weighted line integral convolution (LIC) with the aim at deeply discovering and exploiting the spatial direction relevancy of the tensors for their regularization. At first, the tensors decomposition, computed by direction projection, yields multiple atomic vector fields, from which, for each point in the tensor field we obtain a family of integral curves that are associated with spatial direction related tensors. Then, LIC is employed with the nonlocal means filtering to smooth the tensors relevant to each integral curve, giving rise to curve-level structure tensor (CLST). At last, a weighted average scheme is carried out on the multiple CLSTs, leading to our proposed weighted anisotropic nonlocal structure tensor (WANST). Experimental results demonstrate that the proposed WANST is superior to the current representative nonlinear structure tensors. The proposed WANST can be applied to industrial surveillance system to enable it perceive image contents, such as flat regions, corners, textures, and edges. In addition, WANST can also help monitoring system improve its image quality. Yuhui Zheng, Yahui Sun 0003, Khan Muhammad 0001, Victor Hugo C. de Albuquerque |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Consistency Graph Modeling for Semantic CorrespondenceabstractTo establish robust semantic correspondence between images covering different objects belonging to the same category, there are three important types of information including inter-image relationship, intra-image relationship and cycle consistency. Most existing methods only exploit one or two types of the above information and cannot make them enhance and complement each other. Different from existing methods, we propose a novel end-to-end Consistency Graph Modeling Network (CGMNet) for semantic correspondence by modeling inter-image relationship, intra-image relationship and cycle consistency jointly in a unified deep model. The proposed CGMNet enjoys several merits. First, to the best of our knowledge, this is the first work to jointly model the three kinds of information in a deep model for semantic correspondence. Second, our model has designed three effective modules including cross-graph module, intra-graph module and cycle consistency module, which can jointly learn more discriminative feature representations robust to local ambiguities and background clutter for semantic correspondence. Extensive experimental results show that our algorithm performs favorably against state-of-the-art methods on four challenging datasets including PF-PASCAL, PF-WILLOW, Caltech-101 and TSS. Tianzhu Zhang 0001, Yuhui Zheng, Mingliang Xu 0001, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Deep High-Resolution Representation Learning for Cross-Resolution Person Re-IdentificationabstractPerson re-identification (re-ID) tackles the problem of matching person images with the same identity from different cameras. In practical applications, due to the differences in camera performance and distance between cameras and persons of interest, captured person images usually have various resolutions. This problem, named Cross-Resolution Person Re-identification, presents a great challenge for the accurate person matching. In this paper, we propose a Deep High-Resolution Pseudo-Siamese Framework (PS-HRNet) to solve the above problem. Specifically, we first improve the VDSR by introducing existing channel attention (CA) mechanism and harvest a new module, i.e., VDSR-CA, to restore the resolution of low-resolution images and make full use of the different channel information of feature maps. Then we reform the HRNet by designing a novel representation head, HRNet-ReID, to extract discriminating features. In addition, a pseudo-siamese framework is developed to reduce the difference of feature distributions between low-resolution images and high-resolution images. The experimental results on five cross-resolution person datasets verify the effectiveness of our proposed approach. Compared with the state-of-the-art methods, the proposed PS-HRNet improves the Rank-1 accuracy by 3.4%, 6.2%, 2.5%,1.1% and 4.2% on MLR-Market-1501, MLR-CUHK03, MLR-VIPeR, MLR-DukeMTMC-reID, and CAVIAR datasets, respectively, which demonstrates the superiority of our method in handling the Cross-Resolution Person Re-ID task. Our code is available at https://github.com/zhguoqing. Guoqing Zhang 0002, Zhicheng Dong 0001, Hao Wang 0101, Yuhui Zheng, Shengyong Chen |
IEEE Trans. Image Process. | 5 |
| 2021 | A Serial Image Copy-Move Forgery Localization Scheme With Source/Target DistinguishmentabstractIn this paper, we improve the parallel deep neural network (DNN) scheme BusterNet for image copy-move forgery localization with source/target region distinguishment. BusterNet is based on two branches, i.e., Simi-Det and Mani-Det, and suffers from two main drawbacks: (a) it should ensure that both branches correctly locate regions; (b) the Simi-Det branch only extracts single-level and low-resolution features using VGG16 with four pooling layers. To ensure the identification of the source and target regions, we introduce two subnetworks that are constructed serially: the copy-move similarity detection network (CMSDNet) and the source/target region distinguishment network (STRDNet). Regarding the second drawback, the CMSDNet subnetwork improves Simi-Det by removing the last pooling layer in VGG16 and by introducing atrous convolution into VGG16 to preserve field-of-views of filters after the removal of the fourth pooling layer; double-level self-correlation is also considered for matching hierarchical features. Moreover, atrous spatial pyramid pooling and attention mechanism allow the capture of multiscale features and provide evidence for important information. Finally, STRDNet is designed to determine the similar regions obtained from CMSDNet directly as tampered regions and untampered regions. It determines regions at the image-level rather than at the pixel-level as made by Mani-Det of BusterNet. Experimental results on four publicly available datasets (new synthetic dataset, CASIA, CoMoFoD, and COVERAGE) demonstrate that the proposed algorithm is superior to the state-of-the-art algorithms in terms of similarity detection ability and source/target distinguishment ability. Beijing Chen, Weijin Tan, Gouenou Coatrieux, Yuhui Zheng, Yun Q. Shi 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | JSNet: A simulation network of JPEG lossy compression and restoration for robust image watermarking against JPEG attack
Beijing Chen, Yunqing Wu, Gouenou Coatrieux, Yuhui Zheng |
Comput. Vis. Image Underst. | 5 |
| 2020 | Cost-sensitive joint feature and dictionary learning for face recognition
Guoqing Zhang 0002, Fatih Porikli, Huaijiang Sun, Quan-Sen Sun, Guiyu Xia, Yuhui Zheng |
Neurocomputing | 6 |
| 2020 | Image splicing localization using residual image and residual-based fully convolutional network
Beijing Chen, Xiaoming Qi, Guanyu Yang 0001, Yuhui Zheng, Bin Xiao 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2020 | Two stages double attention convolutional neural network for crowd counting
Zhao Zou, Yuhui Zheng, Shoukun Xu |
Multim. Tools Appl. | 3 |
| 2020 | Low Rank Component Induced Spatial-Spectral Kernel Method for Hyperspectral Image ClassificationabstractKernel methods, e.g., composite kernels (CKs) and spatial-spectral kernels (SSKs), have been demonstrated to be an effective way to exploit the spatial-spectral information nonlinearly for improving the classification performance of hyperspectral image (HSI). However, these methods are always conducted with square-shaped window or superpixel techniques. Both techniques are likely to misclassify the pixels that lie at the boundaries of class, and thus a small target is always smoothed away. To alleviate these problems, in this paper, we propose a novel patch-based low rank component induced spatial-spectral kernel method, termed LRCISSK, for HSI classification. First, the latent low-rank features of spectra in each cubic patch of HSI are reconstructed by a low rank matrix recovery (LRMR) technique, and then, to further explore more accurate spatial information, they are used to identify a homogeneous neighborhood for the target pixel (i.e., the centroid pixel) adaptively. Finally, the adaptively identified homogenous neighborhood which consists of the latent low-rank spectra is embedded into the spatial-spectral kernel framework. It can easily map the spectra into the nonlinearly complex manifolds and enable a classifier (e.g., support vector machine, SVM) to distinguish them effectively. Experimental results on three real HSI datasets validate that the proposed LRCISSK method can effectively explore the spatial-spectral information and deliver superior performance with at least 1.30% higher OA and 1.03% higher AA on average when compared to other state-of-the-art classifiers. Le Sun 0002, Yuhui Zheng, Hiuk Jae Shim, Zebin Wu 0001, Byeungwoo Jeon |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Optimal Discriminative Projection for Sparse Representation-Based Classification via Bilevel OptimizationabstractRecently, sparse representation-based classification (SRC) has been widely studied and has produced state-of-the-art results in various classification tasks. Learning useful and computationally convenient representations from complex redundant and highly variable visual data is crucial for the success of SRC. However, how to find the best feature representation to work with SRC remains an open question. In this paper, we present a novel discriminative projection learning approach with the objective of seeking a projection matrix such that the learned low-dimensional representation can fit SRC well and that it has well discriminant ability. More specifically, we formulate the learning algorithm as a bilevel optimization problem, where the optimization includes an ℓ1-norm minimization problem in its constraints. Through the bilevel optimization model, the relationship between sparse representation and the desired feature projection can be explicitly exploited during the learning process. Therefore, SRC can achieve a better performance in the transformed subspace. The optimization model can be solved by using a stochastic gradient ascent algorithm, and the desired gradient is computed using implicit differentiation. Furthermore, our method can be easily extended to learn a dictionary. The extensive experimental results on a series of benchmark databases show that our method outperforms many state-of-the-art algorithms. Guoqing Zhang 0002, Huaijiang Sun, Yuhui Zheng, Guiyu Xia, Lei Feng 0003, Quan-Sen Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Multi-Task Deep Dual Correlation Filters for Visual TrackingabstractCorrelation filters combined with deep features have delivered impressive results in visual tracking task. However, existing approaches treat deep features produced by different network layers independently, limiting their representation power. To address this issue, this paper proposes a multi-task deep dual correlation filters (MDDCF) based method for robust visual tracking. First, a new multi-task learning scheme is designed to take full advantage of the multi-level features of deep networks, where target representation with individual features is regarded as a single task. As such, the interdependencies between different levels of features can be better explored. Second, we reformulate the objective function of the dual correlation filters and propose a new alternating optimization method, allowing joint training of the correlation filters and network parameters. Third, we design an effective object template update scheme which can well capture the target appearance variations. Extensive experimental evaluations on seven benchmark datasets show that the proposed MDDCF tracker performs favorably against state-ofthe-art methods. Yuhui Zheng, Xinyan Liu 0002, Xu Cheng 0003, Kaihua Zhang 0001, Yi Wu 0001, Shengyong Chen |
IEEE Trans. Image Process. | 1 |
| 2020 | Dynamically Spatiotemporal Regularized Correlation TrackingabstractRecently, due to the high performance, spatially regularized strategy has been widely applied to addressing the issue of boundary effects existed in correlation filter (CF)-based visual tracking. Specifically, it introduces a spatially regularized term to penalize the coefficients of the CFs to be learned depending on their spatial locations. However, the regularization weights are often formed as a fixed Gaussian function, and hence may cause the learned model degenerate due to the inflexible constraints on the ever-changing CFs to be learned over time during tracking. To address this issue, in this paper, we develop a dynamically spatiotemporal regularization model to constrain the CFs to be learned with the ever-changing regularization weights learned from two consecutive frames. The proposed method jointly learns the CFs along with the dynamically spatiotemporal constraint term, which can be efficiently solved in the Fourier domain by the alternative direction method. Extensive evaluations on the popular data sets OTB-100 and VOT-2016 demonstrate that the proposed tracker performs favorably against the baseline tracker and several recently proposed state-of-the-art methods. Yuhui Zheng, Huihui Song 0003, Kaihua Zhang 0001, Jiaqing Fan, Xinyan Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | A Novel Light-Weight Subjective Trust Inference Framework in MANETsabstractThere is an inherent reliance on collaboration among the participants of mobile ad hoc networks in order to achieve the fixed functionalities. However, they are susceptible to the destruction of the malicious attacks or denial of cooperation. Therefore, it becomes obvious that the security issue is urgently needed to be addressed. Over the last few years, many trust-considered countermeasures have been proposed. The design of trust quantification methods is the key of these countermeasures. In this study, we abstract a novel light-weight subjective trust inference framework, which is divided into trust assessment and trust prediction. The process of node trust assessment is based on node's historical behaviours. Then utilizing the obtained trust data sequence, we introduce the SCGM(1,1)-weighted Markov stochastic chain measure to predict node's trust for future decision making. Experimental results have been conducted to evaluate the effectiveness of the proposed trust model. As an important security application, based on the standard On-Demand Multicast Routing Protocol (ODMRP), we make four major improvements which take the issue of trust into consideration, and propose a novel trust-based routing protocol called the On-Demand Trust-Based Multicast Routing protocol (ODTMRP). And finally, convincing experimental results are presented using three routing evaluation metrics. Hui Xia 0001, Zhetao Li, Yuhui Zheng, Anfeng Liu, Young-June Choi, Hiroo Sekiya |
IEEE Trans. Sustain. Comput. | 3 |
| 2019 | Glaucoma Progression Prediction Using Retinal Thickness via Latent Space Linear RegressionabstractPrediction of glaucomatous visual field loss has significant clinical benefits because it can help with early detection of glaucoma as well as decision-making for treatments. Glaucomatous visual loss is conventionally captured through visual field sensitivity (VF ) measurement, which is costly and time-consuming. Thus, existing approaches mainly predict future VF utilizing limited VF data collected in the past. Recently, optical coherence tomography (OCT) has been adopted to measure retinal layers thickness (RT ) for considerably more low-cost treatment assistance. There then arises an important question in the context of ophthalmology: are RT measurements beneficial for VF prediction? In this paper, we propose a novel method to demonstrate the benefits provided by RT measurements. The challenge is management of the two heterogeneities of VF data and RT data as RT data are collected according to different clinical schedules and lie in a different space to VF data. To tackle these heterogeneities, we propose latent progression patterns (LPPs), a novel type of representations for glaucoma progression. Along with LPPs, we propose a method to transform VF series to an LPP based on matrix factorization and a method to transform RT series to an LPP based on deep neural networks. Partial VF and RT information is integrated in LPPs to provide accurate prediction. The proposed framework is named deeply-regularized latent-space linear regression (\em DLLR). We empirically demonstrate that our proposed method outperforms the state-of-the-art technique by 12% for the best case in terms of the mean of the root mean square error on a real dataset. Yuhui Zheng, Linchuan Xu, Taichi Kiwaki, Jing Wang 0023, Hiroshi Murata, Ryo Asaoka, Kenji Yamanishi |
KDD | 1 |
| 2019 | Domain adaptive collaborative representation based classification
Guoqing Zhang 0002, Yuhui Zheng, Guiyu Xia |
Multim. Tools Appl. | 2 |
| 2019 | Multi-Kernel Coupled Projections for Domain Adaptive Dictionary LearningabstractDictionary learning has produced state-of-the-art results in various classification tasks. However, if the training data have a different distribution than the testing data, the learned sparse representation might not be optimal. Recently, several domain-adaptive dictionary learning (DADL) methods and kernels have been proposed and have achieved impressive performance. However, the performance of these single kernel-based methods heavily depends heavily on the choice of the kernel, and the question of how to combine multiple kernel learning (MKL) with the DADL framework has not been well studied. Motivated by these concerns, in this paper, we propose a multi-kernel domain-adaptive sparse representation-based classification (MK-DASRC) and then use it as a criterion to design a multi-kernel sparse representation-based domain-adaptive discriminative projection method, in which the discriminative features of the data in the two domains are simultaneously learned with the dictionary. The purpose of this method is to maximize the between-class sparse reconstruction residuals of data from both domains, and minimize the within-class sparse reconstruction residuals of data in the low-dimensional subspace. Thus, the resulting representations can satisfactorily fit MK-DASRC and simultaneously display discriminability. Extensive experimental results on a series of benchmark databases show that our method performs better than the state-of-the-art methods. Yuhui Zheng, Guoqing Zhang 0002, Baihua Xiao, Fu Xiao 0001, Jianwei Zhang 0005 |
IEEE Trans. Multim. | 1 |
| 2019 | Spatially Regularized Structural Support Vector Machine for Robust Visual TrackingabstractStructural support vector machine (SSVM) is popular in the visual tracking field as it provides a consistent target representation for both learning and detection. However, the spatial distribution of feature is not considered in standard SSVM-based trackers, therefore leading to limited performance. To obtain a robust discriminative classifier, this paper proposes a novel tracking framework that spatially regularizes SSVM, which yields a new spatially regularized SSVM (SRSSVM). We utilize the spatial regularization prior to penalize the learning classifier with the same size as the target region. The location of classifier spatially located far from the center of region is assigned large weight and vice versa. Then, it is introduced into the SSVM model as a regularization factor to learn the robust discriminative model. Furthermore, an optimizing algorithm with dual coordination descent is presented to efficiently solve the SRSSVM tracking model. Our proposed SRSSVM tracking method has low computational cost like the traditional linear SSVM tracker while can significantly improve the robustness of the discriminative classifier. The experimental results on three popular tracking benchmark data sets show that the proposed SRSSVM tracking method performs favorably against the state-of-the-art trackers. Yuhui Zheng, Le Sun 0002, Shunfeng Wang, Jianwei Zhang 0005, Jifeng Ning |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Joint Bayesian guided metric learning for end-to-end face verification
Chunyan Xu, Jian Yang 0003, Jianjun Qian, Yuhui Zheng, LinLin Shen |
Neurocomputing | 5 |
| 2018 | Scalable transfer support vector machine with group probabilities
Tongguang Ni, Xiaoqing Gu, Jun Wang 0024, Yuhui Zheng |
Neurocomputing | 4 |
| 2018 | Sparse regression with output correlation for cardiac ejection fraction estimation
Bin Gu 0001, Yingying Shan, Victor S. Sheng, Yuhui Zheng, Shuo Li 0001 |
Inf. Sci. | 4 |
| 2018 | Quaternion discrete fractional random transform for color image adaptive watermarking
Beijing Chen, Chunfei Zhou, Byeungwoo Jeon, Yuhui Zheng |
Multim. Tools Appl. | 4 |
| 2018 | Few-shot learning for short text classification
Leiming Yan, Yuhui Zheng, Jie Cao 0011 |
Multim. Tools Appl. | 2 |
| 2018 | Selection of regularization parameter in GMM based image denoising method
Yuhui Zheng, Jianwei Zhang 0005, Jin Wang 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Student's t-Hidden Markov Model for Unsupervised Learning Using Localized Feature SelectionabstractRecently, the hidden Markov model (HMM) with student’s t-mixture model (SMM), called student’s t-HMM (SHMM) for short, has received much attention in unsupervised learning of sequential data. However, the current existing SHMMs fail to take into consideration of the relevant features embedded in local subspaces, thus influencing their performances in clustering. To address the problem, a novel SHMM is proposed by combining the measure of localized feature saliency (LFS) with SMM and utilizing two student’s t-distributions as subcomponents to respectively describe the distributions of useful features and non-salient “features,” with the purpose of accurately modeling the hidden state observation emission distributions of SHMM. Moreover, we exploit the variational Bayesian learning technique to simultaneously estimate the LFS, the number of components and other parameters of the herein proposed SHMM. Experimental results on both synthetic and real data sets demonstrate the improved robustness, effectiveness, and accuracy of our model. Yuhui Zheng, Byeungwoo Jeon, Le Sun 0002, Jianwei Zhang 0005, Hui Zhang 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Homogeneous region based low rank representation in hidden field for hyperspectral classificationabstractIn this paper, a new classifier under Bayesian framework is proposed to explore homogeneous region based low rank representation in hidden field for classification of hyperspectral imagery (HSI). This classifier integrates low rank representation and superpixel segmentation simultaneously, in which the HSI data is assumed to be lying in a low rank subspace within each homogeneous region of an estimated hidden field. First, the HSI data is projected into the Principal Component space, then the first principal component image is segmented into hundreds of homogeneous regions. Following, the spectral-only supervised Bayesian classifier, i.e., Sparse Multinomial Logistic Regression (SMLR), is utilized for estimating the likelihood probabilities of testing samples, then spatial information is exploited by low rank representation within each superpixel in a hidden field which is approximated to the pre-estimated likelihood probabilities. The proposed model can be easily solved by alternating direction method of multipliers (ADMM). Experimental results on real hyperspectral data, i.e., AVIRIS Indian Pines and ROSIS University of Pavia, show that the proposed classifier outperforms other state-of-the-art classifiers in terms of quantitative assessment and visual effect. Le Sun 0002, Byeungwoo Jeon, Yuhui Zheng, Yang Xu 0006, Zebin Wu 0001 |
IGARSS | 3 |
| 2017 | Image segmentation using a hierarchical student's-t mixture modelabstractAs a significant tool, finite mixture models (FMMs) have been widely used for image segmentation. However, there are two problems with standard FMMs: first, the conditional probability is sensitive to outliers. Second, the robustness to image noise is inadequate. In this study, the authors present a novel hierarchical Student's‐ t MM (HSMM), which includes standard FMMs as a sub‐problem. Additionally, to incorporate more image spatial information, they apply a mean template not only to the prior/posterior probability, but also to the sub‐conditional distribution. Thus, their HSMM is more robust to outliers and image noise owing to the spatial constraints from the mean template. In the standard SMM, a t ‐distribution is used to calculate the conditional probability. In this study, the authors present a novel hierarchical student's‐ t mixture model (HSMM), which includes the standard FMM as a sub‐problem. Finally, though they use Student's‐ t ‐distribution to solve the image segment problems of this study, their HSMM achieves excellent performance, is elastic and can encompass any other model that is based on FMMs. Experimental results demonstrate that their proposed method is robust and effective. Lingcheng Kong, Hui Zhang 0015, Yuhui Zheng, Jiezhong Zhu, Q. M. Jonathan Wu |
IET Image Process. | 3 |
| 2017 | A robust modified Gaussian mixture model with rough set for image segmentation
Zexuan Ji, Yong Xia 0001, Yuhui Zheng |
Neurocomputing | 4 |
| 2017 | Dynamic dictionary optimization for sparse-representation-based face classification using local difference images
Chang-Bin Shao, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001, Yuhui Zheng |
Inf. Sci. | 5 |
| 2017 | Hyperspectral Image Restoration Using Low-Rank Representation on Spectral Difference ImageabstractThis letter presents a novel mixed noise (i.e., Gaussian, impulse, stripe noises, or dead lines) reduction method for hyperspectral image (HSI) by utilizing low-rank representation (LRR) on spectral difference image. The proposed method is based on the assumption that all spectra in the spectral difference space of HSI lie in the same low-rank subspace. The LRR on the spectral difference space was exploited by nuclear norm of difference image along the spectral dimension. It showed great potential in removing structured sparse noise (e.g., stripes or dead lines located at the same place of each band) and heavy Gaussian noise. To simultaneously solve the proposed model and reduce computational load, alternating direction method of multipliers was utilized to achieve robust reconstruction. The experimental results on both simulated and real HSI data sets validated that the proposed method outperformed many state-of-the-art methods in terms of quantitative assessment and visual quality. Le Sun 0002, Byeungwoo Jeon, Yuhui Zheng, Zebin Wu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Gaussian mixture model learning based image denoising method with adaptive regularization parameters
Jianwei Zhang 0005, Tong Li 0021, Yuhui Zheng, Jin Wang 0001 |
Multim. Tools Appl. | 4 |
| 2017 | Efficient Anonymous Authenticated Key Agreement Scheme for Wireless Body Area NetworksabstractWireless body area networks (WBANs) are widely used in telemedicine, which can be utilized for real-time patients monitoring and home health-care. The sensor nodes in WBANs collect the client’s physiological data and transmit it to the medical center. However, the clients’ personal information is sensitive and there are many security threats in the extra-body communication. Therefore, the security and privacy of client’s physiological data need to be ensured. Many authentication protocols for WBANs have been proposed in recent years. However, the existing protocols fail to consider the key update phase. In this paper, we propose an efficient authenticated key agreement scheme for WBANs and add the key update phase to enhance the security of the proposed scheme. In addition, session keys are generated during the registration phase and kept secretly, thus reducing computation cost in the authentication phase. The performance analysis demonstrates that our scheme is more efficient than the currently popular related schemes. Tong Li 0021, Yuhui Zheng, Ti Zhou |
Secur. Commun. Networks | 2 |
| 2016 | Hyperspectral unmixing based on L1-L2 sparsity and total variationabstractThis paper proposes a novel linear hyperspectral unmixing method based on l1-l2sparsity and total variation (TV) regularization. First, the enhanced sparsity based on l1-l2norm is explored to depict the intrinsic sparse characteristic of the fractional abundances in sparse regression unmixing model. By taking the correlation between hyperspectral pixels into account, total variation is minimized to enforce the spatial smoothness. Finally, the proposed model is solved by the extended alternating direction method of multipliers (ADMM). Experimental results on simulated and real hyperspectral datasets validate the excellent performances of the proposed method. Le Sun 0002, Byeungwoo Jeon, Yuhui Zheng |
ICIP | 3 |
| 2016 | Non-local-based spatially constrained hierarchical fuzzy C-means method for brain magnetic resonance imaging segmentationabstractOwing to the existence of noise and intensity inhomogeneity in brain magnetic resonance (MR) images, the existing segmentation algorithms are hard to find satisfied results. In this study, the authors propose an improved fuzzy C ‐mean clustering method (FCM) to obtain more accurate results. First, the authors modify the traditional regularisation smoothing term by using the non‐local information to reduce the effect of the noise. Second, inspired by the mechanism of the Gaussian mixture model, the distance function of FCM is defined by using the form of certain exponential function consisting of not only the distance but also the covariance and the prior probability to improve the robustness. Meanwhile, the bias field is modelled by using orthogonal basis functions to reduce the effect of intensity inhomogeneity. Finally, they use the hierarchical strategy to construct a more flexibility function, which considers the improved distance function itself as a sub‐FCM, to make the method more robust and accurate. Compared with the state‐of‐the‐art methods, experiment results based on synthetic and real MR images demonstrate its accuracy and robustness. Hui Zhang 0015, Yuhui Zheng, Byeungwoo Jeon, Q. M. Jonathan Wu |
IET Image Process. | 4 |
| 2016 | Efficient data integrity auditing for storage security in mobile health cloud
Yongjun Ren, Jian Shen 0001, Yuhui Zheng, Jin Wang 0001, Han-Chieh Chao |
Peer-to-Peer Netw. Appl. | 3 |
| 2016 | An improved anisotropic hierarchical fuzzy c-means method based on multivariate student t-distribution for brain MRI segmentation
Hui Zhang 0015, Yuhui Zheng, Byeungwoo Jeon, Q. M. Jonathan Wu |
Pattern Recognit. | 3 |
| 2015 | Atomic decomposition based anisotropic non-local structure tensorabstractThe existing non-local structure tensors utilize the isotropic nature of the neighborhoods and compare similarity of tensors by the Euclidean distance for tensor field regularization, thus resulting in limited performances in image analysis. In this paper, we present an anisotropic nonlocal tensor regularization method by using a directional projection based atomic decomposition scheme, which offers two advantages: better exploitation of spatial directional information for anisotropically regularizing tensor field, and straightforward employment of the Euclidean distance to compute smoothing weights without extending the original non-local means filter to tensor field. Experimental results show that the proposed anisotropic structure tensor is superior to existing representative nonlinear structure tensors, in terms of corner detection and image denoising. Yuhui Zheng, Byeungwoo Jeon, Quan-Sen Sun |
ICIP | 1 |
| 2014 | Effective fuzzy clustering algorithm with Bayesian model and mean template for image segmentationabstractFuzzy c‐means (FCMs) with spatial constraints have been considered as an effective algorithm for image segmentation. The well‐known Gaussian mixture model (GMM) has also been regarded as a useful tool in several image segmentation applications. In this study, the authors propose a new algorithm to incorporate the merits of these two approaches and reveal some intrinsic relationships between them. In the authors model, the new objective function pays more attention on spatial constraints and adopts Gaussian distribution as the distance function. Thus, their model can degrade to the standard GMM as a special case. Our algorithm is fully free of the empirically pre‐defined parameters that are used in traditional FCM methods to balance between robustness to noise and effectiveness of preserving the image sharpness and details. Furthermore, in their algorithm, the prior probability of an image pixel is influenced by the fuzzy memberships of pixels in its immediate neighbourhood to incorporate the local spatial information and intensity information. Finally, they utilise the mean template instead of the traditional hidden Markov random field (HMRF) model for estimation of prior probability. The mean template is considered as a spatial constraint for collecting more image spatial information. Compared with HMRF, their method is simple, easy and fast to implement. The performance of their proposed algorithm, compared with state‐of‐the‐art technologies including extensions of possibilistic fuzzy c‐means (PFCM), GMM, FCM, HMRF and their hybrid models, demonstrates its improved robustness and effectiveness. Hui Zhang 0015, Q. M. Jonathan Wu, Yuhui Zheng, Thanh Minh Nguyen 0001, Dingcheng Wang |
IET Image Process. | 3 |