Yongqiang Tang

dblp:51/1668 · DBLP profile ↗
← Back
58ranked-venue papers
6as first author
52since 2021 · last 2026
0000-0001-9333-8200ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 3 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual Activation-Weight Sparsity: A Training-Free Framework for Efficient Large Language Model Compression
abstract
Luoyang Sun, Guangyan Li, Cheng Deng, Haifeng Zhang, Jian Zhao, Yongqiang Tang, Wensheng Zhang, Jun Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Luoyang Sun, Guangyan Li, Cheng Deng 0001, Haifeng Zhang 0002, Jian Zhao 0006, Yongqiang Tang, Wensheng Zhang 0002, Jun Wang 0012
ACL (1)6
2026 Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
abstract
Large Language Models (LLMs) have demonstrated strong capabilities in web search and reasoning. However, their dependence on static training corpora makes them prone to factual errors and knowledge gaps. Retrieval-Augmented Generation (RAG) addresses this limitation by incorporating external knowledge sources, especially structured Knowledge Graphs (KGs), which provide explicit semantics and efficient retrieval. Existing KG-based RAG approaches, however, generally assume that anchor entities are accessible to initiate graph traversal, which limits their robustness in open-world settings where accurate linking between the user query and the KG entity is unreliable. To overcome this limitation, we propose AnchorRAG, a novel multi-agent collaboration framework for open-world RAG without the predefined anchor entities. Specifically, a predictor agent dynamically identifies candidate anchor entities by aligning user query terms with KG nodes and initializes independent retriever agents to conduct parallel multi-hop explorations from each candidate. Then a supervisor agent formulates the iterative retrieval strategy for these retriever agents and synthesizes the resulting knowledge paths to generate the final answer. This multi-agent collaboration framework improves retrieval robustness and mitigates the impact of ambiguous or erroneous anchors. Extensive experiments on four public benchmarks demonstrate that AnchorRAG significantly outperforms existing baselines and establishes new state-of-the-art results on the real-world reasoning tasks.
Jiasheng Xu, Mingda Li 0002, Yongqiang Tang, Wensheng Zhang 0002
WWW3
2026 MaMoE4Rec: Multimodal recommendation with Hop-Aware graph modeling and Mixture-of-Experts fusion
Sirui Zheng, Jin Liu 0016, Bo Huang 0014, Yongqiang Tang, Lan You, Hamido Fujita
Expert Syst. Appl.4
2026 Auto-weighted graph structure learning of multi-dimensional biomarkers for indirect estimation of pediatric reference intervals
Jianguo Zheng, Yongqiang Tang, Yaguang Peng, Rui Chen 0032, Mingda Li 0002, Ruohua Yan, Wensheng Zhang 0002, XiaoXia Peng
Expert Syst. Appl.2
2026 Beyond alignment: Discovering cross-graph triples for knowledge graph integration
Mingda Li 0002, Ao Gao, Yongqiang Tang, Yuanpeng Deng, Wensheng Zhang 0002
Knowl. Based Syst.3
2026 CoT defender: Preemptive chain-of-thought occupation for jailbreak attack mitigation
Jin Liu 0016, Yongqiang Tang, Zhiwen Xie, Xiao Yu 0008, Bo Huang 0014
Neural Networks3
2026 Graph Condensation via Homophily Node Refining and Fine-Grained Distribution Matching
abstract
The remarkable success of GNNs has provoked the challenge of high computational and memory overhead when training with large-scale graphs. As a promising solution, graph condensation is committed to constructing synthetic graphs with significantly smaller size, which are expected to preserve the essential characteristics of the original ones. During this process, a core problem is how to accurately portray and align the data distribution structures between the original graph space and the synthetic graph space. A mainstream idea in existing research is matching the class distributions between the two spaces. Unfortunately, they generally overlook two key issues: 1) heterophilic nodes in original graphs may render the chaotic class distribution patterns; 2) coarse-grained matching of the overall class centroid between original and synthetic spaces is insufficient for data with complex subcategory distributions. In this paper, we propose a novel Graph Condensation method via homophily node Refinement and fine-grained class Distribution matching (GCRD). Given the original large-scale graph, we first distinguish the nodes into advantageous homophilic nodes and detrimental heterophilic nodes, followed by adaptively assigning node weights to refine the generated class distribution patterns of the original graphs. Furthermore, with the refined class distribution patterns, we propose a fine-grained distribution matching objective to more delicately align the local distribution structure of subclasses within each class. The rigorous theoretical analysis confirms the effectiveness of our proposal in precisely learning the class information. Extensive experiments demonstrate our state-of-the-art classification and cross-architecture generalization performance against various baselines.
Ruiwen Yuan, Yongqiang Tang, Wensheng Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 S3Det: Strip steel surface defect detector via enhanced deformable convolution and dual cross-layer pyramid
Yange Sun, Huaping Guo, Yongqiang Tang, Wensheng Zhang 0002
Pattern Recognit. Lett.4
2026 A Zero-Shot Network for Low-Light Image Enhancement With Brightness-Aware Representation and Semantic-Consistency Guidance
abstract
Low-light image enhancement aims to restore visibility, contrast, and semantic fidelity to images captured under poor illumination, which often suffer from uneven lighting, detail loss, and noise. Most existing deep learning methods usually rely on paired datasets, which limits their generalization ability. To overcome these limitations, a new zero-shot low-light enhancement network, termed BRSGNet, is introduced to improve brightness-aware representation and enhance semantic consistency. Specifically, a Brightness-Aware Module (BAM) with multi-scale dilated convolution and spatial attention mechanisms is designed to improve illumination estimation while preserving global-local details. Then, a Semantic-Aware Network (SANet) incorporating a Semantic Feature Consistency Loss (SFC-Loss) is introduced to enforce semantic alignment between the enhanced and original images without requiring ground-truth references. Extensive experiments on three public benchmarks demonstrate that the proposed approach achieves competitive performance in terms of both visual quality and semantic consistency, showcasing strong generalization across diverse low-light conditions.
Huihong Huang, Zhida Ren, Xiaowen Shi, Yongqiang Tang, Wensheng Zhang 0002
IEEE Signal Process. Lett.5
2026 AdaptGCD: Multi-Expert Adapter Tuning for Generalized Category Discovery
abstract
Different from the traditional semi-supervised learning paradigm that is constrained by the close-world assumption, Generalized Category Discovery (GCD) presumes that the unlabeled dataset contains new categories not appearing in the labeled set, and aims to not only classify old categories but also discover new categories in the unlabeled data. Existing studies on GCD typically devote to transferring the general knowledge from the self-supervised pretrained model to the target GCD task via some fine-tuning strategies, such as partial tuning and prompt learning. Nevertheless, these fine-tuning methods fail to make a sound balance between the generalization capacity of pretrained backbone and the adaptability to the GCD task. To fill this gap, in this paper, we propose a novel adapter-tuning-based method named AdaptGCD, which is the first work to introduce the adapter tuning into the GCD task and provides some key insights expected to enlighten future research. Furthermore, considering the discrepancy of supervision information between the old and new classes, a multi-expert adapter structure equipped with a route assignment constraint is elaborately devised, such that the data from old and new classes are separated into different expert groups. Extensive experiments are conducted on 7 widely-used datasets. The remarkable performance improvements highlight the efficacy of our proposal and it can be also combined with other advanced methods like SPTNet for further enhancement.
Yuxun Qu, Yongqiang Tang, Chenyang Zhang 0003, Wensheng Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2026 Multiplex Graph Guided Deep Survival Analysis
abstract
Survival analysis is extensively employed to analyze the probability of the event of interest, particularly in the medical field. Most current research treats patients as isolated entities, neglecting the complex associations among them, which leads to underutilization of valuable information. Recently, several studies address this limitation by incorporating patient graph structures. However, these approaches generally overlook two critical issues: 1) the exploration of heterogeneous inter-patient relationships, and 2) flexible and scalable inductive inference for test samples. To overcome these challenges, this study introduces a novel framework, Multiplex Graph Guided Deep Survival Analysis (MGG-Surv). Specifically, we employ multiplex patient graphs to capture comprehensive inter-patient associative information. Furthermore, we propose a teacher-student dual network architecture, where the teacher network encodes multiplex graphs, and the learned graph knowledge is transferred to the student network via a unidirectional connection termed Graph-Guided Distillation. The student network integrates this graph knowledge to predict survival outcomes without requiring the patient graphs. These innovative designs facilitate comprehensive integration of inter-patient relationships while achieving flexible and scalable graph-free inference. Experiments on four datasets, encompass-ing both single and competing risks, demonstrate the superior performance of our framework.
Chang Cui, Yongqiang Tang, Yuxun Qu, Wensheng Zhang 0002
IEEE Trans. Knowl. Data Eng.2
2025 A Structure-aware Invariant Learning Framework for Node-level Graph OOD Generalization
abstract
Graph Neural Networks (GNNs) have been proven effective in modeling graph data, mostly depending on the in-distribution assumption. While in the out-of-distribution (OOD) scenarios, especially for the more challenging node-level task, the feature and structure distribution shifts between training and test nodes lead to performance degradation. To improve node-level OOD generalization, typical approaches introduce graph augmentation to enrich the training environments and conduct invariant learning to learn stable representations across various augmented environments. However, their graph augmentations emphasize diversity but neglect the preservation of invariant patterns which are fundamental to invariant learning. Moreover, most of them simply conduct the classic invariant learning objective but lack the consideration of the graph-specific structure information. Therefore, to mitigate their weakness, we propose a Structure-aware Invariant learning framework for Node-level Graph OOD generalization (SING). Specifically, we develop the invariance constraint regularization terms during the optimization of augmentations. Additionally, we define the structure embedding to elucidate the structural property and design the structure embedding alignment loss to optimize the augmentations and the invariant representations. By introducing the structure information, we further integrate the unique structural property into invariant learning, thereby boosting the invariant message-passing GNNs. The extensive experiments on the transductive GOOD benchmark and the inductive datasets empirically validate our superior OOD generalization performance to baselines.
Ruiwen Yuan, Yongqiang Tang, Wensheng Zhang 0002
KDD (1)2
2025 KAN-LSTM: A New LSTM Structure for the Prediction of the Stock Market
abstract
ABSTRACT Accurate stock market prediction is crucial for investors to formulate correct investment strategies. However, the non‐linearity, high dimensionality, and volatility of financial data pose significant challenges to existing stock market prediction models. To effectively address the complex datasets faced by stock market prediction, this paper proposes a new and more efficient deep learning hybrid model, KAN‐LSTM, based on the LSTM (long short‐term memory) and integrating the KAN (Kolmogorov–Arnold network). The hybrid architecture improves the learning process by replacing the original MLP (multi‐layer perceptron) with the KAN, overcoming the limitations of poor interpretability and fixed activation functions in LSTM. Prediction experiments conducted on multidimensional financial data in the stock market show that the KAN‐LSTM hybrid model outperforms the original LSTM in all evaluation metrics, demonstrating superior performance and more efficient prediction capabilities. Specifically, the MAE (mean absolute error) decreased by 2.43%, the RMSE (root mean squared error) decreased by 1.92%, the MAPE (mean absolute percentage error) decreased by 2.2%, and the increased by 19.08%.
Weiping Zhu 0004, Jin Liu 0016, Yongqiang Tang, Xiao Liu 0004
Concurr. Comput. Pract. Exp.4
2025 Disentangled and reassociated deep representation for dynamic survival analysis with competing risks
Chang Cui, Yongqiang Tang, Wensheng Zhang 0002
Knowl. Based Syst.2
2025 IBCS: Learning Information Bottleneck-Constrained Denoised Causal Subgraph for Graph Classification
abstract
The significant success of graph learning has provoked a meaningful but challenging task of extracting the precise causal subgraphs that can interpret and improve the predictions. Unfortunately, current works merely center on partially eliminating either the spurious or the noisy parts, while overlook the fact that in more practical and general situations, both the spurious and noisy subgraph coexist with the causal one. This brings great challenges and makes previous methods fail to extract the true causal substructure. Unlike existing studies, in this paper, we propose a more reasonable problem formulation that hypothesizes the graph is a mixture of causal, spurious, and noisy subgraphs. With this regard, an Information Bottleneck-constrained denoised Causal Subgraph (IBCS) learning model is developed, which is capable of simultaneously excluding the spurious and noisy parts. Specifically, for the spurious correlation, we design a novel causal learning objective, in which beyond minimizing the empirical risks of causal and spurious subgraph classification, the intervention is further conducted on spurious features to cut off its correlation with the causal part. On this basis, we further impose the information bottleneck constraint to filter out label-irrelevant noise information. Theoretically, we prove that the causal subgraph extracted by our IBCS can approximate the ground-truth. Empirically, extensive evaluations on nine benchmark datasets demonstrate our superiority over state-of-the-art baselines.
Ruiwen Yuan, Yongqiang Tang, Yanghao Xiao, Wensheng Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Fine-Grained Interactive Transformers for Continuous Dynamic Link Prediction
abstract
Dynamic link prediction (DLP) plays a critical role in understanding and forecasting evolving relationships in real-world systems across various domains. However, accurately predicting future links remains challenging, as existing methods often overlook the independent modeling of dynamic interactions within individual nodes and the fine-grained characterization of latent interactions across node sequences. To address these challenges, we propose FineFormer (Fine-grained Interactive Transformer), a novel framework that alternates between self-attention and cross-attention mechanisms, enhanced with layer-wise contrastive learning. This design enables FineFormer to uncover fine-grained temporal dependencies both within single node sequences and across different node sequences. Specifically, self-attention captures temporal-spatial dynamics within the interaction sequences of individual nodes, while cross-attention focuses on the complex interactions across the sequences of pairs of nodes. Additionally, by strategically applying layer-wise contrastive learning, FineFormer refines node representations and enhances the model's ability to distinguish between connected and unconnected node pairs during feature refinement. FineFormer is evaluated on five challenging and diverse real-world DLP datasets. Experimental results demonstrate that FineFormer consistently outperforms state-of-the-art baselines, particularly in capturing complex, fine-grained interactions in continuous-time dynamic networks.
Yajing Wu, Yongqiang Tang, Wensheng Zhang 0002
IEEE Trans. Cybern.2
2025 A Temporal-Guided Graph Multitask Learning Framework for Multiperiod Optimal Power Flow
Baicheng Chen, Hui Liu 0014, Yongqiang Tang, Hongzhou Li
IEEE Trans. Ind. Informatics4
2025 Upper Limb Motor Sequence Analysis: From Isolated to Sequential
abstract
Motor skills are performed through sequential movements rather than isolated actions. Yet, decoding these sequences from biosignals poses a significant challenge. To address this gap, this study transitions motor decoding from classifying movements in isolated time windows to segmenting sequential movements. The proposed algorithm segments the electromyography (EMG) sequence in a coarse-to-fine manner. It begins with frame-level segmentation and locating the approximate boundaries at the movement-level. A region-growing-inspired fusion strategy is then designed to incorporate the coarse segmentation and localization results for the fined output. Experiments on a self-collected EMG dataset demonstrate impressive results in segmenting movements for participant-dependent/independent setups (accuracy:$94.2\hbox{\%}/74.7\hbox{\%}$; dice coefficient:$92.5\hbox{\%}/61.7\hbox{\%}$; mean Intersection over Union:$80.9\hbox{\%}/51.9\hbox{\%}$). Further analysis shows the algorithm's ability to capture the natural rhythm in participants' movement sequences. This research paves the way for a deep understanding of motor sequences, which benefits various applications, such as rehabilitation engineering.
Tian-Yu Xiang, Xiao-Hu Zhou, Mei-Jiang Gui, Xiaoliang Xie, Shiqi Liu 0004, Hao Li 0077, De-Xing Huang, Jiaxing Wang 0001, Yongqiang Tang, Jiamou Liu, Zeng-Guang Hou
IEEE Trans. Ind. Informatics9
2025 Contrastive Learning With Transformer to Predict the Chronicity of Children With Immune Thrombocytopenia
abstract
Immune thrombocytopenia (ITP) is a typically self-limiting and immune-mediated bleeding disorder in children. Approximately 20% of children with ITP experience chronicity, leading to reduced quality of life and increased treatment burden. The accurate prediction of chronicity would enable clinicians to make personalized treatment plans at an early stage. However, due to the self-limiting nature of ITP and the scarcity of available children patients, the data presents two prominent issues: small data and imbalanced class, which are unfavorable for effectively training a deep learning model. To handle these issues concurrently, we proposed a novel method that integrates contrastive learning with the Transformer. First, we adopt the FT-Transformer as our backbone, which allows our model to flexibly process heterogeneous tabular data. Second, we amplify and balance the original data via random masking and oversampling, respectively. Lastly, we build contrastive pairs according to the latent representations generated by the FT-Transformer encoder, such that the amplified and oversampled synthetic data can be utilized thoroughly. The experimental results on real-world ITP children data show that our proposal outperforms the state-of-the-art methods, and demonstrate the significant advantages of dealing with insufficient and imbalanced problems.
Yuntian Wang, Yongqiang Tang, Jingyao Ma, Zhenping Chen, Chang Cui, Mingda Li 0002, Runhui Wu, Wensheng Zhang 0002
IEEE J. Biomed. Health Informatics2
2025 Cross-Modal Recipe Retrieval With Fine-Grained Prompting Alignment and Evidential Semantic Consistency
abstract
Alignment between the food images and the corresponding recipes is an emerging cross-modal representation learning task. In this task, the recipes are composed of three components, i.e., food title, ingredient lists, and cooking instructions, which require a fine-grained alignment between the features of the two modalities. Existing methods usually aggregate the recipes into global embeddings and then align them with the global image embeddings. Meanwhile, semantic classification is frequently used in these methods to regularize the embeddings of the two modalities. While these methods are efficient, there remain two problems: (1) Forcing the alignment between the global images and recipes embeddings may result in losing the component-specific information. (2) The high diversity of food appearance leads to high uncertainty in the semantic classification of food images and recipes. To solve these problems, we propose a Fine-grained Prompting and Alignment (FPA) model to enhance the feature extraction and bring more component-specific information for fine-grained alignment. Furthermore, to regularize the semantic information contained in the cross-modal features, we design an Evidential Semantic Consistency (ESC) loss to keep the cross-modal semantic consistency. We have conducted comprehensive experiments on the benchmark dataset Recipe1M and the state-of-the-art results on the cross-modal recipe retrieval task demonstrate the effectiveness of our method.
Jin Liu 0016, Zhizhong Zhang 0001, Yuan Xie 0006, Yongqiang Tang, Wensheng Zhang 0002, Xiaohui Cui
IEEE Trans. Multim.5
2025 Constrained Maximum Cross-Domain Likelihood for Domain Generalization
abstract
As a recent noticeable topic, domain generalization aims to learn a generalizable model on multiple source domains, which is expected to perform well on unseen test domains. Great efforts have been made to learn domain-invariant features by aligning distributions across domains. However, existing works are often designed based on some relaxed conditions which are generally hard to satisfy and fail to realize the desired joint distribution alignment. In this article, we propose a novel domain generalization method, which originates from an intuitive idea that a domain-invariant classifier can be learned by minimizing the Kullback-Leibler (KL)-divergence between posterior distributions from different domains. To enhance the generalizability of the learned classifier, we formalize the optimization objective as an expectation computed on the ground-truth marginal distribution. Nevertheless, it also presents two obvious deficiencies, one of which is the side-effect of entropy increase in KL-divergence and the other is the unavailability of ground-truth marginal distributions. For the former, we introduce a term named maximum in-domain likelihood to maintain the discrimination of the learned domain-invariant representation space. For the latter, we approximate the ground-truth marginal distribution with source domains under a reasonable convex hull assumption. Finally, a constrained maximum cross-domain likelihood (CMCL) optimization problem is deduced, by solving which the joint distributions are naturally aligned. An alternating optimization strategy is carefully designed to approximately solve this optimization problem. Extensive experiments on four standard benchmark datasets, i.e., Digits-DG, PACS, Office-Home, and miniDomainNet, highlight the superior performance of our method.
Yongqiang Tang, Wensheng Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 Dual-Space Contrastive Learning for Open-World Semi-Supervised Classification
abstract
Despite recent progress in semi-supervised learning (SSL), its scalability remains limited in realistic scenarios where unseen classes may appear in the unlabeled data. To address this challenge, open-world SSL (OWSSL) is proposed in recent years and attracts much attention. One core difficulty in OWSSL is to enhance the representative ability for unlabeled samples, especially for those in novel classes. More recently, several works introduce contrastive learning into OWSSL and achieve impressive performance. However, they mainly focus on conducting contrastive learning solely in either feature or prediction spaces, while ignoring the thorough exploration of information potentials in dual spaces. In this study, we propose a novel method to handle OWSSL tasks via dual-space contrastive learning (DSCL). DSCL contains two modules: intraspace contrastive learning and interspace contrastive learning. In the intraspace module, we bridge the two spaces with a learnable classifier and impose contrastive learning in the dual spaces, such that the category discriminative information could be effectively utilized to improve the representative ability. In interspace module, to further enhance the utilization of complementary information from dual spaces, we introduce neighborhood information from feature space to enhance predictive learning and meanwhile utilize the cluster structure from the prediction spaces to improve intraclass compactness of the features. Compared with state-of-the-art competitors, the proposed DSCL achieves superior performance on the popular benchmarks, i.e., CIFAR100, Imagenet100, CIFAR10, CUB-200, and Scar.
Yuxun Qu, Yongqiang Tang, Chenyang Zhang 0003, Xiangrui Cai, Xiaojie Yuan, Wensheng Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 Clustering Enhanced Multiplex Graph Contrastive Representation Learning
abstract
Multiplex graph representation learning has attracted considerable attention due to its powerful capacity to depict multiple relation types between nodes. Previous methods generally learn representations of each relation-based subgraph and then aggregate them into final representations. Despite the enormous success, they commonly encounter two challenges: 1) the latent community structure is overlooked and 2) consistent and complementary information across relation types remains largely unexplored. To address these issues, we propose a clustering-enhanced multiplex graph contrastive representation learning model (CEMR). In CEMR, by formulating each relation type as a view, we propose a multiview graph clustering framework to discover the potential community structure, which promotes representations to incorporate global semantic correlations. Moreover, under the proposed multiview clustering framework, we develop cross-view contrastive learning and cross-view cosupervision modules to explore consistent and complementary information in different views, respectively. Specifically, the cross-view contrastive learning module equipped with a novel negative pairs selecting mechanism enables the view-specific representations to extract common knowledge across views. The cross-view cosupervision module exploits the high-confidence complementary information in one view to guide low-confidence clustering in other views by contrastive learning. Comprehensive experiments on four datasets confirm the superiority of our CEMR when compared to the state-of-the-art rivals.
Ruiwen Yuan, Yongqiang Tang, Yajing Wu, Wensheng Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 Shapley value-based class activation mapping for improved explainability in neural networks
Huaiguang Cai, Yang Yang 0056, Yongqiang Tang, Zhengya Sun, Wensheng Zhang 0002
Vis. Comput.3
2024 LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
abstract
Large language models (LLMs) show excellent performance in difficult tasks, but they often require massive memories and computational resources. How to reduce the parameter scale of LLMs has become research hotspots. In this study, we get an important observation that the multi-head self-attention (MHA) sub-layer of Transformer exhibits noticeable low-rank structure, while the feed-forward network (FFN) sub-layer does not. With this regard, we design a novel structured compression method LoRAP, which organically combines Low-Rank matrix approximation And structured Pruning. For the MHA sub-layer, we proposal an input activation weighted singular value decomposition method and allocate different parameter amounts for each weight matrix based on the differences in low-rank properties of matrices.For the FFN sub-layer, we propose a gradient-free structured channel pruning method and save the least important 1% of parameters which actually play a vital role in model performance. Extensive evaluations on zero-shot perplexity and zero-shot task classification indicate that our proposal is superior to previous structured compression rivals under multiple compression ratios. Our code will be released soon.
Guangyan Li, Yongqiang Tang, Wensheng Zhang 0002
ICML2
2024 Addressing Hidden Confounding with Heterogeneous Observational Datasets for Recommendation
abstract
The collected data in recommender systems generally suffers selection bias. Considerable works are proposed to address selection bias induced by observed user and item features, but they fail when hidden features (e.g., user age or salary) that affect both user selection mechanism and feedback exist, which is called hidden confounding. To tackle this issue, methods based on sensitivity analysis and leveraging a few randomized controlled trial (RCT) data for model calibration are proposed. However, the former relies on strong assumptions of hidden confounding strength, whereas the latter relies on the expensive RCT data, thereby limiting their applicability in real-world scenarios. In this paper, we propose to employ heterogeneous observational data to address hidden confounding, wherein some data is subject to hidden confounding while the remaining is not. We argue that such setup is more aligned with practical scenarios, especially when some users do not have complete personal information (thus assumed with hidden confounding), while others do have (thus assumed without hidden confounding). To achieve unbiased learning, we propose a novel meta-learning based debiasing method called MetaDebias. This method explicitly models oracle error imputation and hidden confounding bias, and utilizes bi-level optimization for model training. Extensive experiments on three public datasets validate our method achieves state-of-the-art performance in the presence of hidden confounding, regardless of RCT data availability.
Yanghao Xiao, Yongqiang Tang, Wensheng Zhang 0002
NeurIPS3
2024 Learning the long-tail distribution in latent space for Weighted Link Prediction via conditional Invertible Neural Networks
Yajing Wu, Chenyang Zhang 0003, Yongqiang Tang, Xuebing Yang, Yanting Yin, Wensheng Zhang 0002
Knowl. Based Syst.3
2024 Multiview Subspace Tensor Self-Representation for SAR Image Semi-Supervised Classification
abstract
Synthetic aperture radar (SAR) image classification has proven its significant importance in automatic remote sensing. However, current methods demand a large volume of training data to ensure satisfactory generalization capability. Given the difficulty in obtaining sufficient labeled samples, semi-supervised learning for SAR image classification becomes extremely important. However, existing studies are hindered by sample relationship modeling and single sample descriptions, leading to limited performance. To address this issue, we propose an innovative approach: multiview subspace tensor self-representation learning with label propagation for SAR image classification. Initially, our method extracts global sample relationship matrices of SAR samples using subspace self-representation learning with affine and nonnegative constraints. To enhance the single sample feature description, we construct multiview tensor learning by stacking the subspace representation matrices of different SAR view features. Finally, the SAR image multiview subspace self-representation matrix is treated as the probability indication matrix for classification. The experiments on MSTAR using extreme 10% labeled samples have shown that the proposed method yields performance boosts of 22.3% and 12.2% to the state-of-the-art method under two rigorous evaluation settings, respectively. The remarkably superior experimental results effectively validate the effectiveness of our proposals.
Yang Yang 0056, Yongqiang Tang, Jiangbo Bai, Lu Zhang 0052, Wensheng Zhang 0002
IEEE Geosci. Remote. Sens. Lett.2
2024 Adaptive-weighted deep multi-view clustering with uniform scale representation
Rui Chen 0032, Yongqiang Tang, Wensheng Zhang 0002
Neural Networks2
2024 Dynamic Functional Connectivity Neural Network for Epileptic Seizure Prediction Using Multi-Channel EEG Signal
abstract
Epilepsy, one of the world's most common neurological diseases, impacts over 1% of the global population. Accurate early prediction of epileptic seizure has a great influence on epileptic patients' lives and attracted extensive attention. However, existing methods do not fully consider the complexity of multi-channel electroencephalogram (EEG) signal, which is the most common measurement for epileptic seizures. In this letter, we propose a Dynamic Functional Connectivity neural Network (DynFCNet) for epileptic seizure prediction. The proposed DynFCNet can discover dynamic brain functional connections and generate dynamic functional connectivity graphs, as well as extract non-Euclidean features from multi-channel EEG signals by Graph Convolutional Network (GCN). To take Euclidean features into consideration, Convolutional Neural Network (CNN) as one of the branches is also employed. Further, we incorporated both intra-group loss and inter-group loss to enhance our DynFCNet. Extensive experiments were implemented on a public multi-channel EEG dataset (CHB-MIT). The results confirm that our proposal outperforms the other competitors
Tao Xu 0060, Yajing Wu, Yongqiang Tang, Wensheng Zhang 0002, Zhihua Cui
IEEE Signal Process. Lett.3
2024 Graph Structure Aware Contrastive Multi-View Clustering
abstract
Multi-view clustering has become a research hotspot in recent decades because of its effectiveness in heterogeneous data fusion. Although a large number of related studies have been developed one after another, most of them usually only concern with the characteristics of the data themselves and overlook the inherent connection among samples, hindering them from exploring structural knowledge of graph space. Moreover, many current works tend to highlight the compactness of one cluster without taking the differences between clusters into account. To track these two drawbacks, in this paper, we propose a graph structure aware contrastive multi-view clustering (namely, GCMC) approach. Specifically, we incorporate the well-designed graph autoencoder with conventional multi-layer perception autoencoder to extract the structural and high-level representation of multi-view data, so that the underlying correlation of samples can be effectively squeezed for model learning. Then the contrastive learning paradigm is performed on multiple pseudo-label distributions to ensure that the positive pairs of pseudo-label representations share the complementarity across views while the divergence between negative pairs is sufficiently large. This makes each semantic cluster more discriminative, i.e., jointly satisfying intra-cluster compactness and inter-cluster exclusiveness. Through comprehensive experiments on eight widely-known datasets, we prove that the proposed approach can perform better than the state-of-the-art opponents.
Rui Chen 0032, Yongqiang Tang, Xiangrui Cai, Xiaojie Yuan, Wensheng Zhang 0002
IEEE Trans. Big Data2
2024 Semi-Supervised Graph Structure Learning via Dual Reinforcement of Label and Prior Structure
abstract
Graph neural networks (GNNs) have achieved considerable success in dealing with graph-structured data by the message-passing mechanism. Actually, this mechanism relies on a fundamental assumption that the graph structure along which information propagates is perfect. However, the real-world graphs are inevitably incomplete or noisy, which violates the assumption, thus resulting in limited performance. Therefore, optimizing graph structure for GNNs is indispensable and important. Although current semi-supervised graph structure learning (GSL) methods have achieved a promising performance, the potential of labels and prior graph structure has not been fully exploited yet. Inspired by this, we examine GSL with dual reinforcement of label and prior structure in this article. Specifically, to enhance label utilization, we first propose to construct the prior label-constrained matrices to refine the graph structure by identifying label consistency. Second, to adequately leverage the prior structure to guide GSL, we develop spectral contrastive learning that extracts global properties embedded in the prior graph structure. Moreover, contrastive fusion with prior spatial structure is further adopted, which promotes the learned structure to integrate local spatial information from the prior graph. To extensively evaluate our proposal, we perform sufficient experiments on seven benchmark datasets, where experimental results confirm the effectiveness of our method and the rationality of the learned structure from various aspects.
Ruiwen Yuan, Yongqiang Tang, Yajing Wu, Jinghao Niu, Wensheng Zhang 0002
IEEE Trans. Cybern.2
2024 SASOD: Saliency-Aware Ship Object Detection in High-Resolution Optical Images
abstract
Ship detection in high-resolution optical remote sensing images (ORSI) is an important yet challenging task with extensive applications, such as maritime security and resource conservation. In recent years, bolstered by deep learning, ship detection has also grown by leaps and bounds. Nevertheless, existing methods still suffer from two challenging issues: 1) imprecise localization for low discriminative ships under the complicated background and 2) missed detections for small ships. To solve the above issues, we propose a novel ship detection method equipped with a saliency-guided feature fusion network (SGFFN) and a dynamic IoU-adaptive strategy (DIAS). SGFFN is designed based on a top-down feature pyramid network to introduce saliency information into the ship detection network and optimize the saliency-aware features. It comprises two components: the resolution-matching saliency supervision (RMS) network and the cross-stage saliency integration network (CSIN). RMS is a bimatching mechanism that adopts diverse prediction structures for the saliency maps with different scales, such that the finer saliency-aware features could be obtained. CSIN is a cross-stage cross-channel integration module that is designed to fuse saliency-aware features with low-level features. Furthermore, a customized training strategy for small ships, i.e., DIAS, is devised to assign appropriate intersection over union (IoU) thresholds for anchors around the small ships during the training phase. Experimental results on two datasets demonstrate that our proposed method achieves state-of-the-art performance.
Zhida Ren, Yongqiang Tang, Yang Yang 0056, Wensheng Zhang 0002
IEEE Trans. Geosci. Remote. Sens.2
2024 Deep Survival Analysis With Latent Clustering and Contrastive Learning
abstract
Survival analysis is employed to analyze the time before the event of interest occurs, which is broadly applied in many fields. The existence of censored data with incomplete supervision information about survival outcomes is one key challenge in survival analysis tasks. Although some progress has been made on this issue recently, the present methods generally treat the instances as separate ones while ignoring their potential correlations, thus rendering unsatisfactory performance. In this study, we propose a novel Deep Survival Analysis model with latent Clustering and Contrastive learning (DSACC). Specifically, we jointly optimize representation learning, latent clustering and survival prediction in a unified framework. In this way, the clusters distribution structure in latent representation space is revealed, and meanwhile the structure of the clusters is well incorporated to improve the ability of survival prediction. Besides, by virtue of the learned clusters, we further propose a contrastive loss function, where the uncensored data in each cluster are set as anchors, and the censored data are treated as positive/negative sample pairs according to whether they belong to the same cluster or not. This design enables the censored data to make full use of the supervision information of the uncensored samples. Through extensive experiments on four popular clinical datasets, we demonstrate that our proposed DSACC achieves advanced performance in terms of both C-index (0.6722, 0.6793, 0.6350, and 0.7943) and Integrated Brier Score (IBS) (0.1616, 0.1826, 0.2028, and 0.1120).
Chang Cui, Yongqiang Tang, Wensheng Zhang 0002
IEEE J. Biomed. Health Informatics2
2024 Semisupervised Progressive Representation Learning for Deep Multiview Clustering
abstract
Multiview clustering has become a research hotspot in recent years due to its excellent capability of heterogeneous data fusion. Although a great deal of related works has appeared one after another, most of them generally overlook the potentials of prior knowledge utilization and progressive sample learning, resulting in unsatisfactory clustering performance in real-world applications. To deal with the aforementioned drawbacks, in this article, we propose a semisupervised progressive representation learning approach for deep multiview clustering (namely, SPDMC). Specifically, to make full use of the discriminative information contained in prior knowledge, we design a flexible and unified regularization, which models the sample pairwise relationship by enforcing the learned view-specific representation of must-link (ML) samples (cannot-link (CL) samples) to be similar (dissimilar) with cosine similarity. Moreover, we introduce the self-paced learning (SPL) paradigm and take good care of two characteristics in terms of both complexity and diversity when progressively learning multiview representations, such that the complementarity across multiple views can be squeezed thoroughly. Through comprehensive experiments on eight widely used image datasets, we prove that the proposed approach can perform better than the state-of-the-art opponents.
Rui Chen 0032, Yongqiang Tang, Yuan Xie 0006, Wensheng Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2023 IDO: Instance dual-optimization for weakly supervised object detection
Zhida Ren, Yongqiang Tang, Wensheng Zhang 0002
Appl. Intell.2
2023 Context-aware mutual learning for semi-supervised human activity recognition using wearable sensors
Yuxun Qu, Yongqiang Tang, Xuebing Yang, Yanlong Wen, Wensheng Zhang 0002
Expert Syst. Appl.2
2023 Open set domain adaptation with latent structure discovery and kernelized classifier learning
Yongqiang Tang, Lei Tian 0007, Wensheng Zhang 0002
Neurocomputing1
2023 A multimodal fusion emotion recognition method based on multitask learning and attention mechanism
Jinbao Xie, Qingyan Wang, Dali Yang, Jinming Gu, Yongqiang Tang, Yury I. Varatnitski
Neurocomputing6
2023 Cross-stream contrastive learning for self-supervised skeleton-based action recognition
Ding Li 0006, Yongqiang Tang, Zhizhong Zhang 0001, Wensheng Zhang 0002
Image Vis. Comput.2
2023 Meta-path infomax joint structure enhancement for multiplex network representation learning
Ruiwen Yuan, Yajing Wu, Yongqiang Tang, Wensheng Zhang 0002
Knowl. Based Syst.3
2023 Partial Domain Adaptation by Progressive Sample Learning of Shared Classes
Lei Tian 0007, Yongqiang Tang, Wensheng Zhang 0002
Neural Process. Lett.2
2023 Affine Subspace Robust Low-Rank Self-Representation: From Matrix to Tensor
abstract
Low-rank self-representation based subspace learning has confirmed its great effectiveness in a broad range of applications. Nevertheless, existing studies mainly focus on exploring the global linear subspace structure, and cannot commendably handle the case where the samples approximately (i.e., the samples contain data errors) lie in several more general affine subspaces. To overcome this drawback, in this paper, we innovatively propose to introduce affine and nonnegative constraints into low-rank self-representation learning. While simple enough, we provide their underlying theoretical insight from a geometric perspective. The union of two constraints geometrically restricts each sample to be expressed as a convex combination of other samples in the same subspace. In this way, when exploring the global affine subspace structure, we can also consider the specific local distribution of data in each subspace. To comprehensively demonstrate the benefits of introducing two constraints, we instantiate three low-rank self-representation methods ranging from single-view low-rank matrix learning to multi-view low-rank tensor learning. We carefully design the solution algorithms to efficiently optimize the proposed three approaches. Extensive experiments are conducted on three typical tasks, including single-view subspace clustering, multi-view subspace clustering, and multi-view semi-supervised classification. The notably superior experimental results powerfully verify the effectiveness of our proposals.
Yongqiang Tang, Yuan Xie 0006, Wensheng Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Deep convolutional self-paced clustering
Rui Chen 0032, Yongqiang Tang, Lei Tian 0007, Wensheng Zhang 0002
Appl. Intell.2
2022 Deep multi-view semi-supervised clustering with sample pairwise constraints
Rui Chen 0032, Yongqiang Tang, Wensheng Zhang 0002
Neurocomputing2
2022 One-Step Multiview Subspace Segmentation via Joint Skinny Tensor Learning and Latent Clustering
abstract
Multiview subspace clustering (MSC) has attracted growing attention due to the extensive value in various applications, such as natural language processing, face recognition, and time-series analysis. In this article, we are devoted to address two crucial issues in MSC: 1) high computational cost and 2) cumbersome multistage clustering. Existing MSC approaches, including tensor singular value decomposition (t-SVD)-MSC that has achieved promising performance, generally utilize the dataset itself as the dictionary and regard representation learning and clustering process as two separate parts, thus leading to the high computational overhead and unsatisfactory clustering performance. To remedy these two issues, we propose a novel MSC model called joint skinny tensor learning and latent clustering (JSTC), which can learn high-order skinny tensor representations and corresponding latent clustering assignments simultaneously. Through such a joint optimization strategy, the multiview complementary information and latent clustering structure can be exploited thoroughly to improve the clustering performance. An alternating direction minimization algorithm, which owns low computational complexity and can be run in parallel when solving several key subproblems, is carefully designed to optimize the JSTC model. Such a nice property makes our JSTC an appealing solution for large-scale MSC problems. We conduct extensive experiments on ten popular datasets and compare our JSTC with 12 competitors. Five commonly used metrics, including four external measures (NMI, ACC, F-score, and RI) and one internal metric (SI), are adopted to evaluate the clustering quality. The experimental results with the Wilcoxon statistical test demonstrate the superiority of the proposed method in both clustering performance and operational efficiency.
Yongqiang Tang, Yuan Xie 0006, Changqing Zhang 0002, Zhizhong Zhang 0001, Wensheng Zhang 0002
IEEE Trans. Cybern.1
2022 Ship Detection in High-Resolution Optical Remote Sensing Images Aided by Saliency Information
abstract
Ship detection is a crucial but challenging task in optical remote sensing images. Recently, thanks to the emergence of deep neural networks, significant progress has been made in ship detection. However, there are still two significant issues that must be addressed: 1) The high-resolution optical images may confuse the background with the ship, leading to more false alarms during detection; 2) The detector receives fewer positive samples due to the sparse and uneven distribution of ships in the optical remote sensing images. In this paper, we innovatively propose employing the saliency information to aid the ship detection task to tackle these two issues. To achieve this goal, we devise two novel modules, Feature-Enhanced Structure (FES) and Saliency Prediction Branch (SPB), to boost the capacity of ship detection in complex environments, and propose a new sampling strategy named Salient Screening Mechanism (SSM) to increase the number of positive samples. More specifically, SSM is adopted during the training phase to mine more positive samples from the ignored set. Then, in an end-to-end learning fashion, a neural network that incorporates our carefully designed FES and SPB is trained to gain more discriminative information for distinguishing the foreground and the background. To evaluate the effectiveness of our proposal, two new datasets HRSC-SO and DOTA-isaid-ship are constructed, which possesses the annotation information for both object detection and saliency detection. We conduct extensive experiments on the constructed dataset, and the results demonstrate that our method outperforms the previous state-of-the-art approaches.
Zhida Ren, Yongqiang Tang, Zewen He, Lei Tian 0007, Yang Yang 0056, Wensheng Zhang 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Inductive Spatiotemporal Graph Convolutional Networks for Short-Term Quantitative Precipitation Forecasting
abstract
Short-term quantitative precipitation forecasting (SQPF) using weather radar is an important but challenging problem as one must cope with inherent nonlinearity and spatiotemporal correlation in the data. In this article, we propose a novel deep learning model, named Inductive spatiotemporal Graph Convolutional Networks (InstGCNs), to overcome these issues in SQPF. The proposed InstGCN can learn a nonlinear mapping from historical radar reflectivity to future rainfall amounts and extract informative spatiotemporal representations simultaneously. Specifically, we first provide a formal definition for formulating the SQPF problem from a graph perspective. Then, based on radar reflectivity and rain gauge observation, we propose a novel graph construction approach that utilizes a special elliptic structure to model the spatial dependence of precipitation areas. In addition, a new Node level Differential Block (Node-DB) is introduced to tackle the nonstationary temporal dependence. To execute inductive graph learning for unseen nodes, we design to decompose a whole graph into subgraphs. We conduct extensive experiments on three real-world datasets in East China and a public weather radar dataset in the southeastern parts of France. The experimental results confirm the advantages of InstGCN compared with several state of the arts.
Yajing Wu, Xuebing Yang, Yongqiang Tang, Chenyang Zhang 0003, Wensheng Zhang 0002
IEEE Trans. Geosci. Remote. Sens.3
2022 Constrained Tensor Representation Learning for Multi-View Semi-Supervised Subspace Clustering
abstract
Multi-view subspace clustering is an effective method to partition data into their corresponding categories. Nevertheless, existing multi-view subspace clustering approaches generally operate in a purely unsupervised manner, while ignoring the valuable weakly supervised information that can be readily obtained in many practical applications. In this paper, we consider the weakly supervised form of sample pair constraints, and devote to promoting the performance of multi-view subspace clustering with the aid of such prior knowledge. To achieve this goal, inspired by the intrinsic block diagonal structure of ideal low-rank representation (LRR), we propose a novel regularization to integrate must-link, cannot-link and normalization constraints into a unified formulation. The proposed regularization can be regarded as a general description for sample pairwise constraints, and thus provides a flexible framework for multi-view semi-supervised subspace clustering task. Furthermore, we devise a contrained tensor representation learning (CTRL) model that takes advantage of our proposed regularization to facilitate the learning of the desired representation tensor. An efficient optimization algorithm based on alternating direction minimization strategy is carefully designed to solve the proposed CTRL model. Extensive experiments on eight challenging real-world datasets are conducted, and the results validate the effectiveness of our designed pairwise constraints regularization, as well as the superiority of the proposed CTRL model.
Yongqiang Tang, Yuan Xie 0006, Chenyang Zhang 0003, Wensheng Zhang 0002
IEEE Trans. Multim.1
2021 Graph Convolutional Regression Networks for Quantitative Precipitation Estimation
abstract
Accurate and high-resolution quantitative precipitation estimation (QPE) plays a crucial role in meteorology and hydrology. However, for acquiring a more accurate QPE, how to depict the complex nonlinear relationship between the radar reflectivity and the true rain rates, as well as adaptively explore the spatial dependencies of precipitation, remains extremely challenging. In this letter, we propose to incorporate the merits of graph convolutional regression networks (GCRNs) and address the aforementioned issues simultaneously in the GCRNs framework. Furthermore, in order to tolerate the variabilities of spatial correlation in the practical precipitation, we expand GCRNs with a multiconvolutional mechanism between the center node and its neighbor rain gauges. Thus, the ability to capture more complicated spatial characteristics of precipitation can be enhanced, and the phenomenon of overwhelming by the neighbor nodes can be released. Extensive experiments were implemented on 12 rainfall processes in Hangzhou, China, 2015. The experimental results confirm that our proposal consistently outperforms the state-of-the-art QPE models.
Yajing Wu, Yongqiang Tang, Xuebing Yang, Wensheng Zhang 0002
IEEE Geosci. Remote. Sens. Lett.2
2021 Improving Domain-Adaptive Person Re-Identification by Dual-Alignment Learning With Camera-Aware Image Generation
abstract
Domain adaptation in person re-identification (re-ID) has always been challenging, especially for the lack of supervision information on the target domain. Existing methods generally introduced extra supervision by adversarial learning techniques, then added all the augmented data in the training process to optimize the re-ID model. However, the direct utilization of all the generated data not only increases additional computational cost but also ignores the potential correlation between the origin and generated data. In this article, we propose a novel dual-alignment learning framework (DAL) with camera-aware image generation to efficiently and effectively tackle this issue. Specifically, we propose a camera transfer matching module to generate additional training images with different camera styles, and construct the matching pairs with each containing a origin image and one corresponding camera transferred image. To strengthen the correlation of images for each matching pair, we align the pseudo-labels via clustering algorithm to reduce the pseudo-labels distribution discrepancy between the origin and generated images. Besides, to avoid model degeneration affected by some inaccurate pseudo-labels on unlabelled data, we maximize the mutual information to align the image feature representations of matching pair. The DAL allows us to decrease the camera variance and enhance the discrimination ability of re-ID model. Extensive experiments on three large-scale benchmarks demonstrate the superiority of DAL over state-of-the-art methods.
Chenyang Zhang 0003, Yongqiang Tang, Zhizhong Zhang 0001, Ding Li 0006, Xuebing Yang, Wensheng Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2021 Tensor Multi-Elastic Kernel Self-Paced Learning for Time Series Clustering
abstract
Time series clustering has attracted growing attention due to the abundant data accessible and extensive value in various applications. The unique characteristics of time series, including high-dimension, warping, and the integration of multiple elastic measures, pose challenges for the present clustering algorithms, most of which take into account only part of these difficulties. In this paper, we make an effort to simultaneously address all aforementioned issues in time series clustering under a unified multiple kernels clustering (MKC) framework. Specifically, we first implicitly map the raw time series space into multiple kernel spaces via elastic distance measure functions. In such high-dimensional spaces, we resort to the tensor constraint based self-representation subspace clustering approach, which involves the self-paced learning paradigm, to explore the essential low-dimensional structure of the data, as well as the high-order complementary information from different elastic kernels. The proposed approach can be extended to more challenging multivariate time series clustering scenario in a direct but elegant way. Extensive experiments on 85 univariate and 10 multivariate time series datasets demonstrate the significant superiority of the proposed approach beyond the baseline and several state-of-the-art MKC methods.
Yongqiang Tang, Yuan Xie 0006, Xuebing Yang, Jinghao Niu, Wensheng Zhang 0002
IEEE Trans. Knowl. Data Eng.1
2020 Multi-scale 2D Representation Learning for weakly-supervised moment retrieval
Ding Li 0006, Yongqiang Tang, Zhizhong Zhang 0001, Wensheng Zhang 0002
ICPR3
2020 Learning to Generate Radar Image Sequences Using Two-Stage Generative Adversarial Networks
abstract
While quantitative precipitation estimation (QPE) using weather radar is widely adopted in operation, precipitation data sets are often highly imbalanced. In particular, extreme precipitation usually lacks representation, which may introduce the bottleneck for radar QPE with machine learning models. Discovering the intrinsic characteristic of extreme precipitation with few samples is challenging. In this letter, we focus on the radar reflectivity data and aim to generate synthetic radar image sequences with respect to extreme precipitation. Considering the relatively long interval between continuous radar images due to radar volume scan, traditional methods in video generation are not suitable. In this letter, we propose Two-stage Generative Adversarial Networks (TsGANs) to address the above-mentioned problem. In general, our TsGAN constructs adversarial process between generators and discriminators: the generator produces samples similar to real data, while the discriminator determines whether or not a sample is eligible. In Stage I, we generate an image sequence containing content and motion features. In Stage II, we design an enhanced net structure to enrich the adversarial processes and further improve the motion features. Experimental testing is performed within the radar coverage in Shenzhen, China, on rainfall events in 2014-2016. Results show that our TsGAN is superior to previous works.
Chenyang Zhang 0003, Xuebing Yang, Yongqiang Tang, Wensheng Zhang 0002
IEEE Geosci. Remote. Sens. Lett.3
2020 Domain Adaptation by Class Centroid Matching and Local Manifold Self-Learning
abstract
Domain adaptation has been a fundamental technology for transferring knowledge from a source domain to a target domain. The key issue of domain adaptation is how to reduce the distribution discrepancy between two domains in a proper way such that they can be treated indifferently for learning. In this paper, we propose a novel domain adaptation approach, which can thoroughly explore the data distribution structure of target domain. Specifically, we regard the samples within the same cluster in target domain as a whole rather than individuals and assigns pseudo-labels to the target cluster by class centroid matching. Besides, to exploit the manifold structure information of target data more thoroughly, we further introduce a local manifold self-learning strategy into our proposal to adaptively capture the inherent local connectivity of target samples. An efficient iterative optimization algorithm is designed to solve the objective function of our proposal with theoretical convergence guarantee. In addition to unsupervised domain adaptation, we further extend our method to the semi-supervised scenario including both homogeneous and heterogeneous settings in a direct but elegant way. Extensive experiments on seven benchmark datasets validate the significant superiority of our proposal in both unsupervised and semi-supervised manners.
Lei Tian 0007, Yongqiang Tang, Liangchen Hu, Zhida Ren, Wensheng Zhang 0002
IEEE Trans. Image Process.2
2020 Tensor Multi-Task Learning for Person Re-Identification
abstract
This paper presents a tensor multi-task model for person re-identification (Re-ID). Due to discrepancy among cameras, our approach regards Re-ID from multiple cameras as different but related classification tasks, each task corresponding to a specific camera. In each task, we distinguish the person identity as a one-vs-all linear classification problem, where one classifier is associated with a specific person. By constructing all classifiers into a task-specific projection matrix, the proposed method could utilize all the matrices to form a tensor structure, and jointly train all the tasks in a uniform tensor space. In this space, by assuming the features of the same person under different cameras are generated from a latent subspace, and different identities under the same perspective share similar patterns, the high-order correlations, not only across different tasks but also within a certain task, can be captured by utilizing a new type of low-rank tensor constraint. Therefore, the learned classifiers transform the original feature vector into the latent space, where feature distributions across cameras can be well-aligned. Moreover, this model can be incorporated into multiple visual features to boost the performance, and easily extended to the unsupervised setting. Extensive experiments and comparisons with recent Re-ID methods manifest the competitive performance of our method.
Zhizhong Zhang 0001, Yuan Xie 0006, Wensheng Zhang 0002, Yongqiang Tang, Qi Tian 0001
IEEE Trans. Image Process.4
2020 Inter-Patient ECG Classification With Symbolic Representations and Multi-Perspective Convolutional Neural Networks
abstract
This paper presents a novel deep learning framework for the inter-patient electrocardiogram (ECG) heartbeat classification. A symbolization approach especially designed for ECG is introduced, which can jointly represent the morphology and rhythm of the heartbeat and alleviate the influence of inter-patient variation through baseline correction. The symbolic representation of the heartbeat is used by a multi-perspective convolutional neural network (MPCNN) to learn features automatically and classify the heartbeat. We evaluate our method for the detection of the supraventricular ectopic beat (SVEB) and ventricular ectopic beat (VEB) on MIT-BIH arrhythmia dataset. Compared with the state-of-the-art methods based on manual features or deep learning models, our method shows superior performance: the overall accuracy of 96.4%, F1 scores for SVEB and VEB of 76.6% and 89.7%, respectively. The ablation study on our method validates the effectiveness of the proposed symbolization approach and joint representation architecture, which can help the deep learning model to learn more general features and improve the ability of generalization for unseen patients. Because our method achieves a competitive inter-patient heartbeat classification performance without complex handcrafted features or the intervention of the human expert, it can also be adjusted to handle various other tasks relative to ECG classification.
Jinghao Niu, Yongqiang Tang, Zhengya Sun, Wensheng Zhang 0002
IEEE J. Biomed. Health Informatics2
2018 Radar and Rain Gauge Merging-Based Precipitation Estimation via Geographical-Temporal Attention Continuous Conditional Random Field
abstract
An accurate, high-resolution precipitation estimation based on rain gauge and radar observations is essential in various meteorological applications. Although numerous studies have demonstrated the effectiveness of merging two information sources rather than using separate sources, approaches that simultaneously consider the local radar reflectivity, the neighborhood rain gauge observations, and the temporal information are much less common. In this paper, we present a new framework for real-time quantitative precipitation estimation (QPE). By formulating the QPE as a continuous conditional random field (CCRF) learning problem, the spatiotemporal correlations of precipitation can be explored more thoroughly. Based on the CCRF, we further improve the accuracy of the precipitation estimation by introducing geographical and temporal attention. Specifically, we first present a data-driven weighting scheme to merge the first law of geography into the proposed framework, and hence, the neighborhood sample closer to the estimated grid can receive more attention. Second, the temporal attention penalizes the similarity between two adjacent timestamps via the discrepancy of two-view estimates, which can model the local temporal consistency and tolerate some drastic changes. A sufficient evaluation is conducted on 11 rainfall processes that occurred in 2015, and the results confirm the advantage of our proposal for real-time precipitation estimation.
Yongqiang Tang, Xuebing Yang, Wensheng Zhang 0002
IEEE Trans. Geosci. Remote. Sens.1