VLDB 2026 Research / reviewers in the wild / expert
Taiping Zhang
dblp:41/2029
· DBLP profile ↗
66ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0001-9891-4203ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Calibrated Affinity Distillation and Global-to-Local Prototype Alignment for Image Clustering
Xuejiao Wan, Taiping Zhang, Zuotao Fu |
ICIC (7) | 2 |
| 2026 | Aegis: A domain generalization framework for medical image segmentation by mitigating feature misalignment
Yuheng Xu, Taiping Zhang, Yuqi Fang |
Pattern Recognit. | 2 |
| 2026 | An explicit suppression paradigm for cross-domain medical image segmentation
Yuheng Xu, Taiping Zhang, Yuqi Fang |
Pattern Recognit. | 2 |
| 2025 | A Domain Generalization Framework Based on Wavelet-Driven Structural Enhancement and Contrastive AlignmentabstractDomain generalization trains models on source domain data to generalize effectively to unseen target domains. Existing methods rely on adversarial training or feature alignment for domain-invariant representation learning, often combined with data augmentation or self-supervised tasks to improve robustness. However, these approaches struggle with modeling complex inter-domain relationships and addressing limited data diversity. To tackle these challenges, we propose a distribution modeling-based contrastive learning framework. Unlike point-based and prototype-based models, it directly models class feature distributions, capturing intra-class consistency and inter-class separability. We further introduce a cross-domain distribution alignment mechanism to minimize inter-domain discrepancies, enhancing domain-invariant feature learning. Additionally, a frequency-domain structural enhancement strategy using discrete wavelet transform preserves critical structural information while reducing visual distortions caused by style variations. Experimental results show that the proposed framework significantly improves segmentation performance under domain shifts and large intra-class variability, offering an efficient solution for domain generalization in medical image segmentation. Yuheng Xu, Taiping Zhang, Yang Liu 0271 |
ICME | 2 |
| 2025 | Adaptive Semantic Anchoring for Enhanced Generalized Few-Shot Semantic SegmentationabstractGeneralized Few-Shot Semantic Segmentation (GFSS) extends Few-Shot Segmentation by utilizing abundant base class samples and limited novel class annotations to simultaneously segment both base and novel classes. Unlike traditional FSS, which assumes class-matched query-support pairs, GFSS removes this restriction,providing greater flexibility and practicality for segmentation tasks in real-world scenarios. However, existing GFSS methods often face two major challenges: first, due to the scarcity of novel class annotations, their prototypes tend to overfit to low-level texture features, resulting in poor generalization. Second, fixed prototypes within the prototype-query matching strategies struggle to adapt to diverse semantic contexts. To address these challenges, we propose Adaptive Semantic Anchoring for Enhanced Generalized Few-Shot Semantic Segmentation, which incorporates two complementary modules: Semantic diffusion enhancement(SDE) and anchor-based prototype refinement(APR). Specifically, SDE suppresses low-level texture dependencies through gradient-based diffusion, ensuring robust high-level semantic representations across classes. Meanwhile, APR adaptively refines class prototypes using high-confidence anchor pixels, combining local details with global semantics to enhance segmentation accuracy. Extensive experiments on the PASCAL-5iand COCO-20ibenchmarks demonstrate that our method significantly outperforms state-of-the-art approaches. Xinrao Chen, Taiping Zhang |
IJCNN | 2 |
| 2025 | A Structure and Semantic Aware Framework for Generalized Medical Image Segmentation via Frequency and Probabilistic LearningabstractDomain generalization in medical image segmentation remains challenging due to domain shifts caused by varying imaging protocols and device heterogeneity in clinical datasets. Existing methods relying on global or random augmentations suffer from limited diversity or neglect distribution constraints, while overlooking critical medical imaging characteristics like high semantic similarity and class imbalance. To address this, we propose a Frequency-Aware Multi-Scale Domain Augmentation (FMSDA) strategy that integrates Fourier Transform to dynamically blend amplitude spectrums across domains, generating structurally consistent yet stylistically diverse images at both global and class-specific scales. Complementing this, we introduce Treasure in Feature, which refines intermediate-layer features through focal loss and cross-domain information consistency, enhancing discriminability for small target regions. Further, we devise probabilistic representation learning that models latent-space distributions via statistical covariance estimation and semantic-guided contrastive alignment, bridging domain gaps by leveraging both certainty and uncertainty. Extensive experiments on Fundus and Prostate benchmarks validate the superiority of our framework, achieving state-of-the-art cross-domain generalization performance while maintaining the integrity of critical anatomical structures. Yuheng Xu, Taiping Zhang |
IJCNN | 3 |
| 2025 | Semantic Token Enhancement and Clustering-Guided Activation for Weakly Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) methods based on image-level annotations often employ Class Activation Maps (CAM) to construct pseudo labels for dense prediction. Recently, Transformer-Based WSSS approaches have attracted increasing attention due to the strong capability of Transformer in capturing global context. These methods typically generate object localization maps by modeling attention interactions between class tokens and patch tokens. However, the self-attention mechanism in Transformer tends on focus on only a limited number of tokens, resulting in sparse attention maps and consequently overlooking semantically relevant regions in the generation of pseudo labels. In this paper, we propose two novel attention mechanisms: SCAR (selective cross-attention refinement) and Self-Squared attention, both aimed at uniformly highlighting entire object regions. The SCAR module extracts localized features by restricting cross-attention operations to the predicted attention regions. Self-Squared attention includes a clustering-aware module that groups similar patch tokens from the same object to guide activation. The proposed methods ultimately generate a clustering-guided class activation map, which captures more comprehensive region coverage. Experimental results on standard benchmark datasets demonstrate that our approach consistently surpasses previous WSSS approaches. Jingjing Hou, Yuheng Xu, Taiping Zhang |
SMC | 3 |
| 2025 | Multi-Feature Guided Generalization: Tackling Domain Shifts in Semi-Supervised Medical Image SegmentationabstractMedical image segmentation faces the dual challenges of limited annotated data and domain shifts when generalizing to unseen domains. Existing methods often tackle one challenge at the expense of the other. To tackle both challenges, we propose a segmentation framework that addresses these issues simultaneously. Inspired by class-level representations, we hypothesize that data from unseen target domains can be expressed as linear combinations of source domain data, which can be approximated via data augmentation. Based on this, we introduce a Multi-Feature Balance Fusion (MFBF) data augmentation mechanism. MFBF combines global and class-specific local augmentations, exploring diverse and even extreme appearances in unknown domains under a balanced multi-feature framework. This approach significantly enriches sample diversity and ensures that augmented samples capture potential target domain distributions, effectively mitigating domain shift. Additionally, to address the limitations of traditional sharpening functions, which may lead to overconfident predictions and overfitting, we introduce a Dynamic Uncertainty Estimation (DUE) mechanism. DUE encourages the model to explore more regions during training while avoiding overfitting to noise or outliers. Experimental results validate the effectiveness of our approach, demonstrating its superior segmentation performance in addressing both data scarcity and domain shift challenges. This provides a practical and efficient solution for semi-supervised learning and domain generalization in medical image segmentation. Yuheng Xu, Taiping Zhang |
SMC | 3 |
| 2025 | Combining hierarchical sparse representation with adaptive prompt for few-shot segmentation
Xiaoliu Luo, Ting Xie 0004, Weisen Qin, Zhao Duan, Taiping Zhang |
Expert Syst. Appl. | 6 |
| 2025 | Layer-Wise Mutual Information Meta-Learning Network for Few-Shot SegmentationabstractThe goal of few-shot segmentation (FSS) is to segment unlabeled images belonging to previously unseen classes using only a limited number of labeled images. The main objective is to transfer label information effectively from support images to query images. In this study, we introduce a novel meta-learning framework called layer-wise mutual information (LayerMI), which enhances the propagation of label information by maximizing the mutual information (MI) between support and query features at each layer. Our approach involves the utilization of a LayerMI Block based on information-theoretic co-clustering. This block performs online co-clustering on the joint probability distribution obtained from each layer, generating a target-specific attention map. The LayerMI Block can be seamlessly integrated into the meta-learning framework and applied to all convolutional neural network (CNN) layers without altering the training objectives. Notably, the LayerMI Block not only maximizes MI between support and query features but also facilitates internal clustering within the image. Extensive experiments demonstrate that LayerMI significantly enhances the performance of baseline and achieves competitive performance compared to state-of-the-art methods on three challenging benchmarks: PASCAL- $5^{i}$ , COCO- $20^{i}$ , and FSS-1000. Xiaoliu Luo, Zhao Duan, Anyong Qin, Zhuotao Tian, Ting Xie 0004, Taiping Zhang, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Combining transformers with CNN for multi-focus image fusion
Zhao Duan, Xiaoliu Luo, Taiping Zhang |
Expert Syst. Appl. | 3 |
| 2024 | PFENet++: Boosting Few-Shot Semantic Segmentation With the Noise-Filtered Context-Aware Prior MaskabstractIn this work, we revisit the prior mask guidance proposed in “Prior Guided Feature Enrichment Network for Few-Shot Segmentation”. The prior mask serves as an indicator that highlights the region of interests of unseen categories, and it is effective in achieving better performance on different frameworks of recent studies. However, the current method directly takes the maximum element-to-element correspondence between the query and support features to indicate the probability of belonging to the target class, thus the broader contextual information is seldom exploited during the prior mask generation. To address this issue, first, we propose the Context-aware Prior Mask (CAPM) that leverages additional nearby semantic cues for better locating the objects in query images. Second, since the maximum correlation value is vulnerable to noisy features, we take one step further by incorporating a lightweight Noise Suppression Module (NSM) to screen out the unnecessary responses, yielding high-quality masks for providing the prior knowledge. Both two contributions are experimentally shown to have substantial practical merit, and the new model named PFENet++ significantly outperforms the baseline PFENet as well as all other competitors on three challenging benchmarks PASCAL-5$^{i}$, COCO-20$^{i}$and FSS-1000. The new state-of-the-art performance is achieved without compromising the efficiency, manifesting the potential for being a new strong baseline in few-shot semantic segmentation. Xiaoliu Luo, Zhuotao Tian, Taiping Zhang, Bei Yu 0001, Yuan Yan Tang, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Spatial Similarity Guidance for Few-Shot SegmentationabstractIn this work, we address the challenging issue of few-shot segmentation. Existing methods mainly explore the target object through the semantic similarity between the query and support pixels. However, the semantic similarity often fails to deal well with the target objects with large variations in appearance and the error predictions along the boundary. To this end, we propose a novel spatial similarity guidance network (S2GNet), which adaptively integrates spatial information with semantic information for building a target-aware correlation region to enhance the target object localization. To promote the overall spatial position understanding of the target object, we exploit boundaries as crucial guidance for spatial information. Thus we jointly train a boundary detection task and a segmentation task in an end-to-end way. With that, a target-aware attention module is further proposed to capture the target correlation region by combining the spatial similarity with the semantic similarity for each pair of pixels in the query image, which refines the location of the target object effectively and improves the segmentation performance. Extensive experiments on both PASCAL-5iand COCO-20idatasets show that our approach can achieve state-of-the-art performances. Xiaoliu Luo, Zhao Duan, Taiping Zhang |
ICASSP | 3 |
| 2023 | Multi-focus image fusion via gradient guidance progressive networkabstractIn this paper, we address the problem of fusing multi-focus images in same scenes. We propose a gradient guidance progressive network for multi-focus image fusion. We explicitly extract gradient features of images, and introduce the gradient guidance progressive module to integrate effectively features. In the module, we employ low-resolution features with large receptive fields to detect focused areas far away from boundaries. While for high-resolution features incorporating detailed gradient features, we only focus on optimizing outputs near boundaries. Benefiting from the separate operations on both areas far away from and near boundaries, the proposed method makes accurate focus region detection with detailed boundaries. Experimental results demonstrate the effectiveness and superiority of the proposed method compared with the state-of-the-art methods. Zhao Duan, Xiaoliu Luo, Taiping Zhang |
ICME | 3 |
| 2023 | A Lightweight Multi-Scale Based Attention Network for Image Super-ResolutionabstractIn this paper, we propose a lightweight multi-scale based attention network (MBAN) for single-image super-resolution (SISR). First, a deep feature transform block (DFTB) is designed for multi-scale feature extraction; this block combines group convolution and improved channel attention (ICA) for performance purposes while remaining sufficiently lightweight. Second, a dual multi-scale attention block (DMAB) is proposed for long-range information interaction; this block employs different window sizes for self-attention (SA) and short connections between different branches to achieve multiscale attention interaction. Finally, our MBAN is constructed by cascaded multi-scale based attention blocks (MBABs) that perform detail restoration; these blocks simultaneously extract multi-scale local features and integrate multi-scale global features with the DFTBs and DMABs. Extensive experiments suggest the superiority of our MBAN over the state-of-the-art (SOTA) lightweight SR methods in terms of both quantitative metrics and visual quality. Yanjie Yang, Jun Luo 0006, Huayan Pu, Mingliang Zhou 0001, Xuekai Wei, Taiping Zhang, Zhaowei Shang |
IECON | 6 |
| 2023 | MTSN: Multiscale Temporal Similarity Network for Temporal Action LocalizationabstractTemporal Action Localization (TAL) aims to predict the categories and temporal segments of all action instances in untrimmed videos, which is a critical and challenging task in the video understanding field. The performances of existing TAL methods remain unsatisfactory, due to the lack of highly effective temporal modeling and refined action proposal decoding. In this paper, we propose Multiscale Temporal Similarity Network (MTSN), a novel one-stage method for TAL, which mainly benefits from dynamic complementary modeling and temporal similarity decoding. Specifically, we first design Dynamic Complementary Context Aggregation (DCCA), a Transformer-based encoder. DCCA performs both long-range and short-range temporal modeling through different interaction range types of attention heads at each feature pyramid level, while higher-level semantic representations are effectively complemented with more short-range detail information in a dynamic fashion. Moreover, Temporal Similarity Mask (TSM) is designed to generate masks through an optimized globally-aware decoding process, including similarity cross-modeling, region-aware optimization and multiscale aggregated residual, which leads to high-quality action proposals. We conduct extensive experiments on two major TAL benchmarks: THUMOS14 and ActivityNet-1.3, where our method establishes a new state-of-the-art and significantly outperforms the previous best methods. Without bells and whistles, on THUMOS14, MTSN achieves an average mAP of 72.1% (+5.3%). On ActivityNet-1.3, MTSN reaches an average mAP of 40.7% (+3.1%), which crosses the 40% average mAP for the first time. Xiaodong Jin, Taiping Zhang |
ACM Multimedia | 2 |
| 2023 | Target-Aware Bi-Transformer for Few-Shot Segmentation
Xianglin Wang, Xiaoliu Luo, Taiping Zhang |
PRCV (2) | 3 |
| 2023 | Semantic Segmentation Based on Vision Transformer via Interactive AttentionabstractSemantic segmentation is a fundamental task in the computer vision community that aims to achieve pixel-wise classification of images. Convolutional Neural Networks (CNNs) have been the backbone of typical semantic segmentation methods. However, the recent success of the Transformer architecture in natural language processing has led to its application in the field of image semantic segmentation. These methods mainly focus on learning more effective information through the encoder, while paying less attention to the decoder. In this paper, we propose a novel attention-based decoder module called the Attention In Attention (AIA) module. This module employs interactive attention to extract spatial and channel information and dynamically determine feature importance. Additionally, we propose the Feature Position Offset Estimation Module (FPOEM) to mitigate feature misalignment when features of different scales are fused. Experiments on two datasets, Cityscapes and ADE20K, show that the method proposed in this paper achieves state-of-the-art performance. Tao Qiu, Xinqi Jiang, Taiping Zhang |
SMC | 5 |
| 2023 | Object Detection via Multi-Scale Token Based on Vision TransformerabstractVisual transformers have achieved impressive performance on object detection. Traditional transformers only focus on multi-scale features between tokens and tokens. However, these methods do not pay attention to the fine-grained features inside a single token, which can lead to the loss of semantic information in the object detection task. To address this issue, we propose a novel network for the above problem, which consists of three components, (1) Internal Multiscale Token Module (IMTM) focuses on the receptive field size of each token and transforms the token dimension size to effectively extract more multiscale features within the self-attention layer, thereby improving the performance and generalization ability of the model. (2) Differential Filter Module (DFM) uses a convolutional network to focus on high-frequency information in the image, helping the Transformer to learn edge features and establish local context, while improving the model performance through residual connections. (3) Feature Fusion Module (FFM) enhances the local and global information extracted by the network by fusing information from different dimensions. Extensive experiments on PASCAL VOC shows that our proposed method can achieve a state-of-the-art performance on object detection. Tao Qiu, Xinqi Jiang, Zhaowei Shang, Taiping Zhang |
SMC | 6 |
| 2023 | Low-light image enhancement with geometrical sparse representation
Taiping Zhang, Linchang Zhao, Darong Huang 0002, Zhenyuan Zhang 0002 |
Appl. Intell. | 2 |
| 2023 | Multi-focus image fusion using structure-guided flow
Zhao Duan, Xiaoliu Luo, Taiping Zhang |
Image Vis. Comput. | 3 |
| 2023 | Intermediate prototype network for few-shot segmentation
Xiaoliu Luo, Zhao Duan, Taiping Zhang |
Signal Process. | 3 |
| 2022 | Pseudo-Interacting Guided Network for Few-Shot SegmentationabstractFew-shot segmentation has got a lot of concerns recently. Existing methods mainly locate and recognize the target object based on a cross-guided way that applies masked target object features of support(query) images to make a feature matching with query(support) images. However, there are some differences between support images and query images because of large appearance and scale variation, which will lead to inaccurate and incomplete segmentation. This problem inspired us to explore the local coherence of the image to guide the segmentation. We try to get some target pixels in the query image and apply these pixels to search for more target pixels in the query image. In this work, we propose a novel network that combines a universal cross-guided branch with a new pseudo-interacting guided branch. Specifically, we first employ the universal cross-guided branch to produce a pseudo-labeling that represents the probability of each pixel belonging to the target object. Then we design a pseudo-interacting guided branch, which applies some pixels with high probabilities based on generated pseudo-labeling to segment the target object in the query image and revises the results of the cross-guided branch simultaneously. Extensive experiments show that our approach outperforms state-of-the-art methods on both PASCAL-5iand COCO-20idatasets. Xiaoliu Luo, Zhao Duan, Taiping Zhang |
ICASSP | 5 |
| 2022 | A Novel GAN based on Progressive Growing Transformer with Capsule EmbeddingabstractGenerative Adversarial Networks (GANs) have achieved great improvement after using Convolutional Neural Networks (CNNs) instead of Multi-Layer Perceptrons (MLPs) to build network architecture. Recently, since Transformer architecture has performed well in compute vision, building a Transformer-based image generation network helps solve some of problems caused by CNNs e.g. CNNs-based GANs are difficult to train. On the other hand, the learning of positional encoding in the Transformer structure is often ignored in Transformer-based GANs. Capsule networks are usually considered to be able to learn position information in image features. Therefore, this paper constructs a Progressive Growing Transformer network with Capsule Embedding GAN (PGTCEGAN). The results from the proposed approach are promising with 3.59 FID and 3.92 FID on CelebA and LSUN-Church datasets respectively in image generation task. Xinqi Jiang, Taiping Zhang, Tao Qiu |
SMC | 2 |
| 2022 | Double Feature Pyramid Networks for Classification and Localization on Object DetectionabstractThe decoupled head for classification and localization have been proven powerful in the most of one-stage and two-stage detectors. However, most object detection algorithm share a Feature Pyramid Networks. We perform a thorough analysis about the effectiveness of Feature Pyramid Networks for these two tasks. The decoupled feature pyramid network performs better than the shared network. Going a step further, we found that the two tasks have different preferences for feature pyramid networks. For higher accuracy, we propose a Scene Parsing Pyramid Network for Classification and a Feature Pyramid Transformer Network for Localization. Scene Parsing Pyramid Network exploit the capability of global context information by different region based context aggregation through pyramid pooling module and pyramid attention feature extraction module. Feature Pyramid Transformer Network can capture the suitable contexts of objects residing in different scales. We evaluate our double Feature Pyramid Networks feature pyramid network in the object detection task by integrating it into the FCOS algorithm. The modified algorithm outperforms previous state-of-the-art feature pyramid based methods with a clear margin on both MS-COCO 2017 validation and test datasets. Taiping Zhang, Tao Qiu, Xinqi Jiang |
SMC | 2 |
| 2022 | Efficient residual attention network for single image super-resolution
Fangwei Hao, Taiping Zhang, Linchang Zhao, Yuan Yan Tang |
Appl. Intell. | 2 |
| 2022 | Siamese networks with an online reweighted example for imbalanced data learning
Linchang Zhao, Zhaowei Shang, Mingliang Zhou 0001, Mu Zhang 0010, Dagang Gu, Taiping Zhang, Yuan Yan Tang |
Pattern Recognit. | 7 |
| 2021 | Graph Affinity Network for Few-Shot SegmentationabstractFew-shot segmentation aims to learn a segmentation model that can be generalized to novel classes with a few annotations. Previous methods mainly establish the correspondence between support images and query images with global information. However, human perception does not tend to learn a whole representation in its entirety at once. In this paper, we propose a novel network to build the correspondence from subparts, parts and whole. Our network mainly contain two novel designs: we firstly adopt graph convolutional network to make pixels not only contain the information of each pixel itself but also include its contextual pixels, and then a learnable Graph Affinity Module(GAM) is proposed to mine more accurate relationships as well as common object location inference between the support images and the query images. Experiments on the PASCAL-5idataset show that our method achieves state-of-the-art performance. Xiaoliu Luo, Taiping Zhang |
ICIP | 2 |
| 2021 | Explore Connection Pattern And Attention Mechanism For Lightweightimage Super-ResolutionabstractDespite the great success of CNN (Convolutional Neural Network) in SISR (Single Image Super-Resolution), the increase in network depth leads to higher computational complexity and memory usage, which extremely hinders real-world applications. To solve this problem, we propose a lightweight cascade fusion network (CFNet) by stacking the cascade fusion block (CFB), which adopts both cascade connections and fusion connections to fully utilize the extraction ability of convolution. Specifically, Cascade connections help to transfer low-level features of source, and fusion connections help to obtain hierarchical features produced by intermediate convolution layers. To further improve the performance, we also design an efficient low-dimensional pixel attention (LPA) mechanism for SISR tasks and summarize several design guidelines. Thanks to LPA module, our CFNet improves the final reconstruction quality with little parameter cost. Extensive experimental results show that the proposed CFNet achieves a better trade-off against the state-of-the-art methods in terms of performance and model complexity. Our codes for CFNet are available at https://github.com/knowback/CFNet. Zhu Qin, Taiping Zhang |
ICIP | 2 |
| 2021 | RGB-Infrared Person Re-Identification Via Multi-Modality Relation Aggregation and Graph Convolution NetworkabstractRGB-Infrared person Re-identification (RGB-IR Re-ID) task aims at using RGB (infrared) query image to matching person in infrared (RGB) gallery image, the large modality gap between two modalities makes the task very challenging. Different modality image’s feature lacked modality-shared information essentially, leading to large cross-modality discrepancy, existing method are not using modality relations to handle the discrepancy. In this paper, we proposed a novel Graph-based Modality-aware Relation Network (GMRN) to solve this problem, our method contains two parts: 1) fine-granularity multi-modality feature aggregation module to incorporate cross-modality information, 2) modality-aware graph convolution network utilize intra-modality relations to further learn discriminative features under two different modalities. Extensive experiments have made on two public RGB-IR Re-ID dataset SYSU-MMOI and RegDB, experiment results show that our method outperforms current state-of-the-art methods by a large margin. Our code is available at https://github.com/clsrsun/GMRN-ReID Jiangshan Sun, Taiping Zhang |
ICIP | 2 |
| 2021 | Channel Interaction with Local Enhancement for Few-Shot Semantic SegmentationabstractState-of-the-art semantic segmentation algorithms are based on deep convolutional neural networks, which aim to exploit a great quantity of labeled samples to predict the class label for each pixel in an image. Few-shot segmentation alleviate the data labeling task, learning a model given only a few labeled samples, which adapts well to unseen classes. Previous works only extract spatial pixels relationship to build the guidance, which ignores the important feature channel information that contain abundant local semantic structure in the images. In this paper, we propose a novel channel interaction network (CINet) for few-shot semantic segmentation. It mainly consists of three novel steps: a local enhancement module that aggregates the spatial pixels connection in both query images and support images; a channel interaction module that captures the similar channel from relationship between each pair of images for segmentation guidance; and a residual attention module that exploiting residual connection of feature scales to alleviate spatial inconsistency between query images and corresponding example images. Extensive experiments on PASCAL-5idataset demonstrate that our model outperforms state-of-the-art methods without parameters increase sharply. Xiaoliu Luo, Taiping Zhang |
IJCNN | 3 |
| 2021 | Channel Hourglass Residual Network For Single Image Super-ResolutionabstractDeep convolutional neural networks (CNNs) for Super-Resolution (SR) from low-resolution (LR) images have achieved remarkable reconstruction performance with the utilization of residual networks and visual attention mechanism. However, the existing single image super-resolution (SISR) methods with deeper or wider network architectures encounter module representation bottleneck and neglect module efficiency in real-world applications. To solve these issues, in this paper, we design channel hourglass residual structure (CHRS) consisted of several nested residual modules for reducing parameters and extracting more representational features. Furthermore, we integrate channel attention (CA) mechanism into CHRS to generate channel hourglass residual block (CHRB) which can be easily extended to other methods for improving performance. We also propose channel hourglass residual network (CHRN) which not only pays attention to network learning efficiency but also learns more discriminative expressions. Extensive experiments demonstrate the effectiveness of our CHRN and the generalization ability of our CHRB. Fangwei Hao, XinDi Ma, Taiping Zhang, Yuan Yan Tang |
IJCNN | 3 |
| 2021 | Target-aware for Few-shot SegmentationabstractFew-shot segmentation refers to learn a segmentation model that can be generalized to novel classes with limited labeled images. Establishing the correspondence between support images and query images effectively has a considerable effect on guiding the segmentation of query images. Most existing methods mainly adopt a trained classification network as the backbone, nevertheless, the classification tasks only focus on the most discriminate regions of the target rather than the targets' integrity and the most discriminate regions may not be part of the target we need to segment while multiple classes object included in images. Besides, there exists another question that the most discriminate regions of the target in support image also do not necessarily appear in query images because of occlusion or incomplete object. All these may cause the correspondence between two images inaccurately. To tackle these problems, we propose a Target-aware Network(TaNet). Our network has two objectives: (1) increasing both intra-object similarity and inter-object dissimilarity for query image and support image to make each object more complete rather than highlight the most discriminate regions; (2) adaptively generating target-aware correspondence between support images and query images. Experiments on PASCAL-5iand COCO-20ishow that our method achieves state-of-the-art performance. Xiaoliu Luo, Taiping Zhang, Zhao Duan |
IJCNN | 2 |
| 2021 | Indian Buffet Process-Based on Nonnegative Matrix Factorization with Single Binary ComponentabstractNonnegative matrix factorization seeks to find a basic matrix and a weight matrix to approximate the nonnegative matrix. It has proven to be a powerful low-rank decomposition technique for nonnegative multivariate data. However, its performance largely depends on the assumption of a fixed number of features. In this work, we propose a new probabilistic nonnegative matrix factorization which factorizes a nonnegative matrix into a low-rank factor matrix with {0,1} constraints and a nonnegative weight matrix. In order to automatically learn the potential binary features and feature number. A deterministic Indian buffet process variational inference is introduced to obtain the binary factor matrix. And the weight matrix is set to satisfy the exponential prior. In order to obtain the real posterior distribution of the two factor matrices, a variational Bayesian exponential Gaussian inference model is established. The comparative experiments on both the synthetic and real-world data sets show the efficacy of the proposed method. XinDi Ma, Taiping Zhang, Yuan Yan Tang |
SMC | 4 |
| 2021 | A robust image representation method against illumination and occlusion variations
Taiping Zhang, Linchang Zhao, Xiaoliu Luo, Yuan Yan Tang |
Image Vis. Comput. | 2 |
| 2021 | DCKN: Multi-focus image fusion via dynamic convolutional kernel network
Zhao Duan, Taiping Zhang, Xiaoliu Luo |
Signal Process. | 2 |
| 2021 | Multi-focus image fusion with Geometrical Sparse Representation
Taiping Zhang, Linchang Zhao, Xiaoliu Luo, Yuan Yan Tang |
Signal Process. Image Commun. | 2 |
| 2020 | Adaptive parameter estimation of GMM and its application in clustering
Linchang Zhao, Zhaowei Shang, Xiaoliu Luo, Taiping Zhang, Yuan Yan Tang |
Future Gener. Comput. Syst. | 5 |
| 2019 | Distribution Preserving Network EmbeddingabstractThe deep autoencoder network which is based on constraining non-negative weights, can learn a low dimensional part-based representation. On the other hand, the inherent structure of the each data cluster can be described by the distribution of the intraclass sample. Then one hopes to learn a new low dimensional feature which can preserve the intrinsic structure embedded in the high dimensional data space perfectly. In this paper, by preserving data distribution, a deep part-based representation can be learned, and the novel algorithm is called Distribution Preserving Network Embedding (DPNE). In DPNE, we first need to estimate the distribution of the original data, and then we seek a part-based representation which respects the distribution. The experimental results on real-world data sets show that the proposed algorithm has good performance in terms of cluster accuracy and adjusted mutual information (AMI). Anyong Qin, Zhaowei Shang, Taiping Zhang, Yuan Yan Tang |
ICASSP | 3 |
| 2019 | A cost-sensitive meta-learning classifier: SPFCNN-Miner
Linchang Zhao, Zhaowei Shang, Anyong Qin, Taiping Zhang, Yuan Yan Tang |
Future Gener. Comput. Syst. | 4 |
| 2019 | MSSTResNet-TLD: A robust tracking method based on tracking-learning-detection framework by using multi-scale spatio-temporal residual network feature model
Bing Liu 0019, Qiao Liu 0001, Taiping Zhang |
Neurocomputing | 3 |
| 2019 | Software defect prediction via cost-sensitive Siamese parallel fully-connected neural networks
Linchang Zhao, Zhaowei Shang, Taiping Zhang, Yuan Yan Tang |
Neurocomputing | 4 |
| 2019 | MSST-ResNet: Deep multi-scale spatiotemporal features for robust visual object tracking
Bing Liu 0019, Qiao Liu 0001, Taiping Zhang |
Knowl. Based Syst. | 4 |
| 2019 | Spectral-Spatial Graph Convolutional Networks for Semisupervised Hyperspectral Image ClassificationabstractCollecting labeled samples is quite costly and time-consuming for hyperspectral image (HSI) classification task. Semisupervised learning framework, which combines the intrinsic information of labeled and unlabeled samples, can alleviate the deficient labeled samples and increase the accuracy of HSI classification. In this letter, we propose a novel semisupervised learning framework that is based on spectral-spatial graph convolutional networks (S2GCNs). It explicitly utilizes the adjacency nodes in graph to approximate the convolution. In the process of approximate convolution on graph, the proposed method makes full use of the spatial information of the current pixel. The experimental results on three real-life HSI data sets, i.e., Botswana Hyperion, Kennedy Space Center, and Indian Pines, show that the proposed S2GCN can significantly improve the classification accuracy. For instance, the overall accuracy on Indian data is increased from 66.8% (GCN) to 91.6%. Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Yulong Wang 0002, Taiping Zhang, Yuan Yan Tang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2019 | Multi-view common component discriminant analysis for cross-view classification
Xinge You, Jiamiao Xu, Wei Yuan 0001, Xiaoyuan Jing, Dacheng Tao, Taiping Zhang |
Pattern Recognit. | 6 |
| 2018 | Distribution preserving learning for unsupervised feature selection
Ting Xie 0004, Taiping Zhang, Yuan Yan Tang |
Neurocomputing | 3 |
| 2018 | Multi-view manifold learning with locality alignment
Xinge You, Shujian Yu, Chang Xu 0002, Wei Yuan 0001, Xiaoyuan Jing, Taiping Zhang, Dacheng Tao |
Pattern Recognit. | 7 |
| 2018 | Video-Based Person Re-Identification by Simultaneously Learning Intra-Video and Inter-Video Distance MetricsabstractVideo-based person re-identification (re-id) is an important application in practice. Since large variations exist between different pedestrian videos, as well as within each video, it's challenging to conduct re-identification between pedestrian videos. In this paper, we propose a simultaneous intra-video and inter-video distance learning (SI2DL) approach for video-based person re-id. Specifically, SI2DL simultaneously learns an intravideo distance metric and an inter-video distance metric from the training videos. The intra-video distance metric is used to make each video more compact, and the inter-video one is used to ensure that the distance between truly matching videos is smaller than that between wrong matching videos. Considering that the goal of distance learning is to make truly matching video pairs from different persons be well separated with each other, we also propose a pair separation based SI2DL (P-SI2DL). P-SI2DL aims to learn a pair of distance metrics, under which any two truly matching video pairs can be well separated. Experiments on four public pedestrian image sequence datasets show that our approaches achieve the state-of-the-art performance. Xiaoke Zhu, Xiaoyuan Jing, Xinge You, Xinyu Zhang 0012, Taiping Zhang |
IEEE Trans. Image Process. | 5 |
| 2017 | Learning the Distribution Preserving Semantic Subspace for ClusteringabstractThis paper proposes a new clustering method for images called distribution preserving indexing (DPI). It aims to find a lower dimensional semantic space approximating the original image space in the sense of preserving the distribution of the data. In the theory, the intrinsic structure of the data clusters can be described by the distribution of the data effectively. Therefore, the cluster structure of the data in a lower dimensional semantic space derived by the DPI becomes clear. Unlike these distance-based clustering methods, which reveal the intrinsic Euclidean structure of data, our method attempts to discover the intrinsic cluster structure of the data space that actually is the union of some sub-manifolds. Moreover, we propose a revised kernel density estimator for the case of high-dimensional data, which is a crucial step in DPI. In addition, we provide a theoretical analysis of the bound of our method. Finally, the extensive experiments compared with other algorithms, on COIL20, CBCL, and MNIST demonstrate the effectiveness of our proposed approach. Jinyu Tian 0001, Taiping Zhang, Anyong Qin, Zhaowei Shang, Yuan Yan Tang |
IEEE Trans. Image Process. | 2 |
| 2016 | Learning Proximity Relations for Feature SelectionabstractThis work presents a feature selection method based on proximity relations learning. Each single feature is treated as a binary classifier that predicts for any three objects X, A, and B whether X is close to A or B. The performance of the classifier is a direct measure of feature quality. Any linear combination of feature-based binary classifiers naturally corresponds to feature selection. Thus, the feature selection problem is transformed into an ensemble learning problem of combining many weak classifiers into an optimized strong classifier. We provide a theoretical analysis of the generalization error of our proposed method which validates the effectiveness of our proposed method. Various experiments are conducted on synthetic data, four UCI data sets and 12 microarray data sets, and demonstrate the success of our approach applying to feature selection. A weakness of our algorithm is high time complexity. Taiping Zhang, Yuan Yan Tang, C. L. Philip Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | An approximate closed-form solution to correlation similarity discriminant analysis
Taiping Zhang, Yuan Yan Tang, C. L. Philip Chen, Zhaowei Shang, Bin Fang 0001 |
Neurocomputing | 1 |
| 2013 | Learning orthogonal projections for Isomap
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang, Ruizong Liu |
Neurocomputing | 4 |
| 2012 | Orthogonal Isometric Projection
Yuan Yan Tang, Bin Fang 0001, Taiping Zhang |
ICPR | 4 |
| 2012 | Document Clustering in Correlation Similarity Measure SpaceabstractThis paper presents a new spectral clustering method called correlation preserving indexing (CPI), which is performed in the correlation similarity measure space. In this framework, the documents are projected into a low-dimensional semantic space in which the correlations between the documents in the local patches are maximized while the correlations between the documents outside these patches are minimized simultaneously. Since the intrinsic geometrical structure of the document space is often embedded in the similarities between the documents, correlation as a similarity measure is more suitable for detecting the intrinsic geometrical structure of the document space than euclidean distance. Consequently, the proposed CPI method can effectively discover the intrinsic structures embedded in high-dimensional document space. The effectiveness of the new method is demonstrated by extensive experiments conducted on various data sets and by comparison with existing document clustering methods. Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Yong Xiang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | A Least-Squares Model to Orthogonal Linear Discriminant AnalysisabstractOrthogonal transformation can delete the correlations among candidate features such that the extracted features do not disturb each other. An orthogonal set of discriminant vectors is more powerful than the classical discriminant vectors. In this paper, we present a new orthogonal linear discriminant analysis (OLDA) model based on least-squares approximation called LS-OLDA for pattern classification, which aims to find an orthogonal transformation W and a diagonal matrix D such that the difference between [Formula: see text] and WDWT is minimized in the least-squares sense, and the trace of D is maximized simultaneously. Theoretical analysis shows that the proposed model coincides with classical OLDA criterion. The experimental results on different standard data sets compared with related methods show that LS-OLDA achieves or approximates closely to the best accuracy, and has lower computational cost. Taiping Zhang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2010 | Incremental Embedding and Learning in the Local Discriminant Subspace With Application to Face RecognitionabstractDimensionality reduction and incremental learning have recently received broad attention in many applications of data mining, pattern recognition, and information retrieval. Inspired by the concept of manifold learning, many discriminant embedding techniques have been introduced to seek low-dimensional discriminative manifold structure in the high-dimensional space for feature reduction and classification. However, such graph-embedding framework-based subspace methods usually confront two limitations: (1) since there is no available updating rule for local discriminant analysis with the additive data, it is difficult to design incremental learning algorithm and (2) the small sample size (SSS) problem usually occurs if the original data exist in very high-dimensional space. To overcome these problems, this paper devises a supervised learning method, called local discriminant subspace embedding (LDSE), to extract discriminative features. Then, the incremental-mode algorithm, incremental LDSE (ILDSE), is proposed to learn the local discriminant subspace with the newly inserted data, which applies incremental learning extension to the batch LDSE algorithm by employing the idea of singular value-decomposition (SVD) updating algorithm. Furthermore, the SSS problem is avoided in our method for the high-dimensional data and the benchmark incremental learning experiments on face recognition show that ILDSE bears much less computational cost compared with the batch algorithm. Bin Fang 0001, Yuan Yan Tang, Taiping Zhang |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2010 | Generalized Discriminant Analysis: A Matrix Exponential ApproachabstractLinear discriminant analysis (LDA) is well known as a powerful tool for discriminant analysis. In the case of a small training data set, however, it cannot directly be applied to high-dimensional data. This case is the so-called small-sample-size or undersampled problem. In this paper, we propose an exponential discriminant analysis (EDA) technique to overcome the undersampled problem. The advantages of EDA are that, compared with principal component analysis (PCA) + LDA, the EDA method can extract the most discriminant information that was contained in the null space of a within-class scatter matrix, and compared with another LDA extension, i.e., null-space LDA (NLDA), the discriminant information that was contained in the non-null space of the within-class scatter matrix is not discarded. Furthermore, EDA is equivalent to transforming original data into a new space by distance diffusion mapping, and then, LDA is applied in such a new space. As a result of diffusion mapping, the margin between different classes is enlarged, which is helpful in improving classification accuracy. Comparisons of experimental results on different data sets are given with respect to existing LDA extensions, including PCA + LDA, LDA via generalized singular value decomposition, regularized LDA, NLDA, and LDA via QR decomposition, which demonstrate the effectiveness of the proposed EDA method. Taiping Zhang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | Weightiness image Partition in 3D Face RecognitionabstractIn this paper we present a novel algorithm suitable to improve the accuracy of 3D face recognition. In the proposed algorithm, we represent the 3D points by point signatures and partition the facial data into fifteen regions according to ¿three courtyards and five eyes¿ theory in pencil sketch on facial image in Chinese traditional art. Then in each partition we use ICA getting eigenvalues of feature and structure character and depth information to represent the 3D facial data. We assign different weightiness to each sub-image according to the result of sub-image variety. In order to match incomplete data under structural constraints, we proposed a reformative robust structural Hausdorff distance to handle these possible cases. Experiments on FRGC v2.0 data set show that the proposed algorithm is robust and effective to 3D face with expression, lighting and expression variance. Yuan Yan Tang, Bin Fang 0001, Taiping Zhang |
SMC | 4 |
| 2009 | Combining Eodh and Directional Gradient Density for Offline Signature VerificationabstractThe main problem to identify skilled forgeries for offline signature verification lies in the fact that it is difficult to formalize distinguished feature representation of the signature patterns and design appropriate fusion scheme for various types of feature vectors. To tackle these problems, in this paper, we propose an approach to extract robust Edge Orientation Distance Histogram (EODH) descriptor which effectively reflects signature structure variations. In addition, directional gradient density features are employed for skilled forgery verification attempt. To exploit the full capacity of two sets of features, we designed the multilevel weighted fuzzy classifier and fuse match scores by way of selection priority. Experiments were conducted on a subcorpus of open MCYT signature database which is widely used for performance evaluation. It shows that the proposed method was able to improve verification accuracy. Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang, Taiping Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2009 | Model-based signature verification with rotation invariant features
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang |
Pattern Recognit. | 4 |
| 2009 | Multiscale facial structure representation for face recognition under varying illumination
Taiping Zhang, Bin Fang 0001, Yuan Yuan 0001, Yuan Yan Tang, Zhaowei Shang, Fangnian Lang |
Pattern Recognit. | 1 |
| 2009 | Face Recognition Under Varying Illumination Using GradientfacesabstractIn this correspondence, we propose a novel method to extract illumination insensitive features for face recognition under varying lighting called the Gradientfaces. Theoretical analysis shows Gradientfaces is an illumination insensitive measure, and robust to different illumination, including uncontrolled, natural lighting. In addition, Gradientfaces is derived from the image gradient domain such that it can discover underlying inherent structure of face images since the gradient domain explicitly considers the relationships between neighboring pixel points. Therefore, Gradientfaces has more discriminating power than the illumination insensitive measure extracted from the pixel domain. Recognition rates of 99.83% achieved on PIE database of 68 subjects, 98.96% achieved on Yale B of ten subjects, and 95.61% achieved on Outdoor database of 132 subjects under uncontrolled natural lighting conditions show that Gradientfaces is an effective method for face recognition under varying illumination. Furthermore, the experimental results on Yale database validate that Gradientfaces is also insensitive to image noise and object artifacts (such as facial expressions). Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang |
IEEE Trans. Image Process. | 1 |
| 2008 | Total variation norm-based nonnegative matrix factorization for identifying discriminant representation of image patterns
Taiping Zhang, Bin Fang 0001, Weining Liu, Yuan Yan Tang |
Neurocomputing | 1 |
| 2008 | Topology Preserving Non-negative Matrix Factorization for Face RecognitionabstractIn this paper, a novel topology preserving non-negative matrix factorization (TPNMF) method is proposed for face recognition. We derive the TPNMF model from original NMF algorithm by preserving local topology structure. The TPNMF is based on minimizing the constraint gradient distance in the high-dimensional space. Compared with L(2) distance, the gradient distance is able to reveal latent manifold structure of face patterns. By using TPNMF decomposition, the high-dimensional face space is transformed into a local topology preserving subspace for face recognition. In comparison with PCA, LDA, and original NMF, which search only the Euclidean structure of face space, the proposed TPNMF finds an embedding that preserves local topology information, such as edges and texture. Theoretical analysis and derivation given also validate the property of TPNMF. Experimental results on three different databases, containing more than 12,000 face images under varying in lighting, facial expression, and pose, show that the proposed TPNMF approach provides a better representation of face patterns and achieves higher recognition rates than NMF. Taiping Zhang, Bin Fang 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 1 |
| 2007 | Offline signature verification: A new rotation invariant approachabstractRotation problem is one of the major difficulties to distinguish signature patterns in off-line skilled signature verification. This paper presents a new approach utilizing Ring- Peripheral features to tackle this problem. In principle, Ring- Peripheral features are able to describe internal and external structure of signatures with different phase shift. In order to extract stable and consistent presentation of signature patterns for verification purpose, FFT is used to eliminate phase effects. In the training samples stage, we employ a selection function to pick up reasonable samples for better threshold estimation. Experiment results demonstrated that the proposed method was successful to improve verification accuracy. Bin Fang 0001, Yuan Yan Tang, Taiping Zhang |
SMC | 4 |
| 2003 | Observational and theoretical study of spectrally resolved ocean optical propertiesabstractIn order to verify the COVE (CERES Ocean Validation Experiment) platform measurements and characterize the ocean optical properties, airborne measurements of spectral and broadband shortwave upwelling and downwelling radiative fluxes were made onboard the NASA Langley OV-10 aircraft between July, 2001 and March, 2003. The processed data and derived spectral albedos are compared with other observed data and model simulation results, and the comparisons show consistently good agreement. Taiping Zhang, William L. Smith Jr., Thomas P. Charlock, C. Ken Rutledge, Zhonghai Jin, Glenn Cota, Bryan E. Fabbri |
IGARSS | 1 |