Yangtao Wang

dblp:150/8599 · DBLP profile ↗
← Back
52ranked-venue papers
18as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 14 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language Models
abstract
Knowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines.
Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang
AAAI1
2026 WaveSculpt: Text-to-3D generation with wavelet-guided score distillation
Weilong Peng, Jianhui Huo, Keke Tang, Yangtao Wang, Yan Wang 0022, Meie Fang
Comput. Aided Geom. Des.4
2026 Skew-normal distributions for modeling asymmetric moving tendencies in pedestrian trajectories
Siyuan Chen 0005, Yatie Xiao, Yangtao Wang, Yanzhao Xie, Tong Zhu 0003, Jinbiao Chen
Neurocomputing3
2026 Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002
Pattern Recognit.2
2026 F3-SD: Focal feature fusion with self-distillation on large vision-language models for cross-modal retrieval
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.2
2026 Prompt-affinity multi-modal class centroids for unsupervised domain adaption
abstract
In recent years, the advancements in large vision-language models (VLMs) like CLIP have sparked a renewed interest in leveraging the prompt learning mechanism to preserve semantic consistency between source and target domains in unsupervised domain adaption (UDA). While these approaches show promising results, they encounter fundamental limitations when quantifying the similarity between source and target domain data , primarily stemming from the redundant and modality-missing class centroids . To address these limitations, we propose P rompt-affinity M ulti-modal C lass C entroids for UDA (termed as PMCC). Firstly, we fuse the text class centroids (directly generated from the text encoder of CLIP with manual prompts for each class) and image class centroids (generated from the image encoder of CLIP for each class based on source domain images) to yield the multi-modal class centroids. Secondly, we conduct the cross-attention operation between each source or target domain image and these multi-modal class centroids. In this way, these class centroids that contain rich semantic information of each class will serve as a bridge to effectively measure the semantic similarity between different domains. Finally, we design a logit bias head and employ a multi-modal prompt learning mechanism to accurately predict the true class of each image for both source and target domains. We conduct extensive experiments on 4 popular UDA datasets including Office-31, Office-Home, VisDA-2017, and DomainNet. The experimental results validate our PMCC achieves higher performance with lower model complexity than the state-of-the-art (SOTA) UDA methods. The code of this project is available at GitHub: https://github.com/246dxw/PMCC .
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.2
2026 Cross-domain distillation for unsupervised domain adaptation with large vision-language models
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Maobin Tang, Meie Fang, Wensheng Zhang 0002
Pattern Recognit.2
2026 RDS-Net: Recursive structure refinement network with Dual-awareness Shape transfer for point cloud completion
Xiaocui Li 0001, Weili Chen, Xinyu Zhang 0012, Yangtao Wang, Wei Liang 0005, Keqin Li 0001
Pattern Recognit.4
2026 Adaptive message passing mechanism for graph neural networks
Yangtao Wang, Linruo Liu, Yanzhao Xie, Maobin Tang, Xiaocui Li 0001
Pattern Recognit.1
2026 PTPD: Prototype-Guided Triplet Prompt Distillation with Vision-language models
Yanzhao Xie, Yangtao Wang, Rukai Wei, Dandan Shao, Maobin Tang, Meie Fang, Weilong Peng, Lisheng Fan, Wensheng Zhang 0002
Pattern Recognit.3
2026 MKGPL: graph prompt learning with multi-view knowledge for few-shot recognition
Yanzhao Xie, Man Qiu, Yangtao Wang, Siyuan Chen 0005, Meie Fang, Maobin Tang, Wensheng Zhang 0002
Pattern Recognit.3
2026 Progressive Hybrid Pseudo-Labeling for Unsupervised Domain Adaptation With Ascending Low-Rank Adaptation
abstract
Unsupervised domain adaptation (UDA) based on large vision-language models (VLMs) has recently demonstrated strong generalization ability, yet it remains fundamentally challenged by noisy pseudo-labels and inefficient adaptation under large domain shifts. In this paper, we propose Progressive Hybrid Pseudo-Labeling for UDA with Ascending Low-Rank Adaptation (termed as PHPL), a parameter-efficient paradigm that addresses these challenges from two complementary perspectives. 1) We introduce a progressive hybrid pseudo-labeling strategy that constructs target-domain supervision by fusing predictions from a frozen teacher model and an adaptive student model with a progressive weighting scheme. By gradually transferring predictive responsibility from the teacher to the student during training, PHPL effectively mitigates early-stage pseudo-label noise and stabilizes self-training under large domain shifts. 2) To enable efficient and stable adaptation of large VLMs, we propose an ascending low-rank adaptation strategy that allocates LoRA capacity in a depth-aware manner. Specifically, larger low-rank updates are assigned to deeper, semantically richer layers, while shallow layers remain lightly parameterized, striking a favorable balance between parameter efficiency and representational expressiveness. We conduct extensive experiments on five widely-used UDA benchmarks, including Office-Home, Office-31, VisDA-2017, Mini-DomainNet, and DomainNet. Experimental results verify that PHPL consistently achieves higher performance across various cross-domain scenarios compared with existing CNN, Transformer, and VLMs-based solutions. Notably, PHPL demonstrates strong robustness on highly challenging large-scale conditions while requiring significantly less computational overhead, validating the effectiveness and scalability of the proposed lightweight adaptation paradigm. The code is available at https://github.com/el2k/PHPL.
Yangtao Wang, Mingxin Huang, Xingwei Deng, Yanzhao Xie, Xiaocui Li 0001
IEEE Trans. Image Process.1
2026 High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream Transformers
abstract
Recently, most image-text matching (ITM) approaches have embraced a dual-stream transformer architecture to facilitate the learning and alignment of cross-modal semantic information. Despite the efficacy of this methodology in bridging the semantic disparity between images and texts, it exhibits two primary limitations. Firstly, it falls short in discriminating the nuanced similarities among features, which leads to misleading outcomes or even compromises the overall ITM process. Secondly, the conventional triplet training paradigm relies on a pre-determined, fixed margin coefficient, thereby impeding its capacity to accurately gauge the similarity relationships between positive and negative samples. In this article, we propose high feature D istinguishability for A daptive I mage-text M atching with dual-stream transformers (termed as DAIM). To address the first limitation, we design a feature discriminability module to bring similar features closer together but with a certain degree of distinction and push dissimilar features farther apart, resulting in high feature distinguishability for accurate ITM. To address the second limitation, we devise a margin optimization module to perceive the similarity distribution between positive and negative samples in real-time during training, thereby adaptively adjusting the margin coefficient to minimize the cross-modal semantic gap to the greatest extent possible. Based on this, we align the multi-level (i.e., representations from low-, middle-, and high-layer transformer encoders) semantic information of cross-modal data by adaptively optimizing the semantic distributions of positive and negative samples. We conduct extensive experiments on two commonly used benchmark datasets, including MSCOCO and Flickr30K. Experimental results verify that DAIM can achieve a higher performance (e.g., 4.7% RSUM gain on MSCOCO) than the state-of-the-art ITM methods. The open-sourced code of this project is available at: https://github.com/Hudjkfhdsjfhdjkg/DAIM.git .
Yangtao Wang, Weibin Huang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, Wensheng Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2025 From Pixels to Shapes: Generative AI for 2D Images and 3D Models
Jianhui Huo, Shijian Xu, Weilong Peng, Yangtao Wang, Yan Wang 0022, Meie Fang
ICIC (19)5
2025 Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text Retrieval
abstract
Image-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA.
Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016
ICME2
2025 Incomplete Multi-view Clustering via Local Reasoning and Correlation Analysis
abstract
In recent years, incomplete multi-view clustering (IMVC) has attracted considerable attention for its ability to acheieve effective clustering results through the integration of key information amidst missing view. However, the existing IMVC methods are still faced with 3 limitations: (1) They exhibit deficiencies in considering the weight distribution within views, (2) they ignore the varying contributions of different views to the common consistent representation, and (3) they struggle to sufficiently extract and recover the vital information within incomplete views. To address these limitations, we incorporates local reasoning and correlation analysis to design an incomplete multi-view clustering method(IMVCLRCA), which introduces a new strategy of feature learning and missing view recovery, fully exploiting local similarity and structural continuity within views and performing precise local reasoning recovery on missing data. By maximizing mutual information between views through contrastive learning, we achieve the consistent representation learning of multiple views. Furthermore, based on semantic consistency, we comprehensively consider the correlation between views, utilized a weight matrix to fuse cross-view data, and constructed a view with a correlation structure, ultimately obtaining a common consistent representation. We conduct extensive experiments on 4 public datasets including Caltech101-20, BBCSport, Scene-15, and LandUse-21. Experimental results demonstrate that IMVCLRCA has higher accuracy and robustness compared to the state-of-the-art IMVC methods. The anonymous code of this project is available on GitHub at https://github.com/ggg2111/2025WSDM-IMVCLRCA.
Xiaocui Li 0001, Xinyu Zhang 0012, Yangtao Wang, Qingyu Shi 0001, Wei Liang 0006
WSDM4
2025 Angle Metric Learning for Discriminative Features on Vehicle Re-Identification
abstract
ABSTRACT Vehicle re‐identification (Re‐ID) facilitates the recognition and distinction of vehicles based on their visual characteristics in images or videos. However, accurately identifying a vehicle poses great challenges due to (i) the pronounced intra‐instance variations encountered under varying lighting conditions such as day and night and (ii) the subtle inter‐instance differences observed among similar vehicles. To address these challenges, the authors propose A ngle M etric learning for D iscriminative F eatures on vehicle Re‐ID (termed as AMDF), which aims to maximise the variance between visual features of different classes while minimising the variance within the same class. AMDF comprehensively measures the angle and distance discrepancies between features. First, to mitigate the impact of lighting conditions on intra‐class variation, the authors employ CycleGAN to generate images that simulate consistent lighting (either day or night), thereby standardising the conditions for distance measurement. Second, Swin Transformer was integrated to help generate more detailed features. At last, a novel angle metric loss based on cosine distance is proposed, which organically integrates angular metric and 2‐norm metric, effectively maximising the decision boundary in angular space. Extensive experimental evaluations on three public datasets including VERI‐776, VERI‐Wild, and VEHICLEID, indicate that the method achieves state‐of‐the‐art performance. The code of this project is released at https://github.com/ZnCu‐0906/AMDF .
Yutong Xie 0009, Shuoqi Zhang, Lide Guo, Rukai Wei, Yanzhao Xie, Yangtao Wang, Maobin Tang, Lisheng Fan
IET Comput. Vis.7
2025 Adaptive Multi-Lens Phase Modulation for Scale-Aware Privacy-Preserving Human Pose Recognition
abstract
Recently, optical privacy protection has emerged as a promising approach for safeguarding visual privacy at the physical acquisition stage. However, existing methods often face a trade‐off between privacy strength and human pose recognition accuracy, particularly in long‐range and multi‐scale scenarios. To address this challenge, we propose a novel adaptive optical privacy‐preserving framework that integrates a learnable optical modulation system with a human pose recognition network. The core of our method lies in a sparse‐weighted multi‐lens model, where a lightweight multilayer perceptron (MLP) predicts a sparse set of coefficients to linearly combine predefined lens phase profiles based on facial region geometry. This enables dynamic control over the point spread function (PSF), adapting the degree of image degradation to subject scale in real time. Additionally, we introduce a privacy‐aware loss function that selectively reduces facial localization accuracy while preserving body pose information. Extensive experiments on MSCOCO and FLIC datasets demonstrate that the proposed method achieves a favorable balance between privacy protection and pose estimation, outperforming previous optical‐ and software‐based baselines.
Weilong Peng, Quanwei Deng, Mingjie Li 0004, Yangtao Wang, Yan Wang 0022, Lisheng Fan, Meie Fang
IET Softw.4
2025 HSALC: hard sample aware label correction for medical image classification
Yangtao Wang, Yicheng Ye, Yanzhao Xie, Maobin Tang, Lisheng Fan
Multim. Tools Appl.1
2024 Domain Alignment with Large Vision-language Models for Cross-domain Remote Sensing Image Retrieval
abstract
Cross-domain remote sensing image retrieval has been a hotspot in the past few years. Most of the existing methods focus on combining semantic learning with domain adaptation on well-labeled source domain and unlabeled target domain. However, they face two serious challenges. (1) They cannot deal with practical scenarios where the source domain lacks sufficient label supervision. (2) They suffer from severe performance degradation when the data distribution between the source domain and target domain becomes highly inconsistent. To address these challenges, we propose D omain A lignment with L arge V ision-language models for cross-domain remote sensing image retrieval (termed as DALV). First, we design a dual-modality prototype guided pseudo-labeling mechanism, which leverages the pre-trained large vision-language model (i.e., CLIP) to assign pseudo-labels for all unlabeled source domain images and target domain images. Second, we compute the confidence scores for these pseudo-labels to distinguish their reliability. Next, we devise a loss reweighting strategy, which incorporates the confidence scores as weight values into the contrastive loss to mitigate the impact of noisy pseudo-labels. Finally, the low-rank adaptation fine-tuning means is adapted to update our model and achieve domain alignment to obtain class discriminative features. Extensive experiments on 12 cross-domain remote sensing image retrieval tasks show that our proposed DALV outperforms the state-of-the-art approaches. The source code is available at https://github.com/ptyy01/DALV.
Guocan Cai, Fufang Li, Yangtao Wang, Xin Tan 0002, Xiaocui Li 0001
CIKM4
2024 Image-text Retrieval with Main Semantics Consistency
abstract
Image-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC.
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang
CIKM2
2024 IE-aware Consistency Losses for Detailed 3D Face Reconstruction from Multiple Images in the Wild
abstract
3D face reconstruction from multiple in-the-wild images in an unsupervised manner poses a significant challenge, primarily due to the pervasive presence of Intrinsic and Extrinsic inconsistencies in facial features. To tackle this, we introduce a novel set of IE-aware consistency losses designed to effectively mitigate these inconsistencies. Our Local Alignment Loss employs neighborhood search techniques to identify and optimize consistent pixel information, thereby reducing intrinsic inconsistencies. In parallel, our Region Subset Selection Loss filters out regions where significant discrepancies exist between the input and reconstructed images, effectively alleviating extrinsic inconsistencies. Extensive experimental results validate the effectiveness of our IE-aware consistency losses in reconstructing detailed 3D facial geometry from images captured in uncontrolled environments.
Weilong Peng, Keke Tang, Kongyang Chen, Yangtao Wang, Ping Li 0016, Meie Fang
ICME5
2023 A hash centroid construction method with Swin transformer for multi-label image retrieval
Yanzhao Xie, Yangtao Wang, Rukai Wei, Yu Liu 0040, Ke Zhou 0001, Lisheng Fan
Neural Comput. Appl.2
2023 TokenCut: Segmenting Objects in Images and Videos With Self-Supervised Transformer and Normalized Cut
abstract
In this paper, we describe a graph-based algorithm that uses the features obtained by a self-supervised transformer to detect and segment salient objects in images and videos. With this approach, the image patches that compose an image or video are organised into a fully connected graph, in which the edge between each pair of patches is labeled with a similarity score based on the features learned by the transformer. Detection and segmentation of salient objects can then be formulated as a graph-cut problem and solved using the classical Normalized Cut algorithm. Despite the simplicity of this approach, it achieves state-of-the-art results on several common image and video detection and segmentation tasks. For unsupervised object discovery, this approach outperforms the competing approaches by a margin of 6.1%, 5.7%, and 2.6% when tested with the VOC07, VOC12, and COCO20 K datasets. For the unsupervised saliency detection task in images, this method improves the score for Intersection over Union (IoU) by 4.4%, 5.6% and 5.2%. When tested with the ECSSD, DUTS, and DUT-OMRON datasets. This method also achieves competitive results for unsupervised video object segmentation tasks with the DAVIS, SegTV2, and FBMS datasets.
Yangtao Wang, Xi Shen 0001, Yuan Yuan 0002, Yuming Du, Maomao Li, Shell Xu Hu, James L. Crowley, Dominique Vaufreydaz
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Label-Affinity Self-Adaptive Central Similarity Hashing for Image Retrieval
abstract
Due to the usage of global similarity, the hashing methods based on predefined hash centers have achieved more accurate retrieval results than the pairwise/triplet-based methods. Nevertheless, the fixed hash centers lack the perception of data distribution and are limited by the pre-determined Hadamard matrix, which consider neither the label semantic information nor the object scale size, resulting in sub-optimal retrieval performance and weak generalization ability. In this paper, we (1) adopt the label semantic information to generate self-adaptive hash centers and (2) propose the label-affinity coefficient (lac) that considers the scale size of each label/object appearing in the given image to calculate the real hash centroid for this image. Based on this, we proposeLabel-affinity Self-adaptive Central Similarity Hashing (LSCSH)for image retrieval. LSCSH consists of a hash code generator module and a hash center adapter module. First, we obtain the label word vector (i.e., the word vector representation of each class label) via the Word2Vector technique to generate and update the hash centers that adapt to the distribution of both label word vectors and generated hash codes. Second, we learnlacto indicate the dominance of different labels corresponding to objects in each given image, which considers the unequal scales of each object (corresponding to a label) to calculate a more accurate hash centroid for each image. Last but not least, we design an asynchronous learning mechanism to enable each hash code and its corresponding hash centroid to adapt to each other dynamically. We conduct extensive experiments on 5 image datasets including CIFAR-10, ImageNet, VOC2012, MS-COCO and NUS-WIDE. The experimental results demonstrate that LSCSH can achieve the state-of-the-art visual retrieval performance on both single-label and multi-label image datasets. The code of this work is released at:https://github.com/lzHZWZ/LSCSH_sourcecode.git.
Yanzhao Xie, Rukai Wei, Jingkuan Song, Yu Liu 0040, Yangtao Wang, Ke Zhou 0001
IEEE Trans. Multim.5
2022 Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut
abstract
Transformers trained with self-supervision using selfdistillation loss (DINO) have been shown to produce attention maps that highlight salient foreground objects. In this paper, we show a graph-based method that uses the selfsupervised transformer features to discover an object from an image. Visual tokens are viewed as nodes in a weighted graph with edges representing a connectivity score based on the similarity of tokens. Foreground objects can then be segmented using a normalized graph-cut to group self-similar regions. We solve the graph-cut problem using spectral clustering with generalized eigen-decomposition and show that the second smallest eigenvector provides a cutting solution since its absolute value indicates the likelihood that a token belongs to a foreground object. Despite its simplicity, this approach significantly boosts the performance of unsupervised object discovery: we improve over the recent state-of-the-art LOST by a margin of 6.9%, 8.1%, and 8.1% respectively on the VOC07, VOC12, and COCO20K. The performance can be further improved by adding a second stage class-agnostic detector (CAD). Our proposed method can be easily extended to unsupervised saliency detection and weakly supervised object detection. For unsupervised saliency detection, we improve IoU for 4.9%, 5.2%, 12.9% on ECSSD, DUTS, DUT-OMRON respectively compared to state-of-the-art. For weakly supervised object detection, we achieve competitive performance on CUB and ImageNet. Our code is available at: https://www.m-psi.fr/Papers/TokenCut2022/
Yangtao Wang, Xi Shen 0001, Shell Xu Hu, Yuan Yuan 0002, James L. Crowley, Dominique Vaufreydaz
CVPR1
2022 Label graph learning for multi-label image recognition with cross-modal fusion
Yanzhao Xie, Yangtao Wang, Yu Liu 0040, Ke Zhou 0001
Multim. Tools Appl.2
2022 STMG: Swin transformer for multi-label image recognition with graph convolution network
Yangtao Wang, Yanzhao Xie, Lisheng Fan, Guangxing Hu
Neural Comput. Appl.1
2021 G-CAM: Graph Convolution Network Based Class Activation Mapping for Multi-label Image Recognition
abstract
In most multi-label image recognition tasks, human visual perception keeps consistent for different spatial transforms of the same image. Existing approaches either learn the perceptual consistency with only image-level supervision or preserve the middle-level feature consistency of attention regions but neglect the (global) label dependencies between different objects over the dataset. To address this issue, we integrate graph convolution network (GCN) and propose G-CAM, which learns visual attention consistency via GCN based class attention mapping (CAM) for multi-label image recognition. G-CAM consists of an image feature extraction module to generate the feature maps of the original image and its transformed one and a GCN module to learn weighted classifiers that capture the label dependencies between different objects. Different from previous works which use fully-connected classification layer, G-CAM first fuses weighted classifiers with the feature vector to generate the predicted labels for each input image, then combines weighted classifiers with the feature maps to respectively obtain the transformed attention heatmaps of the original image and the attention heatmaps of its transformed one. We can compute the attention consistency loss according to the distance between these two attention heatmaps. Finally, this loss is combined with the multi-label classification loss to update the whole network in an end-to-end manner. We conduct extensive experiments on three multi-label image datasets including FLICKR25K, MS-COCO and NUS-WIDE. Experimental results demonstrate G-CAM can achieve better performance compared with the state-of-the-art multi-label image recognition methods.
Yangtao Wang, Yanzhao Xie, Yu Liu 0040, Lisheng Fan
ICMR1
2021 Multi-view clustering via neighbor domain correlation learning
Xiaocui Li 0001, Ke Zhou 0001, Chunhua Li 0002, Xinyu Zhang 0012, Yu Liu 0040, Yangtao Wang
Neural Comput. Appl.6
2021 Unsupervised deep hashing with node representation for image retrieval
Yangtao Wang, Jingkuan Song, Ke Zhou 0001, Yu Liu 0040
Pattern Recognit.1
2020 Fast Graph Convolution Network Based Multi-label Image Recognition via Cross-modal Fusion
abstract
In multi-label image recognition, it has become a popular method to predict those labels that co-occur in an image via modeling the label dependencies. Previous works focus on capturing the correlation between labels, but neglect to effectively fuse the image features and label embeddings, which severely affects the convergence efficiency of the model and inhibits the further precision improvement of multi-label image recognition. To overcome this shortcoming, in this paper, we introduce Multi-modal Factorized Bilinear pooling (MFB) which works as an efficient component to fuse cross-modal embeddings and propose F-GCN, a fast graph convolution network (GCN) based multi-label image recognition model. F-GCN consists of three key modules: (1) an image representation learning module which adopts a convolution neural network (CNN) to learn and generate image representations, (2) a label co-occurrence embedding module which first obtains the label vectors via the word embeddings technique and then adopts GCN to capture label co-occurrence embeddings and (3) an MFB fusion module which efficiently fuses these cross-modal vectors to enable an end-to-end model with a multi-label loss function. We conduct extensive experiments on two multi-label datasets including MS-COCO and VOC2007. Experimental results demonstrate the MFB component efficiently fuses image representations and label co-occurrence embeddings and thus greatly improves the convergence efficiency of the model. In addition, the performance of image recognition has also been promoted compared with the state-of-the-art methods.
Yangtao Wang, Yanzhao Xie, Yu Liu 0040, Ke Zhou 0001, Xiaocui Li 0001
CIKM1
2020 Content Sifting Storage: Achieving Fast Read for Large-scale Image Dataset Analysis
abstract
Analyzing large-scale image dataset requires all images to be read from disks first, leading to high read latency. Therefore, we propose a Content Sifting Storage (CSS) system, which aims to reduce the read latency by only reading sifted relevant data. CSS generates embedded content metadata via deep learning and manages the metadata via Semantic Hamming Graph, which achieves fast read based on content similarity meeting the given analysis. Extensive experimental results on image datasets show that compared with conventional semantic storage systems, our CSS can greatly reduce the read latency by 82.21% to 94.8% with more than 98% recall rate.
Yu Liu 0040, Hong Jiang 0001, Yangtao Wang, Ke Zhou 0001, Li Liu 0047
DAC3
2020 Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost
abstract
Sector errors are a common type of error in modern disks. A sector error that occurs during I/O operations might cause inaccessibility of an application. Even worse, it could result in permanent data loss if the data is being reconstructed, and thereby severely affects the reliability of a storage system. Many disk scrubbing schemes have been proposed to solve this problem. However, existing approaches have several limitations. First, schemes use machine learning (ML) to predict latent sector errors (LSEs), but only leverage a single snapshot of training data to make a prediction, and thereby ignore sequential dependencies between different statuses of a hard disk over time. Second, they accelerate the scrubbing at a fixed rate based on the results of a binary classification model, which may result in unnecessary increases in scrubbing cost. Third, they naively accelerate the scrubbing of the full disk which has LSEs based on the predictive results, but neglect partial high-risk areas (the areas that have a higher probability of encountering LSEs). Lastly, they do not employ strategies to scrub these high-risk areas in advance based on I/O accesses patterns, in order to further increase the efficiency of scrubbing.We address these challenges by designing a Tier-Scrubbing (TS) scheme that combines a Long Short-Term Memory (LSTM) based Adaptive Scrubbing Rate Controller (ASRC), a module focusing on sector error locality to locate high-risk areas in a disk, and a piggyback scrubbing strategy to improve the reliability of a storage system. Our evaluation results on realistic datasets and workloads from two real world data centers demonstrate that TS can simultaneously decrease the Mean-Time-To-Detection (MTTD) by about 80% and the scrubbing cost by 20%, compared to a state-of-the-art scrubbing scheme.
Ji Zhang 0010, Yuanzhang Wang, Yangtao Wang, Ke Zhou 0001, Sebastian Schelter, Ping Huang 0001, Yong-guang Ji
DAC3
2020 A Machine Learning Based Write Policy for SSD Cache in Cloud Block Storage
abstract
Nowadays, SSD cache plays an important role in cloud storage systems. The associated write policy, which enforces an admission control policy regarding filling data into the cache, has a significant impact on the performance of the cache system and the amount of write traffic to SSD caches. Based on our analysis on a typical cloud block storage system, approximately 47.09% writes are write-only, i.e., writes to the blocks which have not been read during a certain time window. Naively writing the write-only data to the SSD cache unnecessarily introduces a large number of harmful writes to the SSD cache without any contribution to cache performance. On the other hand, it is a challenging task to identify and filter out those write-only data in a real-time manner, especially in a cloud environment running changing and diverse workloads.In this paper, to alleviate the above cache problem, we propose an ML-WP, Machine Learning Based Write Policy, which reduces write traffic to SSDs by avoiding writing write-only data. The main challenge in this approach is to identify write-only data in a real-time manner. To realize ML-WP and achieve accurate write-only data identification, we use machine learning methods to classify data into two groups (i.e., write-only and normal data). Based on this classification, the write-only data is directly written to backend storage without being cached. Experimental results show that, compared with the industry widely deployed write-back policy, ML-WP decreases write traffic to SSD cache by 41.52%, while improving the hit ratio by 2.61% and reducing the average read latency by 37.52%.
Yu Zhang 0101, Ke Zhou 0001, Ping Huang 0001, Hua Wang 0008, Jianying Hu, Yangtao Wang, Yong-guang Ji
DATE6
2020 Deep Self-Taught Graph Embedding Hashing With Pseudo Labels For Image Retrieval
abstract
It has always been a tricky task to generate image hashing function via deep learning without labels and allocate the relative distance between data through their features. Existing methods can complete this task and prevent the overfitting problem using shallow graph embedding technique. However, they only capture the first-order proximity. To address this problem, we design DSTGeH, a deep self-taught graph embedding hashing framework which learns hash function without labels for image retrieval. DSTGeH introduces deep graph embedding means to capture more complex topological relationships (the second-order proximity) on the graph and maps these relationships into pseudo labels, which enables an end-to-end hash model and helps recognize the samples outside the graph. We present the ablation studies and compare DSTGeH with the state-of-the-art label-free hashing algorithms. Extensive experiments show DSTGeH can achieve the best performances and produce an overwhelming advantage on multi-object datasets.
Yu Liu 0040, Yangtao Wang, Jingkuan Song, Chan Guo, Ke Zhou 0001, Zhili Xiao
ICME2
2020 Label-Attended Hashing for Multi-Label Image Retrieval
abstract
For the multi-label image retrieval, the existing hashing algorithms neglect the dependency between objects and thus fail to capture the attention information in the feature extraction, which affects the precision of hash codes. To address this problem, we explore the inter-dependency between objects through their co-occurrence correlation from the label set and adopt Multi-modal Factorized Bilinear (MFB) pooling component so that the image representation learning can capture this attention information. We propose a Label-Attended Hashing (LAH) algorithm which enables an end-to-end hash model with inter-dependency feature extraction. LAH first combines Convolutional Neural Network (CNN) and Graph Convolution Network (GCN) to separately generate the image representation and label co-occurrence embeddings, then adopts MFB to fuse these two modal vectors, finally learns the hash function with a Cauchy distribution based loss function via back propagation. Extensive experiments on public multi-label datasets demonstrate that (1) LAH can achieve the state-of-the-art retrieval results and (2) the usage of co-occurrence relationship and MFB not only promotes the precision of hash codes but also accelerates the hash learning. GitHub address: https://github.com/IDSM-AI/LAH.
Yanzhao Xie, Yu Liu 0040, Yangtao Wang, Lianli Gao, Peng Wang 0037, Ke Zhou 0001
IJCAI3
2020 Semantic-aware data quality assessment for image big data
Yu Liu 0040, Yangtao Wang, Ke Zhou 0001, Yujuan Yang
Future Gener. Comput. Syst.2
2020 A low cost and un-cancelled laplace noise based differential privacy algorithm for spatial decompositions
Xiaocui Li 0001, Yangtao Wang, Jingkuan Song, Yu Liu 0040, Xinyu Zhang 0012, Ke Zhou 0001, Chunhua Li 0002
World Wide Web2
2020 A framework for image dark data assessment
Ke Zhou 0001, Yangtao Wang, Yu Liu 0067, Yujuan Yang, Guoliang Li 0001, Lianli Gao, Zhili Xiao
World Wide Web2
2019 An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning
abstract
Configuration tuning is vital to optimize the performance of database management system (DBMS). It becomes more tedious and urgent for cloud databases (CDB) due to the diverse database instances and query workloads, which make the database administrator (DBA) incompetent. Although there are some studies on automatic DBMS configuration tuning, they have several limitations. Firstly, they adopt a pipelined learning model but cannot optimize the overall performance in an end-to-end manner. Secondly, they rely on large-scale high-quality training samples which are hard to obtain. Thirdly, there are a large number of knobs that are in continuous space and have unseen dependencies, and they cannot recommend reasonable configurations in such high-dimensional continuous space. Lastly, in cloud environment, they can hardly cope with the changes of hardware configurations and workloads, and have poor adaptability. To address these challenges, we design an end-to-end automatic CDB tuning system, CDBTune, using deep reinforcement learning (RL). CDBTune utilizes the deep deterministic policy gradient method to find the optimal configurations in high-dimensional continuous space. CDBTune adopts a try-and-error strategy to learn knob settings with a limited number of samples to accomplish the initial training, which alleviates the difficulty of collecting massive high-quality samples. CDBTune adopts the reward-feedback mechanism in RL instead of traditional regression, which enables end-to-end learning and accelerates the convergence speed of our model and improves efficiency of online tuning. We conducted extensive experiments under 6 different workloads on real cloud databases to demonstrate the superiority of CDBTune. Experimental results showed that CDBTune had a good adaptability and significantly outperformed the state-of-the-art tuning tools and DBA experts.
Ji Zhang 0010, Yu Liu 0040, Ke Zhou 0001, Guoliang Li 0001, Zhili Xiao, Jiashu Xing, Yangtao Wang, Tianheng Cheng, Li Liu 0047, Minwei Ran, Zekang Li
SIGMOD Conference8
2018 A More Secure Spatial Decompositions Algorithm via Indefeasible Laplace Noise in Differential Privacy
Xiaocui Li 0001, Yangtao Wang, Xinyu Zhang 0012, Ke Zhou 0001, Chunhua Li 0002
ADMA2
2017 Multiple Medoids based Multi-view Relational Fuzzy Clustering with Minimax Optimization
abstract
Multi-view data becomes prevalent nowadays because more and more data can be collected from various sources. Each data set may be described by different set of features, hence forms a multi-view data set or multi-view data in short. To find the underlying pattern embedded in an unlabelled multi-view data, many multi-view clustering approaches have been proposed. Fuzzy clustering in which a data object can belong to several clusters with different memberships is widely used in many applications. However, in most of the fuzzy clustering approaches, a single center or medoid is considered as the representative of each cluster in the end of clustering process. This may not be sufficient to ensure accurate data analysis. In this paper, a new multi-view fuzzy clustering approach based on multiple medoids and minimax optimization called M4-FC for relational data is proposed. In M4-FC, every object is considered as a medoid candidate with a weight. The higher the weight is, the more likely the object is chosen as the final medoid. In the end of clustering process, there may be more than one mediod in each cluster. Moreover, minimax optimization is applied to find consensus clustering results of different views with its set of features. Extensive experimental studies on several multi-view data sets including real world image and document data sets demonstrate that M4-FC not only outperforms single medoid based multi-view fuzzy clustering approach, but also performs better than existing multi-view relational clustering approaches.
Yangtao Wang, Lihui Chen 0001, Xiaoli Li 0001
IJCAI1
2017 Multi-view fuzzy clustering with minimax optimization for effective clustering of data from multiple sources
Yangtao Wang, Lihui Chen 0001
Expert Syst. Appl.1
2017 Large Scale Document Categorization With Fuzzy Clustering
abstract
Clustering documents into coherent categories is a very useful and important step for document processing and understanding. The introducing of fuzzy set theory into clustering provides a favorable mechanism to capture overlapping among document clusters. Document dataset is commonly represented as a collection of high-dimensional vectors, which may not be able to fit into memory entirely, when the dataset is large and with a very high dimensionality. However, most of the existing fuzzy clustering approaches deal with small and static datasets. Some of them may have a good scalability but they are only effective for low dimensional data. The study presented in this paper is about new efforts on fuzzy clustering of large-scale and high-dimensional data-especially suitable for document categorization. To consider both large scale and high dimensionality into the problem formulation, our key idea is to incorporate document-tailored fuzzy clustering into a scheme, which is effective for dealing with a large-scale problem. We first identified three representative schemes in fuzzy clustering for handling large-scale data, namely sampling extension, single pass, and divide ensemble. The limitation of fuzzy C-means (FCM)-based approaches for a large document clustering are then investigated. Based on the study, we propose new approaches by incorporating each of hyperspherical FCM and fuzzy coclustering with the three scale-up schemes, respectively. This enables our new approaches to maintain effectiveness for high-dimensional data with an extended scalability. Extensive experimental studies with real-world large document datasets have been conducted and the results demonstrate that the proposed approaches perform consistently better over existing ones in document categorization.
Jian-Ping Mei, Yangtao Wang, Lihui Chen 0001, Chunyan Miao
IEEE Trans. Fuzzy Syst.2
2016 Hyperspherical Fuzzy clustering for online document categorization
abstract
For document data which are typically represented as high dimensional and sparse vectors, cosine distance based Hyperspherical Fuzzy C-Means(HFCM) has been shown to be more effective than classic Euclidean distance based fuzzy c-means(FCM) for document categorization. The existing HFCM approach assumes a static dataset and performs clustering in a batch mode. This design makes HFCM no more suitable when the dataset keeps increasing in its size or is too large to be loaded wholly. In this paper, we work on fuzzy clustering approaches for online document clustering. Specifically, we propose to perform hyperspherical fuzzy c-means in an online manner based on the stochastic gradient method. In this Online Hyperspherical Fuzzy C-Means(OnHFCM), documents are assumed to come one by one and centroids of clusters are updated immediately when each document is arrived. Such a formulation allows OnHFCM to be applicable to large datasets which may be incrementally changing in size. Two variants of OnHFCM with different objective functions are presented in this paper. In addition to the classic online algorithm which processes documents one by one, mini-batch algorithm that handles a small batch of documents at a time is also provided. Experimental results on real-world benchmarks demonstrate that the proposed OnHFCM algorithms achieved better performance than existing ones in online document categorization.
Jian-Ping Mei, Yangtao Wang
FUZZ-IEEE2
2016 K-MEAP: Multiple Exemplars Affinity Propagation With Specified K Clusters
abstract
Recently, an attractive clustering approach named multiexemplar affinity propagation (MEAP) has been proposed as an extension to the single exemplar-based AP. MEAP is able to automatically identify multiple exemplars for each cluster associated with a superexemplar. However, if the cluster number is a prior knowledge and can be specified by the user, MEAP is unable to make use of such knowledge directly in its learning process. Instead, it has to rely on rerunning the process as many times as it takes by tuning parameters until it generates the desired number of clusters. The process of MEAP rerunning may be very time-consuming. In this paper, we propose a new clustering algorithm called Multiple Exemplars Affinity Propagation with Specified K Clusters which is able to generate specified K clusters directly while retaining the advantages of MEAP. Two kinds of new additional messages are introduced in K-MEAP in order to control the number of clusters in the process of message passing. Detailed problem formulation, derived messages, and in-depth analysis of the proposed K-MEAP are provided. Experimental studies on 11 real-world data sets with different kinds of applications demonstrate that K-MEAP not only generates K clusters directly and efficiently without tuning parameters but also outperforms related approaches in terms of clustering accuracy.
Yangtao Wang, Lihui Chen 0001
IEEE Trans. Neural Networks Learn. Syst.1
2014 Incremental fuzzy clustering for document categorization
abstract
Incremental clustering has been proposed to handle large datasets which can not fit into memory entirely. Single pass fuzzy c-means (SpFCM) and Online fuzzy c-means (OFCM) are two representative incremental fuzzy clustering methods. Both of them extend the scalability of fuzzy c-means (FCM) by processing the dataset chunk by chunk. However, due to the data sparsity and high-dimensionality, SpFCM and OFCM fail to produce reasonable results for document data. In this study, we work on clustering approaches that take care of both the large-scale and high-dimensionality issues. Specifically, we propose two methods for incrementally clustering of document data. The first method is a modification of the existing FCM-based incremental clustering with a step to normalize the centroids in each iteration, while the other method is incremental clustering, i.e., Single-Pass or Online, with weighted fuzzy co-clustering. We use several benchmark document datasets for experimental study. The experimental results show that the proposed approaches achieved significant improvements over existing SpFCM and OFCM in document clustering.
Jian-Ping Mei, Yangtao Wang, Lihui Chen 0001, Chunyan Miao
FUZZ-IEEE2
2014 Stochastic gradient descent based fuzzy clustering for large data
abstract
Data is growing at an unprecedented rate in commercial and scientific areas. Clustering algorithms for large data which require small memory consumption and scalability become increasingly important under this circumstance. In this paper, we propose a new clustering approach called stochastic gradient based fuzzy clustering(SGFC) which achieves the optimization based on stochastic approximation to handle such kind of large data. We derive an adaptive learning rate which can be updated incrementally and maintained automatically in gradient descent approach employed in SGFC. Moreover, SGFC is extended to a mini-batch SGFC to reduce the stochastic noise. Additionally, multi-pass SGFC is also proposed to improve the clustering performance. Experiments have been conducted on synthetic data to show the effectiveness of our derived adaptive learning rate. Experimental studies have been also conducted on several large benchmark datasets including real world image and document datasets. Compared with existing fuzzy clustering approaches for large data, the mini-batch SGFC shows comparable or better accuracy with significant less time consumption. These results demonstrate the great potential of SGFC for large data analysis.
Yangtao Wang, Lihui Chen 0001, Jian-Ping Mei
FUZZ-IEEE1
2014 Multi-exemplar based clustering for imbalanced data
abstract
Clustering is an important unsupervised technique of data analysis to find the underlining information of the unlabelled data. Many clustering approaches have been developed and reported in the literature and some of them are widely applied in real world problems such as k-means and fuzzy k-means. However, when handling imbalanced data in which the classes have very different sizes, the performance of these algorithms may not be very effective. The results of these algorithms always generate clusters with similar sizes which is called "uniform effect". To prevent uniform effect and improve the clustering performance, we proposed a new approach called multi-exemplar merging clustering(MEMC) for imbalanced data in this paper. Our approach is composed of two stages of processing: multiple exemplars identification stage and exemplars merging stage. In the first stage, multiple exemplars which are the data objects selected to represent the data set are identified using MEAP. In the second stage, the exemplars are merged based on the proposed overlapping measure(OM) which reflects the degree of overlapping between clusters. Experimental results on several synthetic and real world data sets are conducted to show the effectiveness of our proposed approach on imbalanced data clustering.
Yangtao Wang, Lihui Chen 0001
ICARCV1
2014 K-MEAP: Generating Specified K Clusters with Multiple Exemplars by Efficient Affinity Propagation
abstract
Recently, an attractive clustering approach named multi-exemplar affinity propagation (MEAP) has been proposed as an extension to the single exemplar based Affinity Propagation (AP). MEAP is able to automatically identify multiple exemplars for each cluster associated with a super exemplar. However, if the cluster number is a prior knowledge and can be specified by the user, MEAP is unable to make use of such knowledge directly in its learning process. Instead it has to rely on re-running the process as many times as it takes by tuning parameters until it generates the desired number of clusters. The process of MEAP re-running may be very time consuming. In this paper, we propose a new clustering algorithm called KMEAP which is able to generate specified K clusters directly while retaining the advantages of MEAP. Two kinds of new additional messages are introduced in MEAP in order to control the number of clusters in the process of message passing. The detailed problem formulation, the derived updating rules for passing messages, and the in-depth analysis of the proposed K-MEAP are provided. Experimental studies demonstrated that K-MEAP not only generates K clusters directly and efficiently without tuning parameters, but also outperforms related approaches in terms of clustering accuracy.
Yangtao Wang, Lihui Chen 0001
ICDM1
2014 Incremental Fuzzy Clustering With Multiple Medoids for Large Data
abstract
As an important technique of data analysis, clustering plays an important role in finding the underlying pattern structure embedded in unlabeled data. Clustering algorithms that need to store all the data into the memory for analysis become infeasible when the dataset is too large to be stored. To handle such large data, incremental clustering approaches are proposed. The key idea behind these approaches is to find representatives (centroids or medoids) to represent each cluster in each data chunk, which is a packet of the data, and final data analysis is carried out based on those identified representatives from all the chunks. In this paper, we propose a new incremental clustering approach called incremental multiple medoids-based fuzzy clustering (IMMFC) to handle complex patterns that are not compact and well separated. We would like to investigate whether IMMFC is a good alternative to capturing the underlying data structure more accurately. IMMFC not only facilitates the selection of multiple medoids for each cluster in a data chunk, but also has the mechanism to make use of relationships among those identified medoids as side information to help the final data clustering process. The detailed problem formulation, updating rules derivation, and the in-depth analysis of the proposed IMMFC are provided. Experimental studies on several large datasets that include real world malware datasets have been conducted. IMMFC outperforms existing incremental fuzzy clustering approaches in terms of clustering accuracy and robustness to the order of data. These results demonstrate the great potential of IMMFC for large-data analysis.
Yangtao Wang, Lihui Chen 0001, Jian-Ping Mei
IEEE Trans. Fuzzy Syst.1