Xiaoyu Zhang 0002

dblp:12/5927-2 · DBLP profile ↗
← Back
79ranked-venue papers
21as first author
40since 2021 · last 2026
0000-0003-1630-6058ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 12 first-author · 18 since 2021Artificial intelligence and machine learning · 32 · 9 first-author · 16 since 2021Security and privacy · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Online Traffic Camouflage Against Network Analyzers via Deep Reinforcement Learning
abstract
Traffic analysis plays a pivotal role in network management. However, despite the prevalence of encryption, attackers are still able to deduce privacy elements such as user behavior and OS identification through advanced learning-based methods that exploit side-channel features. Existing defense strategies, which manipulate feature distribution to evade traffic analyzers, are often hampered by the need for impractical decoder deployment across all routes in symmetric framework methods. Moreover, reversing feature distribution modifications to real-time traffic, especially through dummy packet crafting or padding, is a complex task. In response to these challenges, we propose Veil, a novel and practical defender designed to protect live connections against encrypted network traffic analyzers. Leveraging an asymmetric deployment structure, Veil is capable of reconstructing live streams at the packet-block level, thereby allowing for seamless deployment on any connection node while enforcing transmission constraints. By employing a traffic-customized DQN framework, Veil not only reverses statistical feature perturbations back to the traffic space but also directs the distribution towards a target class. Extensive experiments conducted on real-world datasets validate the efficacy of Veil in efficiently evading analyzers in both targeted and untargeted modes, outperforming existing defense mechanisms. Notably, Veil addresses the key issues of impractical decoder deployment and complex real-time traffic manipulation, offering a more viable solution for network traffic privacy protection. The source code is publicly available at https://github.com/SecTeamPolaris/Veil, facilitating further research and application in the field of network security.
Wenhao Li 0005, Jie Chen 0093, Zhaoxuan Li, Shuai Wang 0079, Huamin Jin, Xiaoyu Zhang 0002
IEEE Trans. Netw. Serv. Manag.6
2025 MUN: Image Forgery Localization Based on M³ Encoder and UN Decoder
abstract
Image forgeries can entirely change the semantic information of an image, and can be used for unscrupulous purposes. In this paper, we propose a novel image forgery localization network named as MUN, which consists of an M^3 encoder and a UN decoder. Firstly, the M^3 encoder is constructed based on a Multi-scale Max-pooling query module to extract Multi-clue forged features. Noiseprint++ is adopted to assist the RGB clue, and its deployment methodology is discussed. A Multi-scale Max-pooling Query (MMQ) module is proposed to integrate RGB and noise features. Secondly, a novel UN decoder is proposed to extract hierarchical features from both top-down and bottom-up directions, reconstructing both high-level and low-level features at the same time. Thirdly, we formulate an IoU-recalibrated Dynamic Cross-Entropy (IoUDCE) loss to dynamically adjust the weights on forged regions according to IoU which can adaptively balance the influence of authentic and forged regions. Last but not least, we propose a data augmentation method, i.e., Deviation Noise Augmentation (DNA), which acquires accessible prior knowledge of RGB distribution to improve the generalization ability. Extensive experiments on publicly available datasets show that MUN outperforms the state-of-the-art works.
Shuhuan Chen, Haichao Shi, Xiaoyu Zhang 0002, Song Xiao 0001, Qiang Cai 0001
AAAI4
2025 Unveiling Deepfakes with Latent Diffusion Counterfactual Explanations
abstract
Deepfake technology, driven by deep learning, produces highly convincing synthetic media, raising concerns about misuse. While DeepFake detection models have achieved impressive accuracy, but due to the difficulty of distinguishing fake from real, interpretability remains challenging that humans cannot understand or trust the detection results. We propose a novel approach to enhance interpretability by generating counterfactual explanations. By integrating ensemble classifier loss and text instructions into the fine-tuning of a Latent Diffusion Model, our method effectively improves the quality and efficiency of generated counterfactual explanations. Experiments on DeepFake datasets validate the effectiveness of our approach, contributing the interpretability of Deepfake detection.
Bo Peng 0002, Jing Dong 0003, Xiaoyu Zhang 0002
ICASSP4
2025 GenFIQA: Generative Face Image Quality Assessment via Identity-conditioned Diffusion Model
abstract
Face recognition (FR) systems are widely deployed but often struggle due to unconstrained image-capturing conditions. Face image quality assessment (FIQA), applied before recognition, mitigates these challenges by filtering out unreliable samples. Current leading FIQA methods evaluate image quality based on the characteristics observed within the FR model pipeline. However, they leave out the inherent differences in identity embeddings between high-and low-quality face images. To this end, we propose Gen-FIQA, which utilizes a generative model to probe and amplify this difference. Specifically, we extract the identity embedding from an input image using a pre-trained FR model, and then use it as a conditioning signal to generate several face images of the same identity. This generation process leverages the inherent prior in the generative model to translate the difference in identity embedding space back to pixel space. To quantify these differences, the quality score is computed as the average cosine similarity between embeddings from the original and generated images. To improve computational efficiency, we further distill GenFIQA into a lightweight regression-based variant, GenFIQA(R). Extensive experiments across five benchmark datasets and four FR models demonstrate the superiority of our methods over thirteen state-of-the-art FIQA methods.
Zheyu Yan, Weisong Zhao, Kai Pang, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001
IJCB6
2025 Straighter Flow Matching via a Diffusion-Based Coupling Prior
Siyu Xing, Jie Cao 0002, Huaibo Huang, Haichao Shi, Xiaoyu Zhang 0002
PRCV (8)5
2025 MPE3: Learning meta-prompt with entity-enhanced semantics for few-shot named entity recognition
Yuwei Xia, Liang Wang 0001, Qiang Liu 0006, Xiaoyu Zhang 0002
Neurocomputing6
2025 Magnifier: Detecting Network Access via Lightweight Traffic-Based Fingerprints
Wenhao Li 0005, Qiang Wang 0059, Huaifeng Bao, Xiaoyu Zhang 0002, Lingyun Ying, Zhaoxuan Li, Huamin Jin, Shuai Wang 0079
IEEE Trans. Inf. Forensics Secur.4
2025 Global Cross-Entropy Loss for Deep Face Recognition
abstract
Contemporary deep face recognition techniques predominantly utilize the Softmax loss function, designed based on the similarities between sample features and class prototypes. These similarities can be categorized into four types: in-sample target similarity, in-sample non-target similarity, out-sample target similarity, and out-sample non-target similarity. When a sample feature from a specific class is designated as the anchor, the similarity between this sample and any class prototype is referred to as in-sample similarity. In contrast, the similarity between samples from other classes and any class prototype is known as out-sample similarity. The terms target and non-target indicate whether the sample and the class prototype used for similarity calculation belong to the same identity or not. The conventional Softmax loss function promotes higher in-sample target similarity than in-sample non-target similarity. However, it overlooks the relation between in-sample and out-sample similarity. In this paper, we propose Global Cross-Entropy loss (GCE), which promotes 1) greater in-sample target similarity over both the in-sample and out-sample non-target similarity, and 2) smaller in-sample non-target similarity to both in-sample and out-sample target similarity. In addition, we propose to establish a bilateral margin penalty for both in-sample target and non-target similarity, so that the discrimination and generalization of the deep face model are improved. To bridge the gap between training and testing of face recognition, we adapt the GCE loss into a pairwise framework by randomly replacing some class prototypes with sample features. We designate the model trained with the proposed Global Cross-Entropy loss as GFace. Extensive experiments on several public face benchmarks, including LFW, CALFW, CPLFW, CFP-FP, AgeDB, IJB-C, IJB-B, MFR-Ongoing, and MegaFace, demonstrate the superiority of GFace over other methods. Additionally, GFace exhibits robust performance in general visual recognition task.
Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Guoying Zhao 0001, Zhen Lei 0001
IEEE Trans. Image Process.4
2024 Poster: Towards Real-Time Intrusion Detection with Explainable AI-Based Detector
abstract
Identifying malicious traffic is crucial for safeguarding internal networks from privacy breaches.Intrusion Detection Systems (IDS) traditionally rely on inefficient and outdated rule-sets, necessitating a shift towards AI-driven, learning-based algorithms for enhanced detection capabilities.Despite their promise, AI-integrated IDS face deployment challenges due to complex, opaque decision-making processes that can lead to latency and an increased risk of false positives.This paper presents the Explainable AI-based Intrusion Detection System (XAI-IDS), addressing the limitations of both rule-based and AI-driven IDS by integrating interpretable deep learning models.XAI-IDS employs tree regularization to transform complex models into efficient, transparent decision trees, facilitating real-time detection with improved accuracy and explainability.Experiments on two benchmark datasets demonstrate XAI-IDS's superior performance, offering a scalable solution to the challenge of identifying malicious traffic with reduced risk of false positives.
Wenhao Li 0005, Duohe Ma, Zhaoxuan Li, Huaifeng Bao, Shuai Wang 0079, Huamin Jin, Xiaoyu Zhang 0002
CCS7
2024 BLIP-Adapter: Bridging Vision-Language Models with Adapters for Generalizable Face Anti-spoofing
abstract
Face anti-spoofing is essential for ensuring the security of facial recognition systems against spoofing attacks. Recent methods have transferred Vision-Language models to face anti-spoofing (e.g., FLIP and CLIPC8), demonstrating that learning perception from supervision in natural language can enhance the model’s detection performance. However, such methods exhibit limited depth in the interaction between images and texts, resulting in poor performance on fine-grained understanding tasks such as face anti-spoofing. Besides, the lack of diversity in image-text pairs for face anti-spoofing further hinders such methods from playing their best. To address these issues, we propose a novel fine-tuning strategy for Vision-Language models in face anti-spoofing. This strategy introduces the Bootstrapping Language-Image Pre-training model (BLIP), known for its novel interaction mechanisms and superior image-text comprehension, to construct a more generalized feature representation for face anti-spoofing. Furthermore, we propose an Adapter module for the text branch to reduce the negative impact of insufficient data diversity and catastrophic forgetting. Extensive experiments conducted on various cross-domain testing benchmarks demonstrate the significant superiority of our method over the state-of-the-art, highlighting its effectiveness and robustness.
Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001
IJCB4
2024 Learning Spatiotemporal Inconsistency via Thumbnail Layout for Face Deepfake Detection
Jian Liang 0001, Lijun Sheng, Xiaoyu Zhang 0002
Int. J. Comput. Vis.4
2024 MetaTKG++: Learning evolving factor enhanced meta-knowledge for temporal knowledge graph reasoning
Yuwei Xia, Mengqi Zhang 0002, Qiang Liu 0006, Liang Wang 0056, Xiaoyu Zhang 0002, Liang Wang 0001
Pattern Recognit.6
2024 Masked Face Transformer
abstract
The COVID-19 pandemic makes wearing masks mandatory. Existing CNN-based face recognition (FR) systems suffer from severe performance degradation as masks occlude the vital facial regions. Recently, Vision Transformers have shown promising performance in various vision tasks with quadratic computation costs. Swin Transformer first proposes a successive window attention mechanism allowing the cross-window connection and more computational efficiency. Despite its potential, the deployment of Swin Transformer in masked face recognition encounters two challenges: 1) the attention range is insufficient to capture locally compatible face regions. 2) Masked face recognition can be defined as an occlusion-robust classification task with a known occlusion position, i.e., the position of the mask is minor-varying, which is overlooked but efficient in improving the model’s recognition accuracy. To alleviate the above problem, we propose a Masked Face Transformer (MFT) with Masked Face-compatible Attention (MFA). The proposed MFA 1) introduces two additional window partition configurations, e.g., row shift and column shift, to enlarge the attention range in Swin with invariant computation costs, and 2) suppresses the interaction between the masked and non-masked regions to retain their discrepancies. Additionally, as mask occlusion leads to a separation between the masked and non-masked samples of the same identity, we propose to explore the relationship between them by a ClassFormer module to enhance intra-class aggregation. Extensive experiments show that MFT outperforms state-of-the-art masked face recognition methods in both simulated and real masked face testing datasets.
Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.5
2024 DyGCN: Efficient Dynamic Graph Embedding With Graph Convolutional Network
abstract
Graph embedding, aiming to learn low-dimensional representations (aka. embeddings) of nodes in graphs, has received significant attention. In recent years, there has been a surge of efforts, among which graph convolutional networks (GCNs) have emerged as an effective class of models. However, these methods mainly focus on the static graph embedding. In the present work, an efficient dynamic graph embedding approach is proposed, called dynamic GCN (DyGCN), which is an extension of the GCN-based methods. The embedding propagation scheme of GCN is naturally generalized to a dynamic setting in an efficient manner, which propagates the change in topological structure and neighborhood embeddings along the graph to update the node embeddings. The most affected nodes are updated first, and then their changes are propagated to further nodes, which in turn are updated. Extensive experiments on various dynamic graphs showed that the proposed model can update the node embeddings in a time-saving and performance-preserving way.
Zeyu Cui, Zekun Li 0001, Xiaoyu Zhang 0002, Qiang Liu 0006, Liang Wang 0001, Mengmeng Ai
IEEE Trans. Neural Networks Learn. Syst.4
2023 Grouped Knowledge Distillation for Deep Face Recognition
abstract
Compared with the feature-based distillation methods, logits distillation can liberalize the requirements of consistent feature dimension between teacher and student networks, while the performance is deemed inferior in face recognition. One major challenge is that the light-weight student network has difficulty fitting the target logits due to its low model capacity, which is attributed to the significant number of identities in face recognition. Therefore, we seek to probe the target logits to extract the primary knowledge related to face identity, and discard the others, to make the distillation more achievable for the student network. Specifically, there is a tail group with near-zero values in the prediction, containing minor knowledge for distillation. To provide a clear perspective of its impact, we first partition the logits into two groups, i.e., Primary Group and Secondary Group, according to the cumulative probability of the softened prediction. Then, we reorganize the Knowledge Distillation (KD) loss of grouped logits into three parts, i.e., Primary-KD, Secondary-KD, and Binary-KD. Primary-KD refers to distilling the primary knowledge from the teacher, Secondary-KD aims to refine minor knowledge but increases the difficulty of distillation, and Binary-KD ensures the consistency of knowledge distribution between teacher and student. We experimentally found that (1) Primary-KD and Binary-KD are indispensable for KD, and (2) Secondary-KD is the culprit restricting KD at the bottleneck. Therefore, we propose a Grouped Knowledge Distillation (GKD) that retains the Primary-KD and Binary-KD but omits Secondary-KD in the ultimate KD loss calculation. Extensive experimental results on popular face recognition benchmarks demonstrate the superiority of proposed GKD over state-of-the-art methods.
Weisong Zhao, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001
AAAI4
2023 Modeling Spoof Noise by De-spoofing Diffusion and its Application in Face Anti-spoofing
abstract
Face anti-spoofing is crucial for ensuring the security and reliability of face recognition systems. Several existing face anti-spoofing methods utilize GAN-like networks to detect presentation attacks by estimating the noise pattern of a spoof image and recovering the corresponding genuine image. But GAN’s limited face appearance space results in the denoised faces cannot cover the full data distribution of genuine faces, thereby undermining the generalization performance of such methods. In this work, we present a pioneering attempt to employ diffusion models to denoise a spoof image and restore the genuine image. The difference between these two images is considered as the spoof noise, which can serve as a discriminative cue for face anti-spoofing. We evaluate our proposed method on several intra-testing and inter-testing protocols, where the experimental results showcase the effectiveness of our method in achieving competitive performance in terms of both accuracy and generalization.
Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001
IJCB3
2023 Rumor Detection with Diverse Counterfactual Evidence
abstract
The growth in social media has exacerbated the threat of fake news to individuals and communities. This draws increasing attention to developing efficient and timely rumor detection methods. The prevailing approaches resort to graph neural networks (GNNs) to exploit the post-propagation patterns of the rumor-spreading process. However, these methods lack inherent interpretation of rumor detection due to the black-box nature of GNNs. Moreover, these methods suffer from less robust results as they employ all the propagation patterns for rumor detection. In this paper, we address the above issues with the proposed Diverse Counterfactual Evidence framework for Rumor Detection (DCE-RD). Our intuition is to exploit the diverse counterfactual evidence of an event graph to serve as multi-view interpretations, which are further aggregated for robust rumor detection results. Specifically, our method first designs a subgraph generation strategy to efficiently generate different subgraphs of the event graph. We constrain the removal of these subgraphs to cause the change in rumor detection results. Thus, these subgraphs naturally serve as counterfactual evidence for rumor detection. To achieve multi-view interpretation, we design a diversity loss inspired by Determinantal Point Processes (DPP) to encourage diversity among the counterfactual evidence. A GNN-based rumor detection model further aggregates the diverse counterfactual evidence discovered by the proposed DCE-RD to achieve interpretable and robust rumor detection results. Extensive experiments on two real-world datasets show the superior performance of our method. Our code is available at https://github.com/Vicinity111/DCE-RD.
Kaiwei Zhang, Junchi Yu, Haichao Shi, Jian Liang 0001, Xiaoyu Zhang 0002
KDD5
2023 Cross-Architecture Distillation for Face Recognition
abstract
Transformers have emerged as the superior choice for face recognition tasks, but their insufficient platform acceleration hinders their application on mobile devices. In contrast, Convolutional Neural Networks (CNNs) capitalize on hardware-compatible acceleration libraries. Consequently, it has become indispensable to preserve the distillation efficacy when transferring knowledge from a Transformer-based teacher model to a CNN-based student model, known as Cross-Architecture Knowledge Distillation (CAKD). Despite its potential, the deployment of CAKD in face recognition encounters two challenges: 1) the teacher and student share disparate spatial information for each pixel, obstructing the alignment of feature space, and 2) the teacher network is not trained in the role of a teacher, lacking proficiency in handling distillation-specific knowledge. To surmount these two constraints, 1) we first introduce a Unified Receptive Fields Mapping module (URFM) that maps pixel features of the teacher and student into local features with unified receptive fields, thereby synchronizing the pixel-wise spatial information of teacher and student. Subsequently, 2) we develop an Adaptable Prompting Teacher network (APT) that integrates prompts into the teacher, enabling it to manage distillation-specific knowledge while preserving the model's discriminative capacity. Extensive experiments on popular face benchmarks and two large-scale verification sets demonstrate the superiority of our method.
Weisong Zhao, Xiangyu Zhu 0001, Zhixiang He, Xiaoyu Zhang 0002, Zhen Lei 0001
ACM Multimedia4
2023 Spoof-Guided Image Decomposition for Face Anti-spoofing
Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Shukai Chen, Peng Li 0035, Zhen Lei 0001
PRCV (5)3
2023 Face Forgery Detection by 3D Decomposition and Composition Search
abstract
Detecting digital face manipulation has attracted extensive attention due to fake media's potential risks to the public. However, recent advances have been able to reduce the forgery signals to a low magnitude. Decomposition, which reversibly decomposes an image into several constituent elements, is a promising way to highlight the hidden forgery details. In this paper, we investigate a novel 3D decomposition based method that considers a face image as the production of the interaction between 3D geometry and lighting environment. Specifically, we disentangle a face image into four graphics components including 3D shape, lighting, common texture, and identity texture, which are respectively constrained by 3D morphable model, harmonic reflectance illumination, and PCA texture model. Meanwhile, we build a fine-grained morphing network to predict 3D shapes with pixel-level accuracy to reduce the noise in the decomposed elements. Moreover, we propose a composition search strategy that enables an automatic construction of an architecture to mine forgery clues from forgery-relevant components. Extensive experiments validate that the decomposed components highlight forgery artifacts, and the searched architecture extracts discriminative forgery features. Thus, our method achieves the state-of-the-art performance.
Xiangyu Zhu 0001, Hongyan Fei, Tianshuo Zhang, Xiaoyu Zhang 0002, Stan Z. Li, Zhen Lei 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 OW-TAL: Learning Unknown Human Activities for Open-World Temporal Action Localization
Yaru Zhang, Xiaoyu Zhang 0002, Haichao Shi
Pattern Recognit.2
2023 Prism: Real-Time Privacy Protection Against Temporal Network Traffic Analyzers
abstract
Traffic analysis is widely used in network monitoring. However, the attackers can sometimes infer sensitive information from the patterns of the encrypted network traffic, which poses a threat to network security. Most existing countermeasures are proposed to obfuscate traffic flows using adversarial examples. However, there are two challenges when adding perturbations to live network traffic. Firstly, the perturbations imposed on the feature space cannot be conveniently projected to original traffic flows in feature-space based methods. Secondly, it is laborious and impractical to apply symmetrical framework to encode/decode the adversarial traffic in traffic-space based approaches. To address the above issues, in this paper, we propose an asymmetric defending scheme, namelyPrism, to protect theliveconnection privacy against attacks of temporal network traffic analyzers. Specifically,Prismfirst extracts standardized temporal features via Power-Law Division (PLD) algorithm, and then employs Time-stacked State Transition Model (TSTM) to obtain the fingerprint of each application. Finally,Prismdefends against the analyzers with online traffic perturbation. Since thePrismis designed as a traffic-space based defender with asymmetric defending structure, the deployment is lightweight and efficient. Experimental results on two real-world datasets demonstrate the effectiveness and generalization of our adversarial perturbations. In particular, it is encouraging to see that our proposed defending scheme outperforms the advanced countermeasures, such as adversarial training and traffic filter.
Wenhao Li 0005, Xiaoyu Zhang 0002, Huaifeng Bao, Zhaoxuan Li, Haichao Shi, Qiang Wang 0059
IEEE Trans. Inf. Forensics Secur.2
2023 Masked Relation Learning for DeepFake Detection
abstract
DeepFake detection aims to differentiate falsified faces from real ones. Most approaches formulate it as a binary classification problem by solely mining the local artifacts and inconsistencies of face forgery, which neglect the relation across local regions. Although several recent works explore local relation learning for DeepFake detection, they overlook the propagation of relational information and lead to limited performance gains. To address these issues, this paper provides a new perspective by formulating DeepFake detection as a graph classification problem, in which each facial region corresponds to a vertex. But relational information with large redundancy hinders the expressiveness of graphs. Inspired by the success of masked modeling, we propose Masked Relation Learning which decreases the redundancy to learn informative relational features. Specifically, a spatiotemporal attention module is exploited to learn the attention features of multiple facial regions. A relation learning module masks partial correlations between regions to reduce redundancy and then propagates the relational information across regions to capture the irregularity from a global view of the graph. We empirically discover that a moderate masking rate (e.g., 50%) brings the best performance gain. Experiments verify the effectiveness of Masked Relation Learning and demonstrate that our approach outperforms the state of the art by 2% AUC on the cross-dataset DeepFake video detection. Code will be available athttps://github.com/zimyang/MaskRelation.
Zimin (Max) Yang, Jian Liang 0001, Xiaoyu Zhang 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.4
2023 StochasticFormer: Stochastic Modeling for Weakly Supervised Temporal Action Localization
abstract
Weakly supervised temporal action localization (WS-TAL) aims to identify the time intervals corresponding to actions of interest in untrimmed videos with video-level weak supervision. For most existing WS-TAL methods, two commonly encountered challenges are under-localization and over-localization, which inevitably bring about severe performance deterioration. To address the issues, this paper proposes a transformer-structured stochastic process modeling framework, namely StochasticFormer, to fully investigate finer-grained interactions among the intermediate predictions to achieve further refined localization. StochasticFormer is built on a standard attention-based pipeline to derive preliminary frame/snippet-level predictions. Then, the pseudo localization module generates variable-length pseudo action instances with the corresponding pseudo labels. Using the pseudo "action instance - action category" pairs as fine-grained pseudo supervision, the stochastic modeler aims to learn the underlying interaction among the intermediate predictions with an encoder-decoder network. The encoder consists of the deterministic and latent path to capture the local and global information, which are subsequently integrated by the decoder to obtain reliable predictions. The framework is optimized with three carefully designed losses, i.e. the video-level classification loss, the frame-level semantic coherence loss, and the ELBO loss. Extensive experiments on two benchmarks, i.e., THUMOS14 and ActivityNet1.2, have shown the efficacy of StochasticFormer compared with the state-of-the-art methods.
Haichao Shi, Xiaoyu Zhang 0002
IEEE Trans. Image Process.2
2023 AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and Localization
abstract
The point process is a solid framework to model sequential data, such as videos, by exploring the underlying relevance. As a challenging problem for high-level video understanding, weakly supervised action recognition and localization in untrimmed videos have attracted intensive research attention. Knowledge transfer by leveraging the publicly available trimmed videos as external guidance is a promising attempt to make up for the coarse-grained video-level annotation and improve the generalization performance. However, unconstrained knowledge transfer may bring about irrelevant noise and jeopardize the learning model. This article proposes a novel adaptability decomposing encoder-decoder network to transfer reliable knowledge between the trimmed and untrimmed videos for action recognition and localization by bidirectional point process modeling, given only video-level annotations. By decomposing the original features into the domain-adaptable and domain-specific ones based on their adaptability, trimmed-untrimmed knowledge transfer can be safely confined within a more coherent subspace. An encoder-decoder-based structure is carefully designed and jointly optimized to facilitate effective action classification and temporal localization. Extensive experiments are conducted on two benchmark data sets (i.e., THUMOS14 and ActivityNet1.3), and the experimental results clearly corroborate the efficacy of our method.
Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035, Jing Dong 0003
IEEE Trans. Neural Networks Learn. Syst.1
2023 Image Quality Assessment-driven Reinforcement Learning for Mixed Distorted Image Restoration
abstract
Due to the diversity of the degradation process that is difficult to model, the recovery of mixed distorted images is still a challenging problem. The deep learning model trained under certain degradation declines significantly in other degradation situations. In this article, we explore ways to use a combination of tools to deal with the mixed distortion. First, we illustrate the limitations of a single deep network in dealing with multiple distortion types and then introduce a hierarchical toolkit with distinguished powerful tools. Second, we investigate how an efficient representation of images combined with a reinforcement learning (RL) paradigm helps to deal with tool noise in continuous restoration. The proposed method can accurately capture the distortion preferences for selecting the optimal recovery tools by RL agent. Finally, to fully utilize random tools for unknown distortion combinations, we adopt the exploration scheme with various quality evaluation methods to achieve more quality improvements. Experimental results demonstrate that the peak signal-to-noise ratio of the proposed method is 3.30 dB higher than other state-of-the-art RL-based methods on the CSIQ single distortion dataset and 0.95 dB higher on the DIV2K mixed distortion dataset.
Xiaoyu Zhang 0002, Wei Gao 0003, Ge Li 0002, Qiuping Jiang, Runmin Cong
ACM Trans. Multim. Comput. Commun. Appl.1
2023 ProGraph: Robust Network Traffic Identification With Graph Propagation
abstract
Network traffic identification is critical for effective network management. Existing methods mostly focus on invariant network environments with stable attribute distributions. Unfortunately, however, they can hardly be adaptive to the variation of practical networks and suffer from significant performance degradation. This problem largely stems from the over-dependence of existing methods on the vulnerable side-channel features. To address this issue, in this paper we propose a graph-based approach, namely ProGraph, to ensure robust network traffic classification among various network environments. The core idea of ProGraph is to construct a correlation graph with session clusters aggregated from different networks, based on which graph propagation can be effectively implemented to predict labels of testing nodes in an iterative manner. ProGraph enhances the correlation between clusters of the same class to provide reliable paths for label dissemination from the labeled clusters to the testing ones. It is encouraging to see that the proposed ProGraph achieves an accuracy of 92.25% in networks with constant attributes, while remaining stable with the accuracy of 90.89% when deployed in different networks, which significantly outperforms the state-of-the-art approaches. Meanwhile, ProGraph can accurately identify the novel classes which do not exist in the training dataset, with an AUC of 95.11. Last but not least, a carefully constructed dataset, namely CrossNet2021, containing network traffic of 20 classes of applications from two distinct networking scenarios, is made publicly available to support further research.
Wenhao Li 0005, Xiaoyu Zhang 0002, Huaifeng Bao, Haichao Shi, Qiang Wang 0059
IEEE/ACM Trans. Netw.2
2022 MetaTKG: Learning Evolutionary Meta-Knowledge for Temporal Knowledge Graph Reasoning
abstract
Reasoning over Temporal Knowledge Graphs (TKGs) aims to predict future facts based on given history.One of the key challenges for prediction is to learn the evolution of facts.Most existing works focus on exploring evolutionary information in history to obtain effective temporal embeddings for entities and relations, but they ignore the variation in evolution patterns of facts, which makes them struggle to adapt to future data with different evolution patterns.Moreover, new entities continue to emerge along with the evolution of facts over time.Since existing models highly rely on historical information to learn embeddings for entities, they perform poorly on such entities with little historical information.To tackle these issues, we propose a novel Temporal Meta-learning framework for TKG reasoning, MetaTKG for brevity.Specifically, our method regards TKG prediction as many temporal meta-tasks, and utilizes the designed Temporal Meta-learner to learn evolutionary metaknowledge from these meta-tasks.The proposed method aims to guide the backbones to learn to adapt quickly to future data and deal with entities with little historical information by the learned meta-knowledge.Specially, in temporal meta-learner, we design a Gating Integration module to adaptively establish temporal correlations between meta-tasks.Extensive experiments on four widely-used datasets and three backbones demonstrate that our method can greatly improve the performance.
Yuwei Xia, Mengqi Zhang 0002, Qiang Liu 0006, Xiaoyu Zhang 0002
EMNLP5
2022 HIRL: Hybrid Image Restoration Based on Hierarchical Deep Reinforcement Learning via Two-Step Analysis
abstract
The restoration of hybrid distorted images in real-world scenarios is still a difficult problem due to the fact that the degrading types and degrees are always unknown. Previous studies typically utilize multiple recovery tools to restore images. However, each tool adopted inevitably introduces additional noise and will affect the subsequent recovery results. To address this issue, in this paper, we propose a hierarchical deep reinforcement learning framework (HIRL), which balance both benefits and noises brought by each tool and select the appropriate type and degree tools. The proposed method endeavors to find a long-term optimal tool sequence, which is better than the greedy strategy that employs the tools with the largest short-term returns. Meanwhile, it benefits from a hierarchical design to reduce time consumption and complexity compared to a brute force strategy. Experiments demonstrate the superiority of the proposed method over other state-of-the-art methods. Furthermore, our framework is highly scalable and can be easily extended to the other recovery pipelines.
Xiaoyu Zhang 0002, Wei Gao 0003
ICASSP1
2022 JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration
abstract
Previous reinforcement learning (RL) based image restoration studies typically train RL agents to search for recovery tools from a constructed toolset and iteratively recover images. However, we argue that these agents rely on pre-trained RL models with fixed-length paths for restoration, which performs poorly in the case of unknown distortions. To address these issues, we propose a joint exploitation and exploration reinforcement learning network (JE2Net). Specifically, we propose a new deep classification network for image feature extraction and tool selection, which serves as a model prior. Second, we design a stochastic strategy to randomly select tools and a dynamic termination strategy to adaptively stop the recovery process. In this way, the model prior and exploration mechanism can be jointly used to expand the search space and obtain more quality gain. Experimental results show that our proposed method is more flexible compared to other state-of-the-art methods and achieves significant quality improvements in the presence of unknown distortions.
Xiaoyu Zhang 0002, Wei Gao 0003, Hui Yuan 0001, Ge Li 0002
ICASSP1
2022 TDRNet: Transformer-Based Dual-Branch Restoration Network for Geometry Based Point Cloud Compression Artifacts
abstract
With the development of 3D point cloud applications, compression plays an important role in lightweight transmission. However, there are barely related works for recovering the compressed point clouds. In this paper, we investigate the geometry-based point cloud compression (G-PCC) artifact removal problem, which is reflected in position offsets and quantity reductions. To address these issues, we propose a novel Transformer-based Dual-branch Restoration Network (TDRNet). First, we design a Transformer Feature Extractor (TFE), which aims to accurately model structural features for position correction and handle inputs of different sizes. Second, a Dual-Branch Restoration (DBR) module is proposed to deeply exploit global shape and local geometry information to restore the quantity reduction distortion. In this way, the proposed TFE and DBR can work cooperatively for the overall compression artifact removal to reconstruct dense point cloud with better completeness and fine-grained details. Experiments show that our proposed TDRNet achieves state-of-the-art results, and our model is expected to provide a baseline for future point cloud compression distortion restoration studies.
Xiaoyu Zhang 0002, Guibiao Liao, Wei Gao 0003, Ge Li 0002
ICME1
2022 Dynamic Graph Modeling for Weakly-Supervised Temporal Action Localization
abstract
Weakly supervised action localization is a challenging task that aims to localize action instances in untrimmed videos given only video-level supervision. Existing methods mostly distinguish action from background via attentive feature fusion with RGB and optical flow modalities. Unfortunately, this strategy fails to retain the distinct characteristics of each modality, leading to inaccurate localization under hard-to-discriminate cases such as action-context interference and in-action stationary period. As an action is typically comprised of multiple stages, an intuitive solution is to model the relation between the finer-grained action segments to obtain a more detailed analysis. In this paper, we propose a dynamic graph-based method, namely DGCNN, to explore the two-stream relation between action segments. To be specific, segments within a video which are likely to be actions are dynamically selected to construct an action graph. For each graph, a triplet adjacency matrix is devised to explore the temporal and contextual correlations between the pseudo action segments, which consists of three components, i.e., mutual importance, feature similarity, and high-level contextual similarity. The two-stream dynamic pseudo graphs, along with the pseudo background segments, are used to derive more detailed video representation. For action localization, a non-local based temporal refinement module is proposed to fully leverage the temporal consistency between consecutive segments. Experimental results on three datasets, i.e., THUMOS14, ActivityNet v1.2 and v1.3, demonstrate that our method is superior to the state-of-the-arts.
Haichao Shi, Xiaoyu Zhang 0002, Lixing Gong, Yong Li 0034, Yongjun Bao
ACM Multimedia2
2022 Robust network traffic identification with graph matching
Wenhao Li 0005, Xiaoyu Zhang 0002, Huaifeng Bao, Qiang Wang 0059, Zhaoxuan Li
Comput. Networks2
2022 Sea Surface Temperature Prediction With Memory Graph Convolutional Networks
abstract
We develop a memory graph convolutional network (MGCN) framework for sea surface temperature (SST) prediction. The MGCN consists of two memory layers: one graph layer and one output layer. The memory layer captures SST temporal changes via temporal convolution units and gate linear units. The graph layer encodes SST spatial changes in terms of characteristics derived from graph Laplacian. The output layer encapsulates information from the previous layers and produces SST prediction results. The MGCN characterizes both the temporal and spatial changes, rendering a comprehensive SST prediction strategy. We use daily mean SST data for two areas near the Bohai Sea and the East China Sea for experimental evaluations and validate that the MGCN performs better than other traditional machine learning methods for nearshore SST prediction. In addition, we test the MGCN on weekly and monthly mean SST datasets and validate that the MGCN is robust and suitable for SST prediction.
Xiaoyu Zhang 0002, Alejandro C. Frery, Peng Ren 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 Consistent Sub-Decision Network for Low-Quality Masked Face Recognition
abstract
The COVID-19 pandemic makes wearing masks mandatory in supermarkets, pharmacies, public transport, etc. Existing facial recognition systems encounter severe performance degradation as the masks occlude key facial regions. Recently, simulation-based methods are proposed to generate masked faces from unmasked faces. However, among simulated faces, there are low-quality samples with negative occlusion, which leads to ambiguous or absent facial features. In this paper, we propose a consistent sub-decision network to obtain sub-decisions that correspond to different facial regions and constrain sub-decisions by weighted bidirectional KL divergence to make the network concentrate on the upper faces without occlusion. In addition, we perform knowledge distillation to drive the masked face embeddings towards an approximation of the original data distribution to mitigate the information loss. Experiments show that the proposed method performs better than the baseline on public masked face recognition datasets, i.e., RMFD, MFR2, and MLFW.
Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001
IEEE Signal Process. Lett.4
2022 Heterogeneous Face Recognition via Face Synthesis With Identity-Attribute Disentanglement
abstract
Heterogeneous Face Recognition (HFR) aims to match faces across different domains (e.g., visible to near-infrared images), which has been widely applied in authentication and forensics scenarios. However, HFR is a challenging problem because of the large cross-domain discrepancy, limited heterogeneous data pairs, and large variation of facial attributes. To address these challenges, we propose a new HFR method from the perspective of heterogeneous data augmentation, named Face Synthesis with Identity-Attribute Disentanglement (FSIAD). Firstly, the identity-attribute disentanglement (IAD) decouples face images into identity-related representations and identity-unrelated representations (called attributes), and then decreases the correlation between identities and attributes. Secondly, we devise a face synthesis module (FSM) to generate a large number of images with stochastic combinations of disentangled identities and attributes for enriching the attribute diversity of synthetic images. Both the original images and the synthetic ones are utilized to train the HFR network for tackling the challenges and improving the performance of HFR. Extensive experiments on five HFR databases validate that FSIAD obtains superior performance than previous HFR approaches. Particularly, FSIAD obtains 4.8% improvement over state of the art in terms of VR@FAR=0.01% on LAMP-HQ, the largest HFR database so far.
Zimin (Max) Yang, Jian Liang 0001, Chaoyou Fu, Mandi Luo, Xiaoyu Zhang 0002
IEEE Trans. Inf. Forensics Secur.5
2022 Action Shuffling for Weakly Supervised Temporal Localization
abstract
Weakly supervised action localization is a challenging task with extensive applications, which aims to identify actions and the corresponding temporal intervals with only video-level annotations available. This paper analyzes the order-sensitive and location-insensitive properties of actions, and embodies them into a self-augmented learning framework to improve the weakly supervised action localization performance. To be specific, we propose a novel two-branch network architecture with intra/inter-action shuffling, referred to as ActShufNet. The intra-action shuffling branch lays out a self-supervised order prediction task to augment the video representation with inner-video relevance, whereas the inter-action shuffling branch imposes a reorganizing strategy on the existing action contents to augment the training set without resorting to any external resources. Furthermore, the global-local adversarial training is presented to enhance the model's robustness to irrelevant noises. Extensive experiments are conducted on three benchmark datasets, and the results clearly demonstrate the efficacy of the proposed method.
Xiaoyu Zhang 0002, Haichao Shi, Xinchu Shi
IEEE Trans. Image Process.1
2021 SAPS: Self-Attentive Pathway Search for weakly-supervised action localization with background-action augmentation
Xiaoyu Zhang 0002, Yaru Zhang, Haichao Shi, Jing Dong 0003
Comput. Vis. Image Underst.1
2021 Weakly-supervised action localization via embedding-modeling iterative optimization
Xiaoyu Zhang 0002, Haichao Shi, Peng Li 0035, Zekun Li 0001, Peng Ren 0001
Pattern Recognit.1
2021 FA-GAN: Face Augmentation GAN for Deformation-Invariant Face Recognition
abstract
Substantial improvements have been achieved in the field of face recognition due to the successful application of deep neural networks. However, existing methods are sensitive to both the quality and quantity of the training data. Despite the availability of large-scale datasets, the long tail data distribution induces strong biases in model learning. In this paper, we present a Face Augmentation Generative Adversarial Network (FA-GAN) to reduce the influence of imbalanced deformation attribute distributions. We propose to decouple these attributes from the identity representation with a novel hierarchical disentanglement module. Moreover, Graph Convolutional Networks (GCNs) are applied to recover geometric information by exploring the interrelations among local regions to guarantee the preservation of identities in face data augmentation. Extensive experiments on face reconstruction, face manipulation, and face recognition demonstrate the effectiveness and generalization ability of the proposed method.
Mandi Luo, Jie Cao 0002, Xin Ma 0031, Xiaoyu Zhang 0002, Ran He 0001
IEEE Trans. Inf. Forensics Secur.4
2020 Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed Videos
abstract
Weakly supervised action recognition and localization for untrimmed videos is a challenging problem with extensive applications. The overwhelming irrelevant background contents in untrimmed videos severely hamper effective identification of actions of interest. In this paper, we propose a novel multi-instance multi-label modeling network based on spatio-temporal pre-trimming to recognize actions and locate corresponding frames in untrimmed videos. Motivated by the fact that person is the key factor in a human action, we spatially and temporally segment each untrimmed video into person-centric clips with pose estimation and tracking techniques. Given the bag-of-instances structure associated with video-level labels, action recognition is naturally formulated as a multi-instance multi-label learning problem. The network is optimized iteratively with selective coarse-to-fine pre-trimming based on instance-label activation. After convergence, temporal localization is further achieved with local-global temporal class activation map. Extensive experiments are conducted on two benchmark datasets, i.e. THUMOS14 and ActivityNet1.3, and experimental results clearly corroborate the efficacy of our method when compared with the state-of-the-arts.
Xiaoyu Zhang 0002, Haichao Shi, Peng Li 0035
AAAI1
2020 Self-Paced Learning with Superpixelwise Features for Hyperspectral Image Classification
abstract
We explore self-paced boost learning (SPBL) with superpix-elwise features for hyperspectral image classification (HIC). Firstly, we conduct feature extraction using superpixelwise principal component analysis (SuperPCA), which reduces the dimensionality of hyperspectral images considering the project discrepancy in different homogeneous regions. Secondly, we perform classification by SPBL on the extracted features, where the learning just focuses on the pixels to be classified and does not make their spatial neighbours involved. SPBL embraces the power of self-paced learning on classifying from simple to complex and that of boost learning on classifying in a robust fashion. Our method is not deep learning grounded and the training does not demand high computing resources. The experimental results on two public hyperspectral image datasets demonstrate that our method is competitive with several prominent ones.
Xiaoxiao Tai, Guangxing Wang 0001, Lirong Han, Xiaoyu Zhang 0002, Peng Ren 0001
IGARSS4
2020 Multi-Scale Deep Residual Learning for Cloud Removal
abstract
This paper proposes a multi-scale deep residual network (MDRN) for removing clouds from remote sensing images. MDRN characterizes learning the residuals of cloud-free images and cloudy images instead of directly mapping the two parts. The learned residuals are subsequently integrated with the input cloudy images to generate the de-clouded results. The advantages of our MDRN are threefold. Firstly, we untangle the sparse details of cloudy images from the bases by a guided filter, making the learning focus on processing the textures and the cloud features in the details. Secondly, we employ multi-scale convolution units (MS-Conv), which have larger receptive fields than plain convolutions and assist in extracting representative features. Thirdly, we leverage deep residual learning to avoid performance degradation and relax learning burdens. Experimental results on the remote sensing image cloud removing dataset (RICE) validate the effectiveness of our MDRN.
Qiaoqiao Yang, Guangxing Wang 0001, Yaxuan Zhao, Xiaoyu Zhang 0002, Guoshuai Dong, Peng Ren 0001
IGARSS4
2020 On Deep Unsupervised Active Learning
abstract
Unsupervised active learning has attracted increasing attention in recent years, where its goal is to select representative samples in an unsupervised setting for human annotating. Most existing works are based on shallow linear models by assuming that each sample can be well approximated by the span (i.e., the set of all linear combinations) of certain selected samples, and then take these selected samples as representative ones to label. However, in practice, the data do not necessarily conform to linear models, and how to model nonlinearity of data often becomes the key point to success. In this paper, we present a novel Deep neural network framework for Unsupervised Active Learning, called DUAL. DUAL can explicitly learn a nonlinear embedding to map each input into a latent space through an encoder-decoder architecture, and introduce a selection block to select representative samples in the the learnt latent space. In the selection block, DUAL considers to simultaneously preserve the whole input patterns as well as the cluster structure of data. Extensive experiments are performed on six publicly available datasets, and experimental results clearly demonstrate the efficacy of our method, compared with state-of-the-arts.
Handong Ma, Zhao Kang 0001, Ye Yuan 0001, Xiaoyu Zhang 0002, Guoren Wang
IJCAI5
2020 A Multi-source Self-adaptive Transfer Learning Model for Mining Social Links
Kai Zhang 0079, Longtao He, Chenglong Li 0001, Xiaoyu Zhang 0002
KSEM (2)5
2020 Malware Classification on Imbalanced Data through Self-Attention
abstract
Malware is an ever-growing threat to the Internet. New and mutated malware are appearing with increasing frequency in recent years. In the real-world scenario, new families of malware are often discovered in the cybersecurity protection system. In order to improve the protection capability of the system, it is necessary to add the newly discovered malware identification characteristics to the online system. However, the samples of newly discovered malware families are usually too small to effectively extract the characteristics of new classes, resulting in a low recognition rate of new families. The essence of this problem is a multi-classification problem based on imbalanced datasets. In this paper, we propose a self-attention based malware classification method to solve the malware classification on imbalanced datasets. An open source dataset is used to simulate the classification of malware on balanced and imbalanced datasets. Our method has reached an accuracy of 98.48%, and the F1-Score of the imbalanced Simda class has reached 89.66% on the Microsoft Kaggle dataset. Experimental results have demonstrated the effectiveness and robustness in malware classification with imbalanced datasets.
Jian Xing, Xiaoyu Zhang 0002, ZiSen Oi, Ge Fu, Qian Qiang, Haoliang Sun
TrustCom4
2020 Attention-aware invertible hashing network with skip connections
Shanshan Li 0005, Qiang Cai 0001, Zhuangzi Li, Hai-Sheng Li 0002, Naiguang Zhang, Xiaoyu Zhang 0002
Pattern Recognit. Lett.6
2020 Hashing Nets for Hashing: A Quantized Deep Learning to Hash Framework for Remote Sensing Image Retrieval
abstract
Fast and accurate remote sensing image retrieval from large data archives has been an important research topic in the remote sensing research literature. Recently, hashing-based remote sensing image retrieval has attracted extreme attention because of its efficient search capabilities. Especially, deep remote sensing image hashing algorithms have been developed based on convolutional neural networks (CNNs) and have shown effective retrieval performance. However, implementing a deep hashing network tends to be highly expensive in terms of storage space and computing resources to be suitable for on-orbit remote sensing image retrieval, which usually operates on resource-limited devices such as satellites and unmanned aerial vehicles (UAVs). To address this limitation, we propose to hash a deep network that in turn hashes remote sensing images. Specifically, we develop a quantized deep learning to hash (QDLH) framework for large-scale remote sensing image retrieval. The weights and activation functions in the QDLH framework are binarized to low-bit representations, which require comparatively much less storage space and computing resources. The QDLH results in a lightweight deep neural network for effective remote sensing image hashing. We conduct extensive experiments on two public remote sensing image data sets by incorporating several state-of-the-art network architectures into our QDLH methodology for remote sensing image hashing. The experimental results demonstrate that the proposed QDLH is effective in saving hardware resources in terms of both storage and computation. Moreover, superior remote sensing image retrieval performance is also achieved by our QDLH, compared with state-of-the-art deep remote sensing image hashing methods.
Peng Li 0035, Lirong Han, Xuanwen Tao, Xiaoyu Zhang 0002, Christos Grecos, Antonio Plaza, Peng Ren 0001
IEEE Trans. Geosci. Remote. Sens.4
2019 Learning Transferable Self-Attentive Representations for Action Recognition in Untrimmed Videos with Weak Supervision
abstract
Action recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video frame/sequence, which is quite costly and time-consuming. In this paper, given only video-level annotations, we propose a novel weakly supervised framework to simultaneously locate action frames as well as recognize actions in untrimmed videos. Our proposed framework consists of two major components. First, for action frame localization, we take advantage of the self-attention mechanism to weight each frame, such that the influence of background frames can be effectively eliminated. Second, considering that there are trimmed videos publicly available and also they contain useful information to leverage, we present an additional module to transfer the knowledge from trimmed videos for improving the classification performance in untrimmed ones. Extensive experiments are conducted on two benchmark datasets (i.e., THUMOS14 and ActivityNet1.3), and experimental results clearly corroborate the efficacy of our method.
Xiaoyu Zhang 0002, Haichao Shi, Kai Zheng 0001, Xiaobin Zhu 0001, Lixin Duan
AAAI1
2019 Residual Invertible Spatio-Temporal Network for Video Super-Resolution
abstract
Video super-resolution is a challenging task, which has attracted great attention in research and industry communities. In this paper, we propose a novel end-to-end architecture, called Residual Invertible Spatio-Temporal Network (RISTN) for video super-resolution. The RISTN can sufficiently exploit the spatial information from low-resolution to high-resolution, and effectively models the temporal consistency from consecutive video frames. Compared with existing recurrent convolutional network based approaches, RISTN is much deeper but more efficient. It consists of three major components: In the spatial component, a lightweight residual invertible block is designed to reduce information loss during feature transformation and provide robust feature representations. In the temporal component, a novel recurrent convolutional model with residual dense connections is proposed to construct deeper network and avoid feature degradation. In the reconstruction component, a new fusion method based on the sparse strategy is proposed to integrate the spatial and temporal features. Experiments on public benchmark datasets demonstrate that RISTN outperforms the state-ofthe-art methods.
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002
AAAI3
2019 A Channel Hopping Strategy Based on the Human Trajectory Similarity for WBANs
abstract
Wireless Body Area Network (WBAN) has permeated in various fields, such as e-health, entertainment, sports and so on. However, Inter-Wban interference makes it rather difficult to ensure the reliability of data transmission. While a WBAN comes into the interference range of others, the collision is inevitable if they are working on the same channel. In this paper, we design a channel hopping strategy based on the human movement trajectory, which allows two WBANs with high meeting probability to hop to different channels, thus decreases the probability of interference. Furthermore, we design a more practical metric, which reflects the asynchronous nature of WBANs' channel hopping status, to measure the probability of two WBANs hopping to the same channel. In addition, the proposed strategy is energy efficiency and low latency without channel sensing or negotiating with other WBANs. Simulation results demonstrate the performance of the proposed strategy.
Xiaoyu Zhang 0002, Bin Liu 0016
BSN1
2019 Fi-GNN: Modeling Feature Interactions via Graph Neural Networks for CTR Prediction
abstract
Click-through rate (CTR) prediction is an essential task in web applications such as online advertising and recommender systems, whose features are usually in multi-field form. The key of this task is to model feature interactions among different feature fields. Recently proposed deep learning based models follow a general paradigm: raw sparse input multi-field features are first mapped into dense field embedding vectors, and then simply concatenated together to feed into deep neural networks (DNN) or other specifically designed networks to learn high-order feature interactions. However, the simple unstructured combination of feature fields will inevitably limit the capability to model sophisticated interactions among different fields in a sufficiently flexible and explicit fashion. In this work, we propose to represent the multi-field features in a graph structure intuitively, where each node corresponds to a feature field and different fields can interact through edges. The task of modeling feature interactions can be thus converted to modeling node interactions on the corresponding graph. To this end, we design a novel model Feature Interaction Graph Neural Networks (Fi-GNN). Taking advantage of the strong representative power of graphs, our proposed model can not only model sophisticated feature interactions in a flexible and explicit fashion, but also provide good model explanations for CTR prediction. Experimental results on two real-world datasets show its superiority over the state-of-the-arts.
Zekun Li 0001, Zeyu Cui, Xiaoyu Zhang 0002, Liang Wang 0001
CIKM4
2019 Semi-Supervised Compatibility Learning Across Categories for Clothing Matching
abstract
Learning the compatibility between fashion items across categories is a key task in fashion analysis, which can decode the secret of clothing matching. The main idea of this task is to map items into a latent style space where compatible items stay close. Previous works try to build such a transformation by minimizing the distances between annotated compatible items, which require massive item-level supervision. However, these annotated data are expensive to obtain and hard to cover the numerous items with various styles in real applications. In such cases, these supervised methods fail to achieve satisfactory performances. In this work, we propose a semi-supervised method to learn the compatibility across categories. We observe that the distributions of different categories have intrinsic similar structures. Accordingly, the better distributions align, the closer compatible items across these categories become. To achieve the alignment, we minimize the distances between distributions with unsupervised adversarial learning, and also the distances between some annotated compatible items which play the role of anchor points to help align. Experimental results on two real-world datasets demonstrate the effectiveness of our method.
Zekun Li 0001, Zeyu Cui, Xiaoyu Zhang 0002, Liang Wang 0001
ICME4
2019 Weakly-Supervised Action Recognition and Localization via Knowledge Transfer
Haichao Shi, Xiaoyu Zhang 0002
PRCV (1)2
2019 Dressing as a Whole: Outfit Compatibility Learning Based on Node-wise Graph Neural Networks
abstract
With the rapid development of fashion market, the customers' demands of customers for fashion recommendation are rising. In this paper, we aim to investigate a practical problem of fashion recommendation by answering the question “which item should we select to match with the given fashion items and form a compatible outfit”. The key to this problem is to estimate the outfit compatibility. Previous works which focus on the compatibility of two items or represent an outfit as a sequence fail to make full use of the complex relations among items in an outfit. To remedy this, we propose to represent an outfit as a graph. In particular, we construct a Fashion Graph, where each node represents a category and each edge represents interaction between two categories. Accordingly, each outfit can be represented as a subgraph by putting items into their corresponding category nodes. To infer the outfit compatibility from such a graph, we propose Node-wise Graph Neural Networks (NGNN) which can better model node interactions and learn better node representations. In NGNN, the node interaction on each edge is different, which is determined by parameters correlated to the two connected nodes. An attention mechanism is utilized to calculate the outfit compatibility score with learned node representations. NGNN can not only be used to model outfit compatibility from visual or textual modality but also from multiple modalities. We conduct experiments on two tasks: (1) Fill-in-the-blank: suggesting an item that matches with existing components of outfit; (2) Compatibility prediction: predicting the compatibility scores of given outfits. Experimental results demonstrate the great superiority of our proposed method over others.
Zeyu Cui, Zekun Li 0001, Xiaoyu Zhang 0002, Liang Wang 0001
WWW4
2019 Active semi-supervised learning based on self-expressive correlation with generative adversarial networks
Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035
Neurocomputing1
2019 A novel framework for semantic segmentation with generative adversarial network
Xiaobin Zhu 0001, Xiaoyu Zhang 0002, Lei Wang 0101
J. Vis. Commun. Image Represent.3
2019 Deep convolutional representations and kernel extreme learning machines for image classification
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Peng Li 0035, Lei Wang 0101
Multim. Tools Appl.3
2019 Hash Code Reconstruction for Fast Similarity Search
abstract
Learning to hash is a popular technique for fast similarity search on a large-scale image database. However, many hashing methods do not achieve satisfactory results because of the quantization loss in the straightforward binary code generation procedure. In order to address this problem, we propose a novel hash code reconstruction framework for existing unsupervised hashing methods. In our proposed approach, the hash codes are generated through reconstructing the original images with relaxed hamming vector representation, such that the final learned codes will be more approximate to characterize the intrinsic image structure. Moreover, our proposed hash code reconstruction algorithm is very efficient for computing, which can be generalized to various hashing methods. Extensive experiments are conducted on four public image datasets by incorporating our proposed scheme with different hashing methods, and the comparison results have shown that significant performance improvements can be achieved with minor additional time cost for fast similarity search task.
Peng Li 0035, Xiaobin Zhu 0001, Xiaoyu Zhang 0002, Peng Ren 0001, Lei Wang 0101
IEEE Signal Process. Lett.3
2018 Generative Adversarial Image Super-Resolution Through Deep Dense Skip Connections
abstract
Abstract Recently, image super‐resolution works based on Convolutional Neural Networks (CNNs) and Generative Adversarial Nets (GANs) have shown promising performance. However, these methods tend to generate blurry and over‐smoothed super‐resolved (SR) images, due to the incomplete loss function and powerless architectures of networks. In this paper, a novel generative adversarial image super‐resolution through deep dense skip connections (GSR‐DDNet), is proposed to solve the above‐mentioned problems. It aims to take advantage of GAN's ability of modeling data distributions, so that GSR‐DDNet can select informative feature representation and model the mapping across the low‐quality and high‐quality images in an adversarial way. The pipeline of the proposed method consists of three main components: 1) The generator of a novel dense skip connection network with the deep structure for learning robust mapping function is proposed to generate SR images from low‐resolution images; 2) The feature extraction network based on VGG‐19 is adopted to capture high frequency feature maps for content loss; and 3) The discriminator with Wasserstein distance is adopted to identify the overall style of SR and ground‐truth images. Experiments conducted on four publicly available datasets demonstrate the superiority against the state‐of‐the‐art methods.
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Hai-Sheng Li 0002, Lei Wang 0101
Comput. Graph. Forum3
2018 A Self-Paced Regularization Framework for Multilabel Learning
abstract
In this brief, we propose a novel multilabel learning framework, called multilabel self-paced learning, in an attempt to incorporate the SPL scheme into the regime of multilabel learning. Specifically, we first propose a new multilabel learning formulation by introducing a self-paced function as a regularizer, so as to simultaneously prioritize label learning tasks and instances in each iteration. Considering that different multilabel learning scenarios often need different self-paced schemes during learning, we thus provide a general way to find the desired self-paced functions. To the best of our knowledge, this is the first work to study multilabel learning by jointly taking into consideration the complexities of both training instances and labels. Experimental results on four publicly available data sets suggest the effectiveness of our approach, compared with the state-of-the-art methods.
Junchi Yan, Xiaoyu Zhang 0002, Qingshan Liu 0001, Hongyuan Zha
IEEE Trans. Neural Networks Learn. Syst.4
2017 Supporting Real-Time Analytic Queries in Big and Fast Data Environments
Guangjun Wu, Xiao-chun Yun, Chao Li 0062, Yipeng Wang 0001, Xiaoyu Zhang 0002, Siyu Jia, Guangyan Zhang
DASFAA (2)6
2017 Compressing deep neural networks for efficient visual inference
abstract
The deployments of deep neural network models on mobile or embedded devices have been challenged due to two main reasons: 1) the large model size for storage, and 2) the large memory bandwidth for inference. To address these issues, this paper develops a deep neural network compression framework to reduce the resource usage for efficient visual inference. By reviewing the trained deep model, we propose a hybrid model compression algorithm via four major modules. Approximation module reduces the number of weights in each fully connected layer with low rank approximation. Then, quantization module analyzes weight distribution in each layer and represents them with low precision fixed point, which reduces the bits for storing each weight. After that, pruning module suppresses small weights to further reduce the number of parameters. Finally, coding module joint optimizes the representation and encoding of the sparse structure of the pruned weights with relative index by Huffman coding. Beyond the compression of model size, we propose an adaptive fixed point memory allocation algorithm to reduce memory footprint in inference. The proposed framework, along with the model compression and memory allocation algorithms, can provide 20-30x compression rate with negligible accuracy loss. We conduct an evaluation on two representative models, AlexNet and VGG-16, for object recognition and face verification tasks, which demonstrate the effectiveness of our proposed compression framework.
Shiming Ge, Zhao Luo, Shengwei Zhao, Xin Jin 0015, Xiaoyu Zhang 0002
ICME5
2017 ListNet-based object proposals ranking
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Qingxiao Guan, Xianfeng Zhao
Neurocomputing2
2016 MC-HOG Correlation Tracking with Saliency Proposal
abstract
Designing effective feature and handling the model drift problem are two important aspects for online visual tracking. For feature representation, gradient and color features are most widely used, but how to effectively combine them for visual tracking is still an open problem. In this paper, we propose a rich feature descriptor, MC-HOG, by leveraging rich gradient information across multiple color channels or spaces. Then MC-HOG features are embedded into the correlation tracking framework to estimate the state of the target. For handling the model drift problem caused by occlusion or distracter, we propose saliency proposals as prior information to provide candidates and reduce background interference. In addition to saliency proposals, a ranking strategy is proposed to determine the importance of these proposals by exploiting the learnt appearance filter, historical preserved object samples and the distracting proposals. In this way, the proposed approach could effectively explore the color-gradient characteristics and alleviate the model drift problem. Extensive evaluations performed on the benchmark dataset show the superiority of the proposed method.
Guibo Zhu, Jinqiao Wang, Yi Wu 0001, Xiaoyu Zhang 0002, Hanqing Lu
AAAI4
2016 Retweeting behavior prediction using probabilistic matrix factorization
abstract
Retweeting is an important mechanism for information diffusion, popular event prediction, and so on. Due to the increasing requirements, in recent years, the task has attracted extensive attentions. In this paper, we propose a novel framework using probabilistic matrix factorization technique to predict retweeting behavior. Our study consists of three components. First, we convert retweeting behavior problem to a matrix factorization problem. Second, following the intuition that a user's social network will affect his retweeting behavior, we extensively study how to model social information to improve the prediction accuracy. Finally, message semantic embedding information is employed in designing a semantic regularization term to constrain the matrix factorization objective function. We also propose a set of metrics to construct the embeddings among messages based on messages' structural and textual features. The empirical results and analysis demonstrate that our methods perform better than the state-of-the-art approaches.
Kai Zhang 0079, Xiao-chun Yun, Jiguang Liang, Xiaoyu Zhang 0002, Chao Li 0062
ISCC4
2016 Collaborative Multi-View Denoising
abstract
In multi-view learning applications, like multimedia analysis and information retrieval, we often encounter the corrupted view problem in which the data are corrupted by two different types of noises, i.e., the intra- and inter-view noises. The noises may affect these applications that commonly acquire complementary representations from different views. Therefore, how to denoise corrupted views from multi-view data is of great importance for applications that integrate and analyze representations from different views. However, the heterogeneity among multi-view representations brings a significant challenge on denoising corrupted views. To address this challenge, we propose a general framework to jointly denoise corrupted views in this paper. Specifically, aiming at capturing the semantic complementarity and distributional similarity among different views, a novel Heterogeneous Linear Metric Learning (HLML) model with low-rank regularization, leave-one-out validation, and pseudo-metric constraints is proposed. Our method linearly maps multi-view data to a high-dimensional feature-homogeneous space that embeds the complementary information from different views. Furthermore, to remove the intra- and inter-view noises, we present a new Multi-view Semi-supervised Collaborative Denoising (MSCD) method with elementary transformation constraints and gradient energy competition to establish the complementary relationship among the heterogeneous representations. Experimental results demonstrate that our proposed methods are effective and efficient.
Lei Zhang 0116, Xiaoyu Zhang 0002, Yong Wang 0032, Binbin Li 0001, Dinggang Shen, Shuiwang Ji
KDD3
2016 Simultaneous optimization for robust correlation estimation in partially observed social network
Xiaoyu Zhang 0002
Neurocomputing1
2016 Weighted hierarchical geographic information description model for social relation estimation
Kai Zhang 0079, Xiao-chun Yun, Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Chao Li 0062
Neurocomputing3
2015 A Markov Random Field Approach to Automated Protocol Signature Inference
Yongzheng Zhang 0002, Yipeng Wang 0001, Jianliang Sun, Xiaoyu Zhang 0002
SecureComm5
2015 Context-aware local abnormality detection in crowded scene
Xiaobin Zhu 0001, Xin Jin 0015, Xiaoyu Zhang 0002, Fugang He, Lei Wang 0101
Sci. China Inf. Sci.3
2015 Update vs. upgrade: Modeling with indeterminate multi-class active learning
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Xiao-chun Yun, Guangjun Wu, Yipeng Wang 0001
Neurocomputing1
2015 Bidirectional Active Learning: A Two-Way Exploration Into Unlabeled and Labeled Data Set
abstract
In practical machine learning applications, human instruction is indispensable for model construction. To utilize the precious labeling effort effectively, active learning queries the user with selective sampling in an interactive way. Traditional active learning techniques merely focus on the unlabeled data set under a unidirectional exploration framework and suffer from model deterioration in the presence of noise. To address this problem, this paper proposes a novel bidirectional active learning algorithm that explores into both unlabeled and labeled data sets simultaneously in a two-way process. For the acquisition of new knowledge, forward learning queries the most informative instances from unlabeled data set. For the introspection of learned knowledge, backward learning detects the most suspiciously unreliable instances within the labeled data set. Under the two-way exploration framework, the generalization ability of the learning model can be greatly improved, which is demonstrated by the encouraging experimental results.
Xiaoyu Zhang 0002, Xiao-chun Yun
IEEE Trans. Neural Networks Learn. Syst.1
2009 Multi-view multi-label active learning for image classification
abstract
Image classification is an important topic in multimedia analysis, among which multi-label image classification is a very challenging task with respect to the large demand for human annotation of multi-label samples. In this paper, we propose a multi-view multi-label active learning strategy, which integrates the mechanism of active learning and multi-view learning. On one hand we explore the sample and label uncertainties within each view; on the other hand we capture the uncertainty over different views based on multi-view fusion. Then the overall uncertainty along the sample, label and view dimensions are obtained to detect the most informative sample-label pairs. Experimental results demonstrate the effectiveness of the proposed scheme.
Xiaoyu Zhang 0002, Jian Cheng 0001, Changsheng Xu, Hanqing Lu, Songde Ma
ICME1
2009 Personalized retrieval of sports video based on multi-modal analysis and user preference acquisition
Yifan Zhang 0001, Changsheng Xu, Xiaoyu Zhang 0002, Hanqing Lu
Multim. Tools Appl.3
2009 Effective Annotation and Search for Video Blogs with Integration of Context and Content Analysis
abstract
In recent years, weblogs (or blogs) have received great popularity worldwide, among which video blogs (or vlogs) are playing an increasingly important role. However, research on vlog analysis is still in the early stage, and how to manage vlogs effectively so that they can be more easily accessible is a challenging problem. In this paper, we propose a novel vlog management model which is comprised of automatic vlog annotation and user-oriented vlog search. For vlog annotation, we extract informative keywords from both the target vlog itself and relevant external resources; besides semantic annotation, we perform sentiment analysis on comments to obtain the overall evaluation. For vlog search, we present saliency-based matching to simulate human perception of similarity, and organize the results by personalized ranking and category-based clustering. An evaluation criterion is also proposed for vlog annotation, which assigns a score to an annotation according to its accuracy and completeness in representing the vlog's semantics. Experimental results demonstrate the effectiveness of the proposed management model for vlogs.
Xiaoyu Zhang 0002, Changsheng Xu, Jian Cheng 0001, Hanqing Lu, Songde Ma
IEEE Trans. Multim.1
2008 Automatic semantic annotation for video blogs
abstract
In recent years, Weblogs (or blogs) have received great popularity worldwide, among which video blogs (or vlogs) are playing an increasingly important role. As vlogs gain in population, how to make them more easily accessible has become a hot research topic. In this paper, we propose a novel automatic annotation model for vlogs. We extract informative keywords from both the target vlog itself and external resources which are semantically and visually relevant to it. We also present a new evaluation criterion, which assigns a score to an annotation according to its accuracy and completeness in representing the vlogpsilas semantics. Experimental results demonstrate the effectiveness of both the annotation model and evaluation criterion.
Xiaoyu Zhang 0002, Changsheng Xu, Jian Cheng 0001, Hanqing Lu, Songde Ma
ICME1
2008 Selective Sampling Based on Dynamic Certainty Propagation for Image Retrieval
Xiaoyu Zhang 0002, Jian Cheng 0001, Hanqing Lu, Songde Ma
MMM1
2007 Weighted Co-SVM for Image Retrieval with MVB Strategy
abstract
In relevance feedback, active learning is often used to alleviate the burden of labeling by selecting only the most informative data. Traditional data selection strategies often choose the data closest to the current classification boundary to label, which are in fact not informative enough. In this paper, we propose the moving virtual boundary (MVB) strategy, which is proved to be a more effective way for data selection. The co-SVM algorithm is another powerful method used in relevance feedback. Unfortunately, its basic assumption that each view of the data be sufficient is often untenable in image retrieval. We present our weighted co-SVM as an extension of co-SVM by attaching weight to each view, and thus relax the view sufficiency assumption. The experimental results show that the weighted co-SVM algorithm outperforms co-SVM obviously, especially with the help of MVB data selection strategy.
Xiaoyu Zhang 0002, Jian Cheng 0001, Hanqing Lu, Songde Ma
ICIP (4)1