Zhiyuan Zhang 0001

dblp:72/1760-1 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
10since 2021 · last 2023
0000-0001-5550-8897ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Fed-FA: Theoretically Modeling Client Data Divergence for Federated Language Backdoor Defense
abstract
Federated learning algorithms enable neural network models to be trained across multiple decentralized edge devices without sharing private data. However, they are susceptible to backdoor attacks launched by malicious clients. Existing robust federated aggregation algorithms heuristically detect and exclude suspicious clients based on their parameter distances, but they are ineffective on Natural Language Processing (NLP) tasks. The main reason is that, although text backdoor patterns are obvious at the underlying dataset level, they are usually hidden at the parameter level, since injecting backdoors into texts with discrete feature space has less impact on the statistics of the model parameters. To settle this issue, we propose to identify backdoor clients by explicitly modeling the data divergence among clients in federated NLP systems. Through theoretical analysis, we derive the f-divergence indicator to estimate the client data divergence with aggregation updates and Hessians. Furthermore, we devise a dataset synthesization method with a Hessian reassignment mechanism guided by the diffusion theory to address the key challenge of inaccessible datasets in calculating clients' data Hessians. We then present the novel Federated F-Divergence-Based Aggregation~(\textbf{Fed-FA}) algorithm, which leverages the f-divergence indicator to detect and discard suspicious clients. Extensive empirical results show that Fed-FA outperforms all the parameter distance-based methods in defending against backdoor attacks among various natural language backdoor attack scenarios.
Zhiyuan Zhang 0001, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
NeurIPS1
2023 ASAT: Adaptively scaled adversarial training in time series
Zhiyuan Zhang 0001, Wei Li 0101, Ruihan Bao, Keiko Harimoto, Yunfang Wu, Xu Sun 0001
Neurocomputing1
2022 GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved Generalization
abstract
Recently, Sharpness-Aware Minimization (SAM) algorithm has shown state-of-the-art generalization abilities in vision tasks.It demonstrates that flat minima tend to imply better generalization abilities.However, it has some difficulty implying SAM to some natural language tasks, especially to models with drastic gradient changes, such as RNNs.In this work, we analyze the relation between the flatness of the local minimum and its generalization ability from a novel and straightforward theoretical perspective.We propose that the shift of the training and test distributions can be equivalently seen as a virtual parameter corruption or perturbation, which can explain why flat minima that are robust against parameter corruptions or perturbations have better generalization performances.On its basis, we propose a Gradient-Strength based Adaptive Sharpness-Aware Minimization (GA-SAM) algorithm to help to learn algorithms find flat minima that generalize better.Results in various language benchmarks validate the effectiveness of the proposed GA-SAM algorithm on natural language tasks.
Zhiyuan Zhang 0001, Ruixuan Luo, Qi Su 0001, Xu Sun 0001
EMNLP1
2022 How to Inject Backdoors with Better Consistency: Logit Anchoring on Clean Data
Zhiyuan Zhang 0001, Lingjuan Lyu, Weiqiang Wang 0002, Lichao Sun 0001, Xu Sun 0001
ICLR1
2022 Stock Trading Volume Prediction with Dual-Process Meta-Learning
Wei Li 0089, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)3
2022 Distributional Correlation-Aware Knowledge Distillation for Stock Trading Volume Prediction
Lei Li 0039, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)2
2021 Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter Corruption
abstract
We argue that the vulnerability of model parameters is of crucial value to the study of model robustness and generalization but little research has been devoted to understanding this matter. In this work, we propose an indicator to measure the robustness of neural network parameters by exploiting their vulnerability via parameter corruption. The proposed indicator describes the maximum loss variation in the non-trivial worst-case scenario under parameter corruption. For practical purposes, we give a gradient-based estimation, which is far more effective than random corruption trials that can hardly induce the worst accuracy degradation. Equipped with theoretical support and empirical validation, we are able to systematically investigate the robustness of different model parameters and reveal vulnerability of deep neural networks that has been rarely paid attention to before. Moreover, we can enhance the models accordingly with the proposed adversarial corruption-resistant training, which not only improves the parameter robustness but also translates into accuracy elevation.
Xu Sun 0001, Zhiyuan Zhang 0001, Xuancheng Ren, Ruixuan Luo, Liangyou Li
AAAI2
2021 Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
abstract
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Wenkai Yang, Lei Li 0039, Zhiyuan Zhang 0001, Xuancheng Ren, Xu Sun 0001
NAACL-HLT3
2021 Neural Network Surgery: Injecting Data Patterns into Pre-trained Models with Minimal Instance-wise Side Effects
abstract
Zhiyuan Zhang, Xuancheng Ren, Qi Su, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhiyuan Zhang 0001, Xuancheng Ren, Qi Su 0001, Xu Sun 0001
NAACL-HLT1
2021 Adversarial parameter defense by multi-step risk minimization
Zhiyuan Zhang 0001, Ruixuan Luo, Xuancheng Ren, Qi Su 0001, Liangyou Li, Xu Sun 0001
Neural Networks1
2020 Rethinking Skip Connection with Layer Normalization
abstract
Skip connection is a widely-used technique to improve the performance and the convergence of deep neural networks, which is believed to relieve the difficulty in optimization due to non-linearity by propagating a linear component through the neural network layers. However, from another point of view, it can also be seen as a modulating mechanism between the input and the output, with the input scaled by a pre-defined value one. In this work, we investigate how the scale factors in the effectiveness of the skip connection and reveal that a trivial adjustment of the scale will lead to spurious gradient exploding or vanishing in line with the deepness of the models, which could by addressed by normalization, in particular, layer normalization, which induces consistent improvements over the plain skip connection. Inspired by the findings, we further propose to adaptively adjust the scale of the input by recursively applying skip connection with layer normalization, which promotes the performance substantially and generalizes well across diverse tasks including both machine translation and image classification datasets.
Xuancheng Ren, Zhiyuan Zhang 0001, Xu Sun 0001, Yuexian Zou
COLING3
2020 Memorized sparse backpropagation
Zhiyuan Zhang 0001, Xuancheng Ren, Qi Su 0001, Xu Sun 0001
Neurocomputing1
2019 Understanding and Improving Layer Normalization
abstract
Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accuracy. However, it is still unclear where the effectiveness stems from. In this paper, our main contribution is to take a step further in understanding LayerNorm. Many of previous studies believe that the success of LayerNorm comes from forward normalization. Unlike them, we find that the derivatives of the mean and variance are more important than forward normalization by re-centering and re-scaling backward gradients. Furthermore, we find that the parameters of LayerNorm, including the bias and gain, increase the risk of over-fitting and do not work in most cases. Experiments show that a simple version of LayerNorm (LayerNorm-simple) without the bias and gain outperforms LayerNorm on four datasets. It obtains the state-of-the-art performance on En-Vi machine translation. To address the over-fitting problem, we propose a new normalization method, Adaptive Normalization (AdaNorm), by replacing the bias and gain with a new transformation function. Experiments show that AdaNorm demonstrates better results than LayerNorm on seven out of eight datasets.
Jingjing Xu 0001, Xu Sun 0001, Zhiyuan Zhang 0001, Guangxiang Zhao, Junyang Lin
NeurIPS3
2019 Automatic Translating Between Ancient Chinese and Contemporary Chinese with Limited Aligned Corpora
Zhiyuan Zhang 0001, Wei Li 0101, Qi Su 0001
NLPCC (2)1
2018 Building an Ellipsis-aware Chinese Dependency Treebank for Web Text
Xuancheng Ren, Xu Sun 0001, Ji Wen, Bingzhen Wei, Weidong Zhan, Zhiyuan Zhang 0001
LREC6