Peng Yang 0013

dblp:57/5443-13 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
13since 2021 · last 2023
0009-0005-3449-5843ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Defending Backdoor Attacks on Vision Transformer via Patch Processing
abstract
Vision Transformers (ViTs) have a radically different architecture with significantly less inductive bias than Convolutional Neural Networks. Along with the improvement in performance, security and robustness of ViTs are also of great importance to study. In contrast to many recent works that exploit the robustness of ViTs against adversarial examples, this paper investigates a representative causative attack, i.e., backdoor. We first examine the vulnerability of ViTs against various backdoor attacks and find that ViTs are also quite vulnerable to existing attacks. However, we observe that the clean-data accuracy and backdoor attack success rate of ViTs respond distinctively to patch transformations before the positional encoding. Then, based on this finding, we propose an effective method for ViTs to defend both patch-based and blending-based trigger backdoor attacks via patch processing. The performances are evaluated on several benchmark datasets, including CIFAR10, GTSRB, and TinyImageNet, which show the proposedds defense is very successful in mitigating backdoor attacks for ViTs. To the best of our knowledge, this paper presents the first defensive strategy that utilizes a unique characteristic of ViTs against backdoor attacks.
Khoa D. Doan, Yingjie Lao, Peng Yang 0013, Ping Li 0001
AAAI3
2023 Power Norm Based Lifelong Learning for Paraphrase Generations
abstract
Lifelong seq2seq language generation models are trained with multiple domains in a lifelong learning manner, with data from each domain being observed in an online fashion. It is a well-known problem that lifelong learning suffers from the catastrophic forgetting (CF). To handle this challenge, existing works have leveraged experience replay or dynamic architecture to consolidate the past knowledge, which however result in incremental memory space or high computational cost. In this work, we propose a novel framework name "power norm based lifelong learning" (PNLLL), which aims to remedy the catastrophic forgetting issues with a power normalization on NLP transformer models. Specifically, PNLLL leverages power norm to achieve a better balance between past experience rehearsal and new knowledge acquisition. These designs enable the knowledge adaptation onto new tasks while memorizing the experience of past tasks. Our experiments on paraphrase generation tasks show that PNLLL not only outperforms SOTA models by a considerable margin and but also largely alleviates forgetting.
Dingcheng Li, Peng Yang 0013, Yue Zhang 0086, Ping Li 0001
SIGIR2
2022 DeepAuth: A DNN Authentication Framework by Model-Unique and Fragile Signature Embedding
abstract
Along with the evolution of deep neural networks (DNNs) in many real-world applications, the complexity of model building has also dramatically increased. Therefore, it is vital to protect the intellectual property (IP) of the model builder and ensure the trustworthiness of the deployed models. Meanwhile, adversarial attacks on DNNs (e.g., backdoor and poisoning attacks) that seek to inject malicious behaviors have been investigated recently, demanding a means for verifying the integrity of the deployed model to protect the users. This paper presents a novel DNN authentication framework DeepAuth that embeds a unique and fragile signature to each protected DNN model. Our approach exploits sensitive key samples that are well crafted from the input space to latent space and then to logit space for producing signatures. After embedding, each model will respond distinctively to these key samples, which creates a model-unique signature as a strong tool for authentication and user identity. The signature embedding process is also designed to ensure the fragility of the signature, which can be used to detect malicious modifications such that an illegitimate user or an altered model should not have the intact signature. Extensive evaluations on various models over a wide range of datasets demonstrate the effectiveness and efficiency of the proposed DeepAuth.
Yingjie Lao, Weijie Zhao 0001, Peng Yang 0013, Ping Li 0001
AAAI3
2022 Feature Fusion Network for Personalized Online Advertising Systems
abstract
Sponsored online advertising delivers many billions of revenues for online ads publishers. The ads systems take userinput query keywords and display ads that are relevant to the query. the task of click-through rate (CTR) prediction aims to estimate the likelihood of a user clicking on the ads, which has become one of the core goals in the ads system. In order to further improve the CTR, user portraits are also considered as an input to make personalized ads display and recommendations, in the current deep learning CTR training platform. The naive combination of user space (~ 109) and feature space (~ 10]12however, would yield a 1021dimensional space. It is not only infeasible to feed the 1021parameters into the embedding layer with any off-the-shelf storage, but also impractical to train the network in such massive-scale dimensional space. In this paper, we design a novel CTR prediction framework for ads systems to tackle the massive-scale user-feature combination challenge. Specifically, we introduce a feature fusion network to explicitly learn user-feature cross embedding in an end-to-end manner. To improve the efficiency, we prune the feature fusion networks to a practical number through a network importance ranking scheme. Extensive empirical experiments on Baidu’s ads data validate the effectiveness of the proposed feature fusion networks.
Weijie Zhao 0001, Peng Yang 0013, Lin Li 0001, Ping Li 0001
IEEE Big Data2
2022 Continual Learning for Natural Language Generations with Transformer Calibration
abstract
Conventional natural language process (NLP) generation models are trained offline with a given dataset for a particular task, which is referred to as isolated learning.Research on sequence-to-sequence language generation aims to study continual learning model to constantly learning from sequentially encountered tasks.However, continual learning studies often suffer from catastrophic forgetting, a persistent challenge for lifelong learning.In this paper, we present a novel NLP transformer model that attempts to mitigate catastrophic forgetting in online continual learning from a new perspective, i.e., attention calibration.We model the attention in the transformer as a calibrated unit in a general formulation, where the attention calibration could give benefits to balance the stability and plasticity of continual learning algorithms through influencing both their forward inference path and backward optimization path.Our empirical experiments, paraphrase generation and dialog response generation, demonstrate that this work outperforms state-of-theart models by a considerable margin and effectively mitigate the forgetting.
Peng Yang 0013, Dingcheng Li, Ping Li 0001
CoNLL1
2022 One Loss for Quantization: Deep Hashing with Discrete Wasserstein Distributional Matching
abstract
Image hashing is a principled approximate nearest neighbor approach to find similar items to a query in a large collection of images. Hashing aims to learn a binary-output function that maps an image to a binary vector. For optimal retrieval performance, producing balanced hash codes with low-quantization error to bridge the gap between the learning stage's continuous relaxation and the inference stage's discrete quantization is important. However, in the existing deep supervised hashing methods, coding balance and low-quantization error are difficult to achieve and involve several losses. We argue that this is because the existing quantization approaches in these methods are heuristically constructed and not effective to achieve these objectives. This paper considers an alternative approach to learning the quantization constraints. The task of learning balanced codes with low quantization error is re-formulated as matching the learned distribution of the continuous codes to a pre-defined discrete, uniform distribution. This is equivalent to minimizing the distance between two distributions. We then propose a computationally efficient distributional distance by leveraging the discrete property of the hash functions. This distributional distance is a valid distance and enjoys lower time and sample complexities. The proposed single-loss quantization objective can be integrated into any existing supervised hashing method to improve code balance and quantization error. Experiments confirm that the proposed approach substantially improves the performance of several representative hashing methods.
Khoa D. Doan, Peng Yang 0013, Ping Li 0001
CVPR2
2022 Identification for Deep Neural Network: Simply Adjusting Few Weights!
abstract
Through the development of powerful algorithms and design tools, deep neural networks (DNNs) have recently approached or even surpassed human-level performance in many real-world applications. Nowadays, since a product-level DNN modeling requires a large amount of training data and expensive computing resources and thus DNN models are considered as valuable data, protecting the intellectual property (IP) of DNN builders becomes an important problem in the security domain. In this paper, we propose a novel watermarking approach that only requires adjusting a few weights, as opposed to prior works that embed watermarks via end-to-end training. The protected model with tiny parameter modifications can output pre-specified labels with carefully selected key samples as inputs, which serves as a strong proof of ownership. Besides, our methodology can be naturally extended to identification, i.e., embedding unique watermarks to identify different users. Watermark embedding is achieved by modifying a very small subset of parameters, guaranteeing a high fidelity while dramatically reducing the computational overhead. The experimental results demonstrate that the proposed algorithm can embed key samples with a high success rate, while well preserving the original functionality of the target model. We show that the proposed method is robust against various transformation attacks.
Yingjie Lao, Peng Yang 0013, Weijie Zhao 0001, Ping Li 0001
ICDE2
2022 Calibrating CNNs for Few-Shot Meta Learning
abstract
Although few-shot meta learning has been extensively studied in machine learning community, the fast adaptation towards new tasks remains a challenge in the few-shot learning scenario. The neuroscience research reveals that the capability of evolving neural network formulation is essential for task adaptation, which has been broadly studied in recent meta-learning researches. In this paper, we present a novel forward-backward meta-learning framework (FBM) to facilitate the model generalization in few-shot learning from a new perspective, i.e., neuron calibration. In particular, FBM models the neurons in deep neural network-based model as calibrated units under a general formulation, where neuron calibration could empower fast adaptation capability to the neural network-based models through influencing both their forward inference path and backward propagation path. The proposed calibration scheme is lightweight and applicable to various feed-forward neural network architectures. Extensive empirical experiments on the challenging few-shot learning benchmarks validate that our approach training with neuron calibration achieves a promising performance, which demonstrates that neuron calibration plays a vital role in improving the few-shot learning performance.
Peng Yang 0013, Shaogang Ren, Ping Li 0001
WACV1
2021 Efficient Learning to Learn a Robust CTR Model for Web-scale Online Sponsored Search Advertising
abstract
Click-through rate (CTR) prediction is crucial for online sponsored search advertising. Several successful CTR models have been adopted in the industry, including the regularized logistic regression (LR). Nonetheless, the learning process suffers from two limitations: 1) Feature crosses for high-order information may generate trillions of features, which are sparse for online learning examples; 2) Rapid changing of data distribution brings challenges to the accurate learning since the model has to perform a fast adaptation on the new data. Moreover, existing adaptive optimizers are ineffective in handling the sparsity issue for high-dimensional features.
Xin Wang 0017, Peng Yang 0013, Shaopeng Chen, Lian Zhao, Jiacheng Guo, Mingming Sun 0001, Ping Li 0001
CIKM2
2021 Adversarial Kernel Sampling on Class-imbalanced Data Streams
abstract
This paper investigates online active learning in the setting of class-imbalanced data streams, where labels are allowed to be queried of with limited budgets. In this setup, conventional learning would be biased towards majority classes and consequently harm the performance. To address this issue, imbalance learning technique adopts both asymmetric losses and asymmetric queries to tackle the imbalance. Although this approach is effective, it may not guarantee the performance in an adversarial setting where the actual labels are unknown, and they may be chosen by the adversary
Peng Yang 0013, Ping Li 0001
CIKM1
2021 Robust Watermarking for Deep Neural Networks via Bi-level Optimization
abstract
Deep neural networks (DNNs) have become state-of-the-art in many application domains. The increasing complexity and cost for building these models demand means for protecting their intellectual property (IP). This paper presents a novel DNN framework that optimizes the robustness of the embedded watermarks. Our method is originated from DNN fault attacks. Different from prior end-to-end DNN watermarking approaches, we only modify a tiny subset of weights to embed the watermark, which also facilities better control of the model behaviors and enables larger rooms for optimizing the robustness of the watermarks.In this paper, built upon the above concept, we pro-pose a bi-level optimization framework where the inner loop phase optimizes the example-level problem to generate robust exemplars, while the outer loop phase proposes a masked adaptive optimization to achieve the robustness of the projected DNN models. Our method alternates the learning of the protected models and watermark exemplars across all phases, where watermark exemplars are not just data samples that could be optimized and adjusted instead. We verify the performance of the proposed methods over a wide range of datasets and DNN architectures. Various transformation attacks including fine-tuning, pruning and overwriting are used to evaluate the robustness.
Peng Yang 0013, Yingjie Lao, Ping Li 0001
ICCV1
2021 Graph-based Adversarial Online Kernel Learning with Adaptive Embedding
abstract
Traditional online learning for vertex classification adapts graph Laplacian regularization into ridge regression, which hardly resolve robustness issue against adversarial examples. To tackle the problem, we propose a more general min-max optimization framework for adversarial online kernel learning (OKL). The derived online algorithm can achieve a min-max regret compared with the optimal model found offline. Nonetheless, optimizing in the reproducing kernel Hilbert space suffers from expensive computational costs. While first-order methods accumulate an optimal $\mathcal{O}(\sqrt{T})$ regret, they only require $\mathcal{O}(t)$ time and space per trial. Second-order methods converge to an optimum much faster with $\mathcal{O}(\log T)$, but suffer an expensive $\mathcal{O}(t^{2})$ per-trial cost. This paper adopts selective sampling to make the OKL scaling to large datasets through constructing a small and accurate embedded space for vertex representation, so that OKL could be performed more efficiently on this sketched kernel space. To achieve this goal, we introduce a novel confidence-aware dictionary selection strategy and a model-based update routine, where for an expected sampling probability p, the computational cost can be reduced by a factor of $p^{2}$ to $\mathcal{O}(p^{2}t^{2})$ space and time per trial, while the regret remains comparable. Numerical experiments are conducted on real-world benchmark datasets to illustrate the efficacy of our method.
Peng Yang 0013, Ping Li 0001
ICDM1
2021 Mitigating Forgetting in Online Continual Learning with Neuron Calibration
abstract
Inspired by human intelligence, the research on online continual learning aims to push the limits of the machine learning models to constantly learn from sequentially encountered tasks, with the data from each task being observed in an online fashion. Though recent studies have achieved remarkable progress in improving the online continual learning performance empowered by the deep neural networks-based models, many of today's approaches still suffer a lot from catastrophic forgetting, a persistent challenge for continual learning. In this paper, we present a novel method which attempts to mitigate catastrophic forgetting in online continual learning from a new perspective, i.e., neuron calibration. In particular, we model the neurons in the deep neural networks-based models as calibrated units under a general formulation. Then we formalize a learning framework to effectively train the calibrated model, where neuron calibration could give ubiquitous benefit to balance the stability and plasticity of online continual learning algorithms through influencing both their forward inference path and backward optimization path. Our proposed formulation for neuron calibration is lightweight and applicable to general feed-forward neural networks-based models. We perform extensive experiments to evaluate our method on four benchmark continual learning datasets. The results show that neuron calibration plays a vital role in improving online continual learning performance and our method could substantially improve the state-of-the-art performance on all~the~evaluated~datasets.
Haiyan Yin, Peng Yang 0013, Ping Li 0001
NeurIPS2
2020 Distributed Primal-Dual Optimization for Online Multi-Task Learning
abstract
Conventional online multi-task learning algorithms suffer from two critical limitations: 1) Heavy communication caused by delivering high velocity of sequential data to a central machine; 2) Expensive runtime complexity for building task relatedness. To address these issues, in this paper we consider a setting where multiple tasks are geographically located in different places, where one task can synchronize data with others to leverage knowledge of related tasks. Specifically, we propose an adaptive primal-dual algorithm, which not only captures task-specific noise in adversarial learning but also carries out a projection-free update with runtime efficiency. Moreover, our model is well-suited to decentralized periodic-connected tasks as it allows the energy-starved or bandwidth-constraint tasks to postpone the update. Theoretical results demonstrate the convergence guarantee of our distributed algorithm with an optimal regret. Empirical results confirm that the proposed model is highly effective on various real-world datasets.
Peng Yang 0013, Ping Li 0001
AAAI1
2020 Adaptive Online Kernel Sampling for Vertex Classification
abstract
This paper studies online kernel learning (OKL) for graph classification problem, since the large approximation space provided by reproducing kernel Hilbert spaces often contains an accurate function. Nonetheless, optimizing over this space is computationally expensive. To address this issue, approximate OKL is introduced to reduce the complexity either by limiting the support vector (SV) used by the predictor, or by avoiding the kernelization process altogether using embedding. Nonetheless, as long as the size of the approximation space or the number of SV does not grow over time, an adversarial environment can always exploit the approximation process. In this paper, we introduce an online kernel sampling (OKS) technique, a new second-order OKL method that slightly improve the bound from $O(d \log(T))$ down to $O(r \log(T))$ where $r$ is the rank of the learned data and is usually much smaller than d. To reduce the computational complexity of second-order methods, we introduce a randomized sampling algorithm for sketching kernel matrix $K_t$ and show that our method is effective to reduce the time and space complexity significantly while maintaining comparable performance. Empirical experimental results demonstrate that the proposed model is highly effective on real-world graph datasets.
Peng Yang 0013, Ping Li 0001
AISTATS1
2020 Efficient Online Multi-Task Learning via Adaptive Kernel Selection
abstract
Conventional multi-task model restricts the task structure to be linearly related, which may not be suitable when data is linearly nonseparable. To remedy this issue, we propose a kernel algorithm for online multi-task classification, as the large approximation space provided by reproducing kernel Hilbert spaces often contains an accurate function. Specifically, it maintains a local-global Gaussian distribution over each task model that guides the direction and scale of parameter updates. Nonetheless, optimizing over this space is computationally expensive. Moreover, most multi-task learning methods require accessing to the entire training instances, which is luxury unavailable in the large-scale streaming learning scenario. To overcome this issue, we propose a randomized kernel sampling technique across multiple tasks. Instead of requiring all inputs’ labels, the proposed algorithm determines whether to query a label or not via considering the confidence from the related tasks over label prediction. Theoretically, the algorithm trained on actively sampled labels can achieve a comparable result with one learned on all labels. Empirically, the proposed algorithm is able to achieve promising learning efficacy, while reducing the computational complexity and labeling cost simultaneously.
Peng Yang 0013, Ping Li 0001
WWW1