Shangyu Chen

dblp:189/9141 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
abstract
Recommendation is crucial for both user experience and company revenue in Meituan as a leading lifestyle company, and generative recommendation models (GRMs) are shown to produce quality recommendations recently. However, existing systems are limited by insufficient functionality support and inefficient implementations for training GRMs in industrial scenarios. As such, we introduce MTGenRec as an efficient and scalable system for GRM training. Specifically, to handle real-time insertions/deletions of sparse embeddings, MTGenRec employs dynamic hash tables to replace static ones. To improve training efficiency, MTGenRec conducts dynamic sequence balancing to address the computation load imbalances among GPUs and adopts feature ID deduplication alongside automatic table merging to accelerate embedding lookup. Extensive experiments show that MTGenRec improves training throughput by 1.6x - 2.4x while achieving good scalability when running over 100 GPUs. MTGenRec has been deployed for many applications in Meituan and is now handling hundreds of millions of requests on a daily basis. On the delivery platform, we observe a 1.22% growth in user order volume and a 1.31% enhancement in online PV_CTR.
Yuxiang Wang 0013, Xiao Yan 0002, Mincong Huang, Ruidong Han, Bin Yin 0004, Shangyu Chen, Xiang Li 0067, Fei Jiang 0009, Wei Lin 0022, Haowei Han, Xiaokai Zhou, Bo Du 0001, Jiawei Jiang 0001
KDD (1)8
2025 MTGR: Industrial-Scale Generative Recommendation Framework in Meituan
abstract
Scaling law has recently been validated in the recommendation system, adopting generative recommendation strategies to achieve scalability. However, these generative approaches require abandoning the meticulously constructed cross features of traditional recommendation models,leading to a significant decline in model performance. To address this challenge, we propose Meituan Generative Recommendation, which is based on the HSTU architecture and is capable of retaining the original deep learning recommendation model (DLRM) features, including cross features. Additionally, MTGR achieves training and inference acceleration through user-level compression to ensure efficient scaling. We also propose Group-Layer Normalization (GLN) to enhance the performance of encoding within different semantic spaces and the dynamic masking strategy to avoid information leakage. We further optimize the training frameworks, enabling support for our models with 10 to 100 times computational complexity compared to the DLRM, without significant cost increases. MTGR achieved 65x FLOPs for single-sample forward inference compared to the DLRM model, resulting in the largest gain in nearly two years both offline and online. This breakthrough was successfully deployed on Meituan, the world's largest food delivery platform, where it has been handling the main traffic.
Ruidong Han, Bin Yin 0004, Shangyu Chen, Fei Jiang 0009, Xiang Li 0067, Mincong Huang, Chunzhen Jing, Yueming Han, MengLei Zhou, Wei Lin 0022
CIKM3
2025 PaRa: Personalizing Text-to-Image Diffusion via Parameter Rank Reduction
abstract
Personalizing a large-scale pretrained Text-to-Image (T2I) diffusion model is chal- lenging as it typically struggles to make an appropriate trade-off between its training data distribution and the target distribution, i.e., learning a novel concept with only a few target images to achieve personalization (aligning with the personalized target) while preserving text editability (aligning with diverse text prompts). In this paper, we propose PaRa, an effective and efficient Parameter Rank Reduction approach for T2I model personalization by explicitly controlling the rank of the diffusion model parameters to restrict its initial diverse generation space into a small and well-balanced target space. Our design is motivated by the fact that taming a T2I model toward a novel concept such as a specific art style implies a small generation space. To this end, by reducing the rank of model parameters during finetuning, we can effectively constrain the space of the denoising sampling trajectories towards the target. With comprehensive experiments, we show that PaRa achieves great advantages over existing finetuning approaches on single/multi-subject generation as well as single-image editing. Notably, compared to the prevailing fine-tuning technique LoRA, PaRa achieves better parameter efficiency (2× fewer learnable parameters) and much better target image alignment.
Shangyu Chen, Zizheng Pan, Jianfei Cai 0001, Dinh Q. Phung
ICLR1
2025 HVQ-VAE: Variational auto-encoder with hyperbolic vector quantization
Shangyu Chen, Pengfei Fang, Mehrtash Harandi, Trung Le 0001, Jianfei Cai 0001, Dinh Q. Phung
Comput. Vis. Image Underst.1
2024 Stereographic Projection for Embedding Hierarchical Structures in Hyperbolic Space
Shangyu Chen, Xiaohao Yang, Pengfei Fang, Mehrtash Harandi, Dinh Q. Phung, Jianfei Cai 0001
ICPR (9)1
2024 Neural Topic Model with Distance Awareness
Shangyu Chen, He Zhao 0001, Viet H. Huynh, Dinh Q. Phung, Jianfei Cai 0001
ICPR (9)1
2024 Generating Semantic Adversarial Examples via Feature Manipulation in Latent Space
abstract
The susceptibility of deep neural networks (DNNs) to adversarial intrusions, exemplified by adversarial examples, is well-documented. Conventional attacks implement unstructured, pixel-wise perturbations to mislead classifiers, which often results in a noticeable departure from natural samples and lacks human-perceptible interpretability. In this work, we present an adversarial attack strategy that implements fine-granularity, semantic-meaning-oriented structural perturbations. Our proposed methodology manipulates the semantic attributes of images through the use of disentangled latent codes. We engineer adversarial perturbations by manipulating either a single latent code or a combination thereof. To this end, we propose two unsupervised semantic manipulation strategies: one based on vector-disentangled representation and the other on feature map-disentangled representation, taking into consideration the complexity of the latent codes and the smoothness of the reconstructed images. Our empirical evaluations, conducted extensively on real-world image data, showcase the potency of our attacks, particularly against black-box classifiers. Furthermore, we establish the existence of a universal semantic adversarial example that is agnostic to specific images.
Shuo Wang 0012, Shangyu Chen, Surya Nepal, Carsten Rudolph, Marthie Grobler
IEEE Trans. Neural Networks Learn. Syst.2
2022 Backdoor Attacks Against Transfer Learning With Pre-Trained Deep Learning Models
abstract
Transfer learning provides an effective solution for feasibly and fast customize accurateStudentmodels, by transferring the learned knowledge of pre-trainedTeachermodels over large datasets via fine-tuning. Many pre-trained Teacher models used in transfer learning are publicly available and maintained by public platforms, increasing their vulnerability to backdoor attacks. In this article, we demonstrate a backdoor threat to transfer learning tasks on both image and time-series data leveraging the knowledge of publicly accessible Teacher models, aimed at defeating three commonly adopted defenses:pruning-based,retraining-basedandinput pre-processing-based defenses. Specifically, ($\mathcal {A}$A) ranking-based selection mechanism to speed up the backdoor trigger generation and perturbation process while defeatingpruning-basedand/orretraining-based defenses. ($\mathcal {B}$B) autoencoder-powered trigger generation is proposed to produce a robust trigger that can defeat theinput pre-processing-based defense, while guaranteeing that selected neuron(s) can be significantly activated. ($\mathcal {C}$C) defense-aware retraining to generate the manipulated model using reverse-engineered model inputs. We launch effective misclassification attacks on Student models over real-world images, brain Magnetic Resonance Imaging (MRI) data and Electrocardiography (ECG) learning systems. The experiments reveal that our enhanced attack can maintain the 98.4 and 97.2 percent classification accuracy as the genuine model on clean image and time series inputs while improving$27.9\%-100\%$27.9%-100%and$27.1\%-56.1\%$27.1%-56.1%attack success rate on trojaned image and time series inputs respectively in the presence of pruning-based and/or retraining-based defenses.
Shuo Wang 0012, Surya Nepal, Carsten Rudolph, Marthie Grobler, Shangyu Chen
IEEE Trans. Serv. Comput.5
2022 Defending Adversarial Attacks via Semantic Feature Manipulation
abstract
Machine learning models have demonstrated vulnerability to adversarial attacks, more specifically misclassification of adversarial examples. In this article, we propose a one-off and attack-agnostic Feature Manipulation (FM)-Defense to detect and purify adversarial examples in an interpretable and efficient manner. The intuition is that the classification result of a normal image is generally resistant to non-significant intrinsic feature changes, e.g., varying the thickness of handwritten digits. In contrast, adversarial examples are sensitive to such changes since the perturbation lacks transferability. To enable manipulation of features, a Combo-variational autoencoder is applied to learn disentangled latent codes that reveal semantic features. The resistance to classification change over the morphs, derived by varying and reconstructing latent codes, is used to detect suspicious inputs. Furthermore, Combo-VAE is enhanced to purify the adversarial examples with good quality by considering class-shared and class-unique features. We empirically demonstrate the effectiveness of detection and quality of purified instances. Our experiments on three datasets show that FM-Defense can detect nearly 100 percent of adversarial examples produced by different state-of-the-art adversarial attacks. It achieves more than 99 percent overall purification accuracy on the suspicious instances that close the manifold of clean examples.
Shuo Wang 0012, Surya Nepal, Carsten Rudolph, Marthie Grobler, Shangyu Chen, Zike An
IEEE Trans. Serv. Comput.5
2020 PART-GAN: Privacy-Preserving Time-Series Sharing
Shuo Wang 0012, Carsten Rudolph, Surya Nepal, Marthie Grobler, Shangyu Chen
ICANN (1)5
2020 OIAD: One-for-all Image Anomaly Detection with Disentanglement Learning
abstract
Anomaly detection aims to recognize samples with anomalous and unusual patterns with respect to a set of normal data. This is significant for numerous domain applications, such as industrial inspection, medical imaging, and security enforcement. There are two key research challenges associated with existing anomaly detection approaches: (1) many approaches perform well on low-dimensional problems however the performance on high-dimensional instances, such as images, is limited; (2) many approaches often rely on traditional supervised approaches and manual engineering of features, while the topic has not been fully explored yet using modern deep learning approaches, even when the well-label samples are limited. In this paper, we propose a One-for-all Image Anomaly Detection system (OIAD) based on disentangled learning using only clean samples. Our key insight is that the impact of small perturbation on the latent representation can be bounded for normal samples while anomaly images are usually outside such bounded intervals, referred to as structure consistency. We implement this idea and evaluate its performance for anomaly detection. Our experiments with three datasets show that OIAD can detect over 90% of anomalies while maintaining a low false alarm rate. It can also detect suspicious samples from samples labeled as clean, coincided with what humans would deem unusual.
Shuo Wang 0012, Shangyu Chen, Surya Nepal, Carsten Rudolph, Marthie Grobler
IJCNN3
2020 Storage Efficient and Dynamic Flexible Runtime Channel Pruning via Deep Reinforcement Learning
abstract
In this paper, we propose a deep reinforcement learning (DRL) based framework to efficiently perform runtime channel pruning on convolutional neural networks (CNNs). Our DRL-based framework aims to learn a pruning strategy to determine how many and which channels to be pruned in each convolutional layer, depending on each individual input instance at runtime. Unlike existing runtime pruning methods which require to store all channels parameters for inference, our framework can reduce parameters storage consumption by introducing a static pruning component. Comparison experimental results with existing runtime and static pruning methods on state-of-the-art CNNs demonstrate that our proposed framework is able to provide a tradeoff between dynamic flexibility and storage efficiency in runtime channel pruning.
Jianda Chen, Shangyu Chen, Sinno Jialin Pan
NeurIPS2
2020 Privacy-Preserving Data Generation and Sharing Using Identification Sanitizer
Shuo Wang 0012, Lingjuan Lyu, Shangyu Chen, Surya Nepal, Carsten Rudolph, Marthie Grobler
WISE (2)4
2019 Deep Neural Network Quantization via Layer-Wise Optimization Using Limited Training Data
abstract
The advancement of deep models poses great challenges to real-world deployment because of the limited computational ability and storage space on edge devices. To solve this problem, existing works have made progress to prune or quantize deep models. However, most existing methods rely heavily on a supervised training process to achieve satisfactory performance, acquiring large amount of labeled training data, which may not be practical for real deployment. In this paper, we propose a novel layer-wise quantization method for deep neural networks, which only requires limited training data (1% of original dataset). Specifically, we formulate parameters quantization for each layer as a discrete optimization problem, and solve it using Alternative Direction Method of Multipliers (ADMM), which gives an efficient closed-form solution. We prove that the final performance drop after quantization is bounded by a linear combination of the reconstructed errors caused at each layer. Based on the proved theorem, we propose an algorithm to quantize a deep neural network layer by layer with an additional weights update step to minimize the final error. Extensive experiments on benchmark deep models are conducted to demonstrate the effectiveness of our proposed method using 1% of CIFAR10 and ImageNet datasets. Codes are available in: https://github.com/csyhhu/L-DNQ
Shangyu Chen, Wenya Wang 0001, Sinno Jialin Pan
AAAI1
2019 Parametric Canonical Correlation Analysis
abstract
Generally, suppose a wave is a linear combination of multiple basis(Not necessarily a sine or cosine waves, it could also be a wavelet, etc.), different types of waves may be similar on some basis, but vary greatly on a certain basis. To address this problem, we introduce a PCCA-based feature extraction method that extends canonical correlation analysis (CCA). The PCCA-based method can train efficient classifiers to rely on only a few samples for periodic signals with support for removing noisy signals. As a demonstration, an efficient system is implemented for the classification of electrocardiogram (ECG) signals by PCCA. The performance is measured using several normal and abnormal ECG signals from the real-world database. These are compared with three commonly-adopted feature extraction techniques using five classes classification tasks related to ECG heartbeats. The AUC(Area under the ROC curve) of the PCCA-based feature extraction technique with two-digits size train dataset for four ECG type-pairs we compared were 0.8805, 0.957, 0.8968 and 1.00 respectively. The experimental results demonstrate that the proposed feature extraction techniques achieve better performance compared to other features extraction techniques with small amount of well-labeled data.
Shangyu Chen, Shuo Wang 0012, Richard O. Sinnott
CloudCom1
2019 Cooperative Pruning in Cross-Domain Deep Neural Network Compression
abstract
The advancement of deep models poses great challenges to real-world deployment because of the limited computational ability and storage space on edge devices. To solve this problem, existing works have made progress to compress deep models by pruning or quantization. However, most existing methods rely on a large amount of training data and a pre-trained model in the same domain. When only limited in-domain training data is available, these methods fail to perform well. This prompts the idea of transferring knowledge from a resource-rich source domain to a target domain with limited data to perform model compression. In this paper, we propose a method to perform cross-domain pruning by cooperatively training in both domains: taking advantage of data and a pre-trained model from the source domain to assist pruning in the target domain. Specifically, source and target pruned models are trained simultaneously and interactively, with source information transferred through the construction of a cooperative pruning mask. Our method significantly improves pruning quality in the target domain, and shed light to model compression in the cross-domain setting.
Shangyu Chen, Wenya Wang 0001, Sinno Jialin Pan
IJCAI1
2019 MetaQuant: Learning to Quantize by Learning to Penetrate Non-differentiable Quantization
abstract
Tremendous amount of parameters make deep neural networks impractical to be deployed for edge-device-based real-world applications due to the limit of computational power and storage space. Existing studies have made progress on learning quantized deep models to reduce model size and energy consumption, i.e. converting full-precision weights ($r$'s) into discrete values ($q$'s) in a supervised training manner. However, the training process for quantization is non-differentiable, which leads to either infinite or zero gradients ($g_r$) w.r.t. $r$. To address this problem, most training-based quantization methods use the gradient w.r.t. $q$ ($g_q$) with clipping to approximate $g_r$ by Straight-Through-Estimator (STE) or manually design their computation. However, these methods only heuristically make training-based quantization applicable, without further analysis on how the approximated gradients can assist training of a quantized network. In this paper, we propose to learn $g_r$ by a neural network. Specifically, a meta network is trained using $g_q$ and $r$ as inputs, and outputs $g_r$ for subsequent weight updates. The meta network is updated together with the original quantized network. Our proposed method alleviates the problem of non-differentiability, and can be trained in an end-to-end manner. Extensive experiments are conducted with CIFAR10/100 and ImageNet on various deep networks to demonstrate the advantage of our proposed method in terms of a faster convergence rate and better performance. Codes are released at: \texttt{https://github.com/csyhhu/MetaQuant}
Shangyu Chen, Wenya Wang 0001, Sinno Jialin Pan
NeurIPS1
2017 Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon
abstract
How to develop slim and accurate deep neural networks has become crucial for real- world applications, especially for those employed in embedded systems. Though previous work along this research line has shown some promising results, most existing methods either fail to significantly compress a well-trained deep network or require a heavy retraining process for the pruned deep network to re-boost its prediction performance. In this paper, we propose a new layer-wise pruning method for deep neural networks. In our proposed method, parameters of each individual layer are pruned independently based on second order derivatives of a layer-wise error function with respect to the corresponding parameters. We prove that the final prediction performance drop after pruning is bounded by a linear combination of the reconstructed errors caused at each layer. By controlling layer-wise errors properly, one only needs to perform a light retraining process on the pruned network to resume its original prediction performance. We conduct extensive experiments on benchmark datasets to demonstrate the effectiveness of our pruning method compared with several state-of-the-art baseline methods. Codes of our work are released at: https://github.com/csyhhu/L-OBS.
Xin Dong 0009, Shangyu Chen, Sinno Jialin Pan
NIPS2