Shang-Tse Chen

dblp:24/9381 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-3441-3471ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorSecurity and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 51% Generative modeling · 12% Efficient and distributed learning · 10%
Network and information security
2 papers
Security and privacy of machine learning · 85% Privacy and data protection · 15%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 28 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
certified robustness
1.622025
Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss · ICML 2025
Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization · AAAI 2024
Machine learning › Trustworthy machine learning
robustness
1.622025
Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss · ICML 2025
Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization · AAAI 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.022024
Annealing Self-Distillation Rectification Improves Adversarial Training · ICLR 2024
Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization · AAAI 2024
Machine learning › Generative modeling
diffusion model
0.912025
DRAG: Data Reconstruction Attack using Guided Diffusion · ICML 2025
Machine learning › Generative modeling › diffusion model
guided diffusion
0.912025
DRAG: Data Reconstruction Attack using Guided Diffusion · ICML 2025
Machine learning › Trustworthy machine learning › robustness › certified robustness
lipschitz-constrained networks
0.912025
Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss · ICML 2025
Security and privacy of machine learning › privacy attack
data reconstruction attack
0.912025
DRAG: Data Reconstruction Attack using Guided Diffusion · ICML 2025
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.812024
Annealing Self-Distillation Rectification Improves Adversarial Training · ICLR 2024
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.812024
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition · EMNLP 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
Annealing Self-Distillation Rectification Improves Adversarial Training · ICLR 2024
Machine learning › Deep learning architectures and training › regularization
label smoothing
0.812024
Annealing Self-Distillation Rectification Improves Adversarial Training · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › certified robustness
randomized smoothing
0.812024
Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization · AAAI 2024
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
robust overfitting
0.812024
Annealing Self-Distillation Rectification Improves Adversarial Training · ICLR 2024
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation
0.812024
Annealing Self-Distillation Rectification Improves Adversarial Training · ICLR 2024
Machine learning › Transfer learning and domain adaptation › domain shift
synthetic-to-real domain gap
0.812024
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition · EMNLP 2024
Mathematical optimization › continuous optimization
quasi-concave optimization
0.812024
Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization · AAAI 2024
Machine learning › Learning theory
online learning
0.322014
Boosting with Online Binary Learners for the Multiclass Bandit Problem · ICML 2014
An Online Boosting Algorithm with Theoretical Justifications · ICML 2012
Security and privacy of machine learning
adversarial defense
0.312018
SHIELD: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression · KDD 2018
Security and privacy of machine learning
adversarial machine learning
0.312018
SHIELD: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression · KDD 2018
Machine learning › Deep learning architectures and training
loss function design
0.312025
Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss · ICML 2025
Smart cities and intelligent transportation › disaster management
fire risk prediction
0.212016
Firebird: Predicting Fire Risk and Prioritizing Fire Inspections in Atlanta · KDD 2016
Smart cities and intelligent transportation › urban computing
urban analytics
0.212016
Firebird: Predicting Fire Risk and Prioritizing Fire Inspections in Atlanta · KDD 2016
Computer vision › Image recognition and object detection
image classification
0.212024
Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization · AAAI 2024
Machine learning › Learning theory › online learning › partial feedback
bandit multiclass learning
0.212014
Boosting with Online Binary Learners for the Multiclass Bandit Problem · ICML 2014
Machine learning › Reinforcement learning
multi-armed bandit
0.212014
Boosting with Online Binary Learners for the Multiclass Bandit Problem · ICML 2014
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.112012
An Online Boosting Algorithm with Theoretical Justifications · ICML 2012
Machine learning › Kernel, tree and ensemble methods › ensemble learning › boosting
online boosting
0.112012
An Online Boosting Algorithm with Theoretical Justifications · ICML 2012
Image and video coding
JPEG compression
0.112018
SHIELD: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression · KDD 2018

Methods — techniques the papers use, named apart from their topics

latent diffusion model · 1.7iterative reconstruction · 1.7randomized smoothing · 1.5logit annealing loss · 0.9block reflector orthogonal layer · 0.9task vectors · 0.8task arithmetic · 0.8quasiconvex optimization · 0.8quasi-convex optimization · 0.8fine-tuning · 0.8annealing · 0.8retraining · 0.7randomization · 0.7ensemble defense · 0.7machine learning · 0.5information visualization · 0.5geocoding · 0.5
YearPublicationVenuePosition
2025 Enhancing Certified Robustness via Block Reflector Orthogonal Layers and Logit Annealing Loss
abstract
Lipschitz neural networks are well-known for providing certified robustness in deep learning. In this paper, we present a novel, efficient Block Reflector Orthogonal (BRO) layer that enhances the capability of orthogonal layers on constructing more expressive Lipschitz neural architectures. In addition, by theoretically analyzing the nature of Lipschitz neural networks, we introduce a new loss function that employs an annealing mechanism to increase margin for most data points. This enables Lipschitz models to provide better certified robustness. By employing our BRO layer and loss function, we design BRONet — a simple yet effective Lipschitz neural network that achieves state-of-the-art certified robustness. Extensive experiments and empirical analysis on CIFAR-10/100, Tiny-ImageNet, and ImageNet validate that our method outperforms existing baselines. The implementation is available at GitHub Link.
Bo-Han Lai, Pin-Han Huang, Bo-Han Kung, Shang-Tse Chen
ICML4
2025 DRAG: Data Reconstruction Attack using Guided Diffusion
abstract
With the rise of large foundation models, split inference (SI) has emerged as a popular computational paradigm for deploying models across lightweight edge devices and cloud servers, addressing data privacy and computational cost concerns. However, most existing data reconstruction attacks have focused on smaller CNN classification models, leaving the privacy risks of foundation models in SI settings largely unexplored. To address this gap, we propose a novel data reconstruction attack based on guided diffusion, which leverages the rich prior knowledge embedded in a latent diffusion model (LDM) pre-trained on a large-scale dataset. Our method performs iterative reconstruction on the LDM’s learned image prior, effectively generating high-fidelity images resembling the original data from their intermediate representations (IR). Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, both qualitatively and quantitatively, in reconstructing data from deep-layer IRs of the vision foundation model. The results highlight the urgent need for more robust privacy protection mechanisms for large models in SI scenarios.
Wa-Kin Lei, Jun-Cheng Chen, Shang-Tse Chen
ICML3
2024 Towards Large Certified Radius in Randomized Smoothing Using Quasiconcave Optimization
abstract
Randomized smoothing is currently the state-of-the-art method that provides certified robustness for deep neural networks. However, due to its excessively conservative nature, this method of incomplete verification often cannot achieve an adequate certified radius on real-world datasets. One way to obtain a larger certified radius is to use an input-specific algorithm instead of using a fixed Gaussian filter for all data points. Several methods based on this idea have been proposed, but they either suffer from high computational costs or gain marginal improvement in certified radius. In this work, we show that by exploiting the quasiconvex problem structure, we can find the optimal certified radii for most data points with slight computational overhead. This observation leads to an efficient and effective input-specific randomized smoothing algorithm. We conduct extensive experiments and empirical analysis on CIFAR-10 and ImageNet. The results show that the proposed method significantly enhances the certified radii with low computational overhead.
Bo-Han Kung, Shang-Tse Chen
AAAI2
2024 Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
abstract
Synthetic data is widely used in speech recognition due to the availability of text-to-speech models, which facilitate adapting models to previously unseen text domains.However, existing methods suffer in performance when they finetune an automatic speech recognition (ASR) model on synthetic data as they suffer from the distributional shift commonly referred to as the synthetic-to-real gap.In this paper, we find that task arithmetic is effective at mitigating this gap.Our proposed method, SYN2REAL task vector, shows an average improvement of 10.03% improvement in word error rate over baselines on the SLURP dataset.Additionally, we show that an average of SYN2REAL task vectors, when we have real speeches from multiple different domains, can further adapt the original ASR model to perform better on the target text domain.
Hsuan Su, Hua Farn, Fan-Yun Sun, Shang-Tse Chen, Hung-yi Lee
EMNLP4
2024 Annealing Self-Distillation Rectification Improves Adversarial Training
abstract
In standard adversarial training, models are optimized to fit invariant one-hot labels for adversarial data when the perturbations are within allowable budgets. However, the overconfident target harms generalization and causes the problem of robust overfitting. To address this issue and enhance adversarial robustness, we analyze the characteristics of robust models and identify that robust models tend to produce smoother and well-calibrated outputs. Based on the observation, we propose a simple yet effective method, Annealing Self-Distillation Rectification (ADR), which generates soft labels as a better guidance mechanism that reflects the underlying distribution of data. By utilizing ADR, we can obtain rectified labels that improve model robustness without the need for pre-trained models or extensive extra computation. Moreover, our method facilitates seamless plug-and-play integration with other adversarial training techniques by replacing the hard labels in their objectives. We demonstrate the efficacy of ADR through extensive experiments and strong performances across datasets.
Yu-Yu Wu, Hung-Jui Wang, Shang-Tse Chen
ICLR3
2020 UnMask: Adversarial Detection and Defense Through Robust Feature Alignment
abstract
Recent research has demonstrated that deep learning architectures are vulnerable to adversarial attacks, high-lighting the vital need for defensive techniques to detect and mitigate these attacks before they occur. We present UnMask, an adversarial detection and defense framework based on robust feature alignment. UnMask combats adversarial attacks by extracting robust features (e.g., beak, wings, eyes) from an image (e.g., "bird") and comparing them to the expected features of the classification. For example, if the extracted features for a "bird" image are wheel, saddle and frame, the model may be under attack. UnMask detects such attacks and defends the model by rectifying the misclassification, re-classifying the image based on its robust features. Our extensive evaluation shows that UnMask detects up to 96.75% of attacks, and defends the model by correctly classifying up to 93% of adversarial images produced by the current strongest attack, Projected Gradient Descent, in the gray-box setting. UnMask provides significantly better protection than adversarial training across 8 attack vectors, averaging 31.18% higher accuracy. We open source the code repository and data with this paper: https://github.com/safreita1/nmask.
Scott Freitas, Shang-Tse Chen, Zijie J. Wang, Polo Chau
IEEE BigData2
2018 SHIELD: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression
abstract
The rapidly growing body of research in adversarial machine learning has demonstrated that deep neural networks (DNNs) are highly vulnerable to adversarially generated images. This underscores the urgent need for practical defense techniques that can be readily deployed to combat attacks in real-time. Observing that many attack strategies aim to perturb image pixels in ways that are visually imperceptible, we place JPEG compression at the core of our proposed SHIELD defense framework, utilizing its capability to effectively "compress away" such pixel manipulation. To immunize a DNN model from artifacts introduced by compression, SHIELD "vaccinates" the model by retraining it with compressed images, where different compression levels are applied to generate multiple vaccinated models that are ultimately used together in an ensemble defense. On top of that, SHIELD adds an additional layer of protection by employing randomization at test time that compresses different regions of an image using random compression levels, making it harder for an adversary to estimate the transformation performed. This novel combination of vaccination, ensembling, and randomization makes SHIELD a fortified multi-pronged defense. We conducted extensive, large-scale experiments using the ImageNet dataset, and show that our approaches eliminate up to 98% of gray-box attacks delivered by strong adversarial techniques such as Carlini-Wagner's L2 attack and DeepFool. Our approaches are fast and work without requiring knowledge about the model.
Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen, Fred Hohman, Michael E. Kounavis, Polo Chau
KDD3
2018 ShapeShifter: Robust Physical Adversarial Attack on Faster R-CNN Object Detector
Shang-Tse Chen, Cory Cornelius, Polo Chau
ECML/PKDD (1)1
2018 ADAGIO: Interactive Experimentation with Adversarial Attack and Defense for Audio
Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen, Michael E. Kounavis, Polo Chau
ECML/PKDD (3)3
2018 Chronodes: Interactive Multifocus Exploration of Event Sequences
abstract
The advent of mobile health (mHealth) technologies challenges the capabilities of current visualizations, interactive tools, and algorithms. We present Chronodes, an interactive system that unifies data mining and human-centric visualization techniques to support explorative analysis of longitudinal mHealth data. Chronodes extracts and visualizes frequent event sequences that reveal chronological patterns across multiple participant timelines of mHealth data. It then combines novel interaction and visualization techniques to enable multifocus event sequence analysis, which allows health researchers to interactively define, explore, and compare groups of participant behaviors using event sequence combinations. Through summarizing insights gained from a pilot study with 20 behavioral and biomedical health experts, we discuss Chronodes's efficacy and potential impact in the mHealth domain. Ultimately, we outline important open challenges in mHealth, and offer recommendations and design guidelines for future research.
Peter J. Polack Jr., Shang-Tse Chen, Minsuk Kahng, Kaya de Barbaro, Rahul C. Basole, Moushumi Sharmin, Polo Chau
ACM Trans. Interact. Intell. Syst.2
2017 Predicting Cyber Threats with Virtual Security Products
abstract
Cybersecurity analysts are often presented suspicious machine activity that does not conclusively indicate compromise, resulting in undetected incidents or costly investigations into the most appropriate remediation actions. There are many reasons for this: deficiencies in the number and quality of security products that are deployed, poor configuration of those security products, and incomplete reporting of product-security telemetry. Managed Security Service Providers (MSSP's), which are tasked with detecting security incidents on behalf of multiple customers, are confronted with these data quality issues, but also possess a wealth of cross-product security data that enables innovative solutions. We use MSSP data to develop Virtual Product, which addresses the aforementioned data challenges by predicting what security events would have been triggered by a security product if it had been present. This benefits the analysts by providing more context into existing security incidents (albeit probabilistic) and by making questionable security incidents more conclusive. We achieve up to 99% AUC in predicting the incidents that some products would have detected had they been present.
Shang-Tse Chen, Yufei Han 0001, Polo Chau, Christopher Gates 0002, Michael Hart, Kevin A. Roundy
ACSAC1
2016 Communication Efficient Distributed Agnostic Boosting
abstract
We consider the problem of learning from distributed data in the agnostic setting, i.e., in the presence of arbitrary forms of noise. Our main contribution is a general distributed boosting-based procedure for learning an arbitrary concept space, that is simultaneously noise tolerant, communication efficient, and computationally efficient. This improves significantly over prior works that were either communication efficient only in noise-free scenarios or computationally prohibitive. Empirical results on large synthetic and real-world datasets demonstrate the effectiveness and scalability of the proposed approach.
Shang-Tse Chen, Maria-Florina Balcan, Polo Chau
AISTATS1
2016 Firebird: Predicting Fire Risk and Prioritizing Fire Inspections in Atlanta
abstract
The Atlanta Fire Rescue Department (AFRD), like many municipal fire departments, actively works to reduce fire risk by inspecting commercial properties for potential hazards and fire code violations. However, AFRD's fire inspection practices relied on tradition and intuition, with no existing data-driven process for prioritizing fire inspections or identifying new properties requiring inspection. In collaboration with AFRD, we developed the Firebird framework to help municipal fire departments identify and prioritize commercial property fire inspections, using machine learning, geocoding, and information visualization. Firebird computes fire risk scores for over 5,000 buildings in the city, with true positive rates of up to 71% in predicting fires. It has identified 6,096 new potential commercial properties to inspect, based on AFRD's criteria for inspection. Furthermore, through an interactive map, Firebird integrates and visualizes fire incidents, property information and risk scores to help AFRD make informed decisions about fire inspections. Firebird has already begun to make positive impact at both local and national levels. It is improving AFRD's inspection processes and Atlanta residents' safety, and was highlighted by National Fire Protection Association (NFPA) as a best practice for using data to inform fire inspections.
Michael A. Madaio, Shang-Tse Chen, Oliver L. Haimson, Xiang Cheng 0002, Matthew Hinds-Aldrich, Polo Chau, Bistra Dilkina
KDD2
2014 Boosting with Online Binary Learners for the Multiclass Bandit Problem
abstract
We consider the problem of online multiclass prediction in the bandit setting. Compared with the full-information setting, in which the learner can receive the true label as feedback after making each prediction, the bandit setting assumes that the learner can only know the correctness of the predicted label. Because the bandit setting is more restricted, it is difficult to design good bandit learners and currently there are not many bandit learners. In this paper, we propose an approach that systematically converts existing online binary classifiers to promising bandit learners with strong theoretical guarantee. The approach matches the idea of boosting, which has been shown to be powerful for batch learning as well as online learning. In particular, we establish the weak-learning condition on the online binary classifiers, and show that the condition allows automatically constructing a bandit learner with arbitrary strength by combining several of those classifiers. Experimental results on several real-world data sets demonstrate the effectiveness of the proposed approach.
Shang-Tse Chen, Hsuan-Tien Lin, Chi-Jen Lu
ICML1
2012 An Online Boosting Algorithm with Theoretical Justifications
Shang-Tse Chen, Hsuan-Tien Lin, Chi-Jen Lu
ICML1