William Aiken

dblp:157/0286 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0003-2421-994XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Lost in the Noise: Evading and Detecting Backdoors in Conditional Diffusion Models
William Aiken, Paula Branco, Guy-Vincent Jourdan
IEA/AIE (1)1
2024 DevilDiffusion: Embedding Hidden Noise Backdoors into Diffusion Models
abstract
Diffusion models represent state-of-the-art deep learning architectures behind many popular and powerful image-synthesizing generative Artificial Intelligence (AI) systems. Their underlying approach, which relies on scheduled noise addition and optimized noise removal, has been applicable to the syn-thesis of sophisticated and nearly photorealistic images across various domains; however, the potential impacts of compro-mised diffusion models remain an underexplored area in the literature. While prior studies have investigated the insertion of explicit triggers into the noise or prompt spaces of diffusion models, it is unrealistic to assume that benign users would intentionally input such triggers into their own diffusion process. In our DevilDiffusion approach, we demonstrate the capability to surreptitiously embed triggers that leverage uncommon but naturally -occurring characteristics of Gaussian noise directly into the noise space of conditional diffusion models. When these specific characteristics occur in naturally -occurring noise, the model will instead construct our target backdoor image. By adjusting the trigger size and the ratio of poisoned images, we can control the trigger rates of the specified target image, ranging from less than 0.01 % to 25 % of all generated images, while still maintaining performance on the target task.
William Aiken, Paula Branco, Guy-Vincent Jourdan
PST1
2023 Going Haywire: False Friends in Federated Learning and How to Find Them
abstract
Federated Learning (FL) promises to offer a major paradigm shift in the way deep learning models are trained at scale, yet malicious clients can surreptitiously embed backdoors into models via trivial augmentation on their own subset of the data. This is especially true in small- and medium-scale FL systems, which consist of dozens, rather than millions, of clients. In this work, we investigate a novel attack scenario for an FL architecture consisting of multiple non-i.i.d. silos of data in which each distribution has a unique backdoor attacker and where the model convergences of adversaries are not more similar than those of benign clients. We propose a new method, dubbed Haywire, as a security-in-depth approach to respond to this novel attack scenario. Our defense utilizes a combination of kPCA dimensionality reduction of fully-connected layers in the network, KMeans anomaly detection to drop anomalous clients, and server aggregation robust to outliers via the Geometric Median. Our solution prevents the contamination of the global model despite having no access to the backdoor triggers. We evaluate the performance of Haywire from model-accuracy, defense-performance, and attack-success perspectives against multiple baselines. Through an extensive set of experiments, we find that Haywire produces the best performances at preventing backdoor attacks while simultaneously not unfairly penalizing benign clients. We carried out additional in-depth experiments across multiple runs that demonstrate the reliability of Haywire.
William Aiken, Paula Branco, Guy-Vincent Jourdan
AsiaCCS1
2023 Measuring Improvement of F1-Scores in Detection of Self-Admitted Technical Debt
abstract
Artificial Intelligence and Machine Learning have witnessed rapid, significant improvements in Natural Language Processing (NLP) tasks. Utilizing Deep Learning, researchers have taken advantage of repository comments in Software Engineering to produce accurate methods for detecting Self-Admitted Technical Debt (SATD) from 20 open-source Java projects’ code. In this work, we improve SATD detection with a novel approach that leverages the Bidirectional Encoder Representations from Transformers (BERT) architecture. For comparison, we re-evaluated previous deep learning methods and applied stratified 10-fold cross-validation to report reliable F1-scores. We examine our model in both cross-project and intra-project contexts. For each context, we use re-sampling and duplication as augmentation strategies to account for data imbalance. We find that our trained BERT model improves over the best performance of all previous methods in 19 of the 20 projects in cross-project scenarios. However, the data augmentation techniques were not sufficient to overcome the lack of data present in the intra-project scenarios, and existing methods still perform better. Future research will look into ways to diversify SATD datasets in order to maximize the latent power in large BERT models.
William Aiken, Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Mehrdad Sabetzadeh, Herna L. Viktor
TechDebt@ICSE1
2021 Neural network laundering: Removing black-box backdoor watermarks from deep neural networks
William Aiken, Hyoungshick Kim, Simon S. Woo, Jungwoo Ryoo
Comput. Secur.1
2019 The Light Will Be with You. Always - A Novel Continuous Mobile Authentication with the Light Sensor
abstract
Existing continuous authentication proposals tend to have two major drawbacks. First, touch-based smartphone authentication approaches typically require explicit user interactions with the smartphone to collect sufficient touch data. These approaches may provide an attacker the opportunity to steal a victim's sensitive data before the system detects the attacker's intrusion. Likewise, an attacker may disable the continuous authentication scheme itself before detection. Second, sensor-based continuous authentication approaches inherently suffer from high energy consumption due to the constant usage of multiple sensors. In this paper, we present a novel continuous authentication system that collects light sensor data from a user's smartphone and analyzes them to authenticate users using support vector machines. We focus on the possibility of collecting light sensor data from users' smartphones while they are conducting daily behaviors to develop an anomaly detection system.
Mohsen Ali Alawami, William Aiken, Hyoungshick Kim
MobiSys2
2018 POSTER: DeepCRACk: Using Deep Learning to Automatically CRack Audio CAPTCHAs
abstract
A Completely Automated Public Turing test to tell Computers and Humans Apart (CAPTCHA) is a defensive mechanism designed to differentiate humans and computers to prevent unauthorized use of online services by automated attacks. They often consist of a visual or audio test that humans can perform easily but that bots cannot solve. However, with current machine learning techniques and open-source neural network architectures, it is now possible to create a self-contained system that is able to solve specific CAPTCHA types and outperform some human users. In this paper, we present a neural network that leverages Mozilla's open source implementation of Baidu's Deep Speech architecture; our model is currently able to solve the audio version of an open-source CATPCHA system (named SimpleCaptcha) with 98.8% accuracy. Our network was trained on 100,000 audio samples generated from SimpleCaptcha and can solve new SimpleCaptcha audio tests in 1.25 seconds on average (with a standard deviation of 0.065 seconds). Our implementation seems additionally promising because it does not require a powerful server to function and is robust to adversarial examples that target Deep Speech's pre-trained models.
William Aiken, Hyoungshick Kim
AsiaCCS1
2018 POSTER: I Can't Hear This Because I Am Human: A Novel Design of Audio CAPTCHA System
abstract
A CAPTCHA (Completely Automated Public Turing test to tell Computers and Humans Apart) provides the first line of defense to protect websites against bots and automatic crawling. Recently, audio-based CAPTCHA systems are started to use for visually impaired people in many internet services. However, with the recent improvement of speech recognition and machine learning system, audio CAPTCHAs have come to struggle to distinguish machines from users, and this situation will likely continue to worsen. Unlike conventional CAPTCHA systems, we propose a new conceptual audio CAPTCHA system, combining certain sound, which is only understandable by a machine. Our experiment results demonstrate that the tested speech recognition systems always provide correct responses for our CAPTCHA samples while humans cannot possibly understand them. Based on this computational gap between the human and machine, we can detect bots with their correct responses, rather than their incorrect ones.
Jusop Choi, Taekkyung Oh, William Aiken, Simon S. Woo, Hyoungshick Kim
AsiaCCS3
2018 An Implementation and Evaluation of Progressive Authentication Using Multiple Level Pattern Locks
abstract
This paper presents a possible implementation of progressive authentication using the Android pattern lock. Our key idea is to use one pattern for two access levels to the device; an abridged pattern is used to access generic applications and a second, extended and higher-complexity pattern is used less frequently to access more sensitive applications. We conducted a user study of 89 participants and a consecutive user survey on those participants to investigate the usability of such a pattern scheme. Data from our prototype showed that for unlocking lowsecurity applications the median unlock times for users of the multiple pattern scheme and conventional pattern scheme were 2824 ms and 5589 ms respectively, and the distributions in the two groups differed significantly (Mann-Whitney U test, p-value less than 0.05, two-tailed). From our user survey, we did not find statistically significant differences between the two groups for their qualitative responses regarding usability and security (t-test, p-value greater than 0.05, two-tailed), but the groups did not differ by more than one satisfaction rating at 90% confidence.
William Aiken, Hyoungshick Kim, Jungwoo Ryoo, Mary Beth Rosson
PST1
2018 A security evaluation framework for cloud security auditing
Syed Rizvi 0001, Jungwoo Ryoo, John Kissell, William Aiken, Yuhong Liu 0003
J. Supercomput.4