Young D. Kwon

dblp:77/5405 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-5216-9057ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
abstract
State-of-the-art text-to-image diffusion models (DMs) achieve remarkable quality, yet their massive parameter scale (8-11B) poses significant challenges for inferences on resource-constrained devices. In this paper, we present HierarchicalPrune, a novel compression framework grounded in a key observation: DM blocks exhibit distinct functional hierarchies, where early blocks establish semantic structures while later blocks handle texture refinements. HierarchicalPrune synergistically combines three techniques: (1) Hierarchical Position Pruning, which identifies and removes less essential later blocks based on position hierarchy; (2) Positional Weight Preservation, which systematically protects early model portions that are essential for semantic structural integrity; and (3) Sensitivity-Guided Distillation, which adjusts knowledge-transfer intensity based on our discovery of block-wise sensitivity variations. As a result, our framework brings billion-scale diffusion models into a range more suitable for on-device inference, while preserving the quality of the output images. Specifically, combined with INT4 weight quantisation, HierarchicalPrune achieves 77.5-80.4% memory footprint reduction (e.g., from 15.8 GB to 3.2 GB) and 27.9-38.0% latency reduction, measured on server and consumer grade GPUs, with the minimum drop of 2.6% in GenEval score and 7% in HPSv2 score compared to the original model. Finally, our comprehensive user study with 85 participants demonstrates that HierarchicalPrune maintains perceptual quality comparable to the original model while significantly outperforming prior works.
Young D. Kwon, Rui Li 0052, Da Li 0001, Sourav Bhattacharya, Stylianos I. Venieris
AAAI1
2026 A cascade framework for on-device uncertainty-aware event detection on microcontrollers
abstract
Pervasive sensing enables diverse wearable event detection (WED) applications, but deploying machine learning models on resource-constrained microcontrollers (MCUs) poses significant challenges, particularly in ensuring prediction reliability under data shifts or out-of-distribution (OOD) inputs. While Uncertainty quantification methods offer a way to assess this reliability, many are computationally prohibitive for MCUs, and detecting multiple events concurrently further exacerbates resource constraints. Addressing these combined challenges, this paper presents an uncertainty and resource-aware framework designed for reliable and efficient multi-event WED on MCUs, significantly extending our preliminary work. The proposed framework achieves this by integrating Evidential Deep Learning (EDL) for efficient, single-pass uncertainty estimation with a novel cascade learning architecture. This architecture promotes resource efficiency via: (i) intra-event sharing using uncertainty-aware early exits within a staged model (shallow, medium, deep), allowing simpler samples to terminate inference earlier; and (ii) inter-event sharing using a multi-head design where multiple event detectors share a common backbone, minimizing overhead. System efficiency is further enhanced through MCU-specific optimizations, including targeted architecture search, quantization, efficient uncertainty operator implementation using standard TensorFlow Lite Micro (TFLM) operations, and library footprint reduction. We conducted extensive experiments on four distinct wearable datasets (Oesense, KWS, ECG5000, and HHAR) and two MCU platforms (STM32F446ZE, STM32H747XI), comparing the proposed framework against strong baselines including Deep Ensembles and Vanilla EDL. Results demonstrate the proposed framework’s effectiveness, achieving competitive accuracy and uncertainty performance (e.g., up to 22% lower NLL than data augmentation) while drastically reducing resource consumption, offering up to 8.64 × faster inference, up to 8.57 × lower energy use, and 55% smaller memory footprint compared to ensemble methods. The proposed framework enables the deployment of reliable, uncertainty-aware multi-event detection on a wider range of low-power MCUs.
Hong Jia, Young D. Kwon, Dong Ma 0001, Nhat Pham, Lorena Qendro, Tam Vu 0001, Cecilia Mascolo
Pervasive Mob. Comput.2
2024 TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce Edge
abstract
On-device training is essential for user personalisation and privacy. With the pervasiveness of IoT devices and microcontroller units (MCUs), this task becomes more challenging due to the constrained memory and compute resources, and the limited availability of labelled user data. Nonetheless, prior works neglect the data scarcity issue, require excessively long training time ($\textit{e.g.}$ a few hours), or induce substantial accuracy loss ($\geq$10%). In this paper, we propose TinyTrain, an on-device training approach that drastically reduces training time by selectively updating parts of the model and explicitly coping with data scarcity. TinyTrain introduces a task-adaptive sparse-update method that $\textit{dynamically}$ selects the layer/channel to update based on a multi-objective criterion that jointly captures user data, the memory, and the compute capabilities of the target device, leading to high accuracy on unseen tasks with reduced computation and memory footprint. TinyTrain outperforms vanilla fine-tuning of the entire network by 3.6-5.0% in accuracy, while reducing the backward-pass memory and computation cost by up to 1,098$\times$ and 7.68$\times$, respectively. Targeting broadly used real-world edge devices, TinyTrain achieves 9.5$\times$ faster and 3.5$\times$ more energy-efficient training over status-quo approaches, and 2.23$\times$ smaller memory footprint than SOTA methods, while remaining within the 1 MB memory envelope of MCU-grade platforms.
Young D. Kwon, Rui Li 0052, Stylianos I. Venieris, Jagmohan Chauhan, Nicholas D. Lane, Cecilia Mascolo
ICML1
2024 TinyTTA: Efficient Test-time Adaptation via Early-exit Ensembles on Edge Devices
abstract
The increased adoption of Internet of Things (IoT) devices has led to the generation of large data streams with applications in healthcare, sustainability, and robotics. In some cases, deep neural networks have been deployed directly on these resource-constrained units to limit communication overhead, increase efficiency and privacy, and enable real-time applications. However, a common challenge in this setting is the continuous adaptation of models necessary to accommodate changing environments, i.e., data distribution shifts. Test-time adaptation (TTA) has emerged as one potential solution, but its validity has yet to be explored in resource-constrained hardware settings, such as those involving microcontroller units (MCUs). TTA on constrained devices generally suffers from i) memory overhead due to the full backpropagation of a large pre-trained network, ii) lack of support for normalization layers on MCUs, and iii) either memory exhaustion with large batch sizes required for updating or poor performance with small batch sizes. In this paper, we propose TinyTTA, to enable, for the first time, efficient TTA on constrained devices with limited memory. To address the limited memory constraints, we introduce a novel self-ensemble and batch-agnostic early-exit strategy for TTA, which enables continuous adaptation with small batch sizes for reduced memory usage, handles distribution shifts, and improves latency efficiency. Moreover, we develop the TinyTTA Engine, a first-of-its-kind MCU library that enables on-device TTA. We validate TinyTTA on a Raspberry Pi Zero 2W and an STM32H747 MCU. Experimental results demonstrate that TinyTTA improves TTA accuracy by up to 57.6\%, reduces memory usage by up to six times, and achieves faster and more energy-efficient TTA. Notably, TinyTTA is the only framework able to run TTA on MCU STM32H747 with a 512 KB memory constraint while maintaining high performance.
Hong Jia, Young D. Kwon, Alessio Orsino, Ting Dang, Domenico Talia, Cecilia Mascolo
NeurIPS2
2024 UR2M: Uncertainty and Resource-Aware Event Detection on Microcontrollers
abstract
Traditional machine learning techniques are prone to generating inaccurate predictions when confronted with shifts in the distribution of data between the training and testing phases. This vulnerability can lead to severe consequences, especially in applications such as mobile healthcare. Uncertainty estimation has the potential to mitigate this issue by assessing the reliability of a model's output. However, existing uncertainty estimation techniques often require substantial computational resources and memory, making them impractical for implementation on microcontrollers (MCUs). This limitation hinders the feasibility of many important on-device wearable event detection (WED) applications, such as heart attack detection. In this paper, we present UR2M, a novel Uncertainty and Resource-aware event detection framework for MCUs. Specifically, we (i) develop an uncertainty-aware WED based on evidential theory for accurate event detection and reliable uncertainty estimation; (ii) introduce a cascade ML framework to achieve efficient model inference via early exits, by sharing shallower model layers among different event models; (iii) optimize the deployment of the model and MCU library for system efficiency. We conducted extensive experiments and compared UR2M to traditional uncertainty baselines using three wearable datasets. Our results demonstrate that UR2M achieves up to 864% faster inference speed, 857% energy-saving for uncertainty estimation, 55% memory saving on two popular MCUs, and a 22% improvement in uncertainty quantification performance. UR2M can be deployed on a wide range of MCUs, significantly expanding real-time and reliable WED applications.
Hong Jia, Young D. Kwon, Dong Ma 0001, Nhat Pham, Lorena Qendro, Tam Vu 0001, Cecilia Mascolo
PerCom2
2023 LifeLearner: Hardware-Aware Meta Continual Learning System for Embedded Computing Platforms
abstract
Continual Learning (CL) allows applications such as user personalization and household robots to learn on the fly and adapt to context. This is an important feature when context, actions, and users change. However, enabling CL on resource-constrained embedded systems is challenging due to the limited labeled data, memory, and computing capacity.
Young D. Kwon, Jagmohan Chauhan, Hong Jia, Stylianos I. Venieris, Cecilia Mascolo
SenSys1
2023 MyoKey: Inertial Motion Sensing and Gesture-Based QWERTY Keyboard for Extended Realities
abstract
Usability challenges and social acceptance of textual input in a context of extended realities (XR) motivate the research of novel input modalities. We investigate the fusion of inertial measurement unit (IMU) control and surface electromyography (sEMG) gesture recognition applied to text entry using a QWERTY-layout virtual keyboard. We design, implement, and evaluate the proposed multi-modal solution named MyoKey. The user can select characters with a combination of arm movements and hand gestures. MyoKey employs a lightweight convolutional neural network classifier that can be deployed on a mobile device with insignificant inference time. We demonstrate the practicality of interruption-free text entry with MyoKey, by recruiting 12 participants and by testing three sets of grasp micro-gestures in three scenarios: empty hand text input, tripod grasp (e.g., pen), and a cylindrical grasp (e.g., umbrella). With MyoKey, users achieve an average text entry rate of 9.33 words per minute (WPM), 8.76 WPM, and 8.35 WPM for the freehand, tripod grasp, and cylindrical grasp conditions, respectively.
Kirill A. Shatilov, Young D. Kwon, Lik-Hang Lee, Dimitris Chatzopoulos, Pan Hui 0001
IEEE Trans. Mob. Comput.2
2022 Causal Analysis on the Anchor Store Effect in a Location-based Social Network
abstract
A particular phenomenon of interest in Retail Eco-nomics is the spillover effect of anchor stores (specific stores with a reputable brand) to non-anchor stores in terms of customer traffic. Prior works in this area rely on small and survey-based datasets that are often confidential or expensive to collect on a large scale. Also, very few works study the underlying causal mechanisms between factors that underpin the spillover effect. In this work, we analyze the causal relationship between anchor stores and customer traffic to non-anchor stores and employ a propensity score matching framework to investigate this effect more efficiently. First of all, to demonstrate the effect, we leverage open and mobile data from London Datastore and Location-Based Social Networks (LBSNs) such as Foursquare. We then perform a large-scale empirical analysis of customer visit patterns from anchor stores to non-anchor stores (e.g., non-chain restaurants) located in the Greater London area as a case study. By studying over 600 neighbourhoods in the Greater London area, we find that anchor stores cause a 14.2-26.5% increase in customer traffic for the non-anchor stores reinforcing the established economic theory Moreover, we evaluate the efficiency of our methodology by studying the confounder balance, dose difference and performance of the matching framework on synthetic data. Through this work, we point decision-makers in the retail industry to a more systematic approach to estimate the anchor store effect and pave the way for further research to discover more complex causal relationships underlying this effect with open data.
Anish K. Vallapuram, Young D. Kwon, Lik-Hang Lee, Fengli Xu, Pan Hui 0001
ASONAM2
2022 YONO: Modeling Multiple Heterogeneous Neural Networks on Microcontrollers
abstract
Internet of Things (IoT) systems provide large amounts of data on all aspects of human behavior. Machine learning techniques, especially deep neural networks (DNN), have shown promise in making sense of this data at a large scale. Also, the research community has worked to reduce the computational and resource demands of DNN to compute on low-resourced micro controllers (MCUs). However, most of the current work in embedded deep learning focuses on solving a single task efficiently, while the multi-tasking nature and applications of IoT devices demand systems that can handle a diverse range of tasks (such as activity, gesture, voice, and context recognition) with input from a variety of sensors, simultaneously. In this paper, we propose YONO, a product quantization (PQ) based approach that compresses multiple heterogeneous models and enables in-memory model execution and model switching for dissimilar multi-task learning on MCUs. We first adopt PQ to learn codebooks that store weights of different models. Also, we propose a novel network optimization and heuristics to maximize the com-pression rate and minimize the accuracy loss. Then, we develop an online component of YONO for efficient model execution and switching between multiple tasks on an MCU at run time without relying on an external storage device. YONO shows remarkable performance as it can compress multiple heterogeneous models with negligible or no loss of accuracy up to 12.37x. Furthermore, YONO's online component enables an efficient execution (latency of 16–159 ms and energy consumption of 3.8-37.9 mJ per operation) and reduces modelloading/switching la-tency and energy consumption by 93.3-94.5% and 93.9-95.0%, respectively, compared to external storage access. Interestingly, YONO can compress various architectures trained with datasets that were not shown during YONO's offline codebook learning phase showing the generalizability of our method. To summarize, YONO shows great potential and opens further doors to enable multi-task learning systems on extremely resource-constrained devices.
Young D. Kwon, Jagmohan Chauhan, Cecilia Mascolo
IPSN1
2022 PROS: an efficient pattern-driven compressive sensing framework for low-power biopotential-based wearables with on-chip intelligence
abstract
While the global healthcare market of wearable devices has been growing significantly in recent years and is predicted to reach $60 billion by 2028, many important healthcare applications such as seizure monitoring, drowsiness detection, etc. have not been deployed due to the limited battery lifetime, slow response rate, and inadequate biosignal quality.
Nhat Pham, Hong Jia, Tuan Dinh, Nam Bui, Young D. Kwon, Dong Ma 0001, Phuc Nguyen 0002, Cecilia Mascolo, Tam Vu 0001
MobiCom6
2021 IAN: interpretable attention network for churn prediction in LBSNs
abstract
With the rise of Location-Based Social Networks (LBSNs) and their heavy reliance on User-Generated Content, it has become essential to attract and keep more users, which makes the churn prediction problem interesting. Recent research focuses on solving the task by utilizing complex neural networks. However, due to the black-box nature of those proposed deep learning algorithms, it is still a challenge for LBSN managers to interpret the prediction results and design strategies to prevent churning behavior. Therefore, in this paper, we perform the first investigation into the interpretability of the churn prediction in LBSNs. We proposed a novel attention-based deep learning network, Interpretable Attention Network (IAN), to achieve high performance while ensuring interpretability. The network is capable to process the complex temporal multivariate multidimensional user data from LBSN datasets (i.e. Yelp and Foursquare) and provides meaningful explanations of its prediction. We also utilize several visualization techniques to interpret the prediction results. By analyzing the attention output, researchers can intuitively gain insights into which features dominate the model's prediction of churning users. Finally, we expect our model to become a robust and powerful tool to help LBSN applications to understand and analyze user churning behavior and in turn remain users.
Young D. Kwon, Youwen Kang, Pan Hui 0001
ASONAM3
2021 Interpretable business survival prediction
abstract
The survival of a business is undeniably pertinent to its success. A key factor contributing to its continuity depends on its customers. The surge of location-based social networks such as Yelp, Diangping, and Foursquare has paved the way for leveraging user-generated content on these platforms to predict business survival. Prior works in this area have developed several quantitative features to capture geography and user mobility among businesses. However, the development of qualitative features is minimal. In this work, we thus perform extensive feature engineering across four feature sets, namely, geography, user mobility, business attributes, and linguistic modelling to develop classifiers for business survival prediction. We additionally employ an interpretability framework to generate explanations and qualitatively assess the classifiers' predictions. Experimentation among the feature sets reveals that qualitative features including business attributes and linguistic features have the highest predictive power, achieving AUC scores of 0.72 and 0.67, respectively. Furthermore, the explanations generated by the interpretability framework demonstrate that these models can potentially identify the reasons from review texts for the survival of a business.
Anish K. Vallapuram, Nikhil Nanda, Young D. Kwon, Pan Hui 0001
ASONAM3
2021 Exploring System Performance of Continual Learning for Mobile and Embedded Sensing Applications
abstract
Continual learning approaches help deep neural network models adapt and learn incrementally by trying to solve catastrophic forgetting. However, whether these existing approaches, applied traditionally to image-based tasks, work with the same efficacy to the sequential time series data generated by mobile or embedded sensing systems remains an unanswered question. To address this void, we conduct the first comprehensive empirical study that quantifies the performance of three predominant continual learning schemes (i.e., regularization, replay, and replay with examples) on six datasets from three mobile and embedded sensing applications in a range of scenarios having different learning complexities. More specifically, we implement an end-to-end continual learning framework on edge devices. Then we investigate the generalizability, trade-offs between performance, storage, computational costs, and memory footprint of different continual learning methods. Our findings suggest that replay with exemplars-based schemes such as iCaRL has the best performance trade-offs, even in complex scenarios, at the expense of some storage space (few MBs) for training examples (1% to 5%). We also demonstrate for the first time that it is feasible and practical to run continual learning on-device with a limited memory budget. In particular, the latency on two types of mobile and embedded devices suggests that both incremental learning time (few seconds - 4 minutes) and training time (1 - 75 minutes) across datasets are acceptable, as training could happen on the device when the embedded device is charging thereby ensuring complete data privacy. Finally, we present some guidelines for practitioners who want to apply a continual learning paradigm for mobile sensing tasks.
Young D. Kwon, Jagmohan Chauhan, Abhishek Kumar 0011, Pan Hui 0001, Cecilia Mascolo
SEC1
2021 FastICARL: Fast Incremental Classifier and Representation Learning with Efficient Budget Allocation in Audio Sensing Applications
abstract
Various incremental learning (IL) approaches have been proposed to help deep learning models learn new tasks/classes continuously without forgetting what was learned previously (i.e., avoid catastrophic forgetting). With the growing number of deployed audio sensing applications that need to dynamically incorporate new tasks and changing input distribution from users, the ability of IL on-device becomes essential for both efficiency and user privacy. However, prior works suffer from high computational costs and storage demands which hinders the deployment of IL on-device. In this work, to overcome these limitations, we develop an end-to-end and on-device IL framework, FastICARL, that incorporates an exemplar-based IL and quantization in the context of audio-based applications. We first employ k-nearest-neighbor to reduce the latency of IL. Then, we jointly utilize a quantization technique to decrease the storage requirements of IL. We implement FastICARL on two types of mobile devices and demonstrate that FastICARL remarkably decreases the IL time up to 78-92% and the storage requirements by 2-4 times without sacrificing its performance. FastICARL enables complete on-device IL, ensuring user privacy as the user data does not need to leave the device.
Young D. Kwon, Jagmohan Chauhan, Cecilia Mascolo
Interspeech1
2020 Enemy at the Gate: Evolution of Twitter User's Polarization During National Crisis
abstract
Social networks are effective platforms to study the real-life behavior of users. In this paper, we study users' political polarization during the times of crisis and its relation to nationalism. To this purpose, we focus on the reaction of Indian and Pakistani Twitter users during February 2019 crisis and the ensuing Indian General Elections in 2019. We show that a national crisis affects the polarization and discourse in both countries. Also, we show that user activities increase during a national crisis, and political discourse strengthens while polarization decreases on critical days. Finally, we highlight the links between this crisis and the Indian elections and show how the political parties discussed the crisis in their campaigns.
Ehsan ul Haq, Tristan Braud, Young D. Kwon, Pan Hui 0001
ASONAM3
2019 Effects of ego networks and communities on self-disclosure in an online social network
abstract
Understanding how much users disclose personal information in Online Social Networks (OSN) has served various scenarios such as maintaining social relationships and customer segmentation. Prior studies on self-disclosure have relied on surveys or users' direct social networks. These approaches, however, cannot represent the whole population nor consider user dynamics at the community level.
Young D. Kwon, Reza Hadi Mogavi, Ehsan ul Haq, Youngjin Kwon, Xiaojuan Ma, Pan Hui 0001
ASONAM1
1997 A stochastic environment modelling method for mobile robot by using 2-D laser scanner
abstract
A method of environment modelling is presented which is based on the stochastic approximation technique. The environment is represented in this paper as stochastic obstacle regions equipped with their own stochastic variables such as mean, variance and eigenvalues. The stochastic variables in each obstacle region are updated by using the distance information of the obstacles, which is acquired from the laser scanner sampling time. The representation of the environment with the stochastic variables enables us to save CPU time and memory consumption in building the map. This technique can also detect if the obstacles are removed and is applicable to the quasi-static environment. If the eigenvalues of the covariance matrix in a certain region are known, then the feature of that region can be extracted as a line or as an ellipse. Thus, the algorithm presented can be used for the navigation and the localization of the mobile robot. The algorithm presented is successfully tested on our mobile robot ARES-II system equipped with the LADAR 2D-laser scanner.
Young D. Kwon, Jin S. Lee
ICRA1