VLDB 2026 Research / reviewers in the wild / expert
Taesik Gong
dblp:206/1779
· DBLP profile ↗
26ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-8967-3652ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Computer networks · 12 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral CuesabstractLarge language models (LLMs) are increasingly used by end users, yet existing personalization methods relying on static profiles or text-only signals fail to capture query-specific expertise variation.We present ExPerT, a querywise personalization framework that adapts LLM responses to users' query domain expertise by combining semantic and behavioral cues.ExPerT consists of two key components: (i) a semantic-behavioral expertise inference module that jointly interprets query text and keystroke dynamics via in-context LLM prompting, and (ii) an expertise-conditioned response generation that adapts the level of detail, terminology, and conceptual complexity.Our user study with 40 participants and 1270 queries demonstrated that ExPerT reduced expertise inference error by 65.7% compared to the strongest baseline (MAE = 0.398 vs. 1.162) and improved response satisfaction by 17.52% (from 3.71 to 4.36) on a 5-point Likert scale.... Yeji Park, Jiwon Tark, Taesik Gong |
ACL (1) | 3 |
| 2025 | Adaptive Camera Sensor for Vision ModelsabstractDomain shift remains a persistent challenge in deep-learning-based computer vision, often requiring extensive model modifications or large labeled datasets to address. Inspired by human visual perception, which adjusts input quality through corrective lenses rather than over-training the brain, we propose Lens, a novel camera sensor control method that enhances model performance by capturing high-quality images from the model’s perspective, rather than relying on traditional human-centric sensor control. Lens is lightweight and adapts sensor parameters to specific models and scenes in real-time (i.e., test-time input adaptation). At its core, Lens utilizes VisiT, a training-free, model-specific quality indicator that evaluates individual unlabeled samples at test time using confidence scores, without additional adaptation costs. To validate Lens, we introduce ImageNet-ES Diverse, a new benchmark dataset capturing natural perturbations from varying sensor and lighting conditions. Extensive experiments on both ImageNet-ES and our new ImageNet-ES Diverse show that Lens significantly improves model accuracy across various baseline schemes for sensor control and model modification, while maintaining low latency in image captures. Lens effectively compensates for large model size differences and integrates synergistically with model improvement techniques. Our code and dataset are available at github.com/Edw2n/Lens.git. Eunsu Baek, Sunghwan Han, Taesik Gong, Hyung-Sin Kim |
ICLR | 3 |
| 2025 | Test-Time Adaptation with Binary FeedbackabstractDeep learning models perform poorly when domain shifts exist between training and test data. Test-time adaptation (TTA) is a paradigm to mitigate this issue by adapting pre-trained models using only unlabeled test samples. However, existing TTA methods can fail under severe domain shifts, while recent active TTA approaches requiring full-class labels are impractical due to high labeling costs. To address this issue, we introduce a new setting of TTA with binary feedback, which uses a few binary feedbacks from annotators to indicate whether model predictions are correct, thereby significantly reducing the labeling burden of annotators. Under the setting, we propose BiTTA, a novel dual-path optimization framework that leverages reinforcement learning to balance binary feedback-guided adaptation on uncertain samples with agreement-based self-adaptation on confident predictions. Experiments show BiTTA achieves substantial accuracy improvements over state-of-the-art baselines, demonstrating its effectiveness in handling severe distribution shifts with minimal labeling effort. Taeckyung Lee, Sorn Chottananurak, Jinwoo Shin, Taesik Gong, Sung-Ju Lee 0001 |
ICML | 5 |
| 2025 | Position: AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm ShiftabstractCurrent AI advances largely rely on scaling neural models and expanding training datasets to achieve generalization and robustness. Despite notable successes, this paradigm incurs significant environmental, economic, and ethical costs, limiting sustainability and equitable access. Inspired by biological sensory systems, where adaptation occurs dynamically at the input (e.g., adjusting pupil size, refocusing vision)—we advocate for adaptive sensing as a necessary and foundational shift. Adaptive sensing proactively modulates sensor parameters (e.g., exposure, sensitivity, multimodal configurations) at the input level, significantly mitigating covariate shifts and improving efficiency. Empirical evidence from recent studies demonstrates that adaptive sensing enables small models (e.g., EfficientNet-B0) to surpass substantially larger models (e.g., OpenCLIP-H) trained with significantly more data and compute. We (i) outline a roadmap for broadly integrating adaptive sensing into real-world applications spanning humanoid, healthcare, autonomous systems, agriculture, and environmental monitoring, (ii) critically assess technical and ethical integration challenges, and (iii) propose targeted research directions, such as standardized benchmarks, real-time adaptive algorithms, multimodal integration, and privacy-preserving methods. Collectively, these efforts aim to transition the AI community toward sustainable, robust, and equitable artificial intelligence systems. Eunsu Baek, Keondo Park, JeongGil Ko, Min-hwan Oh, Taesik Gong, Hyung-Sin Kim |
NeurIPS | 5 |
| 2025 | SNAP: Low-Latency Test-Time Adaptation with Sparse UpdatesabstractTest-Time Adaptation (TTA) adjusts models using unlabeled test data to handle dynamic distribution shifts. However, existing methods rely on frequent adaptation and high computational cost, making them unsuitable for resource-constrained edge environments. To address this, we propose SNAP, a sparse TTA framework that reduces adaptation frequency and data usage while preserving accuracy. SNAP maintains competitive accuracy even when adapting based on only 1\% of the incoming data stream, demonstrating its robustness under infrequent updates. Our method introduces two key components: (i) Class and Domain Representative Memory (CnDRM), which identifies and stores a small set of samples that are representative of both class and domain characteristics to support efficient adaptation with limited data; and (ii) Inference-only Batch-aware Memory Normalization (IoBMN), which dynamically adjusts normalization statistics at inference time by leveraging these representative samples, enabling efficient alignment to shifting target domains. Integrated with five state-of-the-art TTA algorithms, SNAP reduces latency by up to 93.12\%, while keeping the accuracy drop below 3.3\%, even across adaptation rates ranging from 1\% to 50\%. This demonstrates its strong potential for practical use on edge devices serving latency-sensitive applications. The source code is available at https://github.com/chahh9808/SNAP. Hyeongheon Cha, Hye Won Chung, Taesik Gong, Sung-Ju Lee 0001 |
NeurIPS | 4 |
| 2025 | SelfReplay: Adapting Self-Supervised Sensory Models via Adaptive Meta-Task ReplayabstractSelf-supervised learning enables effective model pre-training on large-scale unlabeled data, which is crucial for user-specific fine-tuning in mobile sensing applications. However, pre-trained models often face significant domain shifts during fine-tuning due to user diversity, leading to performance degradation. To address this, we propose SelfReplay, an adaptive approach designed to align self-supervised models to different domains. SelfReplay consists of two stages: MetaSSL, which leverages meta-learning with self-supervised learning to pre-train domain-adaptive weights, and ReplaySSL, which further adapts the pre-trained model to each user's domain by replaying the meta-learned self-supervised task with a few user-specific samples. This produces a personalized model tailored to each user. Evaluations on mobile sensing benchmarks demonstrate that SelfReplay outperforms existing baselines, improving the F1-score by 9.4%p on average. On-device analyses on a commodity smartphone show the efficiency of SelfReplay's adaptation step, required just once after deployment, with SimCLR completing in only 10 seconds while using less than 100MB of memory. Hyungjun Yoon, Jae Hyun Kwak, Biniyam Aschalew Tolera, Gaole Dai, Mo Li 0001, Taesik Gong, Kimin Lee, Sung-Ju Lee 0001 |
SenSys | 6 |
| 2025 | Synergy: Towards On-Body AI via Tiny AI Accelerator Collaboration on WearablesabstractThe advent of tiny artificial intelligence (AI) accelerators enables AI to run at the extreme edge, offering reduced latency, lower power cost, and improved privacy. When integrated into wearable devices, these accelerators open exciting opportunities, allowing various AI apps to run directly on the body. We present Synergy that provides AI apps with besteffort performance via system-driven holistic collaboration over AI accelerator-equipped wearables. To achieve this, Synergy provides device-agnostic programming interfaces to AI apps, giving the system visibility and controllability over the app's resource use. Then, Synergy maximizes the inference throughput of concurrent AI models by creating various execution plans for each app considering AI accelerator availability and intelligently selecting the best set of execution plans. Synergy further improves throughput by leveraging parallelization opportunities over multiple computation units. Our evaluations with 7 baselines and 8 models demonstrate that, on average, Synergy achieves a 23.0× improvement in throughput, while reducing latency by 73.9% and power consumption by 15.8%, compared to the baselines Taesik Gong, Utku Günay Acer, Fahim Kawsar, Chulhong Min |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | From Vision to Motion: Translating Large-Scale Knowledge for Data-Scarce IMU ApplicationsabstractPre-training representations acquired via self-supervised learning could achieve high accuracy on even tasks with small training data. Unlike in vision and natural language processing domains, pre-training for IMU-based applications is challenging, as there are few public datasets with sufficient size and diversity to learn generalizable representations. To overcome this problem, we propose IMG2IMU that adapts pre-trained representation from large-scale images to diverse IMU sensing tasks. We convert the sensor data into visually interpretable spectrograms for the model to utilize the knowledge gained from vision. We further present a sensor-aware pre-training method for images that enables models to acquire particularly impactful knowledge for IMU sensing applications. This involves using contrastive learning on our augmentation set customized for the properties of sensor data. Our evaluation with four different IMU sensing tasks shows that IMG2IMU outperforms the baselines pre-trained on sensor data by an average of 9.6%p F1-score, illustrating that vision knowledge can be usefully incorporated into IMU sensing applications where only limited training data is available. Hyungjun Yoon, Hyeongheon Cha, Hoang C. Nguyen, Taesik Gong, Sung-Ju Lee 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | AETTA: Label-Free Accuracy Estimation for Test-Time AdaptationabstractTest-time adaptation (TTA) has emerged as a viable solution to adapt pretrained models to domain shifts using unlabeled test data. However, TTA faces challenges of adaptation failures due to its reliance on blind adaptation to unknown test samples in dynamic scenarios. Traditional methods for out-of-distribution performance estimation are limited by unrealistic assumptions in the TTA context, such as requiring labeled data or retraining models. To address this issue, we propose AETTA, a label-free accuracy estimation algorithm for TTA. We propose the prediction disagreement as the accuracy estimate, calculated by comparing the target model prediction with dropout inferences. We then improve the prediction disagreement to extend the applicability of AETTA under adaptation failures. Our extensive evaluation with four baselines and six TTA methods demonstrates that AETTA shows an average of 19.8%p more accurate estimation compared with the baselines. We further demonstrate the effectiveness of accuracy estimation with a model recovery case study, showcasing the practicality of our model recovery based on accuracy estimation. The source code is available at https://github.com/taeckyung/AETTA. Taeckyung Lee, Sorn Chottananurak, Taesik Gong, Sung-Ju Lee 0001 |
CVPR | 3 |
| 2024 | By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual PromptingabstractLarge language models (LLMs) have demonstrated exceptional abilities across various domains.However, utilizing LLMs for ubiquitous sensing applications remains challenging as existing text-prompt methods show significant performance degradation when handling long sensor data sequences.We propose a visual prompting approach for sensor data using multimodal LLMs (MLLMs).We design a visual prompt that directs MLLMs to utilize visualized sensor data alongside the target sensory task descriptions.Additionally, we introduce a visualization generator that automates the creation of optimal visualizations tailored to a given sensory task, eliminating the need for prior task-specific knowledge.We evaluated our approach on nine sensory tasks involving four sensing modalities, achieving an average of 10% higher accuracy than text-based prompts and reducing token costs by 15.8×.Our findings highlight the effectiveness and cost-efficiency of visual prompts with MLLMs for various sensory tasks.The source code is available at https://github. com/diamond264/ByMyEyes. ### InstructionYou are an expert in sensor data analysis.Given the sensor data, determine the correct answer from the options listed in the question.Provide the answer with the format of ANSWER , where ANSWER corresponds to one of the options listed in the question.If the answer is not in the options, choose the most possible option.The ECG data is collected from a lead II ECG sensor.The ECG data is recorded over 10 seconds.The data is normalized with the statistics of the user's data.Please refer to the provided examples and use them to answer the following question for the target data.### Examples *Example of normal*: Average heartbeat in the ECG signal (list of ['lead II']): [-0.34, -0.34, -0.35, -0.35, -0.36, … ECG_P_Peaks in the ECG signal (list of (index, value)): [(21, 0.08), (129, 0.19), (239, 0.22), … ECG_Q_Peaks in the ECG signal (list of (index, value)): [(35, -0.61), (137, -0.46), (246, -0.48), … ECG_S_Peaks in the ECG signal (list of (index, value)): [(42, -1.4), (149, -1.35), (260, -1.33), … ECG_T_Peaks in the ECG signal (list of (index, value)): [(63, 2.18), (171, 2.11), (282, 2.31), … *Example of conduction disturbance*: Average heartbeat in the ECG signal (list of ['lead II']): [-0.15, -0.24, -0.29, -0.31, -0.28, … ECG_P_Peaks in the ECG signal (list of (index, value)): [(4, 0.14), (57, 0.3), (103, 0.22), … ECG_Q_Peaks in the ECG signal (list of (index, value)): [(14, -0.05), (65, -0.05), (109, -0.07), … ECG_S_Peaks in the ECG signal (list of (index, value)): [(22, -2.28), (73, -2.1), (124, -2.35), … ECG_T_Peaks in the ECG signal (list of (index, value)): [(82, 0.03), (142, 0.64), (245, 0.25), … ### Question Average heartbeat in the ECG signal (list of ['lead II']): [-0.39, -0.39, -0.39, -0.39, -0.4,… ECG_P_Peaks in the ECG signal (list of (index, value)): [(15, 0.14), (94, -0.26), (173, -0.23), … ECG_Q_Peaks in the ECG signal (list of (index, value)): [(23, -0.27), (102, -0.81), (182, -0.55), … ECG_S_Peaks in the ECG signal (list of (index, value)): [(34, 0.15), (116, -0.51), (192, -0.45), … ECG_T_Peaks in the ECG signal (list of (index, value)): [(50, 1.39), (130, 1.1), (209, 1.31), … *Question*: When the sensor data is used for a task for classifying ECG data into 2 categories: conduction disturbance, normal, what is the most likely answer among ['conduction disturbance', 'normal']?*Answer*: Hyungjun Yoon, Biniyam Aschalew Tolera, Taesik Gong, Kimin Lee, Sung-Ju Lee 0001 |
EMNLP | 3 |
| 2024 | Poster: Time-Efficient Sparse and Lightweight Adaptation for Real-Time Mobile ApplicationabstractWhen deployed in mobile scenarios, deep learning models often suffer from performance degradation due to domain shifts. Test-Time Adaptation (TTA) offers a viable solution, but current approaches face latency issues on resource-constrained mobile devices. We propose TESLA: Time-Efficient Sparse and Lightweight Adaptation strategy for real-time mobile applications, which skips adaptation for specific batches to increase the inference sample rate. Our method balances model accuracy and inference speed by accumulating domain-informative samples from non-adapted batches and sparsely adapting them. Experiments on edge devices demonstrate competitive accuracy even with sparse adaptation rates, highlighting the effectiveness of our approach in real-time mobile applications. Our strategy can seamlessly integrate with existing lightweight adaptation and optimization algorithms, further accelerating inference across diverse mobile systems. Hyeongheon Cha, Taesik Gong, Sung-Ju Lee 0001 |
MobiSys | 2 |
| 2024 | DEX: Data Channel Extension for Efficient CNN Inference on Tiny AI AcceleratorsabstractTiny machine learning (TinyML) aims to run ML models on small devices and is increasingly favored for its enhanced privacy, reduced latency, and low cost. Recently, the advent of tiny AI accelerators has revolutionized the TinyML field by significantly enhancing hardware processing power. These accelerators, equipped with multiple parallel processors and dedicated per-processor memory instances, offer substantial performance improvements over traditional microcontroller units (MCUs). However, their limited data memory often necessitates downsampling input images, resulting in accuracy degradation. To address this challenge, we propose Data channel EXtension (DEX), a novel approach for efficient CNN execution on tiny AI accelerators. DEX incorporates additional spatial information from original images into input images through patch-wise even sampling and channel-wise stacking, effectively extending data across input channels. By leveraging underutilized processors and data memory for channel extension, DEX facilitates parallel execution without increasing inference latency. Our evaluation with four models and four datasets on tiny AI accelerators demonstrates that this simple idea improves accuracy on average by 3.5%p while keeping the inference latency the same on the AI accelerator. The source code is available at https://github.com/Nokia-Bell-Labs/data-channel-extension. Taesik Gong, Fahim Kawsar, Chulhong Min |
NeurIPS | 1 |
| 2023 | LanSER: Language-Model Supported Speech Emotion RecognitionabstractSpeech emotion recognition (SER) models typically rely on costly human-labeled data for training, making scaling methods to large speech datasets and nuanced emotion taxonomies difficult. We present LanSER, a method that enables the use of unlabeled data by inferring weak emotion labels via pre-trained large language models through weakly-supervised learning. For inferring weak labels constrained to a taxonomy, we use a textual entailment approach that selects an emotion label with the highest entailment score for a speech transcript extracted via automatic speech recognition. Our experimental results show that models pre-trained on large datasets with this weak supervision outperform other baseline models on standard SER datasets when fine-tuned, and show improved label efficiency. Despite being pre-trained on labels derived only from text, we show that the resulting representations appear to model the prosodic content of speech. Taesik Gong, Josh Belanich, Krishna Somandepalli, Arsha Nagrani, Brian Eoff, Brendan Jou |
INTERSPEECH | 1 |
| 2023 | SoTTA: Robust Test-Time Adaptation on Noisy Data StreamsabstractTest-time adaptation (TTA) aims to address distributional shifts between training and testing data using only unlabeled test data streams for continual model adaptation. However, most TTA methods assume benign test streams, while test samples could be unexpectedly diverse in the wild. For instance, an unseen object or noise could appear in autonomous driving. This leads to a new threat to existing TTA algorithms; we found that prior TTA algorithms suffer from those noisy test samples as they blindly adapt to incoming samples. To address this problem, we present Screening-out Test-Time Adaptation (SoTTA), a novel TTA algorithm that is robust to noisy samples. The key enabler of SoTTA is two-fold: (i) input-wise robustness via high-confidence uniform-class sampling that effectively filters out the impact of noisy samples and (ii) parameter-wise robustness via entropy-sharpness minimization that improves the robustness of model parameters against large gradients from noisy samples. Our evaluation with standard TTA benchmarks with various noisy scenarios shows that our method outperforms state-of-the-art TTA methods under the presence of noisy samples and achieves comparable accuracy to those methods without noisy samples. The source code is available at https://github.com/taeckyung/SoTTA. Taesik Gong, Taeckyung Lee, Sorn Chottananurak, Sung-Ju Lee 0001 |
NeurIPS | 1 |
| 2022 | MyDJ: Sensing Food Intakes with an Attachable on Your Eyeglass FrameabstractVarious automated eating detection wearables have been proposed to monitor food intakes. While these systems overcome the forgetfulness of manual user journaling, they typically show low accuracy at outside-the-lab environments or have intrusive form-factors (e.g., headgear). Eyeglasses are emerging as a socially-acceptable eating detection wearable, but existing approaches require custom-built frames and consume large power. We propose MyDJ, an eating detection system that could be attached to any eyeglass frame. MyDJ achieves accurate and energy-efficient eating detection by capturing complementary chewing signals on a piezoelectric sensor and an accelerometer. We evaluated the accuracy and wearability of MyDJ with 30 subjects in uncontrolled environments, where six subjects attached MyDJ on their own eyeglasses for a week. Our study shows that MyDJ achieves 0.919 F1-score in eating episode coverage, with 4.03 × battery time over the state-of-the-art systems. In addition, participants reported wearing MyDJ was almost as comfortable (94.95%) as wearing regular eyeglasses. Jaemin Shin 0005, Seungjoo Lee, Taesik Gong, Hyungjun Yoon, Hyunchul Roh, Andrea Bianchi, Sung-Ju Lee 0001 |
CHI | 3 |
| 2022 | NOTE: Robust Continual Test-time Adaptation Against Temporal CorrelationabstractTest-time adaptation (TTA) is an emerging paradigm that addresses distributional shifts between training and testing phases without additional data acquisition or labeling cost; only unlabeled test data streams are used for continual model adaptation. Previous TTA schemes assume that the test samples are independent and identically distributed (i.i.d.), even though they are often temporally correlated (non-i.i.d.) in application scenarios, e.g., autonomous driving. We discover that most existing TTA methods fail dramatically under such scenarios. Motivated by this, we present a new test-time adaptation scheme that is robust against non-i.i.d. test data streams. Our novelty is mainly two-fold: (a) Instance-Aware Batch Normalization (IABN) that corrects normalization for out-of-distribution samples, and (b) Prediction-balanced Reservoir Sampling (PBRS) that simulates i.i.d. data stream from non-i.i.d. stream in a class-balanced manner. Our evaluation with various datasets, including real-world non-i.i.d. streams, demonstrates that the proposed robust TTA not only outperforms state-of-the-art TTA algorithms in the non-i.i.d. setting, but also achieves comparable performance to those algorithms under the i.i.d. assumption. Code is available at https://github.com/TaesikGong/NOTE. Taesik Gong, Jongheon Jeong, Taewon Kim, Jinwoo Shin, Sung-Ju Lee 0001 |
NeurIPS | 1 |
| 2022 | Adapting to Unknown Conditions in Learning-Based Mobile SensingabstractMany applications utilize sensors on mobile devices and apply deep learning for diverse applications. However, they have rarely enjoyed mainstream adoption due to many differentindividual conditionsusers encounter. Individual conditions are characterized by users’ unique behaviors and different devices they carry, which collectively make sensor inputs different. It is impractical to train countless individual conditions beforehand and we thus argue meta-learning is a great approach in solving this problem. We presentMetaSensethat leverages “seen” conditions in training data to adapt to an “unseen” condition (i.e., the target user). Specifically, we design a meta-learning framework that learns “how to adapt” to the target via iterative training sessions of adaptation. MetaSense requires very few training examples from the target (e.g., one or two) and thus requires minimal user effort. In addition, we propose asimilar condition detector(SCD) that identifies when the unseen condition has similar characteristics to seen conditions and leverages this hint to further improve the accuracy. Our evaluation with 10 different datasets shows that MetaSense improves the accuracy of state-of-the-art transfer learning and meta learning methods by 15 and 11 percent, respectively. Furthermore, our SCD achieves additional accuracy improvement (e.g., 15 percent for human activity recognition). Taesik Gong, Yeonsu Kim, Ryuhaerang Choi, Jinwoo Shin, Sung-Ju Lee 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2020 | Messaging Beyond Texts with Real-time Image SuggestionsabstractWhile people primarily communicate with text in mobile chat applications, they are increasingly using visual elements such as images, emojis, and memes. Using such visual elements could help users communicate clearly and make chatting experience enjoyable. However, finding and inserting contextually appropriate images during the chat can be both tedious and distracting. We introduce MilliCat, a real-time image suggestion system that recommends images that match the chat content within a mobile chat application (i.e., autocomplete with images). MilliCat combines natural language processing (e.g., keyword extraction, dependency parsing) and mobile computing (e.g., resource and energy-efficiency) techniques to autonomously make image suggestions when users might want to use images. Through multiple user studies, we investigated the effectiveness of our design choices, the frequency and motivation of image usage by the participants, and the impact of MilliCat on mobile chat experiences. Our results indicate that MilliCat’s real-time image suggestion enables users to quickly and conveniently select and display images on mobile chat by significantly reducing the latency in the image selection process (3.19 × improvement) and consequently more frequent image usage (1.8 ×) than existing solutions. Our study participants reported that they used images more often with MilliCat as the images helped them convey information more effectively, emphasize their opinion, express emotions, and have fun chatting experience. Joon-Gyum Kim, Taesik Gong, Kyungsik Han, Juho Kim 0001, JeongGil Ko, Sung-Ju Lee 0001 |
MobileHCI | 2 |
| 2019 | AudiDoS: Real-Time Denial-of-Service Adversarial Attacks on Deep Audio ModelsabstractDeep learning has enabled personal and IoT devices to rethink microphones as a multi-purpose sensor for understanding conversation and the surrounding environment. This resulted in a proliferation of Voice Controllable Systems (VCS) around us. The increasing popularity of such systems is also prone to attracting miscreants, who often want to take advantage of the VCS without the knowledge of the user. Consequently, understanding the robustness of VCS, especially under adversarial attacks, has become an important research topic. Although there exists some previous work on audio adversarial attacks, their scopes are limited to embedding the attacks onto pre-recorded music clips, which when played through speakers cause VCS to misbehave. As an attack-audio needs to be played, the occurrence of this type of attacks can be suspected by a human listener. In this paper, we focus on audio-based Denial-of-Service (DoS) attack, which is unexplored in the literature. Contrary to previous work, we show that adversarial audio attacks in real-time and overthe-air are possible, while a user interacts with VCS. We show that the attacks are effective regardless of the user's command and interaction timings. In this paper, we present a first-of-itskind imperceptible and always-on universal audio perturbation technique that enables such DoS attack to be successful. We thoroughly evaluate the performance of the attacking scheme across (i) two learning tasks, (ii) two model architectures and (iii) three datasets. We demonstrate that the attack can introduce as high as 78% error rate in audio recognition tasks. Taesik Gong, Alberto Gil C. P. Ramos, Sourav Bhattacharya, Akhil Mathur, Fahim Kawsar |
ICMLA | 1 |
| 2019 | Dissecting 802.11ac Performance - Why You Should Turn Off MU-MIMOabstractWhile the recent Wi-Fi standard 802.11ac achieves Gb/s theoretical capacity with Multi-User MIMO (MU-MIMO) technology, several studies reported that throughput of 802.11ac in practice is far from Gb/s link speed. We investigate the downlink throughput of Wi-Fi systems with commercially available 802.11ac products in multiple indoor environments to reveal the throughput of MU-MIMO system that user experiences in practice. From our experiments, Single-User MIMO (SU-MIMO) outperformed MU-MIMO at every experimental environments. We further provide analysis on our experimental results considering channel sounding overhead, user grouping, environmental impact, and transmission mode selection. Hyunwoo Choi, Taesik Gong, Jaehun Kim, Jaemin Shin 0005, Sung-Ju Lee 0001 |
MobiSys | 2 |
| 2019 | Real-Time Object Identification with a Smartphone KnockabstractWe propose Knocker, a real-time object identification technique with smartphones. Knocker leverages unique impulse signals that are generated by knocking on an object with a smartphone. Knocker does not require any special augmentation for both smartphones and objects. Taesik Gong, Hyunsung Cho, Bowon Lee, Sung-Ju Lee 0001 |
MobiSys | 1 |
| 2019 | Towards Condition-Independent Deep Mobile SensingabstractDeep mobile sensing applications are suffering from various individual conditions in the wild. We propose a meta-learned adaptation technique to adapt to a target condition with a few labeled data. We evaluate our system on a public dataset and it outperforms baselines. Taesik Gong, Yeonsu Kim, Jinwoo Shin, Sung-Ju Lee 0001 |
MobiSys | 1 |
| 2019 | Bringing Context into Emoji RecommendationsabstractWe present Reeboc that combines machine learning and k-means clustering to analyze the conversation of a chat, extract different emotions or topics of the conversation, and recommend emojis that represent various contexts to the user. Instead of simply analyzing a single input sentence, we consider recent sentences exchanged in a conversation. we performed a user study with 17 participants in 8 groups in a realistic mobile chat environment. Participants spent the least amount of time in identifying and selecting the emojis of their choice with Reeboc (38% faster than without emoji recommendation). Joon-Gyum Kim, Taesik Gong, Evey Huang, Juho Kim 0001, Sung-Ju Lee 0001, Bogoan Kim, Jaeyeon Park 0001, Woojeong Kim, Kyungsik Han, JeongGil Ko |
MobiSys | 2 |
| 2019 | MetaSense: few-shot adaptation to untrained conditions in deep mobile sensingabstractRecent improvements in deep learning and hardware support offer a new breakthrough in mobile sensing; we could enjoy context-aware services and mobile healthcare on a mobile device powered by artificial intelligence. However, most related studies perform well only with a certain level of similarity between trained and target data distribution, while in practice, a specific user's behaviors and device make sensor inputs different. Consequently, the performance of such applications might suffer in diverse user and device conditions as training deep models in such countless conditions is infeasible. To mitigate the issue, we propose MetaSense, an adaptive deep mobile sensing system utilizing only a few (e.g., one or two) data instances from the target user. MetaSense employs meta learning that learns how to adapt to the target user's condition, by rehearsing multiple similar tasks generated from our unique task generation strategies in offline training. The trained model has the ability to rapidly adapt to the target user's condition when a few data are available. Our evaluation with real-world traces of motion and audio sensors shows that MetaSense not only outperforms the state-of-the-art transfer learning by 18% and meta learning based approaches by 15% in terms of accuracy, but also requires significantly less adaptation time for the target user. Taesik Gong, Yeonsu Kim, Jinwoo Shin, Sung-Ju Lee 0001 |
SenSys | 1 |
| 2019 | Use MU-MIMO at your own risk - Why we don't get Gb/s Wi-Fi
Hyunwoo Choi, Taesik Gong, Jaehun Kim, Jaemin Shin 0005, Sung-Ju Lee 0001 |
Ad Hoc Networks | 2 |
| 2017 | Enjoy the Silence: Noise Control with SmartphonesabstractWhile a certain sound serves a purpose to someone, for others, the same sound could be noise. Specifically, an alarm sound would be necessary for someone to wake up and start the day, while it would be an unwanted sound for people sharing the same room, who need not wake up as early. Noise cancellation is useful in this scenario, but most existing techniques require costly equipments (e.g., high quality speakers or microphones) or devices that are uncomfortable to wear during sleep. We thus explore the possibility of using only commodity smartphones to achieve active noise control and present Virtual Earplugs. As most people own a smartphone that include a microphone and speakers, and has an alarm clock feature, we believe our system could be easily used in practice. We highlight the technical challenges in realizing our vision and demonstrate the feasibility of our approach using our preliminary prototype. Our results indicate that we can reduce the alarm sound by up to 19~dB in only 2.1~seconds of processing delay. Taesik Gong, Jun Hyuk Chang, Joon-Gyum Kim, Soowon Kang, Donghwi Kim, Sung-Ju Lee 0001 |
ICCCN | 1 |