VLDB 2026 Research / reviewers in the wild / expert
Ngoc Thi Nguyen
dblp:211/2981
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-1011-8937ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMSEG: Semantic Segmentation and Few-Shot Recognition of Human Activity Recognition Data Using Large Language ModelsabstractHuman Activity Recognition (HAR) is vital for proactive health monitoring, but current systems are hindered by fixed-length time windows and the labor-intensive process of collecting annotated data which complicates effective HAR classifier development. To address these challenges, we contribute LLMSEG, a novel two-stage framework that integrates dynamic segmentation and HAR classification. Utilizing an LLM-based segmentation approach with signature activities, LLMSEG dynamically adjusts window lengths based on contextual activity, significantly enhancing accuracy over traditional methods. Additionally, it generates sensor data descriptions enriched with contextual cues, enabling effective few-shot classification without extensive labeled datasets. Systematic evaluations across various HAR datasets demonstrate that LLMSEG outperforms fixed-window methods, balances expressivity and efficiency, and generalizes well across activities. Performance can be further improved by carefully designed contextual prompts. Furthermore, LLMSEG supports robust deployment on both on-device (Raspberry Pi) and edge (laptop) devices. Overall, LLMSEG enhances the generality and reliability of HAR solutions while accommodating diverse deployment scenarios. Jiashu Liu, Kevin Post, Reo Kuchida, Agustin Zuniga, Fatemeh Sarhaddi, Huber Flores, Petteri Nurmi, Ngoc Thi Nguyen |
SenSys | 8 |
| 2026 | AI See What You Did There - The Prevalence of LLM-Generated Answers in MOOC ResponsesabstractLarge language models (LLMs) are reshaping the educational landscape, particularly in online learning environments where student supervision is often limited. Early evidence and anecdotal reports suggest that the use of AI-generated content is highly prevalent among students. However, definitive statistics remain elusive, primarily due to the challenges associated with distinguishing between AI-generated and human-generated responses. Establishing clear evidence and effective mechanisms for identifying AI-generated responses is crucial for understanding the significance of this challenge and for developing policies to address it. To tackle these issues, we present a large-scale empirical study on the prevalence of AI-generated content in online education. Our study analyzes over 4045 student responses from an introductory MOOC on the Internet of Things, employing textual analysis techniques to evaluate various metrics for identifying AI-generated responses and understanding their characteristics. Our findings reveal that a significant majority of student responses (up to 90.1%) exhibit strong similarities to AI-generated content in both wording and contextual meaning, regardless of the specific LLM or similarity metric employed. In terms LLMs usage, DeepSeek, Gemini, and Grok are the three most popular LLMs used to generate responses. Petteri Nurmi, Musfira Khan, Zahra Safaei, Ngoc Thi Nguyen, Fatemeh Sarhaddi, Mika Tompuri, Henrik Nygren, Päivi Kinnunen, Agustin Zuniga |
SIGCSE (1) | 4 |
| 2026 | SARF: Sparsity-Aware Reconstruction Framework for Large-Scale DatasetsabstractLarge-scale datasets, particularly those collected from smart devices and Internet of Things sensors, usually exhibit significant temporal and spatial sparsity, resulting in high amounts of missing data. Unless addressed in the analysis, this sparsity can result in substantial gaps and biases as well as limit the generalizability of conclusions drawn from such data. To address this challenge in data quality, we contribute the Sparsity-Aware Reconstruction Framework (SARF) as a novel and unified data fusion and reconstruction framework that enhances data quality and addresses sparsity. SARF analyzes datasets, partitioning the data into segments with similar characteristics, and reconstructs the data in each segment individually by selecting a reconstruction technique that is tailored to the internal temporal-spatial characteristics of the dataset. Through extensive experiments on two representative datasets - mobile application measurements and IoT sensor data from low-cost air quality sensors - we demonstrate that the targeted adaptation of reconstruction strategies employed by SARF significantly enhances the quality of reconstructed data. Our results show the robustness of SARF's performance across spatiotemporal variations, outperforming current state-of-the-art methods by margins up to 68% on average (74% for compressive sensing, 53% for convolutional sparse coding, 78% for deep learning). These findings underscore SARF's potential to enhance datadriven insights across multiple domains, paving the way for more robust analyses of sparsity-affected datasets. Agustin Zuniga, Huber Flores, Ngoc Thi Nguyen, Pan Hui 0001, Sasu Tarkoma, Petteri Nurmi |
IEEE Trans. Big Data | 3 |
| 2025 | SpikEy: Preventing Drink Spiking using Wearables
Zhigang Yin, Ngoc Thi Nguyen, Agustin Zuniga, Mohan Liyanage, Petteri Nurmi, Huber Flores |
ICMI | 2 |
| 2024 | The Price is Right? The Economic Value of Sharing SensorsabstractWe study user's valuations of smartphone sensing resources and the factors mediating them through a systematic auction study with 108 bids from$N=18$participants, two resource use conditions [fixed battery (FB) and variable battery (VB)] and three sensors (camera, microphone, and GPS) with differing energy and privacy costs. We use a second-price sealed-bid reverse auction as this allows us to elicit the participants’ truthful perceived value for sharing resources. We show that most users would be willing to share even highly-privacy intrusive sensors if they are sufficiently compensated. At the FB level, participants placed much lower value for sharing GPS (€13) than camera (€30) or microphone (€32.5). The values people place on sharing access to resources generally reflect four considerations: 1) the perceived value of the sensor type; 2) the value of the data captured by the sensor; 3) the impact of sharing on the device; and 4) personal variations related to sharing motives, personal tendencies, and the broader sharing context. We address the practical impact of our results by presenting two case studies (collaborative sensing and collaborative AI). Finally, we derive design implications for sharing sensing resources on personal devices. Ngoc Thi Nguyen, Maria Zubair, Agustin Zuniga, Sasu Tarkoma, Pan Hui 0001, Hyowon Lee 0001, Simon T. Perrault, Mostafa H. Ammar, Huber Flores, Petteri Nurmi |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Man and the Machine: Effects of AI-assisted Human Labeling on Interactive Annotation of Real-time Video StreamsabstractAI-assisted interactive annotation is a powerful way to facilitate data annotation—a prerequisite for constructing robust AI models. While AI-assisted interactive annotation has been extensively studied in static settings, less is known about its usage in dynamic scenarios where the annotators operate under time and cognitive constraints, e.g., while detecting suspicious or dangerous activities from real-time surveillance feeds. Understanding how AI can assist annotators in these tasks and facilitate consistent annotation is paramount to ensure high performance for AI models trained on these data. We address this gap in interactive machine learning (IML) research, contributing an extensive investigation of the benefits, limitations, and challenges of AI-assisted annotation in dynamic application use cases. We address both the effects of AI on annotators and the effects of (AI) annotations on the performance of AI models trained on annotated data in real-time video annotations. We conduct extensive experiments that compare annotation performance at two annotator levels (expert and non-expert) and two interactive labeling techniques (with and without AI assistance). In a controlled study with \(N=34\) annotators and a follow-up study with 51,963 images and their annotation labels being input to the AI model, we demonstrate that the benefits of AI-assisted models are greatest for non-expert users and for cases where targets are only partially or briefly visible. The expert users tend to outperform or achieve similar performance as the AI model. Labels combining AI and expert annotations result in the best overall performance as the AI reduces overflow and latency in the expert annotations. We derive guidelines for the use of AI-assisted human annotation in real-time dynamic use cases. Marko Radeta, Rúben Freitas, Claudio Rodrigues, Agustin Zuniga, Ngoc Thi Nguyen, Huber Flores, Petteri Nurmi |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2021 | Intelligent Shifting Cues: Increasing the Awareness of Multi-Device Interaction OpportunitiesabstractThe ever-increasing ubiquity of smart devices is creating new opportunities for people to interact and engage with digital information using multiple devices. In the simplest case this can refer to choosing which device to use for a particular task (e.g., phone, laptop or smartwatch), whereas a more complex example is simultaneously taking advantage of the capabilities of different devices (e.g., laptop and smart TV). Despite these types of opportunities becoming increasing available, currently the full potential of multi-device interactions is not being realized as people struggle to take advantage of them. As our first contribution, we study people’s willingness to engage with multi-device interactions and rank the factors that mediate this response through an online survey (N = 60). Our results show that users are strongly in favour of using multiple devices, but lack the awareness or information to engage with them, or feel that establishing the interactions is too laborious and would disrupt the fluidity of the interactions. Motivated by this result, as our second contribution we design and evaluate intelligent shifting cues, visualizations that offer information about available interaction opportunities and how to establish them, and study how they influence users willingness to engage in multi-device interactions. Results of our study show that the cues can be effective in helping people to engage with multiple devices, but that the suitability of the proposed device and fit with task are important mediating factors. We end the paper by deriving design implications for intelligent systems that can support people in engaging with multi-device interactions. Ngoc Thi Nguyen, Agustin Zuniga, Huber Flores, Hyowon Lee 0001, Simon T. Perrault, Petteri Nurmi |
UMAP | 1 |