Po-Chih Kuo

dblp:165/4431 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-4020-3147ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Benchmarking Reinforcement Learning Algorithms for ICU Ventilator Settings: An Interpretable and Probabilistic Patient Environment for Doctor Agents
abstract
Mechanical ventilation is essential in intensive care units (ICUs), but prolonged use increases patient risk. Reinforcement learning (RL) offers potential for optimizing ventilator management, yet its clinical adoption is limited by the lack of interpretable and realistic simulation environments. We propose an interpretable and probabilistic patient environment simulator based on action-based k-nearest neighbors and empirical transition probabilities, modeling stochastic state transitions grounded in real ICU data (MIMIC-IV and eICU). The simulator supports anomaly detection and provides probabilistic next-state distributions to enhance transparency and safety. Within this environment, we benchmark seven offline RL algorithms under clinically guided reward designs, including five distinct reward function configurations to explore the impact of reward shaping on agent behavior. Our results show that RL agents such as Double DQN and NFQ outperform empirical physician policies in meeting extubation guidelines, especially for high-severity patients. This benchmark enables standardized, interpretable evaluation of RL-based decision support tools for critical care.
Ya-Hsi Chang, Po-Chih Kuo
AAAI2
2026 AIDED: Augmenting Interior Design with Human Experience Data for Designer-AI Co-Design
abstract
Interior design often struggles to capture the subtleties of client experience, leaving gaps between what clients feel and what designers can act upon. We present AIDED, a designer-AI co-design workflow that integrates multimodal client data into generative AI (GAI) design processes. In a within-subjects study with twelve professional designers, we compared four modalities: baseline briefs, gaze heatmaps, questionnaire visualizations, and AI-predicted overlays. Results show that questionnaire data were trusted, creativity-enhancing, and satisfying; gaze heatmaps increased cognitive load; and AI-predicted overlays improved GAI communication but required natural language mediation to establish trust. Interviews confirmed that an authenticity-interpretability trade-off is central to balancing client voices with professional control. Our contributions are: (1) a system that incorporates experiential client signals into GAI design workflows; (2) empirical evidence of how different modalities affect design outcomes; and (3) implications for future AI tools that support human-data interaction in creative practice.
Yang Chen Lin, Chen-Ying Chien, Kai-Hsin Hou, Hung-Yu Chen, Po-Chih Kuo
CHI5
2025 Explainable Detection of Alzheimer's Disease Through Analysis of Human Behavior in Video
abstract
Current diagnostic methods for Alzheimer’s disease (AD), such as brain imaging and cognitive impairment questionnaires, are costly and time-consuming, making early detection challenging. In this work, we develop a computer vision-based method to detect AD using behavioral data collected from the Timed Up and Go (TUG) test and the Cookie Theft (CT) picture description task. By analyzing body joints and facial landmarks through Convolutional Neural Networks and Support Vector Machines, we classified subjects into AD and Non-AD categories across four subtasks: Walking, Sit-Stand, Turning, and Describing. Our approach achieved an F1-score of 0.92±0.03, demonstrating the potential of video-based analysis for AD detection. To enhance the explainability of our model, we applied model explanation methods, identifying key features and symptoms involved in the decision-making process.
Bao-Hsuan Huang, Po-Chih Kuo, Likai Huang, Chaur-Jong Hu
ICASSP2
2025 Mask Augmentation For Tumor Classification In Medical Images
abstract
Tumor detection and classification in medical images are critical for guiding patient management and treatment decisions. However, accurate segmentation and classification of tumors remain challenging due to their small size relative to the overall image. Existing approaches often face difficulties with limited data and potential segmentation errors, resulting in suboptimal performance in real-world applications. To address these challenges, we propose a novel approach leveraging the Segmentation Mask Augmentation (SMA) framework. Our framework enhances the robustness of tumor classification models by generating diverse and imprecise segmentation masks during training, thereby simulating real-world scenarios. Experimental results across four distinct datasets demonstrate the effectiveness of our approach. Our framework presents a promising solution for robust tumor classification, with potential implications for improving clinical diagnosis and patient management.
Chun-Chieh Weng, Huan-Yu Chen, Jing-Tong Tzeng, Ching-Heng Lin, Po-Chih Kuo, Chi-Chun Lee
ICASSP5
2025 Learnable Frequency-Weighting Layer for Improving EEG-Based BCI Performance
abstract
This paper presents a learnable frequency-weighting layer for enhancing EEG-based Brain-Computer Interface (BCI) performance, particularly in higher-order cognitive tasks. The proposed layer, integrating wavelet transform with convolution operations, selectively emphasizes task-relevant frequency components while suppressing those associated with noise or unrelated brain activities. We evaluated our method on two datasets—a local dataset focusing on spatial perspective-taking tasks and the Cog-BCI dataset involving the N-back task, a standard paradigm for assessing working memory load. Empirical results indicate that incorporating the learnable frequency-weighting layer into commonly used deep learning models (EEGNet-V1, EEGNet-V4, ShallowConvNet, and DeepConvNet) consistently yields significant improvements in classification accuracy. Notably, our approach demonstrates effectiveness in tasks exhibiting similar brain-wave patterns, underscoring the adaptability of the method to a wide range of EEG-based BCI applications. Overall, this study provides a robust and generalizable approach to improving EEG-based BCI systems, enhancing their practicality in real-world cognitive applications.
Tsu-Jen Ding, Hui-Yu Hsu, Po-Chih Kuo, Zai-Fu Yao, Cheng-Yu Pan, Bo-Yu Pan
SMC3
2024 Dementia Assessment Using Mandarin Speech with an Attention-Based Speech Recognition Encoder
abstract
Dementia diagnosis requires a series of different testing methods, which is complex and time-consuming. Early detection of dementia is crucial as it can prevent further deterioration of the condition. This paper utilizes a speech recognition model to construct a dementia assessment system tailored for Mandarin speakers during the picture description task. By training an attention-based speech recognition model on voice data closely resembling real-world scenarios, we have significantly enhanced the model’s recognition capabilities. Subsequently, we extracted the encoder from the speech recognition model and added a linear layer for dementia assessment. We collected Mandarin speech data from 99 subjects and acquired their clinical assessments from a local hospital. We achieved an accuracy of 92.04% in Alzheimer’s disease detection and a mean absolute error of 9% in clinical dementia rating score prediction.1
Zih-Jyun Lin, Yi-Ju Chen, Po-Chih Kuo, Likai Huang, Chaur-Jong Hu
ICASSP3
2024 Representing scents: An evaluation framework of scent-related experiences through associations between grounded and psychophysiological data
Yang Chen Lin, Shang-Lin Yu, An-Yu Zhuang, Chiayun Lee, Yao An Ting, Sheng-Kai Lee, Bo-Jyun Lin, Po-Chih Kuo
Int. J. Hum. Comput. Stud.8
2018 Educational Model Based on Hands-on Brain-Computer Interface: Implementation of Music Composition Using EEG
abstract
Technological advancements are producing profound changes in educational models. This has principally manifested as a move away from textbooks to hands-on instruction, which has been shown to make learning more enjoyable and effective. Brain-computer interfaces (BCIs) based on electroencephalography (EEG) have been used in medical applications and as controllers in video games; however, these devices have not yet been used for educational purposes at the high-school level due to the cost of these devices and the lack of teachers trained in their use. In this study, we developed a hands-on educational model in which students learn how to design and build their own BCI system. They also learn to program the device through the analysis of EEG data and transform the signals into suitable outputs. Finally, the students implement their BCI system by measuring their own brain activity. This paper also examines a practical implementation of the system involving an EEG-generated composition of music. A questionnaire survey of students indicated that they enjoyed the hands-on BCI course and were highly motivated to actively engage in it.
Pei-Chi Hu, Po-Chih Kuo
SMC3