VLDB 2026 Research / reviewers in the wild / expert
Richard H. Y. So
dblp:322/5265 · also Richard Hau Yue So
· DBLP profile ↗
12ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-3595-0386ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Novel Weighted Sparse Component Analysis for Underdetermined Blind Speech SeparationabstractSparse component analysis (SCA) is a popular underdetermined blind speech separation (UBSS) method. It models all sources to have an identical distribution. As speeches do not have identical distribution, SCA performs suboptimal. Some studies have improved the performance of SCA by weighting the sources through a reweighting scheme. However, they are not UBSS methods because they assume that the mixing process is known. This paper proposes a novel weighting scheme, called sparse spatial component analysis (SSCA) without the need to know the mixing process. In SSCA, weights, sources, and the parameters for modeling the mixing process are jointly optimized, making it a UBSS method. Simulation experiments show that for instantaneous mixtures, SSCA outperforms SCA and reweighted SCA, improving the source-to-distortion ratio (SDR) by 4 dB and reducing the computational time by 40%. Further, experiments using real-world recordings reveal that SSCA outperforms multichannel non-negative matrix factorization and full-rank covariance analysis (FCA) in terms of SDR. The speed of SSCA is 200% faster than FCA. Yudong He, Baeck Hyun Woo, Richard H. Y. So |
ICASSP | 3 |
| 2025 | Evaluating the Efficacy of Pulse Transit Time between Palm and Forehead in Blood Pressure EstimationabstractTraditional cuff-based blood pressure (BP) measurement techniques, though widely employed, require a cuff that must be fitted and inflated and thus are limited by convenience issues. In contrast, contactless BP monitoring solutions offer a promising alternative. This study explored pulse transit time (PTT) as a feature for accurate BP monitoring using remote photoplethysmography (rPPG). The investigation examined the order of PTT (PTT Order) between the palm and forehead and their impacts on the accuracy of BP estimation. Our findings showed variation in the dominant order between the two sites among the subjects. Nevertheless, the inverse of mean PTT extracted from the two sites in dominant order (PTT Dominant Order) consistently showed a higher linear correlation with systolic blood pressure (SBP). The mean and standard deviation of R-squared derived from the inverse of mean PTT with the dominant order and SBP among the 16 subjects were 0.81 ± 0.13. Additionally, subgroup analysis identified significant differences in SBP across gender and exercise status. Furthermore, our data revealed a hysteresis phenomenon in 25% of the subjects, characterized by SBP returning to baseline levels during post-exercise resting while heart rate (HR) remained persistently elevated. Chuchu Qiu, Jing Wei Chin, Tsz Tai Chan, Kwan Long Wong, Richard H. Y. So |
ICMI | 5 |
| 2025 | Remote Blood Pressure Estimation from Facial Videos Using Transfer Learning: Leveraging PPG to rPPG ConversionabstractBlood pressure (BP) monitoring is crucial for health assessment, but existing contact-based methods face cost and comfort barriers. Remote photoplethysmography (rPPG) offers a promising contactless solution, yet research is hampered by limited rPPG datasets with corresponding BP labels. This paper presents a transfer learning methodology for BP measurement. This approach involves utilizing a base dataset comprising signals produced by PPG (PPG-signals) to acquire knowledge that can be transferred to a target dataset containing signals generated by rP PG (rPPG-signals). In our study, we trained diverse deep-learning models using publicly available datasets containing PPG-signals. Subsequently, these models were fine-tuned and evaluated using a public dataset that specifically consists of rP PG-signals. Additionally, we explored the relationship between BP and heart rate, and examined different loss functions and normalization approaches to optimize the performance of the deep learning models. The findings of our study demonstrate that our best model achieved a better performance than the state-of-the-art model, with mean absolute error (MAE) of 8.721 (reduced by 4.879) mmHg and 8.653 (reduced by 1.647) mmHgfor systolic blood pres-sure (SBP) and diastolic blood pressure (DBP) in a dataset with clinical settings, showing promising potential for remote BP estimation. Chun-Hong Cheng, Jing Wei Chin, Kwan Long Wong, Tsz Tai Chan, Hau Ching Lo, Kwan Lok Pang, Richard H. Y. So, Bryan Yan |
WACV | 7 |
| 2025 | A novel property to modify weighted l1 minimization for improved compressed sensingabstractWeighted l1 minimization schemes are common methods to achieve compressed sensing (CS). However, they fail in the presence of inaccurate prior knowledge or improper scaling of weights due to inappropriately assigned large weights causing large and destructive errors in signal recovery. This paper proposes a theory-based algorithm to identify and correct such destructive weights for each signal entry. The enhancement is achieved through a novel sparsity-inducing property (SIP) which establishes a necessary condition for successful signal recovery. SIP outperforms existing properties such as coherence, restricted isometry property, and nullspace property by indicating which signal entries fail to be recovered. This unique advantage enables us to correct destructive weights that do not satisfy the SIP condition, making signal recovery successful where it previously failed. Results from many numerical experiments demonstrate that our proposed method can improve the signal recovery capability, robustness, and stability of the weighted l1 minimization for a wide range of applications, including sparse and compressive signal recovery, noise-aware recovery, sparse error correction, fast image acquisition, and sub-Nyquist sampling. Yudong He, Baeck Hyun Woo, Fauzan Abdurrahim, Richard H. Y. So |
Signal Process. | 4 |
| 2023 | Predicting Subjective Discomfort Associated With Lens Distortion in VR Headsets During Vestibulo-Ocular Response to VR ScenesabstractWith advances in Virtual Reality (VR) technology, user expectation for a near-perfect experience is also increasing. The push for a wider field-of-view can increase the challenges of correcting lens distortion. Past studies on imperfect VR experiences have focused on motion sickness provoked by vection-inducing VR stimuli and discomfort due to mismatches in accommodation and binocular convergence. Disorientation and discomfort due to unintended optical flow induced by lens distortion, referred to as dynamic distortion (DD), has, to date, received little attention. This study examines and models the effects of DD during head rotations with various fixed gazes stabilized by vestibulo-ocular reflex (VOR). Increases in DD levels comparable to lens parameters from poorly designed commercial VR lenses significantly increase discomfort scores of viewers in relation to disorientation, dizziness, and eye strain. Cross-validated results indicate that the model is able to predict significant differences in subjective scores resulting from different commercial VR lenses and these predictions correlated with empirical data. The present work provides new insights to understand symptoms of discomfort in VR during user interactions with static world-locked / space-stabilized scenes and contributes to the design of discomfort-free VR headset lenses. Tsz Tai Chan, Richard H. Y. So, Jerry Jia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Harvesting Partially-Disjoint Time-Frequency Information for Improving Degenerate Unmixing Estimation TechniqueabstractThe degenerate unmixing estimation technique (DUET) is one of the most efficient blind source separation algorithms tackling the challenging situation when the number of sources exceeds the number of microphones. However, as a time-frequency mask-based method, DUET erroneously results in interference components retention when source signals overlap each other in both frequency and time domains. In this paper, to avoid the erroneous retention, instead of masking, we propose to use multiple linear spatial filters (e.g., the minimum variance distortionless response filter) to extract the desired signals. These filters are constructed based on the information embedded in the detected single-source-points, that is, time-frequency points contributed by a single source. In comparison with the conventional DUET, our method achieved an impressive improvement greater than 5 dB in the source-to-interference ratio and 2 to 5 dB improvement in the source-to-distortion ratio, respectively. Findings are substantiated by unmixing simulation using live-recorded mixture signals from up to four sources. Audio examples can be found on the web page: "https://ydcnanhe.github.io/demoicassp2022/" Yudong He, Qifeng Chen 0001, Richard H. Y. So |
ICASSP | 4 |
| 2021 | Active head rolls enhance sonar-based auditory localization performanceabstractAnimals utilize a variety of active sensing mechanisms to perceive the world around them. Echolocating bats are an excellent model for the study of active auditory localization. The big brown bat (Eptesicus fuscus), for instance, employs active head roll movements during sonar prey tracking. The function of head rolls in sound source localization is not well understood. Here, we propose an echolocation model with multi-axis head rotation to investigate the effect of active head roll movements on sound localization performance. The model autonomously learns to align the bat's head direction towards the target. We show that a model with active head roll movements better localizes targets than a model without head rolls. Furthermore, we demonstrate that active head rolls also reduce the time required for localization in elevation. Finally, our model offers key insights to sound localization cues used by echolocating bats employing active head movements during echolocation. Lakshitha P. Wijesinghe, Melville Wohlgemuth, Richard H. Y. So, Jochen Triesch, Cynthia F. Moss, Bertram E. Shi |
PLoS Comput. Biol. | 3 |
| 2020 | Truth-to-Estimate Ratio Mask: A Post-Processing Method for Speech Enhancement Direct at Low Signal-to-Noise RatiosabstractThis study proposes a bi-directional recurrent neural network (Bi-RNN) post-processing method for speech enhancement (SE) at low signal-to noise ratios (SNR). Current speech enhancement solutions performed badly under low SNR situations. Loizou and Kim proposed a solution to reduce speech distortion errors in time-frequency (T-F) domain but it requires the knowledge of ground truth. As ground truth is unknown in real-life applications, the current study proposes to use a Bi-RNN to implement Loizou and Kim's solution as a post-processing method for SE engines. Our solutions do not require prior knowledge of ground truth. The effectiveness of the proposed method is investigated with a spectral subtraction (SS) SE engine, a non-negative matrix factorization (NMF) SE engine, and a deep neural network ideal ratio mask (DNN-IRM) SE engine, under matched/mis-matched noise and different SNR conditions. Experimental results demonstrate that the proposed post-processing method effectively improved both perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI) for all of these SE engines, especially at low SNR conditions. Richard H. Y. So |
ICASSP | 4 |
| 2019 | Effects of Base-Frequency and Spectral Envelope on Deep-Learning Speech Separation and Recognition Models
Jun Hui, Shutao Chen, Richard H. Y. So |
INTERSPEECH | 4 |
| 2011 | Could OKAN be an objective indicator of the susceptibility to visually induced motion sickness?abstractInternational Workshop Agreement 3 organized by the International Standard Organization calls for more research to determine simple objective ways to assess susceptibility to visually induce motion sickness (VIMS) without making viewers sick (So and Ujike, 2010). This study examines the use of measurable optokinetic afternystagmus (OKAN) parameters to predict susceptibility to VIMS. Eighteen participants were recruited. They were exposed to a sickness provoking virtual rotating drum (210 degrees field-of-view) with striped patterns rotating at 60 degrees per second for 30 minutes (Phase 1). Sickness data were collected before, during, and after the exposure. These participants were invited back for OKAN measurements at least two weeks after Phase 1 was completed to minimize any adaption effect (Phase 2). Out of the 18 participants, 10 participants (i.e., 55%) exhibited consistent patterns of OKAN. Correlations between the time constants of OKAN and levels of VIMS experienced by the same viewers were found. The possibility of using OKAN as an objective indicator of the susceptibility to visually induced motion sickness is discussed. Cuiting Guo, Jennifer T. T. Ji, Richard H. Y. So |
VR | 3 |
| 1999 | Cybersickness: An Experimental Study to Isolate the Effects of Rotational Scene OscillationsabstractHead-coupled virtual reality systems can cause symptoms of sickness (cybersickness). A study has been conducted to investigate the effects of scene oscillations on the level and types of cybersickness. Sixteen male subjects participated in the experiments. They were exposed to four 20-minute virtual simulation sessions, in a balanced order with 10 days separation. The 4 simulation sessions exposed the subjects to similar visual scene oscillation in different axis: pitch axis, yaw axis, roll axis and no oscillation (speed: 30/spl deg//s, range: +/-60/spl deg/). Verbal ratings of nausea level were taken at 5-minute intervals and sickness symptoms were measured before and after the exposure using the Simulator Sickness Questionnaire (SSQ). Significant differences were found between the no oscillation condition and the oscillating conditions. With scene oscillation, nausea ratings increased significantly after 5-minute exposure for all the oscillation axes (pitch, yaw, and roll axes). Total sickness scores were obtained from the SSQ and their profiles with different scene oscillation axes were presented. Richard H. Y. So, W. T. Lo |
VR | 1 |
| 1996 | Experimental studies of the use of phase lead filters to compensate lags in head-coupled visual displaysabstractDisplay lags degrade performance when using the head to track a target presented on a helmet-mounted display. These lags originate from delays in measuring the position of the head and the time required to generate the image of the target. This paper presents two laboratory studies on the use of phase lead filters to improve head tracking performance in the presence of display lags. In the preliminary study, the benefits of lag compensation by a phase lead filter were impeded by associated changes in filter gain. The frequency responses of two phase lead filters were then optimized to have near unity gain at frequencies below 0.7 Hz where there was most head motion. The main study showed that these optimized filters significantly improved head tracking performance with a system having a total lag of 140 ms. At frequencies above about 0.7 Hz, a greater than unity filter gain caused jittery image movement. Although this jittering degraded head tracking performance it was removed by an alternative lag compensation technique involving 'image deflection'. This deflection shifted the displayed image to its correct horizontal and vertical position relative to the head. Image deflection, combined with the phase lead filters, produced a tracking performance unaffected by lag. Richard H. Y. So, Michael J. Griffin |
IEEE Trans. Syst. Man Cybern. Part A | 1 |