VLDB 2026 Research / reviewers in the wild / expert
Linhui Sun
dblp:66/7278
· DBLP profile ↗
12ranked-venue papers
8as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KCAM-SENet: Speech enhancement network with KAN-based channel attention module
Linhui Sun, Zhaowei Ding, Yuhang Qin, Shengchen Li, Xi Shao, Chng Eng Siong |
Speech Commun. | 1 |
| 2024 | Research on the influences of information presentation on information capture and performance in digital control systemsabstractThis study investigates the effects of information presentation on operator performance in digital control systems at different levels of task difficulty in order to resolve the conflict between ‘large amounts of information’ and ‘limited display’. By using software to present figures of varying difficulty as target information, the operator’s behaviour such as page switching and target clicking is simulated to obtain data. Simultaneously, eye-tracking technique was used to analyse how the way information was presented at different levels of difficulty influenced the subjects’ search performance. The results show that subjects’ performance gradually decreases as the difficulty of the numbers grows for both presentations, and it is found that the complexity of the content of the navigation bar may have influenced the subjects to some extent. In addition, the different position of the navigation bar where the target is located will result in a change in the search pattern. It is noticeable that dividing the single page into double pages can effectively alleviate the contradiction of ‘large amount of information’ and ‘limited display’ in designing the system interface. Therefore, this study provides certain guideline value for the interface design of digital control systems. Linhui Sun, Zigu Guo, Xiaofang Yuan, Xinping Wang |
Behav. Inf. Technol. | 1 |
| 2024 | Dual-Branch Modeling Based on State-Space Model for Speech EnhancementabstractTraditional time-frequency domain speech enhancement methods either only enhance the amplitude spectral features without changing the phase that contributes to the naturalness, intelligibility and harmonic structure, or improve the estimation of the complex spectral features including the real and imaginary components, which limits the accuracy of amplitude and phase estimation. To address this issue, we propose a joint dual-branch structured state-space model that leverages the strengths of both branches while keeping computational complexity low. Specifically, we introduce interaction modules between the two branches to facilitate information exchange, enabling features learned from one branch to compensate for missing parts in the other. Furthermore, to reduce model complexity, we introduce the diagonal version of structured state-space sequence (S4D) model for speech feature sequence denoising in both branches. Experimental results show that our low-complexity model achieves significant improvements over previous advanced systems on VoiceBank+DEMAND and TIMIT+NOISE92 datasets. Linhui Sun, Aifei Gong, Chng Eng Siong |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Asynchronous Event Processing with Local-Shift Graph Convolutional NetworkabstractEvent cameras are bio-inspired sensors that produce sparse and asynchronous event streams instead of frame-based images at a high-rate. Recent works utilizing graph convolutional networks (GCNs) have achieved remarkable performance in recognition tasks, which model event stream as spatio-temporal graph. However, the computational mechanism of graph convolution introduces redundant computation when aggregating neighbor features, which limits the low-latency nature of the events. And they perform a synchronous inference process, which can not achieve a fast response to the asynchronous event signals. This paper proposes a local-shift graph convolutional network (LSNet), which utilizes a novel local-shift operation equipped with a local spatio-temporal attention component to achieve efficient and adaptive aggregation of neighbor features. To improve the efficiency of pooling operation in feature extraction, we design a node-importance based parallel pooling method (NIPooling) for sparse and low-latency event data. Based on the calculated importance of each node, NIPooling can efficiently obtain uniform sampling results in parallel, which retains the diversity of event streams. Furthermore, for achieving a fast response to asynchronous event signals, an asynchronous event processing procedure is proposed to restrict the network nodes which need to recompute activations only to those affected by the new arrival event. Experimental results show that the computational cost can be reduced by nearly 9 times through using local-shift operation and the proposed asynchronous procedure can further improve the inference efficiency, while achieving state-of-the-art performance on gesture recognition and object recognition. Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
AAAI | 1 |
| 2023 | Unsupervised Noise Adaptation Using Data SimulationabstractDeep neural network based speech enhancement approaches aim to learn a noisy-to-clean transformation using a supervised learning paradigm. However, such a trained-well transformation is vulnerable to unseen noises that are not included in training set. In this work, we focus on the unsupervised noise adaptation problem in speech enhancement, where the ground truth of target domain data is completely unavailable. Specifically, we propose a generative adversarial network based method to efficiently learn a converse clean-to-noisy transformation using a few minutes of unpaired target domain data. Then this transformation is utilized to generate sufficient simulated data for domain adaptation of the enhancement model. Experimental results show that our method effectively mitigates the domain mismatch between training and test sets, and surpasses the best baseline by a large margin. Chen Chen 0075, Heqing Zou, Linhui Sun, Chng Eng Siong |
ICASSP | 4 |
| 2022 | Action Representing by Constrained Conditional Mutual Information
Haoyuan Gao, Yifaan Zhang, Linhui Sun |
ACCV (4) | 3 |
| 2022 | MENet: A Memory-Based Network with Dual-Branch for Efficient Event Stream Processing
Linhui Sun, Yifan Zhang 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu |
ECCV (24) | 1 |
| 2020 | PEAN: 3D Hand Pose Estimation Adversarial NetworkabstractDespite recent emerging research attention, 3D hand pose estimation still suffers from the problems of predicting inaccurate or invalid poses which conflict with physical and kinematic constraints. To address these problems, we propose a novel 3D hand pose estimation adversarial network (PEAN) which can implicitly utilize such constraints to regularize the prediction in an adversarial learning framework. PEAN contains two parts: a 3D hierarchical estimation network (3DHNet) to predict hand pose, which decouples the task into multiple subtasks with a hierarchical structure; a pose discrimination network (PDNet) to judge the reasonableness of the estimated 3D hand pose, which back-propagates the constraints to the estimation network. During the adversarial learning process, PDNet is expected to distinguish the estimated 3D hand pose and the ground truth, while 3DHNet is expected to estimate more valid pose to confuse PDNet. In this way, 3DHNet is capable of generating 3D poses with accurate positions and adaptively adjusting the invalid poses without additional prior knowledge. Experiments show that the proposed 3DHNet does a good job in predicting hand poses, and introducing PDNet to 3DHNet does further improve the accuracy and reasonableness of the predicted results. As a result, the proposed PEAN achieves the state-of-the-art performance on three public hand pose estimation datasets. Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu |
ICPR | 1 |
| 2020 | Furion: Engineering High-Quality Immersive Virtual Reality on Today's Mobile DevicesabstractDespite the growing market penetration, today's high-end virtual reality (VR) systems remain tethered, which not only limits users' VR experience but also creates a safety hazard. In this paper, we perform a systematic design study of the “elephant in the room” facing the VR industry - is it feasible to enable high-quality VR apps on untethered mobile devices such as smartphones? Our quantitative, performance-driven design study makes two contributions. First, we show that the QoE achievable for high-quality VR applications on today's mobile hardware and wireless networks via local rendering or offloading is about 10X away from the acceptable QoE, yet waiting for future mobile hardware or next-generation wireless networks (e.g., 5G) is unlikely to help, because of power limitation and the higher CPU utilization needed for processing packets under higher data rate. Second, we present Furion, a VR framework that enables high-quality, immersive mobile VR on today's mobile devices and wireless networks. Furion exploits a key insight about the VR workload that foreground interactions and background environment have contrasting predictability and rendering workload, and employs a split renderer architecture running on both the phone and the server. Supplemented with video compression, use of panoramic frames, parallel decoding on multiple cores on the phone, and view-based bitrate adaptation we demonstrate Furion can support high-quality VR apps on today's smartphones over WiFi, with under 14 ms latency and 60 FPS (the phone display refresh rate). Zeqi Lai, Y. Charlie Hu, Yong Cui 0001, Linhui Sun, Ningwei Dai, Hung-Sheng Lee |
IEEE Trans. Mob. Comput. | 4 |
| 2019 | Joint dictionary learning using a new optimization method for single-channel blind source separation
Linhui Sun, Keli Xie, Ting Gu, Zhen Yang 0001 |
Speech Commun. | 1 |
| 2019 | Speech emotion recognition based on DNN-decision tree SVM model
Linhui Sun, Sheng Fu |
Speech Commun. | 1 |
| 2017 | Furion: Engineering High-Quality Immersive Virtual Reality on Today's Mobile DevicesabstractIn this paper, we perform a systematic design study of the "elephant in the room" facing the VR industry -- is it feasible to enable high-quality VR apps on untethered mobile devices such as smartphones? Our quantitative, performance-driven design study makes two contributions. First, we show that the QoE achievable for high-quality VR applications on today's mobile hardware and wireless networks via local rendering or offloading is about 10X away from the acceptable QoE, yet waiting for future mobile hardware or next-generation wireless networks (e.g. 5G) is unlikely to help, because of power limitation and the higher CPU utilization needed for processing packets under higher data rate. Second, we present Furion, a VR framework that enables high-quality, immersive mobile VR on today's mobile devices and wireless networks. Furion exploits a key insight about the VR workload that foreground interactions and background environment have contrasting predictability and rendering workload, and employs a split renderer architecture running on both the phone and the server. Supplemented with video compression, use of panoramic frames, and parallel decoding on multiple cores on the phone, we demonstrate Furion can support high-quality VR apps on today's smartphones over WiFi, with under 14ms latency and 60 FPS (the phone display refresh rate). Zeqi Lai, Y. Charlie Hu, Yong Cui 0001, Linhui Sun, Ningwei Dai |
MobiCom | 4 |