VLDB 2026 Research / reviewers in the wild / expert
Shuangquan Wang
dblp:23/5003
· DBLP profile ↗
26ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-3967-9693ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mobile-StereoHPE: Real-Time Mobile 3D Hand Pose Estimation from Stereo Gray ImagesabstractMobile XR scenarios pose significant challenges for 3D bare-hand interaction, requiring accurate and low-latency 3D hand tracking within strict computational constraints. Monocular RGB-based methods struggle with estimating absolute 3D hand joints, while depth sensor or large-model approaches are unsuitable for mobile devices. We propose Mobile-StereoHPE, a lightweight and efficient framework for absolute 3D hand pose estimation using stereo gray images. Our approach introduces two novel components: a Feature Fusion Module (FFM) that aggregates stereo image features while minimizing computational overhead, and a Cross Feature Attention (CFA) module that enhances inter-view correspondence learning. These innovations empower our compact CNN backbone to achieve state-of-the-art accuracy on the MVHand benchmark with the smallest model size, a 33% improvement in 3D joint accuracy, and more than twice the speed of current large-model methods. Extensive experiments validate Mobile-StereoHPE as an ideal solution for next-generation XR applications. Dongfang Zhao 0017, Menghe Zhang, Yangwen Liang, Shuangquan Wang, Kee-Bong Song |
ICME | 4 |
| 2025 | 3D-AMTA: Occlusion-Aware Real-Time 3D Hand Pose Estimation with Auto Mask and Token-Specific AttentionabstractUnderstanding hand motion from a single RGB image is challenging due to occlusions and high articulation. This paper presents 3D-AMTA, a transformer-based framework with Auto Mask and Token-specific Attention for occlusion-aware 3D hand pose estimation (HPE). We propose two novel architectural enhancements: auto mask for high-occlusion scenarios, and token-specific attention for fine-grained hand articulations. These modules seamlessly integrate into transformer-based architectures that enhance real-time performance in interactive systems. To enable efficient deployment on robotic and embedded platforms, we propose 3D-AMTA-Mobile, a lightweight variant optimized for on-device processing. It achieves 267 FPS on NVIDIA RTX 2080Ti-GPU while maintaining high accuracy, making it well-suited for resource-constrained robotic applications. Extensive evaluations on FreiHAND and HO3D demonstrate that our approach consistently outperforms state-of-the-art methods in terms of accuracy, efficiency, and inference speed. These advancements contribute to robust hand perception for interactive robotics and AR-based teleoperation. Dongfang Zhao 0017, Menghe Zhang, Yangwen Liang, Shuangquan Wang, Kee-Bong Song |
IROS | 4 |
| 2025 | Leveraging Latent Diffusion in 3D Gaussian Splatting for Novel View Synthesis
Yangwen Liang, Shuangquan Wang, Kee-Bong Song |
MMM (5) | 4 |
| 2025 | Towards Recognizing Food Types for Unseen SubjectsabstractRecognizing food types through sensor signals for unseen users remains remarkably challenging despite extensive recent studies. The efficacy of prior machine learning techniques is dwarfed by giant variations of data collected from multiple participants, partly because users have varied chewing habits and wear sensor devices in various manners. This work treats the problem as an instance of the domain adaptation problem, where each user represents a domain. We develop the first multi-source domain adaptation (MSDA) method for food-typing recognition, which consists of three major components: stratified normalization, a multi-source domain adaptor, and adaptive ensemble learning. New techniques are developed for each component. Using a real-world dataset comprised of 15 participants, we demonstrate that our method achieves \(1.33\times\) to \(2.13\times\) improvement in accuracy compared with nine state-of-the-art MSDA baselines. Additionally, we perform an in-depth ablation study to examine the behavior of each component and confirm its efficacy. Jiexiong Guan, Wei Niu 0002, Shuangquan Wang, Zhenming Liu, Gang Zhou 0002, Bin Ren 0002 |
ACM Trans. Comput. Heal. | 5 |
| 2024 | Temporally-Consistent Video Semantic Segmentation with Bidirectional Occlusion-guided Feature PropagationabstractDespite recent progress in static image segmentation, video segmentation is still challenging due to the need for an accurate, fast, and temporally consistent model. Conducting per-frame static image segmentation on a video is not acceptable since it is computationally prohibitive and prone to temporal inconsistency. In this paper, we present bidirectional occlusion-guided feature propagation (BOFP) method with the goal of improving temporal consistency of segmentation results without sacrificing segmentation accuracy, while at the same time keeping the operations at a low computation cost. It leverages temporal coherence in the video by feature propagation from keyframes to other frames along the motion paths in both forward and backward directions. We propose an occlusion-based attention network to estimate the distorted areas based on bidirectional optical flows, and utilize them as cues for correcting and fusing the propagated features. Extensive experiments on benchmark datasets demonstrate that the proposed BOFP method achieves superior performance in terms of temporal consistency while maintaining comparable level of segmentation accuracy at a low computation cost, striking a great balance among the three performance metrics essential to evaluate video segmentation solutions. Razieh Kaviani Baghbaderani, Shuangquan Wang, Hairong Qi 0001 |
WACV | 3 |
| 2022 | 3D Texture Super Resolution via the Rendering LossabstractDeep learning-based methods have made significant impact and demonstrated superior performance for the classical image and video super-resolution (SR) tasks. Yet, deep learning-based approaches to super-resolve the appearance of 3D objects are still sparse. Due to the nature of rendering 3D models, 2D SR methods applied directly to 3D object texture may not be a good approach. In this paper, we propose a rendering loss derived from the rendering of a 3D model and demonstrate its application to the SR task in the context of 3D texturing. Unlike other literature on the 3D appearance SR, no geometry information of the 3D model is required during network inference. Experimental results demonstrate that incorporating the rendering loss during network training outperforms existing state-of-the-art methods for 3D appearance SR. Furthermore, we provide a new 3D dataset consisting of 97 complete 3D models for further research in this field. Rohit Ranade, Yangwen Liang, Shuangquan Wang, Dongwoon Bai |
ICASSP | 3 |
| 2022 | Towards Socially Acceptable Food Type RecognitionabstractAutomatic food type recognition is an essential task of dietary monitoring. It helps medical professionals recognize a user's food contents, estimate the amount of energy intake, and design a personalized intervention model to prevent many chronic diseases, such as obesity and heart disease. Various wearable and mobile devices are utilized as platforms for food type recognition. However, none of them has been widely used in our daily lives and, at the same time, socially acceptable enough for continuous wear. In this paper, we propose a food type recognition method that takes advantage of Airpods Pro, a pair of widely used wireless in-ear headphones designed by Apple, to recognize 20 different types of food. As far as we know, we are the first to use this socially acceptable commercial product to recognize food types. Audio and motion sensor data are collected from Airpods Pro. Then 135 representative features are extracted and selected to construct the recognition model using the lightGBM algorithm. A real-world data collection is conducted to comprehensively evaluate the performance of the proposed method for seven human subjects. The results show that the average f1-score reaches 94.4% for the ten-fold cross-validation test and 96.0% for the self-evaluation test. Jiexiong Guan, Y. Alicia Hong, Shuangquan Wang, Zhenming Liu, Bin Ren 0002, Gang Zhou 0002 |
MSN | 5 |
| 2021 | DVIO: Depth-Aided Visual Inertial Odometry for RGBD SensorsabstractIn past few years we have observed an increase in the usage of RGBD sensors in mobile devices. These sensors provide a good estimate of the depth map for the camera frame, which can be used in numerous augmented reality applications. This paper presents a new visual inertial odometry (VIO) system, which uses measurements from a RGBD sensor and an inertial measurement unit (IMU) sensor for estimating the motion state of the mobile device. The resulting system is called the depth-aided VIO (DVIO) system. In this system we add the depth measurement as part of the nonlinear optimization process. Specifically, we propose methods to use the depth measurement using one-dimensional (1D) feature parameterization as well as three-dimensional (3D) feature parameterization. In addition, we propose to utilize the depth measurement for estimating time offset between the unsynchronized IMU and the RGBD sensors. Last but not least, we propose a novel block-based marginalization approach to speed up the marginalization processes and maintain the real-time performance of the overall system. Experimental results validate that the proposed DVIO system outperforms the other state-of-the-art VIO systems in terms of trajectory accuracy as well as processing time. Abhishek Tyagi, Yangwen Liang, Shuangquan Wang, Dongwoon Bai |
ISMAR | 3 |
| 2020 | Mag-Barcode: Magnet Barcode Scanning for Indoor Pedestrian TrackingabstractIn typical scenarios for indoor localization and tracking, it is essential to accurately track the pedestrians when they are crossing the connections of different spaces. In this paper, we propose a magnet barcode scanning-based solution for indoor pedestrian tracking. We assemble multiple magnet bars into magnet arrays as a unique magnet barcode, and deploy different magnet barcodes at different connections to label them. We embed an inertial measurement unit (IMU) into the pedestrian`s shoes. When the pedestrian crosses these connections, the magnetometer from the IMU scans the magnet barcode and recognize its corresponding ID. In this way, indoor pedestrian tracking can be regarded as a process of continuously scanning different magnet barcodes. By performing correlation analysis on these barcodes, the trace of pedestrian can be effectively depicted in the indoor map. To build a unique magnet barcode based on the magnet bar arrays, we provide an optimized structure for building the magnet barcode. To tackle the diversities of the pedestrian's gait traces in identifying the magnet barcode, we provide a generalized model based on the space axis for magnet barcode identification. As far as we know, this is the first work to use the magnet bar array to construct the magnet barcode for indoor pedestrian tracking. The real experiment results show that our system can achieve an average accuracy of 88.9% in identifying the magnet barcodes and an average accuracy of 93.1 % for indoor pedestrian tracking. Zefan Ge, Lei Xie 0004, Shuangquan Wang, Xinran Lu, Gang Zhou 0002, Sanglu Lu |
IWQoS | 3 |
| 2019 | TennisEye: tennis ball speed estimation using a racket-mounted motion sensorabstractAggressive tennis shots with high ball speed are the key factor in winning a tennis match. Today's tennis players are increasingly focused on improving ball speed. As a result, in recent tennis tournaments, records of tennis shot speeds are broken again and again. The traditional method for calculating the tennis ball speed uses multiple high-speed cameras and computer vision technology. This method is very expensive and hard to set up. Another way to calculate the tennis ball speed is to use motion sensors, which are lower cost and easier to set up. In this paper, we propose an approach for tennis ball speed estimation based on a racket-mounted motion sensor. We divide the tennis strokes into three categories: serve, groundstroke, and volley. For a serve, a regression model is proposed to estimate the ball speed. For a groundstroke or volley, two models are proposed: a regression model and a physical model. We use the physical model to estimate the ball speed for advanced players and the regression model for beginner players. Under the leave-one-subject-out cross-validation test, evaluation results show that TennisEye is 10.8% more accurate than the state-of-the-art work. Hongyang Zhao, Shuangquan Wang, Gang Zhou 0002, Woosub Jung |
IPSN | 2 |
| 2017 | BrainStorm: a psychosocial game suite design for non-invasive cross-generational cognitive capabilities data collectionabstractCurrently available traditional as well as videogame-based cognitive assessment techniques are inappropriate due to several reasons. This paper presents a novel psychosocial game suite, BrainStorm, for non-invasive cross-generational cognitive capabilities data collection, which additionally provides cross-generational social support. A motivation behind the development of presented game suite is to provide an entertaining and exciting platform for its target users in order to collect gameplay-based cognitive capabilities data in a non-invasive manner. An extensive evaluation of the presented game suite demonstrated high acceptability and attraction for its target users. Besides, the data collection process is successfully reported as transparent and non-invasive. Yiqiang Chen 0001, Lisha Hu, Shuangquan Wang, Jindong Wang 0001, Zhenyu Chen 0003, Xinlong Jiang, Jianfei Shen |
J. Exp. Theor. Artif. Intell. | 4 |
| 2017 | Continuous Authentication With Touch Behavioral Biometrics and Voice on Wearable GlassesabstractWearable glasses are on the rising edge of development with great user popularity. However, user data stored on these devices bring privacy risks to the owner. To better protect the owner's privacy, a continuous authentication system is needed. In this paper, we propose a continuous and noninvasive authentication system for wearable glasses, named GlassGuard. GlassGuard discriminates the owner and an impostor with behavioral biometrics from six types of touch gestures (single-tap, swipe forward, swipe backward, swipe down, two-finger swipe forward, and two-finger swipe backward) and voice commands, which are all available during normal user interactions. With data collected from 32 users on Google Glass, we show that GlassGuard achieves 99% detection rate and 0.5% false alarm rate after 3.5 user events on average when all types of user events are available with equal probability. Under five typical usage scenarios, the system has a detection rate above 93% and a false alarm rate below 3% after less than five user events. Ge Peng, Gang Zhou 0002, David T. Nguyen, Xin Qi 0001, Qing Yang 0005, Shuangquan Wang |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2016 | A Study of Players' Experiences During Brain Games Play
Yiqiang Chen 0001, Shuangquan Wang, Zhenyu Chen 0003, Jianfei Shen, Lisha Hu, Jindong Wang 0001 |
PRICAI | 3 |
| 2015 | Unobtrusive Sensing Incremental Social Contexts Using Fuzzy Class Incremental LearningabstractBy utilizing captured characteristics of surrounding contexts through widely used Bluetooth sensor, user-centric social contexts can be effectively sensed and discovered by dynamic Bluetooth information. At present, state-of-the-art approaches for building classifiers can basically recognize limited classes trained in the learning phase; however, due to the complex diversity of social contextual behavior, the built classifier seldom deals with newly appeared contexts, which results in degrading the recognition performance greatly. To address this problem, we propose, an OSELM (online sequential extreme learning machine) based class incremental learning method for continuous and unobtrusive sensing new classes of social contexts from dynamic Bluetooth data alone. We integrate fuzzy clustering technique and OSELM to discover and recognize social contextual behaviors by real-world Bluetooth sensor data. Experimental results show that our method can automatically cope with incremental classes of social contexts that appear unpredictably in the real-world. Further, our proposed method have the effective recognition capability for both original known classes and newly appeared unknown classes, respectively. Zhenyu Chen 0003, Yiqiang Chen 0001, Xingyu Gao 0001, Shuangquan Wang, Lisha Hu, Chenggang Yan 0001, Nicholas D. Lane, Chunyan Miao |
ICDM | 4 |
| 2015 | Poster: A Continuous and Noninvasive User Authentication System for Google GlassabstractNo abstract available. Ge Peng, David T. Nguyen, Gang Zhou 0002, Shuangquan Wang |
MobiSys | 4 |
| 2014 | A Theoretical Analysis of Path Loss Based Activity RecognitionabstractBody area networks are used extensively in the medical field and elderly care. These networks perform a collection of roles including monitoring an individual's activities, through a process known as activity recognition. Current approaches to activity recognition require specialized hardware for advanced sensors which puts stress on battery life. We show that it is theoretically possible to distinguish between different human activities/postures by using radio signal propagation only and provide strategies for doing this. This is immensely beneficial to the field of sensor networks for two reasons: 1) It removes the need for more energy intensive components thus reducing strain on limited power resources and thus 2) reduces form factor for sensor network nodes. We show that activity recognition can be done using only radio signals at low transmit power levels with good accuracy. Lower power consumption and reduced form factor are desirable features for body networks and have been identified as being essential for wide adoption of body networks. Iberedem N. Ekure, Shuangquan Wang, Gang Zhou 0002 |
MASS | 2 |
| 2014 | b-COELM: A fast, lightweight and accurate activity recognition model for mini-wearable devices
Lisha Hu, Yiqiang Chen 0001, Shuangquan Wang, Zhenyu Chen 0003 |
Pervasive Mob. Comput. | 3 |
| 2012 | Extreme learning machine-based device displacement free activity recognition model
Yiqiang Chen 0001, Zhongtang Zhao, Shuangquan Wang, Zhenyu Chen 0003 |
Soft Comput. | 3 |
| 2012 | Instantaneous Mutual Information and Eigen-Channels in MIMO Mobile Rayleigh FadingabstractIn this paper, we study two important metrics in multiple-input multiple-output (MIMO) time-varying Rayleigh flat fading channels. One is the eigen-channel, and the other is the instantaneous mutual information (IMI). Their second-order statistics, such as the correlation coefficient, level crossing rate (LCR), and average fade/outage duration, are investigated, assuming a general nonisotropic scattering environment. Exact closed-form expressions are derived and Monte Carlo simulations are provided to verify the accuracy of the analytical results. For the eigen-channels, we found they tend to be spatio-temporally uncorrelated in large MIMO systems. For the IMI, the results show that its correlation coefficient can be well approximated by the squared amplitude of the correlation coefficient of the channel, under certain conditions. Moreover, we also found the LCR of IMI is much more sensitive to the scattering environment than that of each eigen-channel. Shuangquan Wang |
IEEE Trans. Inf. Theory | 1 |
| 2010 | Wearable accelerometer based extendable activity recognition systemabstractRecognizing the human activities of daily living (ADL) is an important research issue in the pervasive environment. Activity recognition is treated as a classification problem and the multi-class classifier is often used. Though the multi-class classifier can obtain high classification accuracy, it can not detect the noise activities and unknown activities, and the system has no extendable recognition capability. In this paper, we proposed a recognition system which can recognize known activities and detect unknown activities simultaneously. For each known activity, one one-class classification model is built up and the combined one-class classification models are used to judge whether a test sample belongs to known activities. For the known samples, the multi-class classifier is used to recognize their types. For the continuous unknown samples, based on segmentation algorithm, training samples of new activities are extracted and added into the recognition system to extend the system's recognition capability. Jie Yang 0002, Shuangquan Wang, Ningjiang Chen |
ICRA | 2 |
| 2009 | Efficient receiver algorithms for DFT-spread OFDM systemsabstractFor the 3GPP LTE uplink transmissions, the DFTspread OFDM technique has been adopted as the air interface in order to reduce the peak-to-average-power ratio (PAPR). In this scheme, each data symbol is spread over many tones by a discrete Fourier transform (DFT) operation at the transmitter before being sent to the orthogonal frequency division multiplexing (OFDM) modulator. Moreover, more than one user can be scheduled over the same frequency and time resource block (RB) via space-division multiple-access (SDMA). The conventional receiver technique for such DFT-spread OFDM systems involves tone-by-tone single-tap equalization followed by an inverse DFT operation. In this paper, we propose a more powerful receiver technique for DFT-spread OFDM systems that consists of an efficient linear pre-filter and a two-symbol soft output demodulator. The proposed method can be applied to both single-user per RB (DFT-S-OFDMA) and multiple users per RB (DFT-S-OFDMSDMA) systems and it offers significant performance gains over the conventional method, especially in the high-rate regime, with little attendant increase in computational complexity. Narayan Prasad, Shuangquan Wang, Xiaodong Wang 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2008 | Quantized Multi-Rank Beamforming for MIMO-OFDM SystemsabstractWe consider the sum-rate maximization via linear preceding in downlink MIMO-OFDM systems with quantized feedback. We address the preceding codebook design based on the capacity measure by introducing a new distance metric. We propose a codebook structure and its associated design algorithm that allows for significant reduction in the memory requirement and computational complexity in real-time system implementation. We then provide a system design approach comprising of four main ingredients: (i) a multi-rank beamforming (MRBF) scheme, (ii) an efficient CQI-based precoder selection algorithm, (iii) reduced feedback strategies, and (iv) novel channel quality indicator (CQI) combining. Our simulation results show that the proposed MRBF scheme can approach the precoding upper bounds with relatively few feedback bits. Moreover, with the same number of bits, the proposed scheme simultaneously achieves higher throughput and lower computational complexity in comparison to the other existing precoding schemes. Mohammad Ali Amir Khojastepour, Narayan Prasad, Shuangquan Wang, Xiaodong Wang 0001, Mohammad Madihian |
IEEE J. Sel. Areas Commun. | 3 |
| 2007 | Envelope Correlation Coefficient for Logarithmic Diversity Receivers RevisitedabstractThe analysis of the envelope correlation coefficient for logarithmic diversity receivers, given by de Neumann (de Neumann, 1989) for one Rayleigh fading branch in isotropic scattering environments (two independent real Gaussian branches), is extended in this letter to the general case where a maximum ratio combiner (MRC) operates on M ges 1 independent Rayleigh fading branches, with the same temporal correlation coefficient, in nonisotropic scattering environments. The derived exact closed-form expressions include the results by B. de Neumann in 1989 as special cases. In addition, when M is not so small, theoretical analysis and Monte Carlo simulations show that the envelope correlation coefficient can be accurately approximated by the squared amplitude of the channel correlation coefficient. Shuangquan Wang |
IEEE Trans. Commun. | 1 |
| 2006 | Statistical Characterization of Eigen-Channels in Time-Varying Rayleigh Flat Fading MIMO SystemsabstractIn this paper, some important second-order statistics such as the correlation coefficient, level crossing rate, and average fade duration of eigen-channels, are studied, in time-varying Rayleigh multiple-input multiple-output (MIMO) channels, assuming a general non-isotropic scattering environment. Exact closed-form expressions are derived and Monte Carlo simulations are provided to verify the accuracy of the analytical results. These results serve as useful tools for analysis and design of MIMO systems in time-varying channels. Shuangquan Wang |
GLOBECOM | 1 |
| 2006 | Low-Complexity Optimal Estimation of MIMO ISI Channels With Binary Training SequencesabstractIn this letter, a novel low-complexity optimal channel estimator using uncorrelated periodic complementary sets of binary sequences is proposed for multiple-input multiple-output (MIMO) intersymbol interference (ISI) channels. The estimator is optimal since it attains the minimum possible Crameacuter-Rao lower bound (CRLB). Moreover, it can be implemented with very low complexity via ASIC/FPGA, which makes it suitable and ready for practical MIMO systems Shuangquan Wang |
IEEE Signal Process. Lett. | 1 |
| 2005 | MIMO frequency selective channel estimation using aperiodic complementary sets of sequencesabstractAccurate estimation of MIMO frequency selective fading channels is important for reliable communication. In this paper, a new channel estimation scheme which relies on uncorrelated aperiodic complementary sets of sequences is proposed. Theoretical analysis and Monte-Carlo simulation show that the estimator achieves the minimum possible Cramer-Rao lower bound (CRLB). Furthermore, low-complexity algorithms for both ASIC/FPGA and DSP implementations are provided, using the special structure of the uncorrelated aperiodic complementary sets of sequences. Shuangquan Wang |
GLOBECOM | 1 |