Chuang Shi

dblp:49/11146 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-1517-2775ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Multimodal Fusion-Based Human Action Recognition Using Wi-Fi CSI and Smartwatch Sensors
abstract
Human Activity Recognition (HAR) has become increasingly important in healthcare, smart homes, and human–computer interaction applications. However, traditional vision-based approaches suffer from privacy concerns and high deployment costs, while single-sensor methods are often limited in robustness and generalization capability. To address these challenges, this study proposes a multimodal HAR framework that integrates a 2×2 Wi-Fi Channel State Information (CSI) array with wearable inertial sensors. The proposed 2×2 CSI array enables synchronized multi-channel acquisition and fusion, improving signal stability and reducing packet loss in complex indoor environments. Meanwhile, accelerometer and gyroscope data are collected from a smartwatch and combined with CSI signals to construct a comprehensive multimodal representation. A hierarchical deep learning architecture, termed M2HAR-Net, is designed to effectively extract and fuse spatial–temporal features from heterogeneous modalities, capturing complementary motion characteristics. To enhance computational efficiency and real-time responsiveness, an ablation study on PCA-based frequency-domain reduction demonstrates that retaining only the first principal component preserves discriminative information while significantly reducing inference overhead. Experimental results on a dataset collected from four participants show that the proposed method achieves an overall accuracy of 98.77% across nine daily activities. Comparative evaluations with several state-of-the-art models further validate the effectiveness and robustness of the proposed framework. These findings indicate that efficient multimodal fusion can substantially improve HAR performance, while future work will focus on cross-environment generalization and large-scale validation.
Xiangnan Cai, Xinyi Dai, Ming Xia 0009, Chuang Shi, Wu Chen 0001
IEEE Internet Things J.5
2026 Space-Based Autonomous Timescale Realization for LEO Constellations Enabled by Low SWaP-C Heterogeneous Clock Ensembles and Intersatellite Links
Fu Zheng, Dong Zhang 0017, Chuang Shi
IEEE Internet Things J.5
2025 Knowledge-Enhanced Conversational Recommendation via Multi-view Graph Contrastive Learning
Yuru Liu, Yong Xu 0001, Cheng Li 0058, Chuang Shi, Qun Fang
ICIC (20)4
2025 DAGR: DyHGSampler and AdaSGC-Based Knowledge Graph Diffusion for Recommendation
Chuang Shi, Yong Xu 0001, Cheng Li 0058, Yuru Liu, Qun Fang
ICIC (8)1
2025 Integrated Health Monitoring and Pedestrian Navigation: A Hybrid Foot-Worn and Wrist-Worn Multi-Sensor System for Seamless 3D Localization and Vital Sign Tracking
abstract
This paper introduces a comprehensive health monitoring and pedestrian localization solution that combines a foot-mounted inertial measurement unit (IMU) with a wrist-worn health monitoring device. The system uses pedestrian dead reckoning (PDR) along with an advanced gait recognition algorithm to deliver continuous 3D localization, achieving an accuracy of less than 1 meter in both indoor and outdoor settings. The wrist-worn sensor integrates electrocardiography (ECG), photoplethysmography (PPG), and an accelerometer to monitor vital signs in real time and detect falls. This integrated approach provides a reliable solution for healthcare monitoring, location tracking, and emergency response applications.
Nanzhu Liu, Ming Xia 0009, Deyou Zhang, Chuang Shi
INDIN6
2025 WavI2I: Wavelet-Driven Inertial Imaging for Robust Industrial Human Activity Recognition
abstract
Wearable sensor–based human activity recognition (HAR) is essential for navigation and industrial automation, where real-time data processing is critical. However, the inherent nonlinearity and noise in sensor data make accurate and fast activity detection a challenging task, and traditional feature extraction techniques often struggle to capture the underlying dynamic changes. In this paper, we introduce WavI2I, a novel framework that enhances HAR by first applying the Continuous Wavelet Transform (CWT) to convert inertial sensor signals into time-frequency planes, which are then processed by Convolutional Neural Networks (CNNs). To further optimize feature extraction, we incorporate a nonlinear scale generator, ensuring a balanced focus on both high- and low-frequency components. These innovations strengthen the CNN’s ability to identify critical features, thereby improving recognition accuracy. Experiments on the UCI-HAR dataset confirm that WavI2I surpasses existing methods in terms of accuracy, computational efficiency, and model simplicity.
Ming Xia 0009, Deyou Zhang, Zhuoyuan She, Chuang Shi
INDIN5
2025 Step Length Estimation Method Based on Residual Neural Network for Pedestrian Dead Reckoning with Shoulder-Mounted IMU
abstract
Pedestrian Dead Reckoning (PDR) technology demonstrates significant application value in smart city location services and IoT terminal positioning due to its signal-independent operation, autonomous navigation capability, and anti-interference advantages. However, existing shoulder-mounted inertial measurement units (IMUs) encounter gait characteristic modeling errors during practical deployment, particularly manifesting as nonlinear error accumulation caused by limited step-length prediction accuracy. To address this technical challenge, this study proposes a step-length estimation model based on residual neural networks (ResNet) with limited-sample training. The architecture achieves precise step-length prediction across various motion states through temporal feature extraction and multi-rate motion pattern analysis. Experimental results validated by five independent test sets demonstrate that the system achieves a relative displacement estimation error below 0.6%, with the mean absolute error (MAE) of single-step length prediction consistently remaining under 0.045 meters. Analytical verification confirms that the proposed step-length estimation method significantly enhances the step-length measurement accuracy of shoulder-mounted IMUs, providing an effective technical solution for high-precision indoor positioning of IoT devices.
Ziwei Yue, Ming Xia 0009, Deyou Zhang, Zhuoyuan She, Chuang Shi
INDIN6
2025 A Practical TDOA-Based Method for UWB Anchor Localization
abstract
This paper presents a practical TDOA-based method for UWB anchor localization, aiming to simplify the process of determining anchor positions, reduce costs, and improve efficiency. By utilizing a small number of tag position coordinates and the TDOA information between anchors and tags, and by introducing a weighted least squares approach, this method can quickly and effectively solve for UWB anchor coordinates even without any prior knowledge of their initial positions. Experimental results demonstrate that the proposed method can achieve positioning accuracy within 1 meter, with the rectangular trajectory demonstrating the highest stability and accuracy, as indicated by the lowest RMSE (Root Mean Square Error) and HDOP (Horizontal Dilution of Precision) values. Future research will focus on optimizing the algorithm under complex environmental conditions, integrating data from multiple sensors such as LiDAR or cameras, enhancing real-time performance, and developing user-friendly interfaces. These efforts aim to further enhance the method's robustness and practical applicability. Additionally, comparative experiments in diverse scenarios will be conducted to validate the effectiveness and applicability of the proposed approach.
Xinchi Zhang, Ming Xia 0009, Ziwei Yue, Deyou Zhang, Chuang Shi
INDIN6
2025 LSTM-Attention with Multi-Sensor Fusion for High-Accuracy 3D Indoor Localization on Smartphones
abstract
While smartphone-based fingerprinting techniques have emerged as promising solutions for indoor localization, their efficacy remains constrained by suboptimal fingerprint database quality and algorithmic limitations in 2D coordinate estimation. Conventional approaches suffer from laborious data collection processes, constrained accuracy, and diminished reliability over extended periods. To address these challenges, this study proposes a novel attention-enhanced LSTM architecture synergistically integrating heterogeneous sensor data (WiFi, barometric pressure, and magnetometer) to achieve simultaneous planar localization and multi-floor identification. A dedicated foot-mounted inertial measurement system is introduced to streamline fingerprint database construction by enabling efficient sparse data acquisition through collaborative smartphone-device interactions. The developed LSTM-Attention framework demonstrates superior performance in initial position matching precision, accelerated model convergence, and enhanced trajectory consistency through adaptive feature weighting. Comprehensive evaluations across multi-story academic and office environments reveal pedestrian localization accuracy within 1.5 meters (horizontal) and floor discrimination success rates exceeding 95%, thereby advancing the state-of-the-art in smartphone-based 3D indoor positioning systems.
Ming Xia 0009, Deyou Zhang, Shengmao Que, Chuang Shi
INDIN6
2025 Over-the-Air Computation via Reconfigurable Intelligent Surface with Phase-Dependent Amplitude Response
abstract
Over-the-air computation (AirComp) leverages the inherent superposition property of wireless multiple-access channels to enable direct signal aggregation from massive users. However, unfavorable channel conditions can severely degrade the computed mean square error (CMSE). To address this limitation, we introduce a reconfigurable intelligent surface (RIS) into the AirComp system. Unlike prior works that assume full signal reflection by each RIS element (RE) regardless of its phase shift, we adopt a practical model accounting for the coupling between the amplitude and phase of each RE. Based on this model, we formulate an optimization problem to minimize the CMSE by jointly optimizing transceiver design and the RIS reflection matrix. To tackle this highly nonconvex problem, we propose a dual-loop optimization framework, where the outer loop employs the genetic algorithm to obtain a near-optimal RIS reflection matrix and the inner loop uses an alternating optimization approach for transceiver design. Simulation results demonstrate that the proposed dual-loop algorithm outperforms baseline methods in reducing the CMSE.
Deyou Zhang, Wanxi Zhang, Ming Xia 0009, Chuang Shi
INDIN5
2025 3DIO: Low-Drift 3-D Deep-Inertial Odometry for Indoor Localization Using an IMU
abstract
The use of mobile devices for indoor localization has proven to be a convenient solution for pedestrians in Internet of Things (IoTs) applications. Radiofrequency (RF) signals, including Wi-Fi, Bluetooth, and others, are among the most commonly used sources. However, their availability cannot be guaranteed in all scenarios. Although pedestrian dead reckoning (PDR) using an inertial measurement unit (IMU) provides a self-contained positioning solution, it is susceptible to error accumulation due to heading uncertainties and varying motions. This article presents a low-drift 3-D deep-inertial odometry (DIO) method for indoor pedestrian localization using an IMU. The proposed approach employs a neural network to regress speeds within the human body frame, ensuring that the speeds are unaffected by absolute heading. These regressed speeds are integrated with inertial navigation to determine position. To enhance accuracy, the method incorporates an invariant extended Kalman filter (InEKF)-based integration for state estimation. Additionally, a learned height is included in the filter to improve 3-D position estimation. The performance of the proposed method is validated through real-world tests in various environments. Results demonstrate that the proposed method outperforms traditional PDR, robust neural inertial navigation (RONIN), and EKF-based techniques. Furthermore, this article examines the method from multiple perspectives, highlighting its strengths in addressing heading drift and varying motions, as well as the impact of height constraints and behavior-based position corrections. The video (YouTube) or (BiliBili) is shared to showcase our work.
Shiyu Bai, Weisong Wen, Chuang Shi
IEEE Internet Things J.3
2025 BDS/GPS Tens of Picoseconds Time Synchronization Method With Application to Communication Network
abstract
Time synchronization is a critical requirement in sixth-generation (6G) communication systems and Internet of Things (IoT) applications. However, the standard timing services provided by global navigation satellite systems (GNSS) are increasingly unable to meet the growing demands for precision and stability in modern networks. The proposed system leverages a single-difference time transfer (SDPT) method to achieve tens of picoseconds precision, enabling effective clock disciplining for oven-controlled crystal oscillators (OCXO). A clock disciplining within a phase-locked loop (PLL) structure further optimizes the output of time and frequency signals, enhancing OCXO stability. Additionally, a differential precise time synchronization system (DPTSS) by the BeiDou navigation satellite system (BDS)/global positioning system (GPS) is developed to deliver high-precision time reference signals for communication networks. Experimental results demonstrate that the real-time 1 pulse-per-second (1PPS) interval achieved by the BDS/GPS time link is better than 0.1 ns, with long-term frequency stability reaching the$10{^{-}16 }$level. Furthermore, the proposed system achieves an 1PPS signal synchronization precision of 0.73 ns in communication networks, significantly outperforming the other state-of-the-art methods.
Dayu Yan, Chuang Shi, Fu Zheng, Caicong Wu, Guifei Jing, Xiaolin Meng
IEEE Internet Things J.3
2025 Multimodal Spatiotemporal Feature-Based Human Motion Pattern Recognition With CNN-Transformer-Attention Framework
Ming Xia 0009, Nanzhu Liu, Ziwei Yue, Chuang Shi, Wu Chen 0001, Xudong Mou
IEEE Internet Things J.6
2025 HPPS: A Head-Mounted Pedestrian Positioning System Integrating Fisheye Camera/RTK/PDR With Factor Graph Optimization
abstract
Substations are critical infrastructures in modern power systems, requiring accurate and reliable personnel positioning to ensure operational safety and inspection efficiency. Pedestrian dead reckoning (PDR) technology has emerged as a preferred solution for real-time seamless positioning due to its autonomous, anti-jamming and passive characteristics. However, in practical applications, the PDR algorithm faces challenges such as the inability to provide absolute positioning due to unknown starting points and cumulative errors over time. To address these challenges, this paper proposes a head-mounted pedestrian positioning system (HPPS) integrating fisheye camera/real-time kinematic (RTK)/PDR based on factor graph optimization (FGO). First, a skyward-facing fisheye camera identifies non-line-of-sight (NLOS) signals caused by architectural obstructions, improving the absolute positioning accuracy of RTK. Next, an adaptive pedestrian gait detection and step length estimation algorithm based on head-mounted inertial measurement units (IMU) is developed to enhance PDR robustness. Finally, the FGO framework integrates the inputs of the fisheye camera, RTK, and PDR to achieve seamless indoor-outdoor positioning. Experimental results demonstrate that the proposed system achieves horizontal root mean square error (RMSE) values of less than 0.42 m and 0.35 m in two distinct outdoor substation environments and less than 1.36 m in indoor-outdoor scenarios. This work highlights the integration of precise positioning technologies within the Internet of Things (IoT) framework, advancing smart grid applications and enabling effective location tracking in complex environments.
Ming Xia 0009, Yunfeng Shan, Deyou Zhang, Qianhua Yang, Chuang Shi
IEEE Internet Things J.6
2025 RCA-SI: A Rapid Consensus Algorithm for Swarm Intelligence in unstable network environments
Guangquan Zeng, Wan Hu, Yongchao Zhou, Desheng Zheng, Xiaoyu Li 0003, Chuang Shi
J. Netw. Comput. Appl.6
2025 An Optimized Model With Encoder-Decoder ConvLSTM for Global Ionospheric Forecasting
abstract
The ionosphere is vital for satellite navigation and radio communication, but observational limitations necessitate ionospheric forecasting. The least squares collocation method is commonly used for Global Navigation Satellite System (GNSS)-based global ionospheric forecasting, though its accuracy and stability need improvement. This study introduces two optimized models based on the ConvLSTM cell with an Encoder-Decoder structure to enhance forecasting performance. Using seven years of historical data, the model provides stable forecasts for the following year. Tests from 2015–2020 show that optimization reduces root mean square error (RMSE) by 10.159%–16.363% compared to the unoptimized method. The Encoder-Decoder ConvLSTM-B model achieves the best performance, lowering RMSE by 2.031%–8.547% compared to the ConvLSTM-A model. These results highlight the effectiveness of the proposed approach in improving ionospheric forecast accuracy.
Cheng Wang 0007, Kaiyu Xue, Chuang Shi
IEEE Geosci. Remote. Sens. Lett.3
2025 Beamforming Design for Active RIS-Aided Over-the-Air Computation
abstract
Over-the-air computation (AirComp) is emerging as a promising technology for wireless data aggregation. However, its performance is hampered by users with poor channel conditions. To mitigate such a performance bottleneck, this paper introduces an active reconfigurable intelligence surface (RIS) into the AirComp system. We begin by exploring the ideal active RIS model and propose a joint optimization of the transceiver and RIS configuration to minimize the mean squared error (MSE) between the target and estimated function values. To manage the resulting tri-convex optimization problem, we employ the alternating optimization (AO) framework to decompose it into three convex subproblems, each of which can be solved optimally. We then investigate two specific cases and analyze their respective asymptotic performance to reveal the superiority of the active RIS in mitigating the MSE relative to its passive counterpart. Lastly, we adapt our transceiver and RIS configuration optimization approach to account for the self-interference of the active RIS. To handle the resulting highly non-convex problem, we further develop a two-layer AO framework. Simulation results confirm the superiority of the active RIS in enhancing AirComp performance compared to its passive counterpart.
Deyou Zhang, Ming Xiao 0001, Chuang Shi, Mikael Skoglund, H. Vincent Poor
IEEE Trans. Commun.3
2025 Improved 3-D Wet Refractivity Retrieval Using Refined One-Step Tropospheric Tomography Method
abstract
Current Global Navigation Satellite System (GNSS) tropospheric tomography techniques struggle in extreme weather due to the reliance on empirical wet mapping functions (MF). Although the tomography method without wet MF (i.e., the One-Step method) has been proposed and validated, its performance is still susceptible to separate ambiguity treatment. To avoid this defect, this study proposes a refined One-Step tomography method that simultaneously estimates tomographic and other GNSS parameters from raw observations without wet MF. Tomography experiments were conducted during an extreme rainstorm in the Hong Kong region in September 2023. Wet refractivity retrieved from the refined One-Step method shows better performance than the traditional Two-Step method, exhibiting lower error by 5.8/10.1/27.1% (radiosonde reference) and 10.7/31.0/29.4% (ERA5 reference) for stages before/during/after the rainstorm. These results are also superior to those from the original One-Step method, manifesting the advantage of refinement for parameter estimation. Additionally, its capability for characterizing horizontal wet refractivity distributions was explored. Results of the refined One-Step method (and the original One-Step method) exhibit a 7.4% (1.6%) higher accuracy under symmetric atmospheric conditions, and a 57.2% (41.3%) higher accuracy under asymmetric conditions than the Two-Step method. Correlation analysis between accuracy improvement and asymmetric condition of atmosphere also indicates that the refined One-Step tomography method is more suitable for characterizing three-dimensional water vapor field during extreme weather.
Linghao Zhou, Chuang Shi, Yunchang Cao
IEEE Trans. Geosci. Remote. Sens.4
2024 Cognitive Virtual Sensing Technique for Feedforward Active Noise Control
abstract
The virtual sensing (VS) technique enables an active noise control (ANC) system to estimate the virtual error signal for control using remote monitoring microphones. However, instances where noise characteristics and primary paths exhibit variations lead to a noticeable decline in performance for the conventional VS technique. To address this challenge, we propose the cognitive VS technique in this paper. Its objective is to enhance VS performance by providing a more precise estimate of the error signal based on environmental cognition. Differing from the previous selective VS technique, the cognitive VS technique connects both the reference and monitoring microphones to a lightweight classifier. Hence, the cognitive VS technique has the capability to dynamically adjust the VS filter in accordance with the noise and environmental conditions identified by the classifier. Simulation results demonstrate that the cognitive VS technique surpasses conventional and selective VS techniques in terms of adaptivity and generalisation when noise characteristics change and primary paths are time-varying.
Rong Xie 0001, Anqi Tu, Chuang Shi, Stephen Elliott, Huiyong Li 0001
ICASSP3
2024 Improving the performance of the ORB-SLAM3 with low-light image enhancement
abstract
Traditional Visual Simultaneous Localization and Mapping (VSLAM) algorithms demonstrate good accuracy and robustness in well-lighting environments by extracting numerous feature points. However, in low-light conditions, insufficient illumination leads to low-contrast images, which hampers the ability of the front-end to extract adequate feature points, resulting in increased tracking errors or complete tracking failure. To overcome these limitations, this paper improves the performance of the ORB-SLAM3 in low-light environments with image enhancement capability. Specifically, we integrate the Histogram Equalization Prior-based (HEP) image enhancement module into ORB-SLAM3. Furthermore, we compare and analyze the results with two other image enhancement algorithms to ensure the effective fusion of image enhancement and ORB-SLAM3. Experiments were conducted and performance comparisons were made to validate the performance of ORB-SLAM3 with image enhancement in low-light environments. Experimental results indicate that the Root Mean Square Errors (RMSE) of the Absolute Pose Error (APE) are 0.99 cm and 0.76 cm on the two public ETH3D low-light datasets, respectively. Compared to the original ORB-SLAM3, the improvement is over 50%. Similarly, the RMSE of the Relative Pose Error (RPE) are 0.54 cm and 0.65 cm, with an improvement of 29% and 34%, respectively.
Tuan Li, Zhixin Wang, Chuang Shi
IPIN4
2024 3D Indoor Localization via Universal Signal Fingerprinting Powered by LSTM
abstract
Fingerprint localization is a critical method for indoor positioning that has garnered considerable attention. Traditional methods for constructing fingerprint databases and matching algorithms frequently exhibit inefficiencies and limitations, which can undermine both the accuracy and the robustness of localization systems. This paper introduces an innovative indoor pervasive localization method leveraging deep learning. We employ a hybrid system of foot-mounted positioning devices and smartphones to efficiently create a universal fingerprint database, subsequently utilizing LSTM-based deep learning methods for accurate pedestrian location matching. Experiments conducted within a standard academic building demonstrate that our proposed method can more accurately map the indoor movement trajectories of pedestrians. The localization results indicate a horizontal accuracy of 2.5 meters and a vertical accuracy of 0.2 meters. Notably, our method shows a 10% improvement in horizontal accuracy and an 18% improvement in vertical accuracy over WiFi-only approaches. Furthermore, when compared to the classical Random Forest models, our method achieves performance enhancements of 20% in horizontal and 15% in vertical accuracy.
Ming Xia 0009, Weisong Wen, Chuang Shi, Yunfeng Shan, Xinqi Tian
IPIN5
2024 Seamless Indoor-Outdoor Foot-Mounted Inertial Pedestrian Positioning System Enhanced by Smartphone PPP/3-D Map/Barometer
abstract
Foot-mounted inertial pedestrian positioning system (FIPPS) is increasingly important in the Internet of Things (IoT) and smart cities because of the passive, anti-interference, and fully automatic advantages. However, FIPPS faces the problem of unknown initial positions and accumulative errors in practical applications, which makes the positioning trajectory inaccurate and unable to be matched in a well-defined coordinate system. A seamless indoor–outdoor inertial pedestrian positioning method based on smartphone precise point positioning (PPP)/3-D map/barometer augmentation is proposed to solve this problem. The main contributions include the following: 1) a 3-D-mapping-aided PPP (3DMA PPP) method is proposed to detect and eliminate non-line-of-sight (NLOS) signals to provide high-precision initial coordinates for FIPPS, which enables positioning trajectories to be matched in the unified WGS-84 coordinate system; 2) the adaptive zero-velocity update (ZUPT) approach based on neuro-fuzzy inference is used to precisely identify the stance phases of various gait patterns to suppress position and velocity error accumulation; 3) a barometer and 3-D building model-based height constraint algorithm is applied to further improve the height estimation of the FIPPS; and 4) an Android application named FYTECH is developed for data transmission between the smartphone and FIPPS and supports the online display of pedestrian positioning trajectories. In a seamless outdoor–indoor experiment with multiple motion patterns, including walking, running, elevators and stairs, we demonstrate that the proposed method can achieve RMS errors better than 1.17 and 1.32 m in the horizontal and vertical directions, respectively. The performance indicates that our method can conveniently fuse multisource sensor information, such as foot wearables, smartphones, and 3-D maps to greatly enhance pedestrians’ seamless indoor–outdoor navigation experience.
Chuang Shi, Ming Xia 0009, Fu Zheng, Tuan Li, Yunfeng Shan, Guifei Jing, Wu Chen 0001, T. C. Hsia
IEEE Internet Things J.2
2024 ATS-UNet: Attentional 2-D Time Sequence UNet for Global Ionospheric One-Day-Ahead Prediction
abstract
The ionosphere contains many charged free particles that significantly affect radio signals passing through it. Due to limitations in observation technology, high-precision ionospheric forecasting has attracted widespread attention. Deep learning is well-suited to address its high-dimensional nonlinearity. In this study, we propose the Attentional 2-D time sequence UNet (ATS-UNet) model based on the UNet network combined with an attention mechanism. Using this model and global ionospheric observation data from 2008 to 2020, we develop a one-day forecast model for the global ionosphere. We select the 2015 and 2020 as test data, and compared the ATS-UNet model with various state-of-the-art models currently used in global ionosphere forecasting. These models include Adaptive Autoregressive (AAR), Long Short-Term Memory (LSTM), ConvLSTM and the UNet model. In 2015, the RMSE (Root Mean Square Error) value of the ATS-UNet model’s forecast results decreases by 14% compared to the AAR model, 6% compared to the LSTM model, 4% compared to the ConvLSTM model, and 3% compared to the UNet model. In 2020, the RMSE value of the ATS-UNet model decreases by 7%, 8%, 6%, and 2% respectively when compared to the same models. The results demonstrate that the ATS-UNet model can effectively improve prediction accuracy.
Kaiyu Xue, Chuang Shi, Cheng Wang 0007
IEEE Geosci. Remote. Sens. Lett.2
2024 Comments on "Primary-Ambient Extraction Using Ambient Spectrum Estimation for Immersive Spatial Audio Reproduction"
abstract
In the above paper, He et al. propose a primary-ambient extraction method using ambient phase estimation with a sparsity constraint (APES). The primary-ambient extraction problem is formulated as a$L1$-norm minimization of the primary component with respect to the phase of the ambient component, based on a stereo signal model they have assumed. Discrete searching (DS) is proposed as the approach for solving the APES in each time-frequency bin, which is computationally complicated and potentially imprecise. This correspondence provides an analytical solution to the APES by exploring the geometric relation between the primary and ambient components in a rotated coordinate.
Haisheng Lu, Jiangnan Liang, Chuang Shi
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Integration of Anomaly Machine Sound Detection into Active Noise Control to Shape the Residual Sound
abstract
An active noise control (ANC) system generates a secondary sound to destructively interfere with the undesirable noise. Existing ANC algorithms are mainly designed to minimize the power of the residual sound, with few considerations to the listening experience. This results in a pressing issue in practice whereby the residual sound is perceived to be different from the undesirable noise. When the ANC system is deployed to reduce the noise level in a factory environment, workers may feel strange because they are used to detecting the anomaly machine sound by their auditory perception. In order to solve this problem, this paper proposes to integrate anomaly sound detection (ASD) into the ANC system in order for the residual sound to represent the same machine status as the original machine noise. The ASD module is used to simulate human judgement. A homothety constrained ANC algorithm is developed to synchronously reduce the sample-wise power and keep the segment-wise machine status of the residual sound. The experiment results validate the effectiveness of the homothety constrained ANC algorithm in noise reduction, and the subjective test results show that the ASD-integrated ANC system results in less confusing perceptions of the residual sound.
Chuang Shi, Mengjie Huang, Huitian Jiang, Huiyong Li 0001
ICASSP1
2022 FlexPDR: Fully Flexible Pedestrian Dead Reckoning Using Online Multimode Recognition and Time-Series Decomposition
abstract
Smartphone-based pedestrian dead reckoning (PDR) has been widely used indoors for continuous localization. However, the specific tracking solutions under different modes are vulnerable to mode transition and thus degrading performance. The robustness of pedestrian navigation may be weakened due to the mix of smartphone motions and walking patterns. Due to this challenge, most existing PDR methods assume that the smartphone is carried in a certain pose, such as handheld horizontally, swinging, calling, and pocketed, ignoring the short period but negative transition impact on tracking, which limits its flexibility when applying in the Internet of Things (IoT) services. To achieve a fully flexible PDR (i.e., FlexPDR), this article enhances the robustness and smoothness of pedestrian tracking during the transition between several phone poses and regular motion modes for the first time. We propose a Bayesian-based real-time multimode recognition method that does not require any posterior information, together with a time-series decomposition approach for adaptively tracking scheme-switching. The proposed FlexPDR system achieved a real-time smartphone indoor positioning with a high position accuracy of 98.11% on the specific situation that mixed mode-switching happens, which outperformed other state-of-the-art methods.
Dayu Yan, Chuang Shi, Tuan Li, You Li 0001
IEEE Internet Things J.2
2022 Differential Error Feedback Active Noise Control With the Auxiliary Filter Based Mapping Method
abstract
This letter proposes a differential error feedback active noise control (FBANC) system in an open end duct to mitigate interferences from its downstream. The interference waves entering the duct cause disturbance to the conventional FBANC controller and even result in divergence of the control filter. This is due to the omni-directivity of the error microphone. The differential microphone array (DMA) can be constructed in a compact size using only one more omni-directional microphone. The DMA forms a frequency-invariant beampattern that presents configurable nulls. When the output of the DMA is used as the error signal, the null can be designed to enhance the robustness of the FBANC system against interferences. However, this differential error signal is converted from the sound pressure gradient instead of the sound pressure. The DMA’s location does not precisely indicate the control point where optimum noise reduction has been achieved. To solve this problem, an auxiliary filter based mapping (AFMap) method is developed to map the differential error signal to the location of an omni-directional microphone in the DMA. Experiment results demonstrate that the proposed differential error FBANC system is much less sensitive to interferences than the conventional FBANC system, and the AFMap method can ensure optimum noise reduction occurring at the target control point.
Chuang Shi, Feiyu Du, Huiyong Li 0001
IEEE Signal Process. Lett.1
2022 A Digital Twin Architecture for Wireless Networked Adaptive Active Noise Control
abstract
The active noise control (ANC) is a complementary technique to the passive noise control (PNC) to reduce the low frequency noise. The ANC controller can be implemented by pre-trained filters or adaptive filters. The adaptive ANC controller is advantageous in its adaptation to environmental changes. However, the algorithm complexity of the adaptive ANC controller increases with the scale of ANC applications, making it difficult to be carried out on low-cost processors. To resolve this problem, cloud computing should be utilized in ANC systems, and thus the wireless networked ANC system is proposed. Since it is crucial for ANC controllers to generate the anti-noise wave in real time, this paper formulates a digital twin architecture that implements the control filter adaptation in the cloud and the anti-noise signal generation on the local controller, respectively. A digital twin filtered-reference least mean squares (DT-FxLMS) algorithm is proposed to coordinate the digital twin with the local controller. Simulation and experiment results demonstrate the effectiveness and efficiency of the wireless networked ANC system based on the digital twin architecture.
Chuang Shi, Feiyu Du, Qianyang Wu
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Selective Virtual Sensing Technique for Multi-channel Feedforward Active Noise Control Systems
abstract
The virtual sensing technique allows the active noise control (ANC) system to work with error microphones that are placed far from the desired zone of quietness (ZoQ). Conventionally, a training stage is required to obtain the auxiliary filters with the temporary error microphones placed in the ZoQ. When the characteristics of the primary noise changes, the auxiliary filters have to be retrained. As a result, the conventional virtual sensing technique can only be used when the frequency band of the primary noise remains unchanged. In order to solve this limitation, this paper proposes a selective virtual sensing technique for the multi-channel feedforward ANC system. The selective virtual sensing technique obtains a bank of auxiliary filters in the subband structure. Based on the frequency-band-matching mechanism, a linear combination of the auxiliary filters is calculated and used in the real-time control stage. Experimental results show that the selective virtual sensing technique achieves better noise reduction performance than the conventional virtual sensing technique when the frequency band of the primary noise fluctuates.
Chuang Shi, Rong Xie 0001, Nan Jiang 0014, Huiyong Li 0001, Yoshinobu Kajikawa
ICASSP1
2019 Toward Robust Crowdsourcing-Based Localization: A Fingerprinting Accuracy Indicator Enhanced Wireless/Magnetic/Inertial Integration Approach
abstract
The next-generation Internet of Things (IoT) systems have an increasingly demand on intelligent localization which can scale with big data without human perception. Thus, traditional localization solutions without accuracy metric will greatly limit vast applications. Crowd sourcing-based localization has been proven to be effective for mass-market location-based IoT applications. This paper proposes an enhanced crowd sourcing-based localization method by integrating inertial, wireless, and magnetic sensors. Both wireless and magnetic fingerprinting accuracy are predicted in real time through the introduction of fingerprinting accuracy indicators (FAIs) from three levels (i.e., signal, geometry, and database). The advantages and limitations of these FAI factors and their performances on predicting location errors and outliers are investigated. Furthermore, the FAI-enhanced extended Kalman filter (EKF) is proposed, which improved the dead-reckoning (DR)/WiFi, DR/Magnetic, and DR/WiFi/Magnetic integrated localization accuracy by 30.2%, 19.4%, and 29.0%, and reduced the maximum location errors by 41.2%, 28.4%, and 44.2%, respectively. These outcomes confirm the effectiveness of the FAI-enhanced EKF on improving both accuracy and reliability of multisensor integrated localization using crowd sourced data.
You Li 0001, Zhe He 0002, Zhouzheng Gao, Yuan Zhuang 0001, Chuang Shi, Naser El-Sheimy
IEEE Internet Things J.5
2017 Multiple parallel branch with folding architecture for multichannel filtered-x least mean square algorithm
abstract
Multichannel active noise control (MCANC) systems are commonly used in acoustic noise or vibration control, such as large-dimension ventilation ducts, open windows and mechanical structures. However, its computational load far exceeds the capabilities of digital signal processors (DSPs) and microcontrollers. Even the field programmable gate array (FPGA) cannot straightforwardly cope with the exponential increase in the computation load of MCANC systems. A novel architecture, called the multiple parallel branch with folding, is proposed for the J × J × M (J reference microphones, J secondary sources and Merror microphones) MCANC implementation with the floating-point arithmetic. This architecture addresses the tradeoff between throughput and hardware resource consumption by using parallel execution and folding. The proposed architecture is validated in an experimental setup that carries out a 4 × 4 × 4 multichannel filtered-x least mean square (FxLMS) algorithm achieving the sampling rate and throughput of 24 KHz and 18.4 Mbps, respectively.
Dongyuan Shi, Jianjun He 0001, Chuang Shi, Tatsuya Murao, Woon-Seng Gan
ICASSP3
2017 Compensation for Nonlinear Distortion of the Frequency Modulation-Based Parametric Array Loudspeaker
abstract
The parametric array loudspeaker (PAL) modulates the audio signal on an ultrasonic carrier. When the modulated signal is transmitted in air, an audio beam is created based on the nonlinear acoustic principle. Each modulation method has advantages and disadvantages. The frequency modulation (FM) is favorable for its low cost and high volume, but the tradeoff is its complicated nonlinear distortion, which is difficult to be reduced. In this paper, the Volterra filter is adopted to model the nonlinearity of the FM-based PAL. A novel complex inverse system is devised to effectively reduce the nonlinear distortion. Three practical aspects are addressed. First, the computational complexity of the Volterra filter is reduced by the parallel cascade structure with almost no compromise to the model accuracy. Second, Volterra filters are identified at discrete input levels to treat the nonlinearity that keeps changing with the time-varying audio input. Third, when the input level is high, a separation approach is proposed to refine the identified Volterra filters, which eventually improves the performance of the proposed inverse system.
Yuta Hatano, Chuang Shi, Yoshinobu Kajikawa
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Automatic gain control for parametric array loudspeakers
abstract
Parametric array loudspeaker (PAL) modulates the audio input on an ultrasonic carrier and relies on airborne nonlinear acoustic effects to generate the audible sound output. The sound output is mainly confined in the beam of the ultrasonic carrier and thus shows a pronounced directivity. There are three parameters that together influence the output volume of a PAL. They are the input level, modulation index, and ultrasound level. In existing PALs, the volume knob is associated with the ultrasound level, while the modulation index is either fixed in the circuit or rarely adjustable by another knob. In this paper, an automatic gain control is proposed to improve the sound quality of the PAL by minimizing the modulation index, maintaining the output-to-input ratio, and ensuring the ultrasound level within the safety range. Simulation and measurement results validate that the proposed approach leads to a reduction in the average total harmonic distortion (THD) level by more than one third for all the tested modulation methods.
Chuang Shi, Yoshinobu Kajikawa
ICASSP1
2016 Synthesis of Volterra filters for the parametric array loudspeaker
abstract
The ultrasound-to-ultrasound Volterra filter (U2VF), which was previously proposed to represent the nonlinear response of the parametric array loudspeaker (PAL), identifies the PAL as a nonlinear system that uses ultrasonic signals as its input. It has been proven that the U2VF is a more generic model as compared to the audio-to-audio Volterra filter (A2VF), when the modulation method is adaptive or the input is time varying. However, there is no explicit solution to a linearization system based on the U2VF. Therefore, this paper proposes a synthesis method to extract A2VFs from the U2VF based on the parallel cascade structure. The mature technique of building a linearization system based on the A2VF can hence be adapted to consort with the U2VF.
Chuang Shi, Yoshinobu Kajikawa
ICASSP1
2015 Identification of the parametric array loudspeaker with a volterra filter using the sparse NLMS algorithm
abstract
Volterra filters can be applied to a wide range of nonlinear systems, keeping only the low order kernels to yield a good approximation. The parametric array loudspeaker (PAL), as a weak nonlinear acoustic system, is an attractive directional sound reproduction device. Volterra filters have been adopted in the linearization system of the PAL that efficiently reduces the nonlinear distortion with no need of solving the nonlinear acoustic equation. In this paper, the ultrasound-to-ultrasound Volterra filter is proposed, being inspired by the nonlinear acoustic principle, to provide a better systematic representation of the PAL. Experiment results are presented to prove the effectiveness of the proposed approach, where the sparse NLMS algorithm is carried out in the identification.
Chuang Shi, Yoshinobu Kajikawa
ICASSP1
2013 A psychoacoustical preprocessing technique for virtual bass enhancement of the parametric loudspeaker
abstract
The parametric loudspeaker is a novel type of loudspeaker that can project a directional sound beam. It is commonly used in creating personal sound zone and projecting private messages to a targeted audience. However, the parametric loudspeaker possesses a very poor bass (or low-frequency) response due inherently to the nonlinear acoustic principle generating sound from ultrasound in air. A psychoacoustic signal processing method known as “virtual bass” has been successfully implemented in some consumer electronics with miniature or flat loudspeaker unit, aiming to enhance their bass performances. In this paper, we adapt this “virtual bass“ approach for parametric loudspeakers. Unlike conventional loudspeakers, the parametric loudspeaker brings in an added degree of complexity in “virtual bass” enhancement due to its inherent nonlinear acoustic property. Accordingly, a new preprocessing technique is proposed for the parametric loudspeaker to psychoacoustically reproduce the low-frequency components within an octave below its cut-off frequency.
Chuang Shi, Hao Mu, Woon-Seng Gan
ICASSP1
2011 Using network RTK corrections and low-cost GPS receiver for precise mass market positioning and navigation applications
abstract
With the advancement in GPS receiver hardware and data processing technologies in recent years, more and more low-cost GPS receivers could output instant high quality carrier phase measurements to provide navigation solutions with an accuracy of a few meters. It is anticipated that improved GPS positioning performance will open doors for many new ITSS and Location-based Services (LBS) applications. This is the reason why the exploitation of the raw carrier phase measurements generated with low-cost receivers has attracted attention of researchers and manufacturers around the world. The University of Nottingham, The Chinese Academy of Surveying and Mapping, and Wuhan University in China have teamed up in recent years to research and develop new positioning and navigation algorithms. These algorithms are used to decode real-time carrier phase measurements from low-cost single frequency GPS receivers, collect instant network RTK corrections via GPRS/3G connections, resolve integer ambiguities and output real-time positioning solutions. A prototype system has been devised and extensively tested using the advanced facilities at the University of Nottingham and in China. It has been found that better than decimeter real-time on-the-fly positioning precision could be achieved. In this paper, the design and structure of this prototype system is introduced first. It is followed by a brief description of the newly developed ambiguity search and resolution technologies. A series of tests under a controlled environment have been conducted and the results are analyzed through comparison with “ground truth” trajectories. Preliminary conclusions are drawn in the paper and discussions are made to further improve the performance of this system.
Yanhui Cai, Xiaolin Meng, Weiming Tang, Chuang Shi
Intelligent Vehicles Symposium5