EDBT 2026 Demo / reviewers in the wild / expert
Shaohu Zhang
dblp:188/9419
· DBLP profile ↗
11ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0001-8985-515XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 4 first-author · 4 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Cross-Attention Transformer and Multifeature Fusion for Cross-Linguistic Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) is important in improving human-computer interaction. Cross-Linguistic SER (CLSER) has been a challenging research problem due to significant variability in the linguistic and acoustic features of different languages. In this study, we propose a novel approach,HuMP-CAT, which combines HuBERT (Hidden Unit BERT), MFCC (Mel-Frequency Cepstral Coefficients), and Prosodic characteristics. These features are fused using a cross-attention transformer (CAT) mechanism during feature extraction. Transfer learning is applied to gain from a source emotional speech dataset to the target corpus for emotion recognition. We use IEMOCAP as the source data set to train the source model and evaluate the proposed method on seven data sets in five languages (i.e., English, German, Spanish, Italian, and Chinese). We show that, by fine-tuning the source model with a small portion of speech from the target datasets,HuMP-CATachieves an average accuracy of 78.75% across the seven datasets, with notable performance of 88.69% on EMODB in German language and 79.48% on EMOVO in Italian language. Our extensive evaluation demonstrates thatHuMP-CAToutperforms existing methods across multiple target languages. Xiantao Jiang, F. Richard Yu, Victor C. M. Leung, Tao Wang 0077, Shaohu Zhang |
IEEE Internet Things J. | 6 |
| 2024 | Privacy Measurement of Physical Attributes on Voice AnonymityabstractVarious methods have been proposed for protecting the speaker's identity while preserving speech intelligibility. However, existing studies fail to consider the overall tradeoff between speech utility, speaker verification, and inference of voice physical attributes, such as emotion, age, accent, and gender. we propose a tradeoff metric to encapsulate voice biometrics as well as different voice attributes, to study the feasibility of applying cutting-edge voice anonymization solutions to achieve the optimum tradeoff between privacy protection and speech utility. Shaohu Zhang, Zhouyu Li, Anupam Das 0001 |
MobiCom | 1 |
| 2023 | Speaker Orientation-Aware Privacy Control to Thwart Misactivation of Voice AssistantsabstractSmart home voice assistants (VAs) such as Amazon Echo and Google Home have become popular because of the convenience they provide through voice commands. VAs continuously listen to detect the wake command and send the subsequent audio data to the manufacturer-owned cloud service for processing to identify actionable commands. However, research has shown that VAs are prone to replay attack and accidental activations when the wake words are spoken in the background (either by a human or played through a mechanical speaker). Existing privacy controls are not effective in preventing such misactivations. This raises privacy and security concerns for the users as their conversations can be recorded and relayed to the cloud without their knowledge. Recent studies have shown that the visual gaze plays an important role when interacting with conservation agents such as VAs, and users tend to turn their heads or body toward the VA when invoking it. In this paper, we propose a device-free, non-obtrusive acoustic sensing system called HeadTalk to thwart the misactivation of VAs. The proposed system leverages the user's head direction information and verifies that a human generates the sound to minimize accidental activations. Our extensive evaluation shows that HeadTalk can accurately infer a speaker's head orientation with an average accuracy of 96.14% and distinguish human voice from a mechanical speaker with an equal error rate of 2.58%. We also conduct a user interaction study to assess how users perceive our proposed approach compared to existing privacy controls. Our results suggest that HeadTalk can not only enhance the security and privacy controls for VAs but do so in a usable way without requiring any additional hardware. Shaohu Zhang, Aafaq Sabir, Anupam Das 0001 |
DSN | 1 |
| 2023 | INSPIRE: Instance-Level Privacy-Pre Serving Transformation for Vehicular Camera VideosabstractThe wide spread of vehicular cameras has raised broad privacy concerns. Ubiquitous vehicular cameras capture bystanders like people or cars nearby without their awareness. To address privacy concerns, most existing works either blur out direct identifiers such as vehicle license plates and human faces, or obfuscate whole video frames. However, the former solution is vulnerable to re-identification attacks based on general features, and the latter severely impacts utility of the transformed videos. In this paper, we propose an INStance-level PrIvacy-pREserving (INSPIRE) video transformation framework for vehicular camera videos. INSPIRE leverages deep neural network models to detect and replace sensitive object instances in vehicular videos with their non-existent counterparts. We design INSPIRE as a modular framework to enable flexible customization of protected instance categories and their protection modules. An implementation of INSPIRE focused on protecting people and cars is described, which we tested on six re-identification datasets and three real-world vehicular video datasets to evaluate its privacy protection and utility preservation capability. Results show that INSPIRE can thwart 97% of re-identification attacks for people and cars while maintaining a 0.75 object detection mean average precision on transformed instances. We also demonstrate experimentally that INSPIRE is robust against model inversion attacks. Compared to solutions that provide comparable privacy protection, INSPIRE achieves relatively 1.76 times higher counting accuracy and 31.61% higher object detection mean average precision. Zhouyu Li, Ruozhou Yu, Anupam Das 0001, Shaohu Zhang, Huayue Gu, Fangtong Zhou, Aafaq Sabir, Dilawer Ahmed, Ahsan Zafar |
ICCCN | 4 |
| 2023 | VoicePM: A Robust Privacy Measurement on Voice AnonymityabstractVoice-based human-computer interaction has become pervasive in laptops, smartphones, home voice assistants, and Internet of Thing (IoT) devices. However, voice interaction comes with security and privacy risks. Numerous privacy-preserving measures have been proposed for hiding the speaker's identity while maintaining speech intelligibility. However, existing works do not consider the overall tradeoff between speech utility, speaker verification, and inference of voice attributes, including emotional state, age, accent, and gender. In this study, we first develop a tradeoff metric to capture voice biometrics as well as different voice attributes. We then propose VoicePM, a robust Voice Privacy Measurement framework, to study the feasibility of applying different state-of-the-art voice anonymization solutions to achieve the optimum tradeoff between privacy and utility. We conduct extensive experiments using anonymization approaches covering signal processing, voice synthesis, voice conversion, and adversarial techniques on three speech datasets that include both English and Chinese speakers to showcase the effectiveness and feasibility of VoicePM. Shaohu Zhang, Zhouyu Li, Anupam Das 0001 |
WISEC | 1 |
| 2021 | A 2-FA for home voice assistants using inaudible acoustic signalabstractVoice assistants have been shown to be vulnerable to replay attacks, impersonation attacks and inaudible voice commands. Existing defenses do not provide a practical solution as they either rely on external hardware or work under very constrained settings. We introduce a hand gesture-based authentication system for smart home voice assistants called HandLock, which uses built-in microphones and speakers to generate and sense inaudible acoustic signals to detect the presence of a known hand gesture. Our proposed approach can act as a second-factor authentication (2-FA) for performing specific sensitive operations like confirming online purchases through voice assistants. The experiments involving 45 participants show that HandLock can achieve on average 96.51% true-positive-rate at the expense of 0.82% false-acceptance-rate. Shaohu Zhang, Anupam Das 0001 |
MobiCom | 1 |
| 2021 | HandLock: Enabling 2-FA for Smart Home Voice Assistants using Inaudible Acoustic SignalabstractThe use of voice-control technology has become mainstream and is growing worldwide. While voice assistants provide convenience through automation and control of home appliances, the open nature of the voice channel makes voice assistants difficult to secure. As a result voice assistants have been shown to be vulnerable to replay attacks, impersonation attacks and inaudible voice commands. Existing defenses do not provide a practical solution as they either rely on external hardware (e.g., motion sensors) or work under very constrained settings (e.g., holding the device close to a user’s mouth). We introduce the concept of using a gesture-based authentication system for smart home voice assistants called HandLock, which uses built-in microphones and speakers to generate and sense inaudible acoustic signals to detect the presence of a known (i.e., authorized) hand gesture. Our proposed approach can act as a second-factor authentication (2-FA) for performing specific sensitive operations like confirming online purchases through voice assistants. Through extensive experiments involving 45 participants, we show that HandLock can achieve on average 96.51% true-positive-rate (TPR) at the expense of 0.82% false-acceptance-rate (FAR). We perform a comprehensive analysis of HandLock under various settings to showcase its accuracy, stability, resilience to attacks, and usability. Our analysis shows that HandLock can not only successfully thwart impersonation attacks, but can do so while incurring very low overheads and is compatible with modern voice assistants. Shaohu Zhang, Anupam Das 0001 |
RAID | 1 |
| 2020 | A WiFi-based Home Security SystemabstractTypical home security systems monitor homes for intrusions by installing contact sensors on doors and windows and motion sensors inside the house. Unfortunately, due to the high deployment and operational costs of today's home security systems, only a small fraction of homes have security systems installed (e.g., only 17% in the US and 15% in China). In this paper, we propose a WiFi based Home Security system (WiHS) that uses commodity WiFi devices, which most modern households already have, to perform the three primary tasks of typical home security systems: 1) detect when a door/window is opened/closed, 2) identify which door/window has been opened/closed, and 3) detect movements inside the house. The design of WiHS is based on our intuitive and theoretical understanding of the impacts of the movements of doors and windows on WiFi signals, which we will develop and present in this paper. We extensively evaluated WiHS using commodity WiFi devices in 3 different houses. WiHS detected intrusions with over 95% accuracy and identified the exact door/window that moved with just 4.5% average error. Shaohu Zhang, Raghav H. Venkatnarayan, Muhammad Shahzad 0001 |
MASS | 1 |
| 2017 | WiTraffic: Low-Cost and Non-Intrusive Traffic Monitoring System Using WiFiabstractThe traffic monitoring system is an imperative tool for traffic analysis and transportation planning. In this paper, we present WiTraffic: the first WiFi-based traffic monitoring system. Compared with existing solutions, it is non-intrusive, cost- effective, and easy-to-deploy. Unique WiFi Channel State Information (CSI) patterns of passing vehicles are captured and analyzed to effectively perform vehicle classification, lane detection, and speed estimation. A machine learning technique is adopted to train vehicle classification models and efficiently categorize vehicles. An Earth Mover's Distance (EMD)-based vehicle lane detection algorithm and vehicle speed estimation mechanism are proposed to further utilize WiFi CSI to identify the lane in which a vehicle is located and to estimate the vehicle speed. We implemented WiTraffic with off-the-shelf WiFi devices and performed real-world experiments with over a week of field data collection in both local roads and highways. The results show that the mean classification accuracy, lane detection accuracy for both local road and highway settings are around 96%, and 95%, respectively. The average root-mean- square error (RMSE) of the proposed CSI-based speed estimation method on a highway was 5mph in our experimental settings. Myounggyu Won, Shaohu Zhang, Sang Hyuk Son |
ICCCN | 2 |
| 2016 | WiTraffic: Non-intrusive Vehicle Classification Using WiFi: Poster AbstractabstractWe present WiTraffic: the first WiFi-based traffic monitoring system. The unique WiFi Channel State Information (CSI) patterns of passing vehicles are captured and analyzed to perform vehicle classification. We implemented WiTraffic with off-the-shelf WiFi devices and performed real-world experiments with over a week of field data collection. The results show that the classification accuracy is around 96%. Shaohu Zhang, Myounggyu Won, Sang Hyuk Son |
SenSys | 1 |
| 2016 | Low-Cost Realtime Horizontal Curve Detection Using Inertial Sensors of a SmartphoneabstractFatal accidents occur frequently on low-volume rural roads, and the accident rates are up to 4 times higher at curves. It is thus of paramount importance to perform road inventory of rural roads to develop safety plans. However, most states in U.S. face a challenge to maintain a database for low-volume rural roads due to limited funds for road inventory. In this paper, we propose to significantly reduce the cost for road inventory specifically focusing on horizontal curve detection by developing a mobile road inventory system based on off-the-shelf smartphones. The proposed system is capable of accurately detecting various kinds of horizontal curves by synthesizing heterogeneous smartphone sensor data to generate curve models by exploiting a machine learning technique. We implemented the system on iOS-based smartphones and tested with more than 400-miles of field data. We demonstrate that the proposed system achieves a median of 93.8% curve identification accuracy with a median of 5% false positive rates. Shaohu Zhang, Myounggyu Won, Sang Hyuk Son |
VTC Fall | 1 |