VLDB 2026 Research / reviewers in the wild / expert
Shaoying Wang
dblp:06/5306
· DBLP profile ↗
12ranked-venue papers
9as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resilient Percentile-Driven Spectrum Sharing for NTN-TN Coexistence
Shaoying Wang, Beatriz Lorenzo, Ming Li 0006, Linke Guo, Xiaonan Zhang 0001 |
INFOCOM | 1 |
| 2026 | Who Speaks What from Afar: Eavesdropping In-Person Conversations via mmWave SensingabstractMulti-participant meetings occur across various domains, such as business negotiations and medical consultations, during which sensitive information like trade secrets, business strategies, and patient conditions is often discussed. Previous research has demonstrated that attackers with mmWave radars outside the room can overhear meeting content by detecting minute speech-induced vibrations on objects. However, these eavesdropping attacks cannot differentiate which speech content comes from which person in a multi-participant meeting, leading to potential misunderstandings and poor decision-making. In this paper, we answer the question ``who speaks what''. By leveraging the spatial diversity introduced by ubiquitous objects, we propose an attack system that enables attackers to remotely eavesdrop on in-person conversations without requiring prior knowledge, such as identities, the number of participants, or seating arrangements. Since participants in in-person meetings are typically seated at different locations, their speech induces distinct vibration patterns on nearby objects. To exploit this, we design a noise-robust unsupervised approach for distinguishing participants by detecting speech-induced vibration differences in the frequency domain. Meanwhile, a deep learning-based framework is explored to combine signals from objects for speech quality enhancement. We validate the proof-of-concept attack on speech classification and signal enhancement through extensive experiments. The experimental results show that our attack can achieve the speech classification accuracy of up to $0.99$ with several participants in a meeting room. Meanwhile, our attack demonstrates consistent speech quality enhancement across all real-world scenarios, including different distances between the radar and the objects. Shaoying Wang, Hansong Zhou, Yukun Yuan |
INFOCOM | 1 |
| 2025 | DLSAG: Dynamic Load-aware Steiner Aggregation for Large-Scale Network Path OptimizationabstractThe growing device intelligence and distributed apps of the Internet of Things (IoT) across multiple fields have caused a sharp rise in wireless network traffic, posing more intense uplink resource competition and traffic scheduling optimization challenges for traditional wireless networks. To address these challenges and enhance the network’s load-bearing capacity and efficiency, this paper proposes a dynamic load-aware Steiner traffic aggregation algorithm (DLSAG), which integrates load awareness with a multi-candidate subtree strategy and employs a neural network to accelerate the computation process, thereby efficiently aggregating data flows. Theoretical analysis demonstrates that the time complexity of DLSAG is significantly reduced compared to traditional algorithm. Through extensive experiments across varying network scales and four baseline algorithms, it has been shown that DLSAG can reduce maximum link utilization by approximately 4.4% to 28.8%, while achieving a 1.5% to 10.2% improvement in Packet Delivery Ratio (PDR) and maintaining a competitive end-to-end delay, which is only about 7.4% to 10.3% higher than the optimal algorithm. Ruitao Li, Mingzhen Wu, Shaoying Wang, Fei Song 0001 |
GLOBECOM | 6 |
| 2025 | Non-Intrusive Speaker Diarization via mmWave SensingabstractSpeaker diarization refers to identifying who speaks what in a conversation. It is critical in sensitive settings like psychological counseling and legal consultations. However, traditional approaches, such as microphone or video, raise privacy concerns and cause discomfort to participants due to their noticeable deployment. To address this, we propose a non-intrusive speaker diarization system via mmWave sensing. Our approach leverages the spatial diversity of signals from multiple objects to distinguish speakers. Specifically, it isolates speech-induced vibrating objects signals and extracts speaker-related features through a two-stage feature extraction process. Our system achieves over 93% accuracy in real-world scenarios, demonstrating its effectiveness in reliably distinguishing speakers. Shaoying Wang, Hansong Zhou, Yukun Yuan 0001, Xiaonan Zhang 0001 |
SenSys | 1 |
| 2025 | FastTalker: Real-time audio-driven talking face generation with 3D Gaussian
Keliang Chen, Fang Cui, Mao Ni, Shaoying Wang, Junlin Che, Yonggang Qi, Fangwei Zhang, Gan Guo, Yunxia Huang |
Image Vis. Comput. | 5 |
| 2025 | Recursive Quadratic Filter Design for Non-Gaussian Systems Under Random Access Protocol: A Zero-Order Hold Strategy
Shaoying Wang, Zidong Wang 0001, Hongli Dong |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Waste Not, Want Not: Service Migration-Assisted Federated Intelligence for Multi-Modality Mobile Edge ComputingabstractFuture mobile edge computing (MEC) is envisioned to provide federated intelligence to delay-sensitive learning tasks with multimodal data. Conventional horizontal federated learning (FL) suffers from high resource demand in response to complicated multi-modal models. Multi-modal FL (MFL), on the other hand, offers a more efficient approach for learning from multi-modal data. In MFL, the entire multi-modal model is split into several sub-models with each tailored to a specific data modality and trained on a designated edge. As sub-models are considerably smaller than the multi-modal model, MFL requires fewer computation resources and reduces communication time. Nevertheless, deploying MFL over MEC faces the challenges of device mobility and edge heterogeneity, which, if not addressed, could negatively impact MFL performance. In this paper, we investigate an Service Migration-assisted Mobile Multi-modal Federated Learning (SM3FL) framework, where the service migration for sub-models between edges is enabled. To effectively utilize both communication and computation resources without extravagance in SM3FL, we develop the optimal strategies of service migration and data sample collection to minimize the wall-clock time, defined as the required training time to reach the learning target. Our experiment results show that the proposed SM3FL framework demonstrates remarkable performance, surpassing other state-of-art FL frameworks via substantially reducing the computing demand by 17.5% and dramatically decreasing the wall-clock time by 25.3%. Hansong Zhou, Shaoying Wang, Chutian Jiang, Xiaonan Zhang 0001, Linke Guo, Yukun Yuan 0001 |
MobiHoc | 2 |
| 2022 | A Dynamic Event-Triggered Approach to Recursive Nonfragile Filtering for Complex Networks With Sensor Saturations and Switching TopologiesabstractIn this article, the nonfragile filtering issue is addressed for complex networks (CNs) with switching topologies, sensor saturations, and dynamic event-triggered communication protocol (DECP). Random variables obeying the Bernoulli distribution are utilized in characterizing the phenomena of switching topologies and stochastic gain variations. By introducing an auxiliary offset variable in the event-triggered condition, the DECP is adopted to reduce transmission frequency. The goal of this article is to develop a nonfragile filter framework for the considered CNs such that the upper bounds on the filtering error covariances are ensured. By the virtue of mathematical induction, gain parameters are explicitly derived via minimizing such upper bounds. Moreover, a new method of analyzing the boundedness of a given positive-definite matrix is presented to overcome the challenges resulting from the coupled interconnected nodes, and sufficient conditions are established to guarantee the mean-square boundedness of filtering errors. Finally, simulations are given to prove the usefulness of our developed filtering algorithm. Shaoying Wang, Zidong Wang 0001, Hongli Dong, Yun Chen 0008 |
IEEE Trans. Cybern. | 1 |
| 2021 | Know Yourself and Know Others: Efficient Common Representation Learning for Few-shot Cross-modal RetrievalabstractLearning the common representations for various modalities of data is the key component in cross-modal retrieval. Most existing deep approaches learn multiple networks to independently project each sample into a common representation. However, each representation is only extracted from the corresponding data, which totally ignores the relationships between other data. Thus it is challenging to learn efficient common representations when lacking sufficient supervised multi-modal data for training, e.g., few-shot cross-modal retrieval. How to efficiently exploit the information contained in other examples is underexplored. In this work, we present the Self-Others Net, a few-shot cross-modal retrieval model that fully exploits information contained both in its own and other samples. First, we propose a self-network to fully exploit the correlations that lurk in the data itself. It integrates the features at different layers and extracts the multi-level information in the self-network. Second, an others-network is further proposed to model the relationships among all samples, which learns the Mahalanobis tensor and mixes the prototypes of all data to capture the non-linear dependencies for common representation learning. Extensive experiments are conducted on three benchmark datasets, which demonstrate clear improvements of the proposed method over the state-of-the-arts. Shaoying Wang, Hanjiang Lai |
ICMR | 1 |
| 2019 | Deep Policy Hashing Network with Listwise SupervisionabstractDeep-networks-based hashing has become a leading approach for large-scale image retrieval, which learns a similarity-preserving network to map similar images to nearby hash codes. The pairwise and triplet losses are two widely used similarity preserving manners for deep hashing. These manners ignore the fact that hashing is a prediction task on the list of binary codes. However, learning deep hashing with listwise supervision is challenging in 1) how to obtain the rank list of whole training set when the batch size of the deep network is always small and 2) how to utilize the listwise supervision. In this paper, we present a novel deep policy hashing architecture with two systems are learned in parallel: aquery network and a shared and slowly changingdatabase network. The following three steps are repeated until convergence: 1) the database network encodes all training samples into binary codes to obtain whole rank list, 2) the query network is trained based on policy learning to maximize a reward that indicates the performance of the whole ranking list of binary codes, e.g., mean average precision (MAP), and 3) the database network is updated as the query network. Extensive evaluations on several benchmark datasets show that the proposed method brings substantial improvements over state-of-the-art hashing methods. Shaoying Wang, Hanjiang Lai, Jian Yin 0001 |
ICMR | 1 |
| 2017 | Robust estimator design for networked uncertain systems with imperfect measurements and uncertain-covariance noises
Shaoying Wang, Huajing Fang, Xuegang Tian |
Neurocomputing | 1 |
| 2015 | Recursive estimation for nonlinear stochastic systems with multi-step transmission delays, multiple packet dropouts and correlated noises
Shaoying Wang, Huajing Fang, Xuegang Tian |
Signal Process. | 1 |