EDBT 2026 Demo / reviewers in the wild / expert
Yuejiao Wang
dblp:61/7539
· DBLP profile ↗
21ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SegTune: Structured and Fine-Grained Control for Song GenerationabstractYuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li, Haorui Zheng, Zewen Song, Zhongliang Liu, Chen Zhang, Pengfei Wan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuejiao Wang, Zihao Ji, Pengfei Cai, Haorui Zheng, Zewen Song, Zhongliang Liu |
ACL (1) | 1 |
| 2026 | SARLiquid: Through-package Liquid Leakage Detection based on mmWave SAR ImagingabstractLiquid leakage detection is critical for product quality and user safety, yet existing methods require line-of-sight (LoS) or direct contact with the liquid, or could even introduce additional health or safety risks. In this paper, we propose SARLiquid, a novel mmWave-based through-package liquid leakage detection method. SARLiquid leverages mmWave signals to see through packaging and identify leakage by reconstructing mmWave images. We employ synthetic aperture radar (SAR) technique to enhance imaging resolution and tame multipath effects. We also develop a dedicated algorithm to calibrate the discontinuous phase in SAR imaging results, and propose a deep learning model for liquid leakage detection and liquid identification. We implement SARLiquid and evaluate it across a wide range of scenarios. Results show that SARLiquid achieves average accuracies above 93% for both liquid leakage detection and liquid identification, 14.70% and 9.98% higher than the baselines, respectively. Zhanjun Hao 0001, Changlong Zhao, Yimiao Sun, Yuejiao Wang, Yuan He 0004 |
NOSSDAV | 4 |
| 2026 | mm-ARnet: Exploring Millimeter Wave Radar Point Clouds for Human Action RecognitionabstractHuman Action Recognition (HAR) offers a wide range of applications, including smart home, smart health, entertainment, security, and surveillance. Traditional vision-based HAR systems face significant limitations due to privacy concerns, lighting dependency, and poor performance in complex environments. Millimeter-wave radar-based activity recognition systems have attracted considerable attention due to their superior sensing capabilities, device-free deployment, privacy preservation, and robustness to environmental variations. However, existing approaches struggle with the inherent sparsity and noise in mmWave radar data, particularly for diverse activity categories spanning from full-body movements to subtle localized gestures. This study proposes mm-ARnet, a comprehensive millimeter-wave point-cloud-based framework for recognizing 16 diverse human activities across three distinct behavioral categories: full-body movements, posture transitions, and localized body movements. Our approach leverages 4D point cloud sequences and introduces a multi-frame fusion with stochastic sampling strategy to enhance point cloud density and mitigate sparsity effects. The core innovation lies in our lightweight TCN+Bi-LSTM temporal modeling pipeline integrated with a novel Temporal Pattern Attention (TPA) mechanism. Extensive experiments conducted across three real-world scenarios with 10 participants demonstrate that mm-ARnet achieves 97.42% accuracy, outperforming state-of-the-art methods while maintaining superior temporal performance. Zhanjun Hao 0001, Jiaxing Xiao, Yuejiao Wang, Fenfang Li |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | RFAR: Action Recognition Based on Single TagabstractHuman action recognition in classrooms has recently become a research hotspot. Traditional solutions usually rely on sensors or computer vision methods. However, these methods have some disadvantages, such as difficulty in deployment, susceptibility to ambient light, and privacy and security issues. This paper proposes RFAR, a contactless method for classroom action recognition. This method utilizes an RFID tag placed on the desktop to capture various actions and subsequently evaluate the student's learning status. To enhance the reliability of singletag identification, fused data consisting of two or three types of data sequences (RSSI, phase, and Doppler shift) are incorporated. Furthermore, a dynamic antenna system is utilized to identify the optimal angle for tag-antenna alignment. Notably, the single-tagper-person design eliminates severe interference among multiple tags and simplifies device deployment in multi-person scenarios. This method is proposed based on COTS RFID devices and shows high robustness across different environments and equipment. Experimental results show a recognition accuracy of 93.9% in single-person scenarios and 81.5% in five-person scenarios. Zhanjun Hao 0001, Yuejiao Wang, Fenfang Li, Hao Liu 0122, Chengrui Tao |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTCabstractMulti-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Training (SOT) for MTASR. Our visualization reveals that CTC guides the encoder to represent different speakers in distinct temporal regions of acoustic embeddings. Leveraging this insight, we propose a novel Speaker-Aware CTC (SACTC) training objective, based on the Bayes risk CTC framework. SACTC is a tailored CTC variant for multi-talker scenarios, it explicitly models speaker disentanglement by constraining the encoder to represent different speakers’ tokens at specific time frames. When integrated with SOT, the SOT-SACTC model consistently outperforms standard SOT-CTC across various degrees of speech overlap. Specifically, we observe relative word error rate reductions of 10% overall and 15% on low-overlap speech. This work represents an initial exploration of CTC-based enhancements for MTASR tasks, offering a new perspective on speaker disentanglement in multi-talker speech recognition .1 Jiawen Kang 0002, Lingwei Meng, Yuejiao Wang, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 4 |
| 2025 | Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile InstructionsabstractRecent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings1. Lingwei Meng, Shujie Hu, Jiawen Kang 0002, Zhaoqing Li, Yuejiao Wang, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 5 |
| 2025 | Adversarial Imitation Learning Based on Weighted Wasserstein Distance
Zhengzuo Qin, Yuejiao Wang, Dongdong Zhao 0002, Shi Yan 0002 |
ISNN | 2 |
| 2025 | SonicFER: Facial Expressions Tracking Through a Commercial Smartphone SpeakerabstractFacial expression recognition technology plays a significant role in advancing the intelligence and personalization of human-computer interaction. Although ultrasonic-based expression recognition methods already exist, most rely on phase shifts or Doppler effects, which suffer from insufficient resolution to accurately capture subtle facial expression variations. To address this issue, this paper proposes an innovative facial expression recognition system, SonicFER, which utilizes smartphones to emit ultrasonic waves and receive echoes, achieving fine-grained perception of facial expressions through real-time monitoring and analysis of channel impulse response (CIR). To effectively detect facial expression movements while eliminating static interference and minor motion artifacts, this paper employs a differential operation combined with variance calculation, along with setting appropriate thresholds for filtering. Through rigorous experimental evaluation, SonicFER achieves a high accuracy of 91.2% in recognizing six common facial expressions and outperforms state-of-the-art technologies across various real-world scenarios. Zhanjun Hao 0001, Zhuoxuan Yang, Yuejiao Wang, Mengqiao Li, Liang Cui |
IEEE Internet Things J. | 3 |
| 2024 | Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech ReconstructionabstractDysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges, we propose a novel multi-modal framework to utilize visual information, e.g., lip movements, in DSR as extra clues for reconstructing the highly abnormal pronunciations. The multi-modal framework consists of: (i) a multi-modal encoder to extract robust phoneme embeddings from dysarthric speech with auxiliary visual features; (ii) a variance adaptor to infer the normal phoneme duration and pitch contour from the extracted phoneme embeddings; (iii) a speaker encoder to encode the speaker’s voice characteristics; and (iv) a mel-decoder to generate the reconstructed mel-spectrogram based on the extracted phoneme embeddings, prosodic features and speaker embeddings. Both objective and subjective evaluations conducted on the commonly used UASpeech corpus show that our proposed approach can achieve significant improvements over baseline systems in terms of speech intelligibility and naturalness, especially for the speakers with more severe symptoms. Compared with original dysarthric speech, the reconstructed speech achieves 42.1% absolute word error rate reduction for patients with more severe dysarthria levels.1 Xueyuan Chen, Yuejiao Wang, Xixin Wu, Disong Wang, Zhiyong Wu 0001, Xunying Liu, Helen M. Meng |
ICASSP | 2 |
| 2024 | UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit NormalizationabstractDysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected by the neuromotor disorder and enhances their social inclusion. NED-based (Neural Encoder-Decoder) systems have significantly improved the intelligibility of the reconstructed speech as compared with GAN-based (Generative Adversarial Network) approaches, but the approach is still limited by training inefficiency caused by the cascaded pipeline and auxiliary tasks of the content encoder, which may in turn affect the quality of reconstruction. Inspired by self-supervised speech representation learning and discrete speech units, we propose a Unit-DSR system, which harnesses the powerful domain-adaptation capacity of HuBERT for training efficiency improvement and utilizes speech units to constrain the dysarthric content restoration in a discrete linguistic space. Compared with NED approaches, the Unit-DSR system only consists of a speech unit normalizer and a Unit HiFi-GAN vocoder, which is considerably simpler without cascaded sub-modules or auxiliary tasks. Results on the UASpeech corpus indicate that Unit-DSR outperforms competitive baselines in terms of content restoration, reaching a 28.2% relative average word error rate reduction when compared to original dysarthric speech, and shows robustness against speed perturbation and noise1. Yuejiao Wang, Xixin Wu, Disong Wang, Lingwei Meng, Helen M. Meng |
ICASSP | 1 |
| 2024 | Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
Lingwei Meng, Jiawen Kang 0002, Yuejiao Wang, Zengrui Jin, Xixin Wu, Xunying Liu, Helen M. Meng |
INTERSPEECH | 3 |
| 2024 | Large Language Model-based FMRI Encoding of Language Functions for Subjects with Neurocognitive Disorder
Yuejiao Wang, Xianmin Gong, Lingwei Meng, Xixin Wu, Helen M. Meng |
INTERSPEECH | 1 |
| 2024 | Synergy-Payoff-Maximization-Based Rechargeable Adaptive Energy-Efficient Dual-Mode Data Gathering Using Renewable Energy SourcesabstractIntegrating wireless energy transfer (WET) and data gathering based on the mobile platforms, such as the unmanned aerial vehicle (UAV) has been recognized as a promising technique to prolong the battery lifetime of resource-constrained wireless sensors in the Internet of Things era. However, it is challenging to jointly schedule dynamic renewable energy sources and communications resources to coordinate heterogeneous performance requirements in rechargeable wireless sensor networks (RWSNs). Hence, this article researches rechargeable adaptive energy-efficient dual-mode data gathering (AED2G) using renewable energy sources. First, considering the limited endurance of UAV and the uncertainty of renewable energy harvesting, a life-expectancy-balance-based AED2G strategy is proposed for optimizing the communication energy efficiency of the fixed data gathering (FDG) and mobile data gathering (MDG). Then, considering WET and MDG, the synergy payoff function of rechargeable MDG (RMDG) is designed, and the corresponding synergy payoff maximization problem is established. The problem is nonconvex due to the coupling of MDG and WET, so it is decomposed into two layers to be quickly solved by the designed hierarchical decomposition framework. The simulation results prove that our algorithm can efficiently use renewable energy sources, whether in FDG or RMDG mode, thereby improving the sustainability of RWSN. Haobo Guo, Yijia Ma, Shumin Sun, Yuejiao Wang, Bing Qi 0001, Juan Gao, Chen Xu 0002 |
IEEE Internet Things J. | 5 |
| 2024 | EarHear: Enabling the Deaf to Hear the World via Smartphone Speakers and MicrophonesabstractSign language plays a vital role in communication and learning for individuals with hearing and speech disabilities, serving as a common language for the deaf. Current state-of-the-art sign language recognition methods primarily rely on computer vision techniques, but they have certain limitations, including susceptibility to light interference and privacy concerns. Ubiquitous acoustic sensing provides new possibilities for sign language recognition, leveraging its high resistance to interference and cost effectiveness. However, existing methods face challenges in achieving satisfactory results due to environmental interference and the complexity of sign language recognition contexts. In this work, we propose EarHear, a robust contactless Chinese Sign Language Recognition and translation system. EarHear adopts a differential-Doppler data preprocessing method to cleverly mitigate the interference caused by the environment. To further identify differences in the morphology, speed, and direction of sign language actions and distinguish similar gestures, we propose the vision transformer for sign language recognition, which is able to model the context dependence of long-range features and output indeterminate long sign language sequences using an attention mechanism. As a result, computational speed and recognition accuracy are improved. Moreover, we explore a large-scale language-model-based sign language translation, which enables sign language recognition results to follow natural language standards, thus realizing a true sense of sign language recognition. The evaluation results based on 15 Chinese sentences show that our system achieves an average recognition rate of 93.38% and a BLEU-1 score of 80.73% for sign language translation, reaching the most advanced level in terms of accuracy and robustness. Zhanjun Hao 0001, Yuejiao Wang, Xiaochao Dang 0001 |
IEEE Internet Things J. | 2 |
| 2023 | A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker OneabstractAlthough automatic speech recognition (ASR) can perform well in common non-overlapping environments, sustaining performance in multi-talker overlapping speech recognition remains challenging. Recent research revealed that ASR model’s encoder captures different levels of information with different layers – the lower layers tend to have more acoustic information, and the upper layers more linguistic. This inspires us to develop a Sidecar separator to empower a well-trained ASR model for multi-talker scenarios by separating the mixed speech embedding between two suitable layers. We experimented with a wav2vec 2.0-based ASR model with a Sidecar mounted. By freezing the parameters of the original model and training only the Sidecar (8.7 M, 8.4% of all parameters), the proposed approach outperforms the previous state-of-the-art by a large margin for the 2-speaker mixed LibriMix dataset, reaching a word error rate (WER) of 10.36%; and obtains comparable results (7.56%) for LibriSpeechMix dataset when limited training. Lingwei Meng, Jiawen Kang 0002, Yuejiao Wang, Xixin Wu, Helen M. Meng |
ICASSP | 4 |
| 2022 | UltrasonicG: Highly Robust Gesture Recognition on Ultrasonic Devices
Zhanjun Hao 0001, Yuejiao Wang, Daiyang Zhang, Xiaochao Dang 0001 |
WASA (2) | 2 |
| 2022 | Flexible Gas-Permeable and Resilient Bowtie Antenna for Tensile Strain and Temperature SensingabstractAs a wireless basic unit, flexible antennas hold a wide range of applications in wearable electronics, soft robotics, and Internet of Things (IoT). However, most of the current flexible antennas are encapsulated by silicone elastomers with poor gas permeability, which severely hinders the evaporation of skin moisture and sweat. In addition, conventional rigid metals as high-frequency conductors are limited by poor elasticity and susceptibility to oxidation for on-skin application. Here, we developed a highly permeable and stretch-resistant flexible bowtie antenna that can capture changes in tensile strain and temperature. A low-impedance flexible carbon nanotube-silver (CNT-Ag) substrate was fabricated as the conductor of the antenna. By optimizing the multibeam bowed geometry and wrapping it in porous thermoplastic polyurethane (TPU) fibers, the final five-beam antenna was obtained and was able to withstand a relatively large tensile stress of 25.2 MPa, yet achieve a high vapor transmission rate of 48.2 mg cm−2 h−1. The antenna obtained an ideal impedance match at 2.28 GHz with doughnut-like radiation and a high radiation efficiency of over 85%. Furthermore, the antenna was successfully used to capture the strain in the wrist epidermis during bending and to detect thermal changes in the beaker of hot water, respectively. Finally, demonstrations of the antenna, such as permeability, radiation to the human body, and integrality in connection with flexible circuits, were carefully developed to reveal its feasibility in the real world. We expect this work to pave the way for the future establishment of epidermally flexible antennas for soft electronics. Hongcheng Xu, Weihao Zheng, Yangbo Yuan, Dandan Xu, Yuxin Qin, Ningjuan Zhao, Qikai Duan, Yujian Jin, Yuejiao Wang, Yang Lu 0002, Libo Gao |
IEEE Internet Things J. | 9 |
| 2022 | Role of Asymptomatic COVID-19 Cases in Viral Transmission: Findings From a Hierarchical Community Contact Network ModelabstractAs part of ongoing efforts to contain the coronavirus disease (COVID-19) pandemic, understanding the role of asymptomatic patients in the transmission system is essential for infection control. However, the optimal approach to risk assessment and management of asymptomatic cases remains unclear. This study proposed a Susceptible, Exposed, Infectious, No symptoms, Hospitalized and reported, Recovered, Death (SEINRHD) epidemic propagation model. The model was constructed based on epidemiological characteristics of COVID-19 in China and accounting for the heterogeneity of social contact networks. The early community outbreaks in Wuhan were reconstructed and fitted with the actual data. We used this model to assess epidemic control measures for asymptomatic cases in three dimensions. The impact of asymptomatic cases on epidemic propagation was examined based on the effective reproduction number, abnormally high transmission events, and type and structure of transmission. Management of asymptomatic cases can help flatten the infection curve. Tracing 75% of the asymptomatic cases corresponds to a 32.5% overall reduction in new cases (compared with tracing no asymptomatic cases). Regardless of population-wide measures, household transmission is higher than other types of transmission, accounting for an estimated 50% of all cases. The magnitude of tracing of asymptomatic cases is more important than the timing; when all symptomatic patients were traced, tested, and isolated in a timely manner, the overall epidemic was not sensitive to the time of implementing the measures to trace asymptomatic patients. Disease control and prevention within families should be emphasized during an epidemic.Note to Practitioners—This article addresses the urgent need to assess the risk of another COVID-19 outbreak caused by asymptomatic cases and to find the optimal, most practical approach to asymptomatic case management. Previous studies mostly focused on the clinical and statistical characteristics of asymptomatic cases; few have evaluated the impact of asymptomatic case measures using mathematical modeling at the community scale. This study proposed a Susceptible, Exposed, Infectious, No symptoms, Hospitalized and reported, Recovered, Death (SEINRHD) propagation model based on local community structures and social contact networks, according to the development characteristics and trend of COVID-19 in a Chinese community. The conclusion provides theoretical support for emergency work of relevant departments in different periods of an epidemic. In the early stages of the epidemic, timely detection and isolation of symptomatic patients should be a priority. Where there are surplus resources for epidemic prevention, the authorities should consider increasing the proportion of asymptomatic patients being traced. Epidemic prevention measures among family members should be a primary focus of attention. This combination of strategies can help reduce the rate of viral transmission and result in extinguishing the epidemic. Tianyi Luo, Zhidong Cao, Yuejiao Wang, Daniel Dajun Zeng, Qingpeng Zhang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2020 | Joint beamforming and power allocation using deep learning for D2D communication in heterogeneous networksabstractDevice‐to‐device (D2D) communication plays a significant role in cellular networks as it can increase the capacity, spectrum efficiency and energy efficiency of the system. However, the large computational complexity of D2D resource management optimisation algorithms creates a serious gap between theoretical design and real‐time processing, which leads to the limited use of D2D communication technology. In this study, a novel deep learning‐based optimisation method is proposed to overcome the high computational complexity of joint beamforming design and power allocation optimisation algorithms in D2D communication. Unlike existing approaches, the authors design a convolutional neural network based end‐to‐end network structure to solve complex computing problems for channel state information under a limited feedback scenario. The Max‐SE loss function which indicates quality‐of‐service (QoS) constraint and interference constraint, together with the mean squared error (MSE) function, are designed to maximise the spectral efficiency of the system while minimising the total transmit power. The simulation results show that the proposed approach can achieve performance comparable to the weighted minimum MSE scheme with low computation time. Yuejiao Wang, Shenghui Wang 0003, Lu Liu 0009 |
IET Commun. | 1 |
| 2019 | Social Cognition Construction of the Avian Flu based on Social Media Big DataabstractDuring the high incidence of avian flu, the mainstream media and social media report a lot on the epidemic, mobilizing the people to prevent and control avian flu. This paper collects reports on avian flu from News, Forums, Apps, WeChat and Microblog and forms five data sets. We extract agenda-settings from the News dataset and build agenda-setting networks of the five datasets. Then we use the QAP test to verify the relevance of these agenda-setting networks. We also project the agenda-setting dissimilarity matrices into a two-dimensional space using the MDS method to form cognitive maps, analyzing the cognitive drift of media platforms relative to News. Results show that the agenda-setting networks of Apps and News have the highest correlation coefficient of 0.9193, while Microblog and News have the lowest correlation coefficient of 0.5611. The cognitive maps of Apps, Forum and WeChat have a slight translation and rotation relative to the cognitive map of News. But their relative positional relationship among agenda-settings are similar with News, expect Microblog. Yuejiao Wang, Zhidong Cao |
ISI | 1 |
| 2009 | Analyzing the evolution of user-visible features: A case study with EclipseabstractIntegrated Development Environments (IDEs) help increase programmer productivity by automating much clerical and administrative work. Thus, it is of great research and practical interest to learn about the characteristics on how IDE features change and mature. To this end, we have conducted an empirical study, analyzing a total of 645 ldquoWhat's Newrdquo release note entries in 7 releases of the Eclipse IDE both quantitatively and qualitatively. It is found that majority of the changes are refinements or incremental additions to the feature architecture set up in early releases (1.0 and 2.0). Motivated by this, a further analysis on usability is performed to characterize how these changes impact programmers effectiveness in using the IDE. We summarize our study methodology and lessons learned. Daqing Hou, Yuejiao Wang |
ICSM | 2 |