VLDB 2026 Research / reviewers in the wild / expert
Pai Zhu
dblp:185/6950
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0005-1628-2991ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Soft Asynchronous Actor-Critic Framework for UAV-Assisted IoV Task Offloading With Vehicle Trajectory PredictionabstractWith the rapid increase in computing demand for the Internet of Vehicles (IoV), the limited on-board computing resources have gradually become a bottleneck restricting the efficient processing of tasks. Unmanned aerial vehicles (UAVs), with their high mobility and flexible deployment capabilities, have become an ideal platform for assisting IoV task offloading. However, the dynamic environment and the mobility of vehicles pose significant challenges for task offloading and resource allocation in UAV-assist IoV networks. This paper proposes a soft asynchronous actor-critic (SAAC) framework with vehicle trajectory prediction based on deep reinforcement learning (DRL) for UAV-assisted IoV task offloading. Firstly, in order to solve the problem of poor stability of traditional multi-agent deep reinforcement learning algorithms in dynamic environments, this paper combines the asynchronous training of multi-agent systems with the soft update and target network mechanism to improve the anti-interference and stability of the algorithm. Secondly, a module based on a recurrent neural network (RNN) integrated with a sliding window mechanism is proposed to address the uncertainty arising from vehicle mobility. This module enables precise prediction of vehicle trajectories, thereby facilitating forward-looking flight strategy planning for UAVs and reducing task transmission delays. Simulation results show that the proposed algorithm is effective, therefore this paper provides a new framework for the UAV-assisted IoV task offloading. Ying Yuan 0001, Pai Zhu, Cong Wang 0009, Guorui Li, Zhengmao Yao |
IEEE Internet Things J. | 2 |
| 2025 | Personalizing Keyword Spotting with Speaker InformationabstractKeyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows models to learn from different data inputs and features. We explore both Text-Dependent and Text-Independent speaker recognition systems to extract speaker information, and we experiment on extracting this information from both the input audio and pre-enrolled user audio. Evaluating our systems on a diverse dataset, our primary approach yields a notable 2.6% relative improvement on Equal Error Rate overall, particularly improving performance by 5.9% for children under 12 years old and up to 24% for underrepresented speaker groups. Moreover, our proposed approach only requires a small 1% increase in the number of parameters, with a minimum impact on latency and computational cost, which makes it a practical solution for real-world applications. Beltran Labrador, Pai Zhu, Guanlong Zhao, Angelo Scorza Scarpati, Alicia Lozano-Diez, Ignacio López-Moreno |
ICASSP | 2 |
| 2025 | GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
Harry Zhang, Kurt Partridge, Pai Zhu, Neng Chen, Hyun Jin Park, Dhruuv Agarwal |
INTERSPEECH | 3 |
| 2025 | LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword SpottingabstractCustom keyword spotting (KWS) allows detecting user-defined spoken keywords from streaming audio. This is achieved by comparing the embeddings from voice enrollments and input audio. State-of-the-art custom KWS models are typically trained contrastively using utterances whose keywords are randomly sampled from training dataset. These KWS models often struggle with confusing keywords, such as "blue" versus "glue". This paper introduces an effective way to augment the training with confusable utterances where keywords are generated and grouped from large language models (LLMs), and speech signals are synthesized with diverse speaking styles from text-to-speech (TTS) engines. To better measure user experience on confusable KWS, we define a new northstar metric using the average area under DET curve from confusable groups (c-AUC). Featuring high scalability and zero labor cost, the proposed method improves AUC by 3.7% and c-AUC by 11.3% on the Speech Commands testing set. Pai Zhu, Dhruuv Agarwal, Kurt Partridge |
INTERSPEECH | 1 |
| 2024 | GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-Shot Keyword SpottingabstractWe propose GE2E-KWS—a generalized end-to-end training and evaluation framework for customized keyword spotting. Specifically, enrollment utterances are separated and grouped by keywords from the training batch and their embedding centroids are compared to all other test utterance embeddings to compute the loss. This simulates runtime enrollment and verification stages, and improves convergence stability and training speed by optimizing matrix operations compared to SOTA triplet loss approaches. To benchmark different models reliably, we propose an evaluation process that mimics the production environment and compute metrics that directly measure keyword matching accuracy. Trained with GE2E loss, our 419KB quantized conformer model beats a 7.5 GB ASR encoder by 23.6% relative AUC, and beats a same size triplet loss model by 60.7% AUC. Our KWS models are natively streamable with low memory footprints, and designed to continuously run on-device with no retraining needed for new keywords (zero-shot). Pai Zhu, Jacob W. Bartel, Dhruuv Agarwal, Kurt Partridge, Hyun Jin Park |
SLT | 1 |
| 2023 | Locale Encoding for Scalable Multilingual Keyword Spotting ModelsabstractA Multilingual Keyword Spotting (KWS) system detects spoken keywords over multiple locales. Conventional monolingual KWS approaches do not scale well to multilingual scenarios because of high development/maintenance costs and lack of resource sharing. To overcome this limit, we propose two locale-conditioned universal models with locale feature concatenation and feature-wise linear modulation (FiLM). We compare these models with two baseline methods: locale-specific monolingual KWS, and a single universal model trained over all data. Experiments over 10 localized language datasets show that locale-conditioned models substantially improve accuracy over baseline methods across all locales in different noise conditions. FiLM performed the best, improving on average FRR by 61% (relative) compared to monolingual KWS models of similar sizes. Pai Zhu, Hyun Jin Park, Alex Park 0001, Angelo Scorza Scarpati, Ignacio López-Moreno |
ICASSP | 1 |
| 2021 | Noisy Student-Teacher Training for Robust Keyword SpottingabstractWe propose self-training with noisy student-teacher approach for streaming keyword spotting, that can utilize large-scale unlabeled data and aggressive data augmentation. The proposed method applies aggressive data augmentation (spectral augmentation) on the input of both student and teacher and utilize unlabeled data at scale, which significantly boosts the accuracy of student against challenging conditions. Such aggressive augmentation usually degrades model performance when used with supervised training with hard-labeled data. Experiments show that aggressive spec augmentation on baseline supervised training method degrades accuracy, while the proposed self-training with noisy student-teacher training improves accuracy of some difficult-conditioned test sets by as much as 60%. Hyun-Jin Park, Pai Zhu, Ignacio López-Moreno, Niranjan Subrahmanya |
Interspeech | 2 |
| 2020 | Training Keyword Spotting Models on Non-IID Data with Federated LearningabstractWe demonstrate that a production-quality keyword-spotting model can be trained on-device using federated learning and achieve comparable false accept and false reject rates to a centrally-trained model. To overcome the algorithmic constraints associated with fitting on-device data (which are inherently non-independent and identically distributed), we conduct thorough empirical studies of optimization algorithms and hyperparameter configurations using large-scale federated simulations. To overcome resource constraints, we replace memory intensive MTR data augmentation with SpecAugment, which reduces the false reject rate by 56%. Finally, to label examples (given the zero visibility into on-device data), we explore teacher-student training. Andrew Hard, Kurt Partridge, Cameron Nguyen, Niranjan Subrahmanya, Aishanee Shah, Pai Zhu, Ignacio López-Moreno, Rajiv Mathews |
INTERSPEECH | 6 |
| 2016 | Machine learning techniques with probability vector for cooperative spectrum sensing in cognitive radio networksabstractWe study cooperative spectrum sensing in cognitive radio networks (CRN) using machine learning techniques in this paper. A low-dimensional probability vector is proposed as the feature vector for machine learning based classification, instead of the N-dimensional energy vector in a CRN with a single primary user (PU) and N secondary users (SUs). This proposed method down-converts a high-dimensional feature vector to a constant two-dimensional feature vector for machine learning techniques while keeping the same spectrum sensing performance if not better. Due to its lower dimension, the probability vector based classification is capable of having a smaller training duration and a shorter classification time for testing vectors. Yingqi Lu, Pai Zhu, Michel Fattouche |
WCNC | 2 |