VLDB 2026 Research / reviewers in the wild / expert
Boyu Zhu
dblp:243/4419
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone ArraysabstractSpeech separation using ad-hoc microphone arrays has been explored, but there is still significant room for improvement, especially in complex scenarios with varying channel conditions. Co-attention, a feature fusion mechanism, is widely used in multimodal fusion to capture the cooperation between modalities and enhance the representation of extracted features. In this paper, we propose a co-attention-based multi-channel model for speech separation with ad-hoc microphone arrays. The co-attention mechanism is integrated into the model to enhance the interaction between different speakers across multiple channels, enabling efficient channel fusion. To the best of our knowledge, this is the first work to apply co-attention for speech separation. Experimental results demonstrate that the proposed method significantly outperforms existing approaches, underscoring the importance of co-attention in optimizing channel fusion for speech separation in challenging acoustic environments. Hongmei Guo, Linfeng Feng, Xueqing Li 0003, Boyu Zhu, Xiao-Lei Zhang 0001, Xuelong Li 0001 |
ICASSP | 5 |
| 2025 | Biometric Confusion Matrix and Inter ZooPlot: Two Novel Visualizations for Biometric Verification EvaluationabstractBiometric Verification Systems (BVS) often suffer from misclassification errors, which are frequently concentrated around a small subset of users for whom the system performs poorly. The Biometric Menagerie was introduced to categorize such users based on their biometric behavior. However, its main representation, the ZooPlot, fails to accurately distinguish all categories defined in the original Menagerie, particularly "Lambs" (easily impersonated) and "Wolves" (frequently impersonate others).This paper proposes two new visualizations for evaluating BVS: Inter ZooPlot and the Biometric Confusion Matrix (BCM). A user study was conducted to assess the effectiveness of different visualizations. Inter ZooPlot and BCM achieved average accuracies of 90.0% and 89.4%, respectively, in distinguishing the four original categories defined in the Biometric Menagerie, outperforming the baseline ZooPlot at 73.9%. Furthermore, we show that BCM can reveal sources of user-specific errors and highlight system imbalances, making it a promising post-hoc explainability method for biometric system analysis. All additional materials are available at: https://github.com/Boyu1998/BCM. Boyu Zhu, Romain Giot |
IJCB | 1 |
| 2025 | Augment Mandarin to Cantonese Speech Databases via Retrieval-Augmented Generation and Speech Synthesis
Boyu Zhu, Ruihao Jing, Chunyu Qiang, Tianrui Wang |
INTERSPEECH | 3 |
| 2025 | EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated CodeabstractExisting code generation benchmarks primarily evaluate functional correctness, with limited attention to code efficiency, and they are often restricted to a single language such as Python. To address this gap, we introduce EffiBench‑X, the first large‑scale multi‑language benchmark specifically designed for robust efficiency evaluation of LLM‑generated code. EffiBench‑X supports Python, C++, Java, JavaScript, Ruby, and Go, and comprises competitive programming tasks paired with human‑expert solutions as efficiency baselines. Evaluating state‑of‑the‑art LLMs on EffiBench‑X reveals that while models frequently generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM‑generated solutions (e.g., Qwen3‑32B) achieve only around 62% of human efficiency on average, with significant language‑specific variation: models tend to perform better in Python, Ruby, and JavaScript than in Java, C++, and Go (e.g., DeepSeek‑R1’s Python code is markedly more efficient than its Java code). These findings highlight the need for research into optimization‑oriented methods to improve the efficiency of LLM‑generated code across diverse languages. The dataset and evaluation infrastructure are publicly available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x. Yuhao Qing, Boyu Zhu, Mingzhe Du, Zhijiang Guo, Terry Yue Zhuo, Qianru Zhang, Jie Zhang 0050, Heming Cui, Siu-Ming Yiu, Dong Huang 0005, See-Kiong Ng, Anh Tuan Luu |
NeurIPS | 2 |
| 2024 | Differentially Private K-Means Publishing with Distributed DimensionsabstractIn this paper, we address the critical concerns related to dataset privacy in the context of k-means clustering publishing within a distributed dimension setting. By leveraging differential privacy mechanisms, we propose a novel framework that integrates a differentially private classifier, constructed through voting based on raw clustering results, and an enhanced generative adversarial network (GAN) simulating the classifier’s behavior in inferring class labels for a public dataset. Our approach generates synthetic clustering results that mimic real outcomes in classification tasks, ensuring differential privacy and minimizing noise. Our contributions include a comprehensive exploration of privacy issues, the introduction of a novel privacy-preserving k-means clustering framework, and theoretical analyses demonstrating sensitivity and differential privacy guarantees. Evaluation on the MNIST dataset demonstrates the effectiveness of the framework, achieving 82.22% accuracy with a (10.48, 10−9)-differential-privacy guarantee, compared to 83.45% accuracy without privacy-preserving. Boyu Zhu, Yuan Zhang 0004, Tingting Chen 0001, Sheng Zhong 0002 |
CSCWD | 1 |
| 2024 | Toward Universal Detection of Adversarial Examples via Pseudorandom ClassifiersabstractAdversarial examples that can fool neural network classifiers have attracted much attention. Existing approaches to detect adversarial examples leverage a supervised scheme in generating attacks (either targeted or non-targeted) for training the detectors, which means the detectors are geared to the attacks chosen at the training time and could be circumvented if the adversary does not act as expected. In this paper, we borrow ideas from cryptography and present a novel approach called pseudorandom classifier. In a nutshell, a pseudorandom classifier is a classifier equipped with a mapping to encode the category labels into random multi-bit labels, and a keyed pseudorandom injective function to transform the input to the classifier. The multi-bit labels enable attack-independent and probabilistic detection if the input sample is adversarial. The pseudorandom injection makes the existing white-box adversarial example generation methods, largely based on back-propagation, no longer applicable. We empirically evaluate our method on MNIST, CIFAR10, Imagenette, CIFAR100, and GTSRB. The results suggest that its performance against adversarial examples is comparable to the state-of-the-art. Boyu Zhu, Changyu Dong, Yuan Zhang 0004, Yunlong Mao, Sheng Zhong 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Secure Deep Neural Network Models Publishing Against Membership Inference Attacks Via Training Task ParallelismabstractVast data and computing resources are commonly needed to train deep neural networks, causing an unaffordable price for individual users. Motivated by the increasing demands of deep learning applications, sharing well-trained models becomes popular. The owner of a pre-trained model can share it by publishing the model directly or providing a prediction interface. Either way, individual users can benefit from deep learning without much cost, and computing resources can be saved. However, recent studies of machine learning security have identified severe threats to these model publishing approaches. This paper will focus on the privacy leakage issue of publishing well-trained deep neural network models. To tackle this problem, we propose a series of secure model publishing solutions based on training task parallelism. Specifically, we show how to estimate private model parameters through parallel model training and generate new model parameters in a privacy-preserving manner to replace the original ones for publishing. Based on data parallelism and parameter generating techniques, we design another two solutions concentrating on model quality and parameter privacy, respectively. Through privacy leakage analysis and experimental attack evaluation, we conclude that deep neural network models published with our solutions can provide on-demand model quality guarantees and resist membership inference attacks. Yunlong Mao, Wenbo Hong, Boyu Zhu, Zhifei Zhu, Yuan Zhang 0004, Sheng Zhong 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Private Deep Neural Network Models Publishing for Machine Learning as a ServiceabstractMachine learning as a service has emerged recently to relieve tensions between heavy deep learning tasks and increasing application demands. A deep learning service provider could help its clients to benefit from deep learning techniques at an affordable price instead of huge resource consumption. However, the service provider may have serious concerns about model privacy when a deep neural network model is published. Previous model publishing solutions mainly depend on additional artificial noise. By adding elaborated noises to parameters or gradients during the training phase, strong privacy guarantees like differential privacy could be achieved. However, this kind of approach cannot give guarantees on some other aspects, such as the quality of the disturbingly trained model and the convergence of the modified learning algorithm. In this paper, we propose an alternative private deep neural network model publishing solution, which caused no interference in the original training phase. We provide privacy, convergence and quality guarantees for the published model at the same time. Furthermore, our solution can achieve a smaller privacy budget when compared with artificial noise based training solutions proposed in previous works. Specifically, our solution gives an acceptable test accuracy with privacy budget ϵ = 1. Meanwhile, membership inference attack accuracy will be deceased from nearly 90% to around 60% across all classes. Yunlong Mao, Boyu Zhu, Wenbo Hong, Zhifei Zhu, Yuan Zhang 0004, Sheng Zhong 0002 |
IWQoS | 2 |
| 2020 | Secure Inter-Domain Forwarding Loop Test in Software Defined NetworksabstractDebugging a traditional network is notoriously difficult due to network devices' heterogeneity and protocols' decentralized nature, but Software-Defined Networking (SDN) is changing this predicament. Recent works have provided very nice approaches for an administrator to perform several fundamental network tests in a single-domain SDN network. However, how to perform these tests securely in multi-domain networks still remains open. In this paper, we study the highly challenging problem of inter-domain forwarding loop test in a SDN environment. We present two novel testing protocols that can be used for inter-domain loop tests. Both protocols are secure in the sense that they protect each domain's private information about its topology and configuration. The first protocol, based on random sampling, is highly efficient with a small error probability diminishing exponentially in the sample size. The second protocol, based on secure set intersection test, guarantees 100 percent accuracy of the result, although not as efficient as the first one. We provide rigorous proofs for the security and accuracy guarantees, and show our protocols have very good efficiency by testing them with real-world network data. Yuan Zhang 0004, Boyu Zhu, Yixin Fang, Suxin Guo, Aidong Zhang 0001, Sheng Zhong 0002 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2019 | Fast Chosen-Key Distinguish Attacks on Round-Reduced AES-192
Chunbo Zhu, Gaoli Wang, Boyu Zhu |
ACISP | 3 |