EDBT 2026 Demo / reviewers in the wild / expert
Haoyang Li 0018
dblp:118/0004-18
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-8235-6753ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction AttacksabstractMachine learning models constitute valuable intellectual property, yet remain vulnerable to model extraction attacks (MEA), where adversaries replicate their functionality through black-box queries. Model watermarking counters MEAs by embedding forensic markers for ownership verification. Current black-box watermarks prioritize MEA survival through representation entanglement, yet inadequately explore resilience against sequential MEAs and removal attacks. Our study reveals that this risk is underestimated because existing removal methods are weakened by entanglement. To address this gap, we propose Watermark Removal attacK (WRK), which circumvents entanglement constraints by exploiting decision boundaries shaped by prevailing sample-level watermark artifacts. WRK effectively reduces watermark success rates by ≥88.79% across existing watermarking benchmarks. For robust protection, we propose Class-Feature Watermarks (CFW), which improve resilience by leveraging class-level artifacts. CFW constructs a synthetic class using out-of-domain samples, eliminating vulnerable decision boundaries between original domain samples and their artifact-modified counterparts (watermark samples). CFW concurrently optimizes both MEA transferability and post-MEA stability. Experiments across multiple domains show that CFW consistently outperforms prior methods in resilience, maintaining a watermark success rate of ≥70.15% in extracted models even under the combined MEA and WRK distortion, while preserving the utility of protected models. Yaxin Xiao, Qingqing Ye 0001, Zi Liang, Haoyang Li 0018, Ronghua Li 0002, Huadi Zheng, Haibo Hu 0001 |
AAAI | 4 |
| 2026 | Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
Dingjie Song, Tianlong Xu, Yifan Zhang 0004, Hang Li 0007, Zhiling Yan, Haoyang Li 0018, Lichao Sun 0001, Qingsong Wen |
AIED (1) | 7 |
| 2026 | TACE-Net: Two-Stage Asymmetric Conditional Enhancement for Weak-Source Recovery in Co-Channel FMabstractFor co-channel FM reception with two simultaneously active sources, two-pass constant modulus algorithm (CMA) can provide a coarse decomposition of the overlapped signals, but the weak branch often remains severely distorted after demodulation. We propose TACE-Net, a two-stage asymmetric conditional enhancement framework for weak-source recovery. Stage I refines the dominant CMA branch, and Stage II enhances the weak branch using the pre-CMA mixture, the weak branch, and the dominant branch refined in Stage I. To benchmark weakbranch recovery, we construct VCTK-Radio, a dataset simulating FM modulation, co-channel mixing, CMA-based separation, and demodulation using the VCTK corpus. On VCTK-Radio, TACENet improves DNSMOS-OVRL from 1.164 to 2.901 and PESQ from 1.169 to 1.901, while reducing WER from 71.25% to 28.72% on the weak branch, outperforming competitive baselines. Haoyang Li 0018, Ritesh Chandra Tewari, Wei Rao 0002, Sirajudeen Gulam Razul, Chng Eng Siong |
IEEE Signal Process. Lett. | 2 |
| 2025 | A Sample-Level Evaluation and Generative Framework for Model Inversion AttacksabstractModel Inversion (MI) attacks, which reconstruct the training dataset of neural networks, pose significant privacy concerns in machine learning. Recent MI attacks have managed to reconstruct realistic label-level private data, such as the general appearance of a target person from all training images labeled on him. Beyond label-level privacy, in this paper we show sample-level privacy, the private information of a single target sample, is also important but under-explored in the MI literature due to the limitations of existing evaluation metrics. To address this gap, this study introduces a novel metric tailored for training-sample analysis, namely, the Diversity and Distance Composite Score (DDCS), which evaluates the reconstruction fidelity of each training sample by encompassing various MI attack attributes. This, in turn, enhances the precision of sample-level privacy assessments. Leveraging DDCS as a new evaluative lens, we observe that many training samples remain resilient against even the most advanced MI attack. As such, we further propose a transfer learning framework that augments the generative capabilities of MI attackers through the integration of entropy loss and natural gradient descent. Extensive experiments verify the effectiveness of our framework on improving state-of-the-art MI attacks over various metrics including DDCS, coverage and FID. Finally, we demonstrate that DDCS can also be useful for MI defense, by identifying samples susceptible to MI attacks in an unsupervised manner. Haoyang Li 0018, Li Bai 0004, Qingqing Ye 0001, Haibo Hu 0001, Yaxin Xiao, Huadi Zheng, Jianliang Xu |
AAAI | 1 |
| 2025 | What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
Jiamin Chang, Haoyang Li 0018, Hammond A. Pearce, Ruoxi Sun 0001, Bo Li 0026, Minhui Xue 0001 |
CCS | 2 |
| 2025 | Speech Enhancement Using Continuous Embeddings of Neural Audio CodecabstractRecent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a novel, efficient SE approach by leveraging the pre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based SE methods, which process discrete speech tokens using Language Models (LMs), we perform SE within the continuous embedding space of the pretrained NAC, which is highly compressed along the time dimension for efficient representation. Our lightweight SE model, optimized through an embedding-level loss, delivers results comparable to SE baselines trained on larger datasets, with a significantly lower real-time factor of 0.005. Additionally, our method achieves a low GMAC of 3.94, reducing complexity 18-fold compared to Sepformer in a simulated cloud-based audio transmission environment. This work highlights a new, efficient NAC-based SE solution, particularly suitable for cloud applications where NAC is used to compress audio before transmission. Haoyang Li 0018, Jia Qi Yip, Tianyu Fan, Chng Eng Siong |
ICASSP | 1 |
| 2025 | LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation GenerationabstractPrevious fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining valuable information for both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that the current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.49% equal error rate. Hieu-Thi Luong, Haoyang Li 0018, Lin Zhang 0054, Kong-Aik Lee, Chng Eng Siong |
ICASSP | 2 |
| 2025 | Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for Privacy
Yaxin Xiao, Qingqing Ye 0001, Huadi Zheng, Haibo Hu 0001, Zi Liang, Haoyang Li 0018, Yijie Jiao |
ICCV | 7 |
| 2025 | From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based MethodologyabstractDeep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND show that GR-KAN requires up to 4× fewer parameters while improving PESQ by up to 0.1. In contrast, KAN, facing scalability issues, outperforms MLP on a small-scale signal modeling task but fails to improve MP-SENet. We demonstrate the first successful use of KAN-based methods for consistent improvement in both time- and SoTA TF-domain SE, establishing GR-KAN as a promising alternative for SE. Haoyang Li 0018, Chen Chen 0075, Sabato Marco Siniscalchi, Songting Liu, Chng Eng Siong |
INTERSPEECH | 1 |
| 2023 | 3DFed: Adaptive and Extensible Framework for Covert Backdoor Attack in Federated LearningabstractFederated Learning (FL), the de-facto distributed machine learning paradigm that locally trains datasets at individual devices, is vulnerable to backdoor model poisoning attacks. By compromising or impersonating those devices, an attacker can upload crafted malicious model updates to manipulate the global model with backdoor behavior upon attacker-specified triggers. However, existing backdoor attacks require more information on the victim FL system beyond a practical black-box setting. Furthermore, they are often specialized to optimize for a single objective, which becomes ineffective as modern FL systems tend to adopt in-depth defense that detects backdoor models from different perspectives. Motivated by these concerns, in this paper, we propose 3DFed, an adaptive, extensible, and multi-layered framework to launch covert FL backdoor attacks in a black-box setting. 3DFed sports three evasion modules that camouflage backdoor models: backdoor training with constrained loss, noise mask, and decoy model. By implanting indicators into a backdoor model, 3DFed can obtain the attack feedback in the previous epoch from the global model and dynamically adjust the hyper-parameters of these backdoor evasion modules. Through extensive experimental results, we show that when all its components work together, 3DFed can evade the detection of all state-of-the-art FL backdoor defenses, including Deepsight, Foolsgold, FLAME, FL-Detector, and RFLBAT. New evasion modules can also be incorporated in 3DFed in the future as it is an extensible framework. Haoyang Li 0018, Qingqing Ye 0001, Haibo Hu 0001, Jin Li 0002, Leixia Wang, Chengfang Fang, Jie Shi 0005 |
SP | 1 |