Qi Li 0033

dblp:181/2688-33 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-1720-2664ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
Yutong Jin, Qi Li 0033, Lingshuang Liu, Jianbing Ni
ACISP (3)2
2026 Bridging Black-Box and No-Box: Embedding Reconstruction Attacks on Deep Recognition Systems
abstract
Deep Neural Network (DNN)-based recognition systems are widely deployed for face and speaker authentication, yet remain vulnerable to Embedding Reconstruction Attacks (ERAs), in which adversaries recover biometric data from embeddings. Prior work assumes white-box or black-box access, requiring stronger adversarial knowledge than many real-world deployments provide. We introduce the first ERA framework that systematically characterizes settings withlessknowledge than black-box access. Our four-tier taxonomy progressively reduces adversarial capabilities, ranging from score-only and decision-only interfaces to no-query/no-feedback scenarios, mirroring the spectrum of commercial recognition APIs. To conduct ERAs under these constraints, we design high-fidelity reconstructors using Stable Diffusion for faces and flow-matching transformers for voices, trained via adaptive knowledge distillation. We formalize per-tier feasibility, proving Tiers 1–3 are practically exploitable and Tier 4 is infeasible under our formal threat model. Experiments on face and voice benchmarks show that our methods outperform existing black-box attacks under identical query budgets, achieving 93.27%, 82.60%, and 62.29% success rates for Tiers 1–3, respectively. We further evaluate compressed DNNs (pruned, quantized, and distilled models), providing the first systematic evidence that restricted-access recognition models remain at high risk in production.
Qi Li 0033, Jianbing Ni, Mohammad Zulkernine, Rongxing Lu
IEEE Trans. Dependable Secur. Comput.1
2025 SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
Xiangman Li, Qi Li 0033, Jianbing Ni, Rongxing Lu
ESORICS (1)3
2025 Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
abstract
Recent advances in LLM watermarking methods such as SynthID-Text by Google DeepMind offer promising solutions for tracing the provenance of AI-generated text. However, our robustness assessment reveals that SynthID-Text is vulnerable to meaning-preserving attacks, such as paraphrasing, copy-paste modifications, and back-translation, which can significantly degrade watermark detectability. To address these limitations, we propose SynGuard, a hybrid framework that combines the semantic alignment strength of Semantic Invariant Robust (SIR) with the probabilistic watermarking mechanism of SynthID-Text. Our approach jointly embeds watermarks at both lexical and semantic levels, enabling robust provenance tracking while preserving the original meaning. Experimental results across multiple attack scenarios show that SynGuard improves watermark recovery by an average of 11.1% in F1 score compared to SynthID-Text. These findings demonstrate the effectiveness of semantic-aware watermarking in resisting real-world tampering. All code, datasets, and evaluation scripts are publicly available at: https://github.com/githshine/SynGuard.
Xia Han, Qi Li 0033, Jianbing Ni, Mohammad Zulkernine
TrustCom2
2024 Proactive Audio Authentication Using Speaker Identity Watermarking
abstract
Generative AI, particularly through “deep fake” technology, stands at the crossroads of innovation and ethical dilemma. On one hand, it brings unprecedented advancements, transforming how we interact with digital content. On the other hand, it significantly compromises privacy and security, casting a shadow over the reliability of speaker recognition systems and fueling misuse in telecommunication fraud and manipulation of public opinion. This stark contrast not only raises legitimate concerns over the safety of sharing personal audio and video but also questions the very authenticity of digital media. To address the challenges of traceability in deepfake content and guarantee the integrity of audio, we propose a new solution specifically designed to counteract voice conversion and synthetic speech attacks. Leveraging cutting-edge deep learning technology, three extension strategies and ensemble learning of synthesis layer, this approach not only overcomes the inherent limitations of existing forensic methods but also resolves the issues associated with high-capacity watermarks. It achieves exceptionally high accuracy and imperceptibility across multiple speech datasets, various synthetic forgery methods, and numerous speech processing algorithms.
Qi Li 0033, Xiaodong Lin 0001
PST1
2023 Secure and Efficient Online Fingerprint Authentication Scheme Based On Cloud Computing
abstract
Privacy protection of biometrics-based on cloud computing is attracting increasing attention. In 2018, Zhuet al.proposed an efficient and privacy-preserving online fingerprint authentication scheme for data outsourcing e-Finga. Under the premise of ensuring user's fingerprint data privacy and message security authentication, the e-Finga scheme can provide accurate and efficient fingerprint identity authentication services. However, our analysis shows that the temporary fingerprint in this scheme uses the deterministic encryption algorithm, which has the risk of leaking the user's fingerprint characteristics. Therefore, we propose a temporary fingerprint attack method for the e-Finga scheme. Experiments demonstrate that an adversary can analyze specific secret parameters and fingerprint features when eavesdropping on a user's temporary fingerprint ciphertext. To counter the temporary fingerprint attack, we propose a secure e-fingerprint scheme– Secure e-finger that uses the learning with errors samples, which has the homomorphic addition property, to encrypt user's temporary fingerprints. Experiments show that the secure e-finger scheme can resist the temporary fingerprint attack. Compared with the unprotected e-Finga scheme, the client running time is increased by about 6% percent, the communication cost on the user side only increased by 0.3125% percent. As a result, our solution can realize secure online fingerprint authentication without losing efficiency. Single user authentication is likely to cause the problem of excessive authority. Based on the Secure e-finger scheme, we propose a threshold scheme based on biological characteristics.
Tanping Zhou, Zelun Yue, Wenchao Liu 0002, Yiliang Han, Qi Li 0033, Xiaoyuan Yang 0002
IEEE Trans. Cloud Comput.6
2021 Voxstructor: Voice Reconstruction from Voiceprint
Panpan Lu, Qi Li 0033, Hui Zhu 0001, Giuliano Sovernigo, Xiaodong Lin 0001
ISC2
2021 Efficient and Privacy-Preserving Speaker Recognition for Cybertwin-Driven 6G
abstract
With the introduction of cybertwin, a new approach to represent human or things in the cyberspace, it is foreseeable that vehicles will be able to offer more and more services in the future. Naturally, considering the safety of drivers, speaker recognition will be widely used in vehicle scenarios. Speaker recognition technologies are experiencing increasing popularity due to the unique and indissoluble link between individuals and their voices. However, the coming cybertwin-driven 6G brings speaker recognition technologies unprecedented challenges, especially in preventing the disclosure of voiceprint. To address these challenges, an efficient and privacy-preserving speaker recognition scheme for cybertwin-driven 6G, referred to as NEATEN, is proposed in this article. With NEATEN, the speaker identity can be recognized at multiple security levels without leaking the voiceprint data. More concretely, based on the random projection data perturbation, voiceprint perturbation algorithms in two phases and the corresponding ciphertext-based similarity computation algorithm are proposed. By using these algorithms, our efficient and accurate speaker recognition scheme can be achieved. Orthogonal to the previous works of biometric identification based on the Euclidean distance, NEATEN makes progress on the non-Euclidean distance, such as cosine distance and complicated distance. Detailed analysis shows that NEATEN can resist various known security threats. Experiments conducted on TIMIT and Voxceleb data sets have demonstrated that NEATEN is highly accurate and efficient, and can be flexibly deployed in a real cybertwin-driven 6G vehicle environment.
Qi Li 0033, Xiaodong Lin 0001
IEEE Internet Things J.1