VLDB 2026 Research / reviewers in the wild / expert
Sichu Liang
dblp:327/3713
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0009-6798-1118ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM JudgesabstractMulti-agent LLM systems routinely generate multiple candidate responses that are aggregated by an LLM judge.To reduce the dominant prefill cost in such pipelines, recent work advocates KV cache reuse across partially shared contexts and reports substantial speedups for generation agents.In this work, we show that these efficiency gains do not transfer uniformly to judge-centric inference.Across GSM8K, MMLU, and HumanEval, we find that reuse strategies that are effective for execution agents can severely perturb judge behavior: end-task accuracy may appear stable, yet the judge's selection becomes highly inconsistent with dense prefill.We quantify this risk using Judge Consistency Rate (JCR) and provide diagnostics showing that reuse systematically weakens cross-candidate attention, especially for later candidate blocks.Our ablation further demonstrates that explicit crosscandidate interaction is crucial for preserving dense-prefill decisions.Overall, our results identify a previously overlooked failure mode of KV cache reuse and highlight judge-centric inference as a distinct regime that demands dedicated, risk-aware system design.1 Sichu Liang, Zhenglin Wang, Jiajia Chu, Hui Zang |
ACL (1) | 1 |
| 2026 | Auditing Partial Dataset Usage in Large Language Models via Fuzzy Membership AggregationabstractThe remarkable capabilities of Large Language Models (LLMs) are fueled by massive internet-scale corpora. However, scraped data owners often do not consent to its use for training, raising significant legal and ethical concerns over copyright and privacy.Data auditingtechniques seek to verify whether a protected dataset was used in training a target LLM, typically framing the task as membership inference: estimating binary sample-level membership and aggregating to a dataset-level decision. In this paper, we identify a fundamental limitation of this crisp binary paradigm: in realistic training pipelines, datasets are rarely used in full. Instead, models are trained on mixtures of partial subsets drawn from multiple sources. Existing auditing techniques, built upon anall-or-noneassumption—declaring a dataset either entirely present or absent from training—collapse inpartial dataset usagescenarios. Their predictions fluctuate unpredictably with the member ratio, causing unstable performance and high false-negative rates. Inspired byfuzzy set theory, we relax the crisp notion of binary membership to a continuousfuzzy membershipin [0,1], quantifying each sample's degree of inclusion in the model's training set. We establish a theoretical bridge between sample-level fuzzy memberships and the dataset-level usage ratio, facilitating inference of the proportion of a protected dataset used during training. Aneural network fuzzifierfirst estimates sample-level fuzzy memberships from binary labels in a reference set, then refines them using dataset-level member ratios as higher-order supervision. Finally, adefuzzificationstage aggregates calibrated memberships to determine partial usage. Across LLMs of varying scales and multiple auditing datasets, ourFuzzy Auditorsubstantially outperforms state-of-the-art crisp binary techniques in detecting partial usage, estimating member proportions, and identifying individual member samples. Hongyu Zhu 0004, Sichu Liang, Bofan Chen, Shi-Lin Wang, Zhuosheng Zhang 0001, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | Stealing Knowledge from Auditable DatasetsabstractThe success of modern deep learning hinges on vast training data, much of which is scraped from the web and may include copyrighted or private content—raising serious legal and ethical concerns when used without authorization. Dataset provenance seeks to identify whether a model has been trained on specific data collections, thus protecting copyright holders while preserving data utility. Existing techniques either watermark datasets to embed distinctive behaviors, or directly infer usage from discrepancies in model outputs between seen and unseen samples. These approaches exploit the fundamental problem of empirical risk minimization to overfit to seen features. Hence, provenance signals are considered inherently hard to erase, while the adversary’s perspective remains largely overlooked, limiting our ability to assess reliability in real-world scenarios. In this work, we present a unified framework that interprets both watermarking and inference-based provenance as manifestations of output divergence, modeling the interaction between auditor and adversary as a min-max game over such divergences. This perspective motivates DivMin, a simple yet effective learning strategy that minimizes the relevant divergence to suppress provenance cues. Experiments across diverse image datasets demonstrate that, starting from a pretrained vision-language model, DivMin retains over 93% of the full fine-tuning performance gain relative to a zero-shot baseline, while evading all six state-of-the-art auditing methods. Our findings establish divergence minimization as a direct and practical path to obfuscating provenance, offering a realistic simulation of potential adversary strategies to guide the development of more robust auditing techniques. Code and Appendix will be available at https://github.com/GradOpt/DivMin. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Shi-Lin Wang, Zhuosheng Zhang 0001 |
ECAI | 2 |
| 2025 | Efficient and Effective Model ExtractionabstractModel extraction aims to steal a functionally similar copy from a machine learning as a service (MLaaS) API with minimal overhead, typically for illicit profit or as a precursor to further attacks, posing a significant threat to the MLaaS ecosystem. However, recent studies have shown that model extraction is highly inefficient, particularly when the target task distribution is unavailable. In such cases, even substantially increasing the attack budget fails to produce a sufficiently similar replica, reducing the adversary’s motivation to pursue extraction attacks. In this paper, we revisit the elementary design choices throughout the extraction lifecycle. We propose an embarrassingly simple yet dramatically effective algorithm, Efficient and Effective Model Extraction (E3), focusing on both query preparation and training routine. E3achieves superior generalization compared to state-of-the-art methods while minimizing computational costs. For instance, with only 0.005× the query budget and less than 0.2× the runtime, E3outperforms classical generative model based data-free model extraction by an absolute accuracy improvement of over 50% on CIFAR-10. Our findings underscore the persistent threat posed by model extraction and suggest that it could serve as a valuable benchmarking algorithm for future security evaluations. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Shi-Lin Wang |
ICASSP | 3 |
| 2025 | Evading Data Provenance in Deep Neural NetworksabstractModern over-parameterized deep models are highly data-dependent, with large scale general-purpose and domain-specific datasets serving as the bedrock for rapid advancements. However, many datasets are proprietary or contain sensitive information, making unrestricted model training problematic. In the open world where data thefts cannot be fully prevented, Dataset Ownership Verification (DOV) has emerged as a promising method to protect copyright by detecting unauthorized model training and tracing illicit activities. Due to its diversity and superior stealth, evading DOV is considered extremely challenging. However, this paper identifies that previous studies have relied on oversimplistic evasion attacks for evaluation, leading to a false sense of security. We introduce a unified evasion framework, in which a teacher model first learns from the copyright dataset and then transfers task-relevant yet identifier-independent domain knowledge to a surrogate student using an out-of-distribution (OOD) dataset as the intermediary. Leveraging Vision-Language Models and Large Language Models, we curate the most informative and reliable subsets from the OOD gallery set as the final transfer set, and propose selectively transferring task-oriented knowledge to achieve a better trade-off between generalization and evasion effectiveness. Experiments across diverse datasets covering eleven DOV methods demonstrate our approach simultaneously eliminates all copyright identifiers and significantly outperforms nine state-of-the-art evasion attacks in both generalization and effectiveness, with moderate computational overhead. As a proof of concept, we reveal key vulnerabilities in current DOV methods, highlighting the need for long-term development to enhance practicality. Hongyu Zhu 0004, Sichu Liang, Zhuomeng Zhang, Fangqi Li 0001, Shi-Lin Wang |
ICCV | 2 |
| 2025 | Revisiting Data Auditing in Large Vision-Language ModelsabstractWith the surge of large language models (LLMs), Large Vision-Language Models (VLMs)-which integrate vision encoders with LLMs for accurate visual grounding-have shown great potential in tasks like generalist agents and robotic control. However, VLMs are typically trained on massive web-scraped images, raising concerns over copyright infringement and privacy violations, and making data auditing increasingly urgent. Membership inference (MI), which determines whether a sample was used in training, has emerged as a key auditing technique, with promising results on open-source VLMs like LLaVA (AUC > 80%). In this work, we revisit these advances and uncover a critical issue: current MI benchmarks suffer from distribution shifts between member and non-member images, introducing shortcut cues that inflate MI performance. We further analyze the nature of these shifts and propose a principled metric based on optimal transport to quantify the distribution discrepancy. To evaluate MI in realistic settings, we construct new benchmarks with i.i.d. member and non-member images. Existing MI methods fail under these unbiased conditions, performing only marginally better than chance. Further, we explore the theoretical upper bound of MI by probing the Bayes Optimality within the VLM's embedding space and find the irreducible error rate remains high. Despite this pessimistic outlook, we analyze why MI for VLMs is particularly challenging and identify three practical scenarios-fine-tuning, access to ground-truth texts, and set-based inference-where auditing becomes feasible. Our study presents a systematic view of the limits and opportunities of MI for VLMs, providing guidance for future efforts in trustworthy data auditing. Code and data will be available at https://github.com/GradOpt/Revisiting-VLM-MIA\faGithub. Hongyu Zhu 0004, Sichu Liang, Boheng Li, Tongxin Yuan, Fangqi Li 0001, Shi-Lin Wang, Zhuosheng Zhang 0001 |
ACM Multimedia | 2 |
| 2024 | Improve Deep Forest with Learnable Layerwise Augmentation Policy SchedulesabstractAs a modern ensemble technique, Deep Forest (DF) employs a cascading structure to construct deep models, providing stronger representational power compared to traditional decision forests. However, its greedy multi-layer learning procedure is prone to overfitting, limiting model effectiveness and generalizability. This paper presents AugDF, an optimized Deep Forest featuring learnable, layerwise data augmentation policy schedules. Specifically, We introduce the Cut Mix for Tabular data (CMT) augmentation technique to mitigate overfitting and develop a population-based search algorithm to tailor augmentation intensity for each layer. Additionally, we propose to incorporate outputs from intermediate layers into a checkpoint ensemble for more stable performance. Experimental results show that AugDF sets new state-of-the-art (SOTA) benchmarks in various tabular classification tasks, outperforming shallow tree ensembles, deep forests, deep neural network, and AutoML competitors. The learned policies also transfer effectively to Deep Forest variants, underscoring its potential for enhancing non-differentiable deep learning modules in tabular signal processing. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Yali Yuan, Shi-Lin Wang, Guang Cheng 0001 |
ICASSP | 2 |
| 2024 | Reliable Model Watermarking: Defending against Theft without Compromising on EvasionabstractWith the rise of Machine Learning as a Service (MLaaS) platforms, safeguarding the intellectual property of deep learning models is becoming paramount. Among various protective measures, trigger set watermarking has emerged as a flexible and effective strategy for preventing unauthorized model distribution. However, this paper identifies an inherent flaw in the current paradigm of trigger set watermarking: evasion adversaries can readily exploit the shortcuts created by models memorizing watermark samples that deviate from the main task distribution, significantly impairing their generalization in adversarial settings. To counteract this, we leverage diffusion models to synthesize unrestricted adversarial examples as trigger sets. By learning the model to accurately recognize them, unique watermark behaviors are promoted through knowledge injection rather than error memorization, thus avoiding exploitable shortcuts. Furthermore, we uncover that the resistance of current trigger set watermarking against removal attacks primarily relies on significantly damaging the decision boundaries during embedding, intertwining unremovability with adverse impacts. By optimizing the knowledge transfer properties of protected models, our approach conveys watermark behaviors to extraction surrogates without aggressive decision boundary perturbation. Experimental results on CIFAR-10/100 and Imagenette datasets demonstrate the effectiveness of our method, showing not only improved robustness against evasion adversaries but also superior resistance to watermark removal attacks compared to state-of-the-art solutions. Hongyu Zhu 0004, Sichu Liang, Fangqi Li 0001, Ju Jia, Shi-Lin Wang |
ACM Multimedia | 2 |