VLDB 2026 Research / reviewers in the wild / expert
Yanru He
dblp:321/2754
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-9546-5043ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InvertVC: High-Fidelity Visual Cryptography via Invertible Diffusion-Based Image Synthesis
Yanyan Han, Suheng Wei, Chaofeng Liu, Yanru He |
ICIC (11) | 5 |
| 2026 | MLAW-HSI: Multi-level Adaptive Watermarking for Hyperspectral Image Classification Models
Hengbo Chen, Yanru He, Chuce He |
ICIC (11) | 3 |
| 2026 | RAD-Gov: A Privacy-Preserving RAG Framework for Few-Shot Government Data Classification
Yufeng Zhu, Yanru He, Xianwei Gao |
ICIC (3) | 4 |
| 2026 | How Far Are We from Automatically Identifying Violations of the Data Minimization Principle in Privacy Policies?abstractData protection laws and regulations require service providers to disclose data practices in privacy policies, specifying what personal information is processed and for what purposes. For compliance, these data practices must adhere to the data minimization principle, limiting the processing of personal information to what is directly relevant and necessary for the service purposes. However, data minimization is context-dependent, making violations difficult to define and quantify in privacy policies. Meanwhile, privacy policies are semantically complex and unstructured, hindering accurate extraction of fine-grained data practices and large-scale automated evaluation. To address these issues, we propose DataMini, a human--LLM collaborative evaluation framework for identifying violations of the data minimization principle in privacy policies. First, DataMini categorizes data minimization violations into two dimensions: inherent violations and contextual violations, establishing fine-grained evaluation criteria. Second, we construct a compliance baseline by mining high-frequency patterns from large-scale privacy policies and integrating expert knowledge to derive compliance mappings for human--LLM collaborative evaluation. Finally, the compliance baseline can automatically verify data practices that satisfy the data minimization principle, enabling the framework to focus exclusively on identifying suspected violations to improve efficiency and accuracy. Extensive evaluations demonstrate that DataMini exhibits superior data practice extraction accuracy of 83.46% and achieves an F1-score of 0.8180 for identifying data minimization violations in privacy policies, reducing manual evaluation effort by approximately 80%. Ziyan Zhou 0001, Yanru He, Yunchuan Guo, Liang Fang 0009, Fenghua Li 0001 |
SIGIR | 2 |
| 2025 | HBS Algorithmic Database Construction: A Chain-of-Thought-Driven Approach
Kaili Dou, Yanyan Han, Yanru He |
Inscrypt (3) | 6 |
| 2025 | A Watermarking Framework for Secure Distribution of Meteorological ImagesabstractMeteorological images typically contain sensitive and critical information, demanding effective security mechanisms to prevent unauthorized copying and distribution. Digital watermarking, a widely adopted protection technique, imperceptibly embeds user identity information into images to enable subsequent traceability. However, existing methods fail to consider the impact of watermark embedding on critical meteorological features, as well as the limitations of high computational overhead and low embedding efficiency in multi-user distribution scenarios. To address these issues, we propose a watermarking framework tailored for the secure distribution of meteorological images. Specifically, we first design a content-adaptive region selection scheme that avoids embedding watermarks in critical meteorological regions such as cloud features. We then develop a two-stage decoupled embedding strategy to enhance distribution efficiency, where the preprocessing stage performs region selection and frequency transformation, while the distribution stage reuses these cached results to rapidly generate user-specific watermarked copies. Extensive experimental results demonstrate that our framework effectively avoids critical regions in meteorological images and maintains high visual quality (average PSNR of 38.10 dB and SSIM of 0.9815) while achieving better extraction accuracy under various attacks. Furthermore, the two-stage decoupled embedding strategy reduces server response latency by over 60.8% in concurrent distribution scenarios. Fenghua Li 0001, Zifu Li, Yanru He |
TrustCom | 6 |
| 2025 | AutoPT: How Far Are We From the Fully Automated Web Penetration Testing?abstractPenetration testing is essential for ensuring Web security by identifying and mitigating vulnerabilities in advance, and the rapid progress of large language models (LLMs) shows great potential to revolutionize this process through intelligent, automated agents. In this work, we establish a comprehensive end-to-end penetration testing benchmark using a real-world penetration testing environment to explore the capabilities of LLM-based agents in this domain. Our results reveal that the agents are familiar to procedures of penetration testing tasks, but they still face limitations in generating accurate commands and executing complete processes. Accordingly, we summarize the current challenges, including the difficulty of maintaining the entire message history and the tendency for the agent to become stuck. Based on the above insights, we propose a Penetration testing State Machine (PSM) that utilizes the Finite State Machine (FSM) methodology to address these limitations. Then, we introduce AutoPT, an automated penetration testing agent based on the principle of PSM driven by LLMs, which utilizes the inherent inference ability of LLM and the constraint framework of state machines. Our evaluation results show that AutoPT outperforms the the ReAct-based baseline and improves the task completion rate from 22% to 41% on the benchmark target. Compared with the baseline and manual work, AutoPT also reduces time and economic costs further. In general, our AutoPT has facilitated the development of automated penetration testing and bring new findings and insights for both academia and industry. Benlong Wu, Kejiang Chen, Xiuwei Shang, Jiapeng Han, Yanru He, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | ProTegO: Protect Text Content against OCR Extraction AttackabstractOnline documents greatly improve the efficiency of information interaction but also cause potential security hazards, such as the ability to copy and reuse text content without authorization readily. To address copyright concerns, recent works have proposed converting reproducible text content into non-reproducible formats, making digital text content observable but not duplicable. However, as the Optical Character Recognition (OCR) technology develops, adversaries can still take screenshots of the target text region and use OCR to extract the text content. None of the existing methods can be well adapted to this kind of OCR extraction attack. In this paper, we propose "ProTegO'', a novel text content protection method against the OCR extraction attack, which generates adversarial underpaintings that do not affect human reading but can interfere with OCR after taking screenshots. Specifically, we design a text-style universal adversarial underpaintings generation framework, which can mislead both text recognition models and commercial OCR services. For invisibility, we take full advantage of the fusion property of human eyes and create complementary underpaintings to display alternatively on the screen. Experimental results demonstrate that ProTegO is a one-size-fits-all method that can ensure good visual quality while simultaneously achieving a high protection success rate on text recognition models with different architectures, outperforming the state-of-the-art methods. Furthermore, we validate the feasibility of ProTegO on a wide range of popular commercial OCR services, including Microsoft, Tencent, Alibaba, Huawei, Baidu, Apple, and Xiaomi. Codes will be available at https://github.com/Ruby-He/ProTegO. Yanru He, Kejiang Chen, Zehua Ma, Jie Zhang 0073, Huanyu Bian, Han Fang 0004, Weiming Zhang 0001, Nenghai Yu |
ACM Multimedia | 1 |
| 2023 | Investigating Neural-based Function Name Reassignment from the Perspective of Binary Code RepresentationabstractBuilding a model to reassign descriptive names for binary functions is considerable assistance for reverse engineering. Existing methods proposed for this issue are based on the low-level representation of binary code (e.g., assembly code), and especially the recent approaches employed neural-based models on instruction sequences. However, their performance is still unsatisfactory. Meanwhile, modern decompilers provide lifted representations of binary code, and their effectiveness has not been adequately studied. This paper further explores the issue of function name reassignment from the perspective of binary code representation. Specifically, we present a general and flexible NEural-based function name Reassignment framework NER, which leverages a decompiler to obtain a specific representation and applies the corresponding serialization strategy on it. NER then uses an alternative neural network to make predictions. Three levels of representation are investigated, including assembly code, Intermediate Representation (IR), and pseudo-code. We observe the binary code representations are significant for the final performance. It demonstrates that the pseudo-code is the most effective one. Based on these findings, we leverage the framework to implement a reassignment model NER-pc, which has 25% and 10% F1 score improvements against the state-of-the-art methods. Besides, more experiments are conducted to verify the design of NER and the effectiveness of NER-pc. Han Gao 0014, Jie Zhang 0073, Yanru He, Shaoyin Cheng, Weiming Zhang 0001 |
PST | 4 |