Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yidan Wang 0001

dblp:191/4607-1 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
0009-0006-2986-5333ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Security and privacy of machine learning · 38% Digital forensics and information hiding · 38% Privacy and data protection · 24%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization · ACL (1) 2025
Digital forensics and information hiding › watermarking › text watermarking
LLM watermarking
0.912025
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models · ACL (1) 2025
Security and privacy of machine learning › privacy attack
personally identifiable information extraction
0.912025
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization · ACL (1) 2025
Privacy and data protection › information leakage
privacy leakage
0.912025
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization · ACL (1) 2025
Digital forensics and information hiding
watermarking
0.912025
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › text generation
text quality preservation
0.312025
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models · ACL (1) 2025
Privacy and data protection
personally identifiable information
0.312025
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

sampling-based watermarking · 1.7logits-based watermarking · 1.7entropy-based embedding · 1.7in-context learning · 0.9gradient-based iterative optimization · 0.9
YearPublicationVenuePosition
2025 PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
abstract
Large Language Models (LLMs) excel in various domains but pose inherent privacy risks.Existing methods to evaluate privacy leakage in LLMs often use memorized prefixes or simple instructions to extract data, both of which well-alignment models can easily block.Meanwhile, Jailbreak attacks bypass LLM safety mechanisms to generate harmful content, but their role in privacy scenarios remains underexplored.In this paper, we examine the effectiveness of jailbreak attacks in extracting sensitive information, bridging privacy leakage and jailbreak attacks in LLMs.Moreover, we propose PIG, a novel framework targeting Personally Identifiable Information (PII) and addressing the limitations of current jailbreak methods.Specifically, PIG identifies PII entities and their types in privacy queries, uses in-context learning to build a privacy context, and iteratively updates it with three gradient-based strategies to elicit target PII.We evaluate PIG and existing jailbreak methods using two privacy-related datasets.Experiments on four white-box and two blackbox LLMs show that PIG outperforms baseline methods and achieves state-of-the-art (SoTA) results.The results underscore significant privacy risks in LLMs, emphasizing the need for stronger safeguards.
Yidan Wang 0001, Yanan Cao 0001, Yubing Ren, Fang Fang 0009, Zheng Lin 0001, Binxing Fang
ACL (1)1
2025 From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
abstract
The rise of Large Language Models (LLMs) has heightened concerns about the misuse of AI-generated text, making watermarking a promising solution.Mainstream watermarking schemes for LLMs fall into two categories: logits-based and sampling-based.However, current schemes entail trade-offs among robustness, text quality, and security.To mitigate this, we integrate logits-based and sampling-based schemes, harnessing their respective strengths to achieve synergy.In this paper, we propose a versatile symbiotic watermarking framework with three strategies: serial, parallel, and hybrid.The hybrid framework adaptively embeds watermarks using token entropy and semantic entropy, optimizing the balance between detectability, robustness, text quality, and security.Furthermore, we validate our approach through comprehensive experiments on various datasets and models.Experimental results indicate that our method outperforms existing baselines and achieves state-of-the-art (SOTA) performance.We believe this framework provides novel insights into diverse watermarking paradigms.Our code is available at https://github.com/redwyd/SymMark.
Yidan Wang 0001, Yubing Ren, Yanan Cao 0001, Binxing Fang
ACL (1)1