Yihao Zhang 0012

dblp:42/1023-12 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0002-0284-1367ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 On Mutation Testing of In-Context Learning Systems
Zeming Wei, Guanzhang Yue, Yihao Zhang 0012, Meng Sun 0002
J. Syst. Archit.3
2025 Automata-Based Steering of Large Language Models for Diverse Structured Generation
Xiaokun Luan, Zemin Wei, Yihao Zhang 0012, Meng Sun 0002
ICFEM3
2025 Component Composition in MedTiny: Multi-Level Constructs and Operational Semantics
abstract
MedTiny is a multi-level, component-based modeling language designed for specifying software systems and provides support for arbitrary-depth automaton hierarchies.This paper presents the high-level operational semantics of MedTiny's core components-Function and Automaton-which enforce strict encapsulation.A port-and-link mechanism enables arbitrary-depth component composition in MedTiny, thereby enhancing model reusability.To demonstrate MedTiny's expressiveness, we establish a formal correspondence between MedTiny and labeled transition systems (LTS).From a model interaction perspective, we contrast MedTiny's port synchronization with the label synchronization, which inherently supports only two-level composition, typically employed in LTS-based models.
Yihao Zhang 0012, Meng Sun 0002
SEKE2
2025 Robust and Efficient Watermarking of Large Language Models Using Error Correction Codes
abstract
Large language models (LLMs) have demonstrated remarkable performance in various tasks, but they also face challenges in intellectual property (IP) protection. Traditional training-based watermarking techniques are computationally expensive, while function invariant transformations (FITs) offer a lightweight alternative. Nevertheless, FIT-based watermarking methods are vulnerable to adaptive attacks, where adversaries can exploit the same transformation to remove or forge watermarks. We propose a novel white-box watermarking scheme that combines error correction codes (ECCs) with weight permutations. By encoding model identifiers using ECCs, our approach guarantees reliable watermark extraction under various attacks. Additionally, we develop a linear assignment-based extraction algorithm to enhance its efficiency. Evaluations on six LLMs show that our method offers robust watermarking capabilities. It has a minimal impact on model performance while effectively defending against removal and forgery attacks. Overall, our approach provides a scalable and secure solution for safeguarding the copyrights of LLMs.
Xiaokun Luan, Zeming Wei, Yihao Zhang 0012, Meng Sun 0002
Proc. Priv. Enhancing Technol.3
2024 Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
abstract
Since the rapid development of Large Language Models (LLMs) has achieved remarkable success, understanding and rectifying their internal complex mechanisms has become an urgent issue. Recent research has attempted to interpret their behaviors through the lens of inner representation. However, developing practical and efficient methods for applying these representations for general and flexible model editing remains challenging. In this work, we explore how to leverage insights from representation engineering to guide the editing of LLMs by deploying a representation discriminator as an editing oracle. We first identify the importance of a robust and reliable discriminator during editing, then propose an \textbf{A}dversarial \textbf{R}epresentation \textbf{E}ngineering (\textbf{ARE}) framework to provide a unified and interpretable approach for conceptual model editing without compromising baseline performance. Experiments on multiple tasks demonstrate the effectiveness of ARE in various model editing scenarios. Our code and data are available at \url{https://github.com/Zhang-Yihao/Adversarial-Representation-Engineering}.
Yihao Zhang 0012, Zeming Wei, Jun Sun 0001, Meng Sun 0002
NeurIPS1
2024 MILE: A Mutation Testing Framework of In-Context Learning Systems
Zeming Wei, Yihao Zhang 0012, Meng Sun 0002
SETTA2
2024 Weighted automata extraction and explanation of recurrent neural networks for natural language tasks
Zeming Wei, Xiyue Zhang 0001, Yihao Zhang 0012, Meng Sun 0002
J. Log. Algebraic Methods Program.3
2023 Using Z3 for Formal Modeling and Verification of FNN Global Robustness (S)
abstract
While Feedforward Neural Networks (FNNs) have achieved remarkable success in various tasks, they are vulnerable to adversarial examples.Several techniques have been developed to verify the adversarial robustness of FNNs, but most of them focus on robustness verification against the local perturbation neighborhood of a single data point.There is still a large research gap in global robustness analysis.The global-robustness verifiable framework DeepGlobal has been proposed to identify all possible Adversarial Dangerous Regions (ADRs) of FNNs, not limited to data samples in a test set.In this paper, we propose a complete specification and implementation of DeepGlobal utilizing the SMT solver Z3 for more explicit definition, and propose several improvements to DeepGlobal for more efficient verification.To evaluate the effectiveness of our implementation and improvements, we conduct extensive experiments on a set of benchmark datasets.Visualization of our experiment results shows the validity and effectiveness of the approach.
Yihao Zhang 0012, Zeming Wei, Xiyue Zhang 0001, Meng Sun 0002
SEKE1