Yihao Peng

dblp:295/3458 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-9190-531XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 APIECHO: Training-Less Anomaly Detection via Intra-API Behavioral Comparison for Web Applications
Yihao Peng, Yiming Wu 0009, Du Wu, Shouling Ji, Hai Wan, Xibin Zhao
SP1
2025 AutoLabel: Automated Fine-Grained Log Labeling for Cyber Attack Dataset Generation
Yihao Peng, Tongxin Zhang, Jieshao Lai, Hai Wan, Xibin Zhao
USENIX Security Symposium1
2025 CPM-based Hierarchical Text Classification
abstract
In the field of natural language processing, hierarchical text classification (HTC) has emerged as a critical task for organizing and analyzing large volumes of text data. The previous work of HTC often falls short in fully leveraging the hierarchical structure of labels, resulting in suboptimal performance. In addition, it is difficult to capture nuanced relationships between parent and child classes, leading to inaccurate predictions and insufficient differentiation between sibling classes under the same parent category. This gap underscores the need for approaches that can more effectively integrate and utilize both hierarchical and corpus-specific information to improve HTC performance. To address these issues, Concept-aware Prompt Mechanism (CPM) is proposed for HTC, which leverages concept information embedded within hierarchical labels to enhance the representation of these labels and improve classification accuracy. Specifically, we introduce a concept initialization module that extracts concept features from hierarchical labels and a novel concept prompt template to integrate these features into the classification process. Our experimental results demonstrate that the proposed CPM achieves state-of-the-art performance on two benchmark datasets, improving Micro-F1 and Macro-F1 scores to varying degrees, particularly in datasets with complex label hierarchies.
Yihao Peng, Peilin Hong
J. Artif. Intell. Res.2
2023 TeSec: Accurate Server-side Attack Investigation for Web Applications
abstract
The user interface (UI) of web applications is usually the entry point of web attacks against enterprises and organizations. Finding the UI elements utilized by the intruders is of great importance both for attack interception and web application fixing. Current attack investigation methods targeting web UI either provide rough analysis results or have poor performance in high concurrency scenarios, which leads to heavy manual analysis work. In this paper, we propose TeSec, an accurate attack investigation method for web UI applications. TeSec makes use of two kinds of correlations. The first one, built from annotated audit log partitioned by PID/TID and delimiter-logs, captures the correspondence between audit log entries and web requests. The second one, modeled by an Aho-Corasick automaton built during system testing period, captures the correspondence between requests and the UI elements/events. Leveraging these two correlations, TeSec can accurately and automatically locate the UI elements/events (i.e., the root cause of the alarm) from an alarm, even in high concurrency scenarios. Furthermore, TeSec only needs to be deployed in the server and does not need to collect logs from the client-side browsers. We evaluate TeSec on 12 web applications. The experimental results show that the matching accuracy between UI events/elements and the alarm is above 99.6%. And security analysts only need to check no more than 2 UI elements on average for each individual forensics analysis. The maximum overhead of average response time and audit log space overhead are low (4.3% and 4.6% respectively).
Yihao Peng, Yilun Sun, Xuancheng Zhang, Hai Wan, Xibin Zhao
SP2
2021 Machine Learning Based Acceleration Method for Ordered Escape Routing
abstract
Escape routing, especially ordered escape routing, is a critical design stage for both printed circuit boards (PCBs) and integrated fan-out (InFO) wafer-level chip-scale packages. Previous works formulate ordered escape routing as boolean satisfiability (SAT) or integer linear programming (ILP) problems. Although optimal routing solutions can be obtained by above-mentioned approaches, the runtime is unacceptable for large-scale designs due to the exponential time complexity of SAT and ILP solvers. In this paper, we first attempt to address ordered escape routing problems with machine learning. We propose a learning-based method to accelerate existing solvers by reducing the solution space of the original problem. The proposed method is flexible, which can be combined with different ordered escape routing algorithms. Specifically, a fully convolutional neural network is trained to predict the probability of each routing grid to be occupied by routing paths. Thus, routing grids with low-probability usage can be removed to reduce the solution space. Experimental results show that the proposed method is effective for both SAT and ILP solvers of ordered escape routing. It achieves an acceleration of 4∼ 370x on average, with a slight increase in the total wirelength. Also, our model has a strong generalization ability. Although it is trained on $10\times 10$ pin array problems, it works well on larger problem sizes such as 14 x 14.
Zhiyang Chen 0006, Weiqing Ji, Yihao Peng, Datao Chen, Hailong Yao 0002
ACM Great Lakes Symposium on VLSI3