Mingzhe Gao

dblp:237/6062 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0005-9672-7400ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HIPPO: A Hierarchy-Preserving and Noise-Tolerant Pre-HLS Power Modeling Framework for FPGA
abstract
Power estimation for customized accelerators, especially those derived from high-level programming languages, entails the invocation of a long electronic design automation (EDA) tool chain, thus incurring large timing overhead that hinders early design optimization. To mitigate this problem, in this paper, we propose HIPPO, an architecture-level power modeling framework for field-programmable gate arrays (FPGAs). HIPPO operates directly on C/C++ programs, whose execution is prior to and independent of any EDA tool including the very front-end, high-level synthesis (HLS). During power modeling, HIPPO exploits the intrinsic C/C++ code hierarchies, including nested loops and operations, and enables multi-level power estimation that aligns with different code hierarchies. Specifically, HIPPO can be decomposed into (1) a code transformation flow that directly converts a C/C++ program with HLS pragmas into hardware-oriented and power-aware control and dataflow graph, (2) a hierarchy-preserving power modeling methodology that combines analytical modeling and data-driven learning approaches to effectively orchestrate different code hierarchies, and (3) an adaptive dataflow coarsening strategy which ensures modeling accuracy, efficiency and robustness by suppressing noise of onboard measurement. Experimental results demonstrate that HIPPO effectively decomposes and accurately predicts both dynamic and total power consumption, achieving average errors of 8.89% (dynamic) and 6.31% (total) for nested loops, and 9.86% (dynamic) and 3.41% (total) for single loops, respectively. These results prove that HIPPO paves the way for power-efficient high-level architecture exploration.
Zefan Lin, Zedong Peng, Mingzhe Gao, Jieru Zhao, Zhe Lin 0007
ICCAD3
2025 Distilling Benign Knowledge with Fine-Grained AST Fragments for Precise Real-World Web Shell Detection
abstract
Web shell detection has become increasingly crucial with the expansion of cloud computing, where automated malware analysis serves as a foundational approach. A key challenge in malware detection lies in balancing the reduction of false positives with maintaining detection accuracy amid rapid software ecosystem evolution. Existing methods require substantial expert intervention to mitigate false positives and often neglect the resource-intensive measures required to address model degradation caused by software updates. This study introduces ASTBAR, a novel method that extracts fine-grained AST fragments to distill benign behavioral knowledge from webserver software. By leveraging program structure and semantic analysis, ASTBAR generates fragment-level representations of benign samples and employs fragment matching to identify malware. Unlike prior techniques, ASTBAR achieves simultaneous improvements in precision, recall, and adaptability to software evolution. The evaluation results demonstrate that ASTBAR achieves an F1 score of$\mathbf{6 5. 3 5 \%}$, outperforming the state-of-theart methods by$\mathbf{1 0. 3 9 \%}$. In a$\mathbf{1 2}$-month industrial deployment spanning over one million users, ASTBAR maintained a 97.63% recall rat while reducing false positives by 700+ cases daily (equivalent to 30 expert hours).
Mingzhe Gao, Ligeng Chen, Yiling He, Lingyun Ying
IWQoS1
2025 PFortifier: Mitigating PHP Object Injection Through Automatic Patch Generation
abstract
PHP Object Injection (POI) vulnerabilities enable unexpected execution of class methods in PHP applications, resulting in various attacks. In the meanwhile, designing effective patches for POI vulnerabilities demands substantial engineering efforts. Existing research mostly focused on the detection of POI gadget chains, whereas the automatic patch generation remains an under-explored problem. In this work, we empirically study known gadget chains, and discover that adversaries usually construct gadget chains by diverging the execution to paths that developers never considered. The methods that get unexpectedly jump into (i.e., executed) are referred to as possible methods (PM). Based on the observation, we propose PFortifier, a framework for automatic POI patch generation. PFortifier operates in two stages: (i) the gadget chain detection phase, in which PFortifier simulates the execution of PHP applications, and detects gadget chains that pass attacker controlled objects to dangerous sinks, and (ii) the patch generation phase, in which PFortifier automatically generates POI patches by restricting PM jumps detected in the first phase. We evaluate PFortifier on 31 PHP applications and frameworks. The experiment results demonstrate the effectiveness of PFortifier: it generates precise patches for 52.53% of gadget chains, and suggests potential patches for 45.45% chains, resulting in a total chain coverage of 97.98%.
Mingzhe Gao, Ligeng Chen, Mingxue Zhang 0001, Gang Liang
SP3
2024 Swift-Mapping: Online Neural Implicit Dense Mapping in Urban Scenes
abstract
Online dense mapping of urban scenes is of paramount importance for scene understanding of autonomous navigation. Traditional online dense mapping methods fuse sensor measurements (vision, lidar, etc.) across time and space via explicit geometric correspondence. Recently, NeRF-based methods have proved the superiority of neural implicit representations by high-fidelity reconstruction of large-scale city scenes. However, it remains an open problem how to integrate powerful neural implicit representations into online dense mapping. Existing methods are restricted to constrained indoor environments and are too computationally expensive to meet online requirements. To this end, we propose Swift-Mapping, an online neural implicit dense mapping framework in urban scenes. We introduce a novel neural implicit octomap (NIO) structure that provides efficient neural representation for large and dynamic urban scenes while retaining online update capability. Based on that, we propose an online neural dense mapping framework that effectively manages and updates neural octree voxel features. Our approach achieves SOTA reconstruction accuracy while being more than 10x faster in reconstruction speed, demonstrating the superior performance of our method in both accuracy and efficiency.
Ke Wu 0021, Kaizhao Zhang, Mingzhe Gao, Jieru Zhao, Zhongxue Gan 0001, Wenchao Ding 0001
AAAI3
2024 Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
abstract
High-level synthesis (HLS) notably speeds up the hardware design process by avoiding RTL programming. However, the turnaround time of HLS increases significantly when post-route quality of results (QoR) are considered during optimization. To tackle this issue, we propose a hierarchical post-route QoR prediction approach for FPGA HLS, which features: (1) a modeling flow that directly estimates latency and post-route resource usage from C/C++ programs; (2) a graph construction method that effectively represents the control and data flow graph of source code and effects of HLS pragmas; and (3) a hierarchical GNN training and prediction method capable of capturing the impact of loop hierarchies. Experimental results show that our method presents a prediction error of less than 10% for different types of QoR metrics, which gains tremendous improvement compared with the state-of-the-art GNN methods. By adopting our proposed methodology, the runtime for design space exploration in HLS is shortened to tens of minutes and the achieved ADRS is reduced to 6.91% on average. Code and models are available at https://github.com/sjtu-zhao-lab/hierarchical-gnn-for-hls.
Mingzhe Gao, Jieru Zhao, Zhe Lin 0007, Minyi Guo
DATE1
2024 AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs
abstract
Recently, the use of large language models (LLMs) for software code generation, e.g., C/C++ and Python, has proven a great success. However, LLMs still suffer from low syntactic and functional correctness when it comes to the generation of register-transfer level (RTL) code, such as Verilog. To address this issue, in this paper, we develop AutoVCoder, a systematic open-source framework that significantly improves the LLMs' correctness of generating Verilog code and enhances the quality of its output at the same time. Our framework integrates three novel techniques, including a high-quality hardware dataset generation approach, a two-round LLM fine-tuning method and a domain-specific retrieval-augmented generation (RAG) mechanism. Experimental results demonstrate that AutoVCoder outperforms both industrial and academic LLMs in Verilog code generation. Code and models are available at https://github.com/sjtu-zhao-lab/AutoVCoder.
Mingzhe Gao, Jieru Zhao, Zhe Lin 0007, Wenchao Ding 0001, Xiaofeng Hou, Yu Feng 0007, Chao Li 0009, Minyi Guo
ICCD1
2023 RecMaL: Rectify the malware family label via hybrid analysis
Mingzhe Gao, Ligeng Chen, Zhengxuan Liu, Lingyun Ying
Comput. Secur.2
2019 An improved kernel-based incremental extreme learning machine with fixed budget for nonstationary time series prediction
Aiqiang Xu, Dianfa Ping, Mingzhe Gao
Neural Comput. Appl.4