Ligeng Chen

dblp:247/6549 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-0076-3708ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 2 first-author · 7 since 2021Security and privacy · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert Substitution
Guoying Zhu, Meng Li 0010, Haipeng Dai 0001, Weijun Wang 0001, Ligeng Chen
ISCA8
2025 MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization
abstract
Deploying large language models (LLMs) with low-rank adaptation (LoRA) on mobile devices is promising due to their capability to complete diverse domain-specific tasks while ensuring privacy and accessibility. In this paper, we introduce MobiLoRA to accelerate LoRA-based LLM inference on mobile devices. MobiLoRA focuses on optimizing the key-value (KV) caches due to the limited computing and memory resources of mobile devices. The key insight of MobiLoRA lies in the utilization of two contexts for on-device LoRA serving: semantic-level contexts, such as prompts with shared prefixes, and system-level contexts, such as the application status (e.g., foreground or killed) of LLM requests. Specifically, for semantic-level contexts, MobiLoRA proposes similarity-aware delta encoding, which leverages token-wise similarity in KV caches across LoRA adapters for efficient storage and reuse. Furthermore, MobiLoRA advocates context-aware KV cache management to optimize cache retention and eviction considering the system-level contexts. We fully implement MobiLoRA and compare it with state-of-the-art LLM serving frameworks using real-world mobile device traces. Results show that MobiLoRA accelerates LoRA-based LLM inference by 57.6% on mobile devices.
Borui Li 0001, Haoran Ma 0006, Ligeng Chen
ACL (1)4
2025 MIINT: Infuse Intuitive Data Correspondence for Model Interpretation
abstract
To help humans understand the mechanism of machine learning models seamlessly, various works are dedicated to opening the black boxes of current learning methods. The mainstream of existing work attempts to calculate feature importance directly from input to output but still suffers from infidelity, inconsistency, instability, and complexity. To improve the performance of interpretable machine learning methods, we find that there are natural correlations between the prediction results and the features under specific tasks, called intuitive data correspondence. Specifically, in image classification tasks, image segmentation can be regarded as a kind of intuitive selection to sketch the relationship between the object and the classification result. Based on this heuristic observation, we propose MIINT to integrate intuitive data correspondence into the traditional post-hoc interpretable machine learning techniques with partial dependency. We conduct experiments on 2 classic image classification datasets, and the results show that MIINT has better fidelity, consistency, stability, and sparsity compared with the baseline method LIME: MIINT reduces infidelity and complexity by more than 10%, and increases consistency and stability by more than 10% and 25% respectively.
Ligeng Chen, Bing Mao 0001
ICME2
2025 Distilling Benign Knowledge with Fine-Grained AST Fragments for Precise Real-World Web Shell Detection
abstract
Web shell detection has become increasingly crucial with the expansion of cloud computing, where automated malware analysis serves as a foundational approach. A key challenge in malware detection lies in balancing the reduction of false positives with maintaining detection accuracy amid rapid software ecosystem evolution. Existing methods require substantial expert intervention to mitigate false positives and often neglect the resource-intensive measures required to address model degradation caused by software updates. This study introduces ASTBAR, a novel method that extracts fine-grained AST fragments to distill benign behavioral knowledge from webserver software. By leveraging program structure and semantic analysis, ASTBAR generates fragment-level representations of benign samples and employs fragment matching to identify malware. Unlike prior techniques, ASTBAR achieves simultaneous improvements in precision, recall, and adaptability to software evolution. The evaluation results demonstrate that ASTBAR achieves an F1 score of$\mathbf{6 5. 3 5 \%}$, outperforming the state-of-theart methods by$\mathbf{1 0. 3 9 \%}$. In a$\mathbf{1 2}$-month industrial deployment spanning over one million users, ASTBAR maintained a 97.63% recall rat while reducing false positives by 700+ cases daily (equivalent to 30 expert hours).
Mingzhe Gao, Ligeng Chen, Yiling He, Lingyun Ying
IWQoS2
2025 Unleashing the Power of LLM to Infer State Machine From the Protocol Implementation
abstract
State machines are essential for enhancing protocol analysis to identify vulnerabilities. However, inferring state machines from network protocol implementations is challenging due to complex code syntax and semantics. Traditional dynamic analysis methods often miss critical state transitions due to limited coverage, while static analysis faces path explosion issues. To overcome these challenges, we introduce a novel state machine inference approach utilizing Large Language Models (LLMs), named ProtocolGPT. This method employs retrieval augmented generation technology to enhance a pre-trained model with specific knowledge from protocol implementations. Through effective prompt engineering, we accurately identify and infer state machines. To the best of our knowledge, our approach represents the first state machine inference that leverages the source code of protocol implementations. Our evaluation of six protocol implementations shows that our method achieves a precision of over 90 %, outperforming the baselines by more than 30 %. Furthermore, integrating our approach with protocol fuzzing improves coverage by more than 20 % and uncovers two 0-day vulnerabilities compared to baseline methods.
Haiyang Wei, Ligeng Chen, Zhengjie Du, Haohui Huang, Guang Cheng 0001, Fengyuan Xu, Linzhang Wang, Bing Mao 0001
IWQoS2
2025 Uncovering Prompt Elements: Cloning System Prompts from Behavioral Traces
abstract
We introduce prompt cloning, a new black-box attack that reconstructs functionally equivalent system prompts rather than extracts original system prompts. Unlike prompt stealing, prompt cloning exploits the insight that system prompts leave persistent behavioral traces in outputs, even under strong alignment and prompt-level defenses. Our method decomposes system behavior into semantically interpretable elements, selectively elicits them through carefully designed queries, and aggregates representative traces to synthesize high-fidelity cloned prompts. Extensive evaluations show that cloned prompts replicate functional behavior with up to 85% semantic similarity, outperforming base LLMs by up to 8%, and even exceeding original system prompts when transferred to different back-end models. We also conduct a large-scale study on GitHub repositories, revealing that single-prompt architectures remain widespread in open-source LLM applications, reinforcing the real-world relevance of our threat model. Our findings reveal that prompt cloning enables unauthorized replication of confidential LLM behavior and underscore the urgent need for defenses that go beyond hiding prompt text.
Hao Wu 0067, Ligeng Chen, Bing Mao 0001
ASE4
2025 PFortifier: Mitigating PHP Object Injection Through Automatic Patch Generation
abstract
PHP Object Injection (POI) vulnerabilities enable unexpected execution of class methods in PHP applications, resulting in various attacks. In the meanwhile, designing effective patches for POI vulnerabilities demands substantial engineering efforts. Existing research mostly focused on the detection of POI gadget chains, whereas the automatic patch generation remains an under-explored problem. In this work, we empirically study known gadget chains, and discover that adversaries usually construct gadget chains by diverging the execution to paths that developers never considered. The methods that get unexpectedly jump into (i.e., executed) are referred to as possible methods (PM). Based on the observation, we propose PFortifier, a framework for automatic POI patch generation. PFortifier operates in two stages: (i) the gadget chain detection phase, in which PFortifier simulates the execution of PHP applications, and detects gadget chains that pass attacker controlled objects to dangerous sinks, and (ii) the patch generation phase, in which PFortifier automatically generates POI patches by restricting PM jumps detected in the first phase. We evaluate PFortifier on 31 PHP applications and frameworks. The experiment results demonstrate the effectiveness of PFortifier: it generates precise patches for 52.53% of gadget chains, and suggests potential patches for 45.45% chains, resulting in a total chain coverage of 97.98%.
Mingzhe Gao, Ligeng Chen, Mingxue Zhang 0001, Gang Liang
SP5
2025 CATI++: empirical study and evaluation for adjacent instruction enhanced type inference
abstract
Abstract Variable-type information is fundamental, and it greatly helps in understanding the program semantics. Previous work applies rule-based and machine learning-based methods to recover variable types from commercial off-the-shelf binaries, heavily relying on the data flow or control flow. However, according to our study, about half of the variables lacked or even had no data flow; this problem has not received much attention from previous work. We empirically explore the severity of this problem to the type inference task and analyze its root causes. Based on compilation properties, we find that the instructions surrounding the instructions that operate on variables provide good contextual information that can be used for co-encoding to overcome the above problem. In this paper, we present an effective machine learning-based method to infer variable types and overcome the challenge of limited data dependency via adjacent instructions co-encoding. Therefore, we implement a system called CATI++, which locates variables from stripped binaries and infers 19 types of variables. We evaluate CATI++ on different compilation options, all of which outperforms state-of-the-art methods. The ablation experiments verify that our scheme is not sensitive to compilation conditions, while our designed method effectively alleviates the problems caused by missing data dependency.
Ligeng Chen, Zhongling He, Bing Mao 0001
Comput. J.1
2024 MFPNet: A Multi-scale Feature Propagation Network for Lightweight Semantic Segmentation
Guoan Xu, Wenjing Jia, Ligeng Chen, Guangwei Gao
ICANN (3)4
2024 Image recoloring for color vision deficiency compensation using Swin transformer
abstract
Abstract People with color vision deficiency (CVD) have difficulty in distinguishing differences between colors. To compensate for the loss of color contrast experienced by CVD individuals, a lot of image recoloring approaches have been proposed. However, the state-of-the-art methods suffer from the failures of simultaneously enhancing color contrast and preserving naturalness of colors [without reducing the Quality of Vision (QOV)], high computational cost, etc. In this paper, we propose an image recoloring method using deep neural network, whose loss function takes into consideration the naturalness and contrast, and the network is trained in an unsupervised manner. Moreover, Swin transformer layer, which has long-range dependency mechanism, is adopted in the proposed method. At the same time, a dataset, which contains confusing color pairs to CVD individuals, is newly collected in this study. To evaluate the performance of the proposed method, quantitative and subjective experiments have been conducted. The experimental results showed that the proposed method is competitive to the state-of-the-art methods in contrast enhancement and naturalness preservation and has a real-time advantage. The code and model will be made available at https://github.com/Ligeng-c/CVD_swin .
Ligeng Chen, Zhenyang Zhu, Wangkang Huang, Kentaro Go, Xiaoyang Mao
Neural Comput. Appl.1
2024 HAFormer: Unleashing the Power of Hierarchy-Aware Features for Lightweight Semantic Segmentation
abstract
Both Convolutional Neural Networks (CNNs) and Transformers have shown great success in semantic segmentation tasks. Efforts have been made to integrate CNNs with Transformer models to capture both local and global context interactions. However, there is still room for enhancement, particularly when considering constraints on computational resources. In this paper, we introduce HAFormer, a model that combines the hierarchical features extraction ability of CNNs with the global dependency modeling capability of Transformers to tackle lightweight semantic segmentation challenges. Specifically, we design a Hierarchy-Aware Pixel-Excitation (HAPE) module for adaptive multi-scale local feature extraction. During the global perception modeling, we devise an Efficient Transformer (ET) module streamlining the quadratic calculations associated with traditional Transformers. Moreover, a correlation-weighted Fusion (cwF) module selectively merges diverse feature representations, significantly enhancing predictive accuracy. HAFormer achieves high performance with minimal computational overhead and compact model size, achieving 74.2% mIoU on Cityscapes and 71.1% mIoU on CamVid test datasets, with frame rates of 105FPS and 118FPS on a single 2080Ti GPU. The source codes are available at https://github.com/XU-GITHUB-curry/HAFormer.
Guoan Xu, Wenjing Jia, Ligeng Chen, Guangwei Gao
IEEE Trans. Image Process.4
2023 RecMaL: Rectify the malware family label via hybrid analysis
Mingzhe Gao, Ligeng Chen, Zhengxuan Liu, Lingyun Ying
Comput. Secur.3
2023 Nimbus++: Revisiting Efficient Function Signature Recovery with Depth Data Analysis
abstract
Function signature recovery is vital for many binary analysis tasks, led by control-flow integrity enhancement. To minimize human effort, existing works attempt to replace rule-based methods with learning-based methods. These works put a lot of work into improving the system’s performance, but this had the unintended consequence of increasing resource usage. However, recovering the function signature is more about providing information for subsequent tasks, e.g. reverse engineering, so both efficiency and performance are significant. To identify the fundamental factors that increase efficiency, we attempt to optimize data-driven systems throughout their lifecycle from a data perspective. To this end, we perform detailed data analysis on a carefully collected dataset. After analysis and exploration, selective input is adopted and a multi-task learning (MTL) structure is introduced for function feature recovery to make full use of mutual information, and the computing resource overhead is optimized based on the observation of information deviation and sub-task relationship. The resource usage of the entire process is significantly reduced by our suggested solution, named Nimbus++ for efficient function signature recovery, without sacrificing performance. Our test findings demonstrate that we even surpass the state-of-the-art method’s prediction accuracy across all function signature recovery tasks by about 1% with just about 12.5% of the processing time.
Ligeng Chen, Bing Mao 0001
Int. J. Softw. Eng. Knowl. Eng.1
2023 NISe: Non-Invasive Secure Framework for Multi-Access Edge Computing
abstract
To address the emerging security challenges in Multi-Access Edge Computing (MEC), it is imperative that solutions go beyond the current infrastructure-centric measures. These methods, including authentication and access control, are insufficient to combat malware that conceals itself within ME applications. The acknowledged flaws in the ME application layer necessitate an immediate call for creative solutions. In this work, we propose a non-invasive security architecture for MEC, meticulously designed to strike a balance between performance burden and security protection capabilities. The objective of the design contains three major aspects, i.e. user experience, service density and serviceability. We conduct a thorough evaluation that enables us to quantify the significance of high bandwidth, low user experience latency and MEC serviceability. The experimental results and ablation studies indicate that our proposed method effectively balances user experience and security capabilities. This not only provides a practical and cost-effective solution but also establishes a strong precedent for the community to develop a secure MEC with superior performance in real-world production environments.
Xuguo Wang, Ligeng Chen, Yu Liang 0001
Int. J. Softw. Eng. Knowl. Eng.2
2023 EVaDe: Efficient and Lightweight Mirai Variants Detection via Approximate Largest Submatrix Search
abstract
The Mirai botnet, notorious for launching significant Distributed Denial of Service (DDoS) attacks and crippling portions of internet services in late 2016, has emerged as a significant threat. Its threat is magnified by the open-source nature of the original Mirai code, which enables a propagation and evolution rate that surpasses traditional malware and frequently defies common sense. As the primary targets of Mirai attacks, Internet of Things (IoT) devices must promptly adapt to the evolving variations of the Mirai threat scenario. In practice, however, IoT devices are frequently constrained by insufficient security detection resources. Therefore, there is an urgent need for a lightweight framework capable of handling Mirai variants and dynamically updating its rule set in order to effectively counter the threat. In response to these challenges, we present Efficient and lightweight Mirai Variants Detection (EVaDe), a novel, lightweight framework for detecting Mirai. EVaDe unleashes the power of sample function mining to efficiently automate the generation of detection rules, requiring limited hardware resources while maintaining effectiveness against Mirai and its numerous variants. In addition, to improve the efficacy of rule generation, we propose a sophisticated algorithm designed to optimize the maximum submatrix problem, thereby facilitating the efficient and rapid extraction of malicious rules from the sample group. We validated the experiments on actual IoT devices with significantly compressed performance overheads. An average sample detection time of 5 ms to make sure the system can be deployed in real production. According to the result, the approach has an average detection rate of 95% for Mirai and its variants, which beats every other well-known piece of commercial antivirus software on the market by 3% to 56%.
Xuguo Wang, Ligeng Chen, Bing Mao 0001
Int. J. Softw. Eng. Knowl. Eng.2
2022 Nimbus: Toward Speed Up Function Signature Recovery via Input Resizing and Multi-Task Learning
abstract
Function signature recovery is important for many binary analysis tasks such as control-flow integrity enforcement, clone detection, and bug finding. Existing works try to substitute learning-based methods with rule-based methods to reduce human effort.They made considerable efforts to enhance the system’s performance, which also bring the side effect of higher resource consumption. However, recovering the function signature is more about providing information for subsequent tasks, and both efficiency and performance are significant.In this paper, we first propose a method called Nimbus for efficient function signature recovery that furthest reduces the whole-process resource consumption without performance loss. Thanks to information bias and task relation (i.e., the relation between parameter count and parameter type recovery), we utilize selective inputs and introduce multi-task learning (MTL) structure for function signature recovery to reduce computational resource consumption, and fully leverage mutual information. Our experimental results show that, with only about the one-eighth processing time of the state-of-the-art method, we even achieve about 1% more prediction accuracy over all function signature recovery tasks.
Ligeng Chen, Bing Mao 0001
QRS2
2022 AVMiner: Expansible and Semantic-Preserving Anti-Virus Labels Mining Method
abstract
With the increase in the variety and quantity of malware, there is an urgent need to speed up the diagnosis and analysis of malware. Extracting the malware family-related tokens from AV (Anti-Virus) labels, provided by online antivirus engines, paves the way for pre-diagnosing the malware. Automatically extracting vital information from AV labels will greatly enhance the detection ability of security enterprises and equip the research ability of security analysts. Recent works like AVCLASS and AVCLASS2 try to extract the attributes of malware from AV labels and establish the taxonomy based on expert knowledge. However, due to the uncertain trend of complicated malicious behaviors, the system needs the following abilities to face the challenge: preserving vital semantics, being expansible, and being free from expert knowledge. In this work, we present AVMiner, an expansible malware tagging system that can mine the most vital tokens from AV labels. AVMiner adopts natural language processing techniques and clustering methods to generate a sequence of tokens without expert knowledge ranked by importance. AVMiner can self-update when new samples come. Finally, we evaluate AVMiner on over 8,000 samples from well-known datasets with manually labeled ground truth, which outperforms previous works.
Ligeng Chen, Zhongling He, Hao Wu 0067, Yuhang Gong, Bing Mao 0001
TrustCom1
2022 DIComP: Lightweight Data-Driven Inference of Binary Compiler Provenance with High Accuracy
abstract
Binary analysis is pervasively utilized to assess software security and test vulnerabilities without accessing source codes. The analysis validity is heavily influenced by the inferring ability of information related to the code compilation. Among the compilation information, compiler type and optimization level, as the key factors determining how binaries look like, are still difficult to be inferred efficiently with existing tools. In this paper, we conduct a thorough empirical study on the binary's appearance under various compilation settings and propose a lightweight binary analysis tool based on the simplest machine learning method, called DIComP to infer the compiler and optimization level via most relevant features according to the observation. Our comprehensive evaluations demonstrate that DIComP can fully recognize the compiler provenance, and it is effective in inferring the optimization levels with up to 90% accuracy. Also, it is efficient to infer thousands of binaries at a millisecond level with our lightweight machine learning model (1MB).
Ligeng Chen, Zhongling He, Hao Wu 0067, Fengyuan Xu, Bing Mao 0001
SANER1
2022 Image recoloring for Red-Green dichromats with compensation range-based naturalness preservation and refined dichromacy gamut
Wangkang Huang, Zhenyang Zhu, Ligeng Chen, Kentaro Go, Xiaoyang Mao
Vis. Comput.3
2021 RoBin: Facilitating the Reproduction of Configuration-Related Vulnerability
abstract
Vulnerability reproduction paves a way in debugging software failures, which need intensive manual efforts. However, some key factors (e.g., software configuration, trigger method) are often missing, so we can not directly reproduce the failure without extra attempts. Even worse, highly customized configuration options of programs create a barrier for reproducing the vulnerabilities that only appear under some specific combinations of configurations. In this paper, we address the problem mentioned above - reproducing the configuration-related vulnerability. We try to solve it by proposing a binary similarity-based method to infer the specific building configurations via the binary from crash report. The main challenges are as follows: precise compilation option inference, program configuration inference, and source-code-to-binary matching. To achieve the goal, we implement RoBin, a binary similarity-based building configuration inference tool. To demonstrate the effectiveness, we test RoBin on 21 vulnerable cases upon 4 well-known open-source programs. It shows a strong ability in pinpointing the building configurations causing the vulnerability. The result can help developers reproduce and diagnose the vulnerability, and finally, patch the programs.
Ligeng Chen, Zhongling He, Dongliang Mu, Bing Mao 0001
TrustCom1
2020 CATI: Context-Assisted Type Inference from Stripped Binaries
abstract
Code analysis is a powerful way to eliminate vulnerabilities. Closed-source programs lack crucial information vital for code analysis because that information is stripped on compilation to achieve smaller executable size. Restoration has always been a challenge for experts. Variable type information is fundamental in this process because it helps to provide a perspective on program semantic. In this paper, we present an efficient approach for inferring types, and we overcome the challenge of scattered information provided by static analysis on stripped binaries. We discover that neighboring instructions are likely to operate the same type of variables, which are leveraged to enrich the features that we rely on. Therefore, we implement a system called CATI, which locates variables from stripped binaries and infers 19 types from variables. Experiments show that it infers variable type with 71.2% accuracy on unseen binaries. Meanwhile, it takes approximately 6 seconds to process a typical binary.
Ligeng Chen, Zhongling He, Bing Mao 0001
DSN1
2019 Building Adversarial Defense with Non-invertible Data Transformations
Wenbo Guo 0002, Dongliang Mu, Ligeng Chen, Jinxuan Gai
PRICAI (3)3