Hongliang Liang

dblp:92/283 · DBLP profile ↗
← Back
48ranked-venue papers
28as first author
23since 2021 · last 2026
0000-0001-6877-780XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 21 · 10 first-author · 12 since 2021Security and privacy · 18 · 12 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 BiMarker: Enhancing text watermark detection for large language models with bipolar watermarks
Qiuping Yi, Zongcheng Ji, Yijian Lu, Shun Zou, Yanqi Li, Keyang Xiao, Hongliang Liang
Neurocomputing8
2026 An end-to-end approach for fixing concurrency bugs via SHB-based context extractor
Qiuping Yi, Keyang Xiao, Zongcheng Ji, Hongliang Liang
J. Syst. Softw.5
2026 BinSleuth: Scalable vulnerability discovery in binaries via automated source identification and path optimization
Shun Zou, Hongliang Liang
J. Syst. Softw.3
2026 One perturbation fools all: An adversarial perturbation can attack different vision models
Jinyan Cai, Tianhao Yu, Hongliang Liang, Qiuping Yi
Knowl. Based Syst.3
2026 Efficient Directed Hybrid Fuzzing via Target-Centric Seed Selection and Generation
abstract
Software vulnerabilities pose severe security threats, highlighting the need for effective automated detection. Directed hybrid fuzzing, which combines the rapid exploration of fuzz testing with the precise constraint solving of symbolic execution, has made notable advancements in vulnerability discovery. However, existing directed hybrid fuzzing approaches still face two key challenges: (1) inefficient seed selection, leading to inadequate prioritization of optimal inputs for symbolic execution, and (2) inefficient seed generation, resulting in suboptimal seed generation. To address these issues, we propose TACO-Fuzz, TArget-Centric cOncolic Fuzzing, which introduces a two-phase target-centric seed selection strategy to prioritize under-explored paths and a target-centric seed generation approach based on constructing extended path conditions, thereby improving seed quality. Our evaluation on a selected set of public benchmarks shows that TACO-Fuzz can outperform several representative state-of-the-art directed fuzzing tools, achieving up to an average speedup of nearly 10x in reaching target locations, along with comparable improvements in reproducing real-world vulnerabilities. Moreover, TACO-Fuzz contributed to the discovery of 17 previously unknown vulnerabilities, each assigned a CVE, and demonstrated faster vulnerability discovery and reproduction in most cases.
Shenghan Liu, Qiuping Yi, Pengbo Du, Hongliang Liang
Proc. ACM Program. Lang.5
2026 LARTS: Language Abstractions for Real-Time and Secure Systems
abstract
Real-time systems must simultaneously deliver predictable timing, fault isolation, and memory safety, yet current operating systems expose only low-level primitives that force developers to manually balance concurrency, isolation, and performance. This paper presents LARTS, a language-aided runtime system that elevates these requirements into language abstractions with enforceable semantics. LARTS introduces execution domain, a unified process–thread abstraction that combines thread-level responsiveness with process-level isolation. Memory is managed through deterministic memory contracts, which bind allocation at load time to eliminate runtime failures and unpredictable latencies. Domain interactions are expressed via deterministic communication channels that integrate efficient transfer, type safety, and priority inheritance, ensuring analyzable end-to-end bounds. Moreover, LARTS enforces secure-by-construction semantics, making classes of bugs such as double fetch and use-after-free unrepresentable in the programming model. We formalize the core semantics of LARTS and show how they guarantee determinism and safety by design. A prototype built on RTEMS demonstrates that LARTS preserves competitive real-time performance while substantially reducing programming complexity and eliminating vulnerabilities in realistic case studies. Our results suggest that high-assurance real-time programming can be treated not as an ad-hoc engineering problem, but as a first-class abstraction with verifiable semantics.
Yanqi Li, Hongliang Liang, Qiuping Yi
Proc. ACM Program. Lang.2
2026 Dynamic Binary Translation of VLIW Code With Software Pipeline
abstract
Dynamic binary translation serves as a pivotal technique for instruction set simulation, yet encounters critical challenges when handling explicit instruction-level parallelism and operational latency inherent in VLIW architectures. Existing approaches demonstrate limited capability in handling non-arithmetic operations, particularly branch and memory access instructions. The complexity intensifies when translating software-pipelined loops featuring architecture-specific instructions due to two inherent characteristics: 1) instruction reordering, overlapping, and masking within loop bodies, and 2) absence of explicit conditional branches and loop counter manipulation instructions. This work presents three original strategies for VLIW code translation. First, a strategy constrains translation block length to resolve branch instruction challenges. Second, an approach manages store-load dependencies through strategic postponement of store instruction processing. Third, a novel software-pipelined loop translation methodology ensures correct execution semantics by serializing parallel iterations, generating state-specific translation blocks, and synchronizing inner and outer loops translations. We implemented these techniques in VEMU and evaluated it against dsplib and Polybench benchmarks. VEMU successfully translates all benchmark programs. Our first strategy does not degrade performances and the second strategy brings 0.2% time overhead. Comparative analysis reveals a speedup ratio between 3.25× and 7× over the Texas Instruments simulator for the same benchmarks, validating the efficacy of the proposed techniques.
Hongliang Liang, Kailai Liao, Guohao Wu, Qiuping Yi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 FENSE: Feedback-Driven Incremental Symbolic Execution for Redundant Path Elimination
abstract
Incremental symbolic execution aims to address the scalability challenges of traditional symbolic execution by concentrating on behavioural differences between program versions introduced during program evolution. Despite progress in the field, existing techniques often struggle to explore these behaviors both efficiently and accurately. In this paper, we introduceFENSE, a novel approach for incremental symbolic execution that improves efficiency by identifying and eliminating redundant paths.FENSEachieves this by summarizing previously explored paths and monitoring variables that may induce divergent incremental behaviors at each branching point. This summarization process enablesFENSEto detect whether a newly explored path subsumes distinct incremental behaviors compared to prior explorations. By effectively pruning redundant paths that exhibit identical incremental behaviors as those previously explored,FENSEachieves a potentially exponential reduction in the number of explored paths. We implemented a prototype ofFENSEand evaluated it on a diverse set of real-world applications. Experimental results demonstrate thatFENSEoutperforms state-of-the-art techniques by significantly reducing both path exploration and execution time. When applied to real-world commits from the GNU Coreutils project,FENSEachieved an average of 76% reduction in explored paths and a 139× speedup overKLEE.
Pengbo Du, Qiuping Yi, Hongliang Liang, Guowei Yang 0001
IEEE Trans. Software Eng.3
2026 CosFormer: A Code Semantic-Aware Transformer for Vulnerability Detection
abstract
Deep learning-based vulnerability detection has made significant strides, surpassing traditional static and dynamic analysis methods. However, existing approaches, including Graph Neural Networks (GNNs) and Transformer-based models, still struggle to fully capture complex code semantics. In this paper, we proposeCosFormer, a novel Code Semantic-aware Transformer tailored for vulnerability detection.CosFormerintroduces two key components: Code Semantic-aware Embedding, which enhances semantic representation at both the token and line levels, and Spatial Dependency-aware Encoding, which integrates structural dependencies from Control Flow Graphs (CFGs) and Program Dependency Graphs (PDGs) to guide attention toward vulnerability-relevant code. We evaluateCosFormeron four benchmark datasets, including a real-world dataset, and demonstrate its superior performance.CosFormerachieves the highest F1 scores across all tasks, outperforming state-of-the-art GNN-based and Transformer-based models, as well as large language models (LLMs). Notably,CosFormerachieves a Cross F1 of 69.37 and a Mixed F1 of 58.42 in generalization evaluation, surpassing all baselines. These results highlightCosFormer’s effectiveness and robustness in detecting vulnerabilities across diverse and previously unseen codebases.
Qianyue Wei, Qiuping Yi, Zongcheng Ji, Hongliang Liang
IEEE Trans. Software Eng.5
2025 Benchmarking and Understanding Compositional Relational Reasoning of LLMs
abstract
Compositional relational reasoning (CRR) is a hallmark of human intelligence, but we lack a clear understanding of whether and how existing transformer large language models (LLMs) can solve CRR tasks. To enable systematic exploration of the CRR capability of LLMs, we first propose a new synthetic benchmark called Generalized Associative Recall (GAR) by integrating and generalizing the essence of several tasks in mechanistic interpretability (MI) study in a unified framework. Evaluation shows that GAR is challenging enough for existing LLMs, revealing their fundamental deficiency in CRR. Meanwhile, it is easy enough for systematic MI study. Then, to understand how LLMs solve GAR tasks, we use attribution patching to discover the core circuits reused by Vicuna-33B across different tasks, and a set of vital attention heads. Intervention experiments show that the correct functioning of these heads significantly impacts task performance. Especially, we identify two classes of heads whose activations represent the abstract notion of true and false in GAR tasks respectively. They play fundamental roles in CRR across various models and tasks.
Ruikang Ni, Da Xiao 0001, Qingye Meng, Shihui Zheng, Hongliang Liang
AAAI6
2025 Symbolic testing of floating-point bugs and exceptions
Dongyu Ma, Luming Yin, Hongliang Liang
J. Syst. Softw.4
2025 Data-independent generation of targeted universal adversarial perturbations
Jinyan Cai, Hongliang Liang, Qiuping Yi
Knowl. Based Syst.3
2025 Toward Automatic Heap Exploit Generation by Using Heap Layout Constraints on Binary Programs
Lianda Yao, Yanqi Li, Hongliang Liang
IEEE Trans. Inf. Forensics Secur.4
2025 Unified and Split Symbolic Execution for Exposing Semantic Differences
abstract
Software evolution is an important activity during a software development lifecycle. Understanding semantic differences between two versions of a software system is a crucial yet challenging task, especially in many safety critical sectors. Consequently, various techniques have been proposed to check semantic differences between a program and its evolution. But, many current techniques are still far from being satisfactory in terms of the accuracy and efficiency. In this article, we propose a novel framework, called US 2 E , which can efficiently and effectively generate the minimal number of test cases that reveal as many semantic differences across two versions as possible. Specifically, given a unified control flow graph that denotes two versions of a program, US 2 E executes as many common nodes as possible and leaves execution of non-common nodes separately in a single concolic execution instance. We evaluate US 2 E on 86 pairs of C programs from 4 benchmarks, and experimental results show that US 2 E can efficiently and effectively generate test cases demonstrating the semantic differences across 2 versions, with better performance than 6 baseline tools.
Hongliang Liang, Luming Yin, Wenying Hu, Wuwei Shen
ACM Trans. Softw. Eng. Methodol.1
2024 Multiple Targets Directed Greybox Fuzzing: From Reachable to Exploited
abstract
Directed Greybox Fuzzing (DGF) is a technique used to efficiently cover specific program locations, making it suitable for vulnerability discovery and regression testing. However, previous DGF tools often suffer from the issue of getting trapped in local optima, where some targets are adequately covered while other targets remain unexploited. We propose the concept of Target-Directed Basic Blocks to efficiently guide the fuzzer, a multi-dimensional energy scheduling algorithm to guide the generation of seeds, and an adaptive exploration-exploitation strategy to prevent falling into local optima. Moreover, we optimize the mutation algorithm by employing a branch-sensitive strategy and a reinforcement learning algorithm based on the multi-armed bandit model. We implemented these algorithms in a tool called SAFuzz and conducted extensive experiments and evaluations in real-world scenarios. The results demonstrate that SAFuzz is effective in bug reproduction, guiding different targets, and validating static analysis results compared to state-of-the-art benchmark tools, i.e., AFLGo, Windranger, AFL++, and Lolly. Moreover, SA Fuzz has detected seven new vulnerabilities in real-world programs, and one of them could not be found by other baseline fuzzers within the time budget.
Xinglin Yu, Hongliang Liang
SANER2
2024 Generative Pre-Trained Transformer-Based Reinforcement Learning for Testing Web Application Firewalls
abstract
Web Application Firewalls (WAFs) are widely deployed to protect key web applications against multiple security threats, so it is important to test WAFs regularly to prevent attackers from bypassing them easily. Machine-learning-based black-box WAF testing is gaining more attention, though existing learning-based approaches have strict requirements on the source and scale of payload data and suffer from the local optimum problem, limiting their effectiveness and practical application. We propose GPTFuzzer, apracticalandeffectivegeneration-based approach to test WAFs by generating attack payloads token-by-token. Specifically, we fine-tune a Generative Pre-trained Transformer language model with reinforcement learning to make GPTFuzzer have the least restrictions on payload data and thus more applicable in practice, and we use reward modeling and KL-divergence penalty to improve the effectiveness of our approach and mitigate the local optimum issue. We implement GPTFuzzer and evaluate it on two well-known open-source WAFs against three kinds of common attacks. Experimental results show that GPTFuzzer significantly outperforms state-of-the-art approaches,i.e.ML-Driven and RAT, finding up to 7.8× (3.2× on average) more bypassing payloads within 1,250,000 requests, or finding out all bypassing payloads using up to 8.1× (3.3× on average) fewer requests.
Hongliang Liang, Da Xiao 0001, Yanjie Zhou, Aibo Wang
IEEE Trans. Dependable Secur. Comput.1
2024 Multiple Targets Directed Greybox Fuzzing
abstract
Directed greybox fuzzing (DGF) can quickly discover or reproduce bugs in programs by seeking to reach a program location or explore some locations in order. However, due to their static stage division and coarse-grained energy scheduling, prior DGF tools perform poorly when facing multiple target locations (targets for short). In this paper, we present multiple targets directed greybox fuzzing which aims to reach multiple programs locations in a fuzzing campaign. Specifically, we propose a novel strategy to adaptively coordinate exploration and exploitation stages, and a novel energy scheduling strategy by considering more relations between seeds and target locations. We implement our approaches in a tool called LeoFuzz and evaluate it on crash reproduction, true positives verification, and vulnerability exposure in real-world programs. Experimental results show that LeoFuzz outperforms six state-of-the-art fuzzers, i.e., QYSM, AFLGo, Lolly, Berry, Beacon and WindRanger in terms of effectiveness and efficiency. Moreover, LeoFuzz has detected 23 new vulnerabilities in real-world programs, and 12 of them have been assigned CVE IDs.
Hongliang Liang, Xinglin Yu, Xianglin Cheng
IEEE Trans. Dependable Secur. Comput.1
2023 Value Peripheral Register Values for Fuzzing MCU Firmware
abstract
Analyzing the security of MCU firmware is important. Fuzzing with peripheral model emulation is proven successful due to its independence of hardware and practicality. However, prior efforts such as DICE and P2IM expose two issues: insufficient exploration of peripheral states and inadequate handling of abnormal behaviors. Our key insight is that peripheral register values are vital for fuzzing MCU firmware. We present a novel approach called VeRa which introduces several techniques: peripheral state triage, rollback explorative execution, low-frequency-first schedule for control-status register values, and random selection for status register values. We integrated VeRa to a state-of-the-art firmware analyzer DICE and evaluated on 122 firmware covering 11 MCU platforms. Evaluation results show that 1) on modeling peripherals, VeRa passed more 45 out of 99 sample firmware than DICE; 2) VeRa outperforms DICE on fuzzing 23 real-world firmware, with 1.4× basic block coverage and 3.1× path coverage on average (up to 13.2× and 33.7×, respectively); 3) The overhead of VeRa is fairly low, adding 5.6% and 1.2% on average to modeling time and fuzzing time respectively; 4) VeRa discovered 3 unique new bugs that DICE cannot find.
Hongliang Liang
ISSRE2
2022 Detection and Privacy Leakage Analysis of Third-Party Libraries in Android Apps
Xiantong Hao, Dandan Ma, Hongliang Liang
SecureComm3
2022 Modeling function-level interactions for file-level bug localization
Hongliang Liang, Dengji Hang
Empir. Softw. Eng.1
2022 Path context augmented statement and network for learning programs
Da Xiao 0001, Dengji Hang, Lu Ai, Shengping Li, Hongliang Liang
Empir. Softw. Eng.5
2021 AST-path Based Compare-Aggregate Network for Code Clone Detection
abstract
Code clone detection remains one of the main challenges in maintaining software projects. Recently, state-of-the-art researches have shown that neural models based on abstract syntax trees (ASTs) can better represent code fragment. However, existing tree-based models are prone to gradient vanishing problems due to the large size of ASTs. In this paper, we represent a code fragment as the set of compositional paths in its abstract syntax tree (AST) and use this code representation to train a classifier to detect clone pairs. Unlike the siamese based model that obtains the embeddings of code fragments separately and then computes the similarity in vector space, our compare-aggregate based network takes two code fragments as a whole to obtain the vectors for classification. To validate our model's ability to detect code clones, we evaluated it on the publicly available dataset BigCloneBench, and the experimental results show our model outperforms the state-of-the-art model ASTNN.
Hongliang Liang, Lu Ai
IJCNN1
2021 Architectural Protection of Trusted System Services for SGX Enclaves in Cloud Computing
abstract
Data security and privacy are of great concern for users of cloud computing. In order to provide such guarantees in public clouds, hardware manufacturers have designed trusted execution environments such as Intel’s Software Guard eXtensions (SGX). Intel SGX supports privacy-preserving, tamper-proof containments called enclaves. Regrettably, an SGX enclave has to rely on the untrusted operating system or hypervisor for underlying services, which contradicts the threat model of Intel SGX. Whereas much of the previous work concentrates on protecting trusted applications by means of modifying a hypervisor, we tackle the problem by reusing existing drivers and leveraging processor-enforced protection. We propose a novel approach, named SMK, to provide trusted system services for SGX enclaves. SMK leverages existing Intel architecture features, i.e., System Management Mode (SMM) and Uniform Extensible Firmware Interface (UEFI). Specifically, we retrofit UEFI firmware and design an isolated micro-kernel inside SMM to securely provision critical system services for enclaves. To highlight the effectiveness and extensibility of SMK, we implement two system services: trusted clock and trusted network. Furthermore, we harden two real-world security-sensitive applications, OpenSSL and OpenVPN, with SMK’s system services. Our evaluation indicates that SMK can supply trusted system services for enclaves with modest runtime overheads.
Hongliang Liang, Yixiu Chen, Tianqi Yang 0003, Zhuosi Xie, Lin Jiang 0002
IEEE Trans. Cloud Comput.1
2020 A Practical Concolic Execution Technique for Large Scale Software Systems
abstract
The state explosion problem faced by concolic execution is significantly serious when detecting program bugs especially in large scale software systems. To mitigate the issue, we propose a practical concolic execution approach to detect vulnerabilities in this paper. First, we identify the correlation between symbolic memory and control flow by statically analyzing the software under test, and distinguish the critical symbolic memory and ordinary symbolic memory according to the above correlation. Then, we design different strategies for two kinds of symbolic memory in order to generate state when performing concolic execution of the target software. We present these ideas in a prototype system, Pracolic, and evaluate it with four file systems in Linux, i.e., ext4, XFS, Btrfs, and ReiserFS. Experimental results show that Pracolic can effectively mitigate the problem of state explosion in concolic execution, and outperform S2E, a state-of-the-art analysis system, on state reduction and code coverage for large scale software systems.
Hongliang Liang, Wenqing Yu, Lu Ai, Lin Jiang 0002
EASE1
2020 GTFuzz: Guard Token Directed Grey-Box Fuzzing
abstract
Directed grey-box fuzzing is an effective technique to find bugs in programs with the guidance of user-specified target locations. However, it can hardly reach a target location guarded by certain syntax tokens (Guard Tokens for short), which is often seen in programs with string operations or grammar/lexical parsing. Only the test inputs containing Guard Tokens are likely to reach the target locations, which challenges the effectiveness of mutation-based fuzzers. In this paper, a Guard Token directed grey-box fuzzer called GTFuzz is presented, which extracts Guard Tokens according to the target locations first and then exploits them to direct the fuzzing. Specifically, to ensure the new test cases generated from mutations contain Guard Tokens, new strategies of seed prioritization, dictionary generation, and seed mutation are also proposed, so as to make them likely to reach the target locations. Experiments on real-world software show that GTFuzz can reach the target locations, reproduce crashes, and expose bugs more efficiently than the state-of-the-art grey-box fuzzers (i.e., AFL, AFLGO and FairFuzz). Moreover, GTFuzz identified 23 previously undiscovered bugs in LibXML2 and MJS.
Hongliang Liang, Xutong Ma, Rong Qu, Jun Yan 0009, Jian Zhang 0001
PRDC2
2020 Sequence Directed Hybrid Fuzzing
abstract
Existing directed grey-box fuzzers are effective compared with coverage-based fuzzers. However, they fail to achieve a balance between effectiveness and efficiency, and it is difficult to cover complex paths due to random mutation. To mitigate the issue, we propose a novel approach, sequence directed hybrid fuzzing (SDHF), which leverages a sequence-directed strategy and concolic execution technique to enhance the effectiveness of fuzzing. Given a set of target statement sequences of a program, SDHF aims to generate inputs that can reach the statements in each sequence in order and trigger potential bugs in the program. We implement the proposed approach in a tool called Berry and evaluate its capability on crash reproduction, true positive verification, and vulnerability detection. Experimental results demonstrate that Berry outperforms four state-of-the-art fuzzers, including directed fuzzers BugRedux, AFLGo and Lolly, and undirected hybrid fuzzer QSYM. Moreover, Berry found 7 new vulnerabilities in real-world programs such as UPX and GNU Libextractor, and 3 new CVEs were assigned.
Hongliang Liang, Lin Jiang 0002, Lu Ai, Jinyi Wei
SANER1
2020 FIT: Inspect vulnerabilities in cross-architecture firmware by deep learning and bipartite matching
Hongliang Liang, Zhuosi Xie, Yixiu Chen, Hua Ning, Jianli Wang
Comput. Secur.1
2020 Establishing Trusted I/O Paths for SGX Client Systems With Aurora
abstract
Today users' private data in edge computing devices (desktops, laptops, and tablets, etc.) is at high risk because they run applications on potentially compromised or malicious systems. To address this problem, hardware vendors propose Trusted Execution Environment (TEE). Particularly, Intel has released a new processor feature called Software Guard eXtension (SGX), and provisions shielded executions (i.e., enclaves) for security-sensitive computations. Regrettably, Intel SGX's design objectives omit trusted I/O paths. Without such guarantees, it is unlikely for an enclave to fulfill its security and privacy purposes because the source or sink of data may have been corrupted. To this end, we propose a novel architecture called Aurora to provide trusted I/O paths for enclave programs even in the presence of untrusted system software. Specifically, Aurora exploits two commercial-off-the-shelf features (System Management Mode, SMM and SGX) and establishes a secure channel between an enclave program and target device. Furthermore, we design and implement trusted paths for HID keyboard, serial port printer, hardware clocks, and USB mass storage, respectively. Leveraging these trusted paths, we protect real-world applications including OpenSSH client, OpenSSL server/client and SQLite database. Security and performance evaluations show that Aurora mitigates several kinds of I/O related attacks and introduces acceptable overheads. Our framework has been open-sourced and is available to the security community.
Hongliang Liang, Yixiu Chen, Lin Jiang 0002, Zhuosi Xie, Tianqi Yang 0003
IEEE Trans. Inf. Forensics Secur.1
2019 Witness: Detecting Vulnerabilities in Android Apps Extensively and Verifiably
abstract
Existing studies on detecting vulnerabilities in apps have two main disadvantages: one is that some studies are limited to detecting a certain vulnerability and lack comprehensive analysis; the other is the lack of valid evidence for vulnerability verification, which leads to high false alarms rate and requires massive manual efforts. We propose the concept of vulnerability pattern to abstract the characteristics of different attacks, e.g., their prerequisites and attack paths, so as to support detecting multiple kinds of vulnerabilities. Also, we present a zero false alarms framework which can find vulnerability instances precisely and generate test cases and triggers to validate the findings, by combing static analysis and dynamic binary instrumentation techniques. We implement our method in a tool named Witness, which currently can detect 8 different types of vulnerabilities and is extensible to support more. Evaluated on 3211 popular apps, Witness successfully detected 243 vulnerability instances, with better precision and more proofs than four existing tools.
Hongliang Liang, Tianqi Yang 0003, Lin Jiang 0002, Yixiu Chen, Zhuosi Xie
APSEC1
2019 JSAC: A Novel Framework to Detect Malicious JavaScript via CNNs over AST and CFG
abstract
JavaScript (JS) is a dominant programming language in web/mobile development, while it is also notoriously abused by attackers due to its powerful characteristics, e.g., dynamic, prototype-based and multi-paradigm, which foil most static and dynamic analysis approaches. To detect malicious JS instances, several machine learning-based methods have been developed recently. However, these methods took JS as a natural language instead of a programming one, which can not capture its syntactic and semantic features. In this paper, we present JSAC, a novel framework to detect JS malware. It combines deep learning and program analysis techniques to capture the syntactic and semantic features of JS programs. Specifically, to get a JS program's syntactic information, we build its abstract syntax tree and employ a tree-based convolutional neural network (CNN) to extract features from it. To get its semantic information, we construct its control flow graph and feed it to another graph-based CNN. Last, the features extracted from two CNNs are fused for final detection. Evaluation on a corpus of 69,523 JS files indicates that JSAC outperforms 4 other models with 98.73% F1-score in detecting JS malware.
Hongliang Liang, Yuxing Yang, Lin Jiang 0002
IJCNN1
2019 Toward Migration of SGX-Enabled Containers
abstract
Containers are becoming the de facto platform for cloud computing. While cloud security has been a major concern, Intel SGX provisions powerful protection guarantees that can be used for containers. However, this technology does not come for free. For example, limited Enclave Page Cache (EPC) challenges the migration design of SGX-enabled containers.We note that previous security protocols are problematic concerning migration of SGX-enabled containers, which will lead to the failure of measures to prevent fork/fallback attacks. In this paper, we propose the migration of SGX-enabled containers and explore the challenges of deploying and migrating SGX-enabled containers considering both EPC resources and persistent storage. To our best knowledge, we are the first to design and implement such a framework for the SGX-enabled container migration that is easy, flexible and lightweight to deploy. We evaluate the proposed framework by migrating SGX-enabled Sqlite3 container and the experimental result shows that the proposed framework has about 15% overhead, which is acceptable due to its security advantage.
Hongliang Liang, Jianqiang Li 0001
ISCC1
2019 Sequence coverage directed greybox fuzzing
abstract
The following topics are dealt with: software maintenance; source code (software); program diagnostics; data mining; Java; software quality; program debugging; learning (artificial intelligence); software metrics; public domain software.
Hongliang Liang, Yini Zhang, Zhuosi Xie, Lin Jiang 0002
ICPC1
2018 Fuzzing: State of the Art
abstract
As one of the most popular software testing techniques, fuzzing can find a variety of weaknesses in a program, such as software bugs and vulnerabilities, by generating numerous test inputs. Due to its effectiveness, fuzzing is regarded as a valuable bug hunting method. In this paper, we present an overview of fuzzing that concentrates on its general process, as well as classifications, followed by detailed discussion of the key obstacles and some state-of-the-art technologies which aim to overcome or mitigate these obstacles. We further investigate and classify several widely used fuzzing tools. Our primary goal is to equip the stakeholder with a better understanding of fuzzing and the potential solutions for improving fuzzing methods in the spectrum of software testing and security. To inspire future research, we also predict some future directions with regard to fuzzing.
Hongliang Liang, Xiaoxiao Pei, Wuwei Shen, Jian Zhang 0001
IEEE Trans. Reliab.1
2017 A Novel Method Makes Concolic System More Effective
abstract
Fuzzing is attractive for finding vulnerabilities in binary programs. However, when the application's input space is huge, fuzzing cannot deal with it well. For discovering vulnerabilities more effective, researchers came up concolic testing, and there are much researches on it recently. A common limitation of concolic systems designed to create inputs is that they often concentrate on path-coverage and struggle to exercise deeper paths in the executable under test, but ignore to find those test cases which can trigger the vulnerabilities. In this paper, we present TSM, a novel method for finding potential vulnerabilities in concolic systems, which can help concolic systems more effective for hunting vulnerabilities. We implemented TSM method on a wide-used concolic testing tool-Fuzzgrind, and the evaluation experiments show that TSM can make Fuzzgrind hunt bugs quickly in real-world software, which are hardly found ever before.
Hongliang Liang, Minhuan Huang, Xiaoxiao Pei
CSCloud1
2017 Fuzzing the Font Parser of Compound Documents
abstract
Currently, complex software (e.g. PDF readers) usually takes various inputs embedded with multiple objects (e.g. fonts, pictures), which may result in bugs. It is a challenge to generate suitable test cases to support fine-grained test to the PDF readers. Compared with the traditional blind fuzzing which does not utilize the information of input grammars, fuzzing with the model of the file format is an effective technique. In this paper, we leverage the structure information of the font files to select seed files among the heterogeneous fonts. A general construction method for generating suitable test cases is proposed. By this means, we can obtain test cases with low overhead. Moreover, to improve the expression ability of the font template in fuzzing PDF readers, we combine file reconstruction and template description. Our methods are evaluated on five common-used PDF readers, and proved effective in triggering crashes.
Hongliang Liang, Huayang Cao
CSCloud1
2017 An end-to-end model for Android malware detection
abstract
Malware detection has been a difficult problem for a very long time. Since the wide use of smart devices in recent years, the number of malwares is increasing rapidly. Most existing methods for malware detection rely too much on manual interventions (e.g. pre-defined features and patterns), which can be easily deceived. In this paper, we propose a novel end-to-end deep learning model to detect Android malwares. Our model takes the raw system call sequence, which is generated during the application's runtime, as input and decides whether the sequence is malicious without any manual intervention. We evaluate the model on 14231 Android applications and obtain a detection accuracy of 93.16%, which is 2.81% higher than the contrast experiment in which we implement the method proposed by other researchers.
Hongliang Liang, Da Xiao 0001
ISI1
2017 Improving the precision of static analysis: Symbolic execution based on GCC abstract syntax tree
abstract
To improve the accuracy of analysis results is one of the hard challenges for static analysis. Especially, static analyzers generally analyze all paths of a program, including infeasible paths, which undoubtedly decreases the analysis accuracy. To mitigate the issue, we design and implement a static analyzer, called ABAZER-SE, which is based on the meta-compilation and the GCC abstract syntax tree. ABAZER-SE combines symbolic execution and static analysis techniques to detect bugs in the source code. In addition, it allows users to write a custom checker for a specific bug.
Hongliang Liang, Shirun Liu, Yini Zhang, Meilin Wang
SNPD1
2017 vmOS: A virtualization-based, secure desktop system
Hongliang Liang, Wenying Hu, Xiaoxiao Pei
Comput. Secur.1
2016 A Correctness Verification Method for C Programs Based on VCC
abstract
The correctness of implementation codes is important especially for safety-critical software usually written in C programming language. We present a correctness verification method (CVM for short) for C codes based on an automatic theorem proving tool-VCC, and propose a specification simplification method to im-prove the correctness and readability of verification specification codes. Using CVM method, the scheduling module of a real-time operating system FreeRTOS6.1.1 is verified, which shows the feasibility and effectiveness when CVM method is applied to the real production software. Experiments show that the CVM method is feasible and effective in verifying the correctness the C codes, and the specification simplification method is also effective.
Hongliang Liang, Daijie Zhang, Xiaoxiao Pei, Jiuyun Xu
CSCloud1
2016 MLSA: A static bugs analysis tool based on LLVM IR
abstract
Program bugs may result in unexpected software error, crash or serious security attack. Static program analysis is one of the most common methods to find program bugs. In this paper we present MLSA - a static analysis tool based on LLVM Intermediate Representation (IR), which can analyze programs written in multiple programming languages. MLSA combines symbolic execution with Z3 SMT solver to find bugs. At present, MLSA can detect some kinds of bugs, such as divide zero error, pointer overflow and dead code. Moreover, as a framework, MLSA follows the scalability and extensibility principles, which can help detect other types of bugs. Experiments show that MLSA is effective in finding bugs in real world software.
Hongliang Liang, Dongyang Wu, Jiuyun Xu
SNPD1
2016 Verifying RTuinOS using VCC: From approach to practice
abstract
The correctness of implementation codes is important especially for safety-critical software which is usually written in C programming language. In this paper, we present a C code correctness verification framework (CVF for short) which is based on an automatic theorem proving tool-VCC, and pro-pose a specification checking method to improve the correctness of verification specification codes. We use a real-time OS, RTuinOS1.0.2, to evaluate our CVF framework. And the result shows it is feasible and effective to apply CVF to a real system software. Besides, the experiments also show that the specification checking method is effective.
Hongliang Liang, Daijie Zhang, Xiaoxiao Pei
SNPD1
2016 Understanding and detecting performance and security bugs in IOT OSes
abstract
Operating system (OS) plays an important role in the efficiency and security of service in Internet of Things (IOT). Considering the limited storage resources and power utilization mechanism in IOT OS, we explore the performance bugs and security bugs of three open source IOT OS (Contiki, TinyOS and RIOT OS), including the features of bugs cause, bugs mitigation, bugs detection rules and bugs fixing in them. We present Rulede, a tool built on LLVM compiler framework to find bugs. Experimental results show that Rulede can effectively detect performance bugs and security bugs in target IOT OSes.
Hongliang Liang
SNPD1
2015 Survey on Privacy Protection of Android Devices
abstract
Nowadays, the ubiquity of smart phones make them carry large amounts of personal sensitive information, but at the same time, there are also many Apps in Android APP market that target to collect users' sensitive data. So it becomes quite important to prevent users from the threat of privacy leakage. In this paper, we analyze the Android's privacy protection mechanism, and describe various threats to users' different types of privacy data. After that, we enumerate two ways that can leak sensitive information, and discuss the current solutions and techniques from aspects of privacy protection enhancement and privacy leakage detection. We also make a fine-grained classification for these two aspects, and study the difference between solutions in each category. Finally, we summarize the deficiency of existing research of Android privacy protection and propose the future research direction.
Hongliang Liang, Dongyang Wu, Jiuyun Xu, Hengtai Ma
CSCloud1
2015 An Uneven Distributed System for Dynamic Taint Analysis Framework
abstract
Dynamic taint analysis has been widely used in software testing, debugging, vulnerability detection and other fields. A popular idea is that we can combine dynamic taint analysis with symbolic execution techniques or fuzz techniques forming the testing framework to test automatically. When testing large applications which costs longer time, a distributed system can be very practical. However, the common distributed system is load balancing which distributes tasks without considering the various performance of each machine, resulting that some machines with poor configuration will burden too much load. In this paper, we present an uneven distributed system, which splits the dynamic taint analysis framework into some modules, and then distributes the modules to different machines classified by their performance. The design and distribution method are all based on the feature of each module. In the studies, we applied the system to test 5 applications compared with the load balancing distributed system, and the results shows it can indeed distribute tasks uneven according to different performance.
Hengtai Ma, Hongliang Liang
CSCloud4
2014 A Lightweight Security Isolation Approach for Virtual Machines Deployment
Hongliang Liang, Changyao Han, Daijie Zhang, Dongyang Wu
Inscrypt1
2013 EAdroid: Providing Environment Adaptive Security for Android System
Hongliang Liang, Shuchang Liu 0005
Inscrypt1
2009 Static Analysis of a Class of Memory Leaks in TrustedBSD MAC Framework
Xinsong Wu, Zhouyi Zhou, Yeping He, Hongliang Liang
ISPEC4
2005 Security On-demand Architecture with Multiple Modules Support
Wenchang Shi, Hongliang Liang, Qinghua Shang, Chunyang Yuan, Bin Liang 0002
ISPEC3