Guojun Peng

dblp:30/4066 · DBLP profile ↗
← Back
39ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0001-5731-8958ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 20 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Computer networks · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OSmartPro: a large language model-assisted option fuzzing approach
abstract
Abstract Program options provide flexible software functionality control but complicate fuzz testing, as triggering many behaviors require specific option combinations. Although existing option-aware fuzzing approaches attempt to mutate options as inputs or leverage AI technologies to extract option relationships from documentation, these methods have limitations. Documentation is often incomplete, and some option dependencies are embedded deeply within program logic via data or control flows, making these methods challenging to detect all possible dependencies. This paper introduces OSmartPro , an advanced option-fuzzing approach that directly extracts options and infers option dependencies from source code. Given LLM’s capabilities to interpret program semantics, OSmartPro employs LLM-assisted static analysis to handle diverse option-parsing structures and extract comprehensive options. Through control and data dependency analysis, it constructs option impact graph , which it uses to guide fuzzing strategies. The tool successfully extracted complete options from all 59 programs in our test set, uncovering undocumented options in over 66% of them. Additionally, OSmartPro inferred 14,701 option combinations, identified 45.03% more execution paths compared to AFL++, and uncovered 54 zero-day vulnerabilities, of which 18 awarded CVE IDs. Lastly, in a benchmark comparison against four option-aware fuzzers, OSmartPro achieved higher line coverage in 66.7% (20 out of 30) of the programs.
Kelin Wang, Mengda Chen, Liang He 0011, Purui Su, Jiongyi Chen, Yan Cai 0001, Chao Feng 0002, Chaojing Tang, Guojun Peng
Cybersecur.11
2026 CTI-Thinker: an LLM-driven system for CTI knowledge graph construction and attack reasoning
abstract
Abstract With the increasing frequency of APT attacks, cyber defense urgently demands high-quality threat intelligence support. Cyber threat intelligence (CTI) knowledge graphs have demonstrated significant potential in aiding threat detection and behavioral reasoning. However, existing CTI data often suffer from unstructured formats, fragmented knowledge, a reliance on manual annotation, and limited semantic mapping to attack techniques. These limitations hinder the robustness and accuracy of downstream reasoning tasks (e.g., attack attribution and intent inference). Moreover, traditional information extraction methods struggle to generalize in scenarios involving cross-paragraph dependencies, emerging threats, and low-resource samples, exhibiting weaknesses in context awareness and sensitivity to prompt variations. To this end, we propose CTI-Thinker, a novel system that integrates large language models with semantic alignment to the ATT&CK framework for CTI knowledge graph construction and threat reasoning. First, CTI-Thinker leverages in-context learning and LoRA-based fine-tuning to extract structured threat entities and relations. Then, it adopts vector-based alignment strategies to unify heterogeneous expressions, enabling entity normalization and knowledge fusion for constructing a high-quality CTI knowledge graph. Finally, a GraphRAG-based reasoning engine is built by incorporating the structured knowledge graph and external ATT&CK resources into a retrieval-augmented generation (RAG) framework, enabling tactical-level inference and CTI-driven question answering. Experimental results demonstrate that CTI-Thinker accurately extracts threat entities and relations and constructs a reliable CTI knowledge graph. It also effectively infers attack intent and supports intelligent reasoning. The system outperforms state-of-the-art methods in precision, robustness, and generalizability, offering a scalable and semantically enriched solution for cyber threat analysis and defense. Graphical abstract
Xiuzhang Yang, Ruijie Zhong, Yuling Chen 0002, Guojun Peng, Dongni Zhang
Cybersecur.4
2026 TSGDroid: Trigger Semantic Graph Modeling for Detecting Suspicious Hidden Sensitive Operations
Dongni Zhang, Xiuzhang Yang, Side Liu, Jinwen Xin, Jianming Fu, Guojun Peng
IEEE Internet Things J.7
2026 Backdoor samples detection based on perturbation discrepancy consistency in pre-trained language models
Zuquan Peng, Jianming Fu, Lixin Zou, Yanzhen Ren, Guojun Peng
Neural Networks6
2025 Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model
abstract
Malicious PDF files have emerged as a persistent threat and become a popular attack vector in web-based attacks. While machine learning-based PDF malware classifiers have shown promise, these classifiers are often susceptible to adversarial attacks, undermining their reliability. To address this issue, recent studies have aimed to enhance the robustness of PDF classifiers. Despite these efforts, the feature engineering underlying these studies remains outdated. Consequently, even with the application of cutting-edge machine learning techniques, these approaches fail to fundamentally resolve the issue of feature instability. To tackle this, we propose a novel approach for PDF feature extraction and PDF malware detection. We introduce the PDFObj IR (PDF Object Intermediate Representation), an assembly-like language framework for PDF objects, from which we extract semantic features using a pretrained language model. Additionally, we construct an Object Reference Graph to capture structural features, drawing inspiration from program analysis. This dual approach enables us to analyze and detect PDF malware based on both semantic and structural features. Experimental results demonstrate that our proposed classifier achieves strong adversarial robustness while maintaining an exceptionally low false positive rate of only 0.07% on baseline dataset compared to state-of-the-art PDF malware classifiers.
Side Liu, Jiang Ming 0002, Guodong Zhou 0002, Jianming Fu, Guojun Peng
CCS6
2025 Beyond Tag Collision: Cluster-based Memory Management for Tag-based Sanitizers
Mengfei Xie, Yan Lin 0003, Jianming Fu, Chenke Luo, Guojun Peng
CCS6
2025 BCDTrack: Bidirectional Constraint-Driven Online Multi-Object Tracking
abstract
Multi-object tracking (MOT) is crucial for video analysis and various computer vision applications. Traditional MOT methods primarily rely on unidirectional trajectory prediction, which can be severely affected by occlusions and short-term object losses, leading to tracking failures or incorrect associations. A significant challenge in MOT is handling these issues while maintaining accurate and consistent tracking over long durations. To address this issue, we propose a novel Bidirectional Constraint-Driven Online Multi-Object Tracking (BCDTrack) method that improves long-term trajectory association. The proposed method performs forward and backward tracking on video sequences using a sliding window approach. In each window, forward tracking is executed first, followed by backward tracking, and the motion information is used to fuse the forward and backward trajectories. Furthermore, to ensure identity (ID) consistency during the fusion process, we design a trajectory fusion strategy that utilizes Kalman filter to perform forward and backward predictions. Extensive experiments and ablation studies on the MOT20 datasets demonstrate that the proposed approach significantly enhances long-term tracking performance, particularly in dynamic and occlusion-prone scenarios, offering superior robustness against object loss and tracking failures.
Guojun Peng, Mithun Mukherjee 0001, Constandinos X. Mavromoustakis
GLOBECOM4
2025 Retrofitting XoM for Stripped Binaries without Embedded Data Relocation
Chenke Luo, Jiang Ming 0002, Mengfei Xie, Guojun Peng, Jianming Fu
NDSS4
2025 MetaBox-v2: A Unified Benchmark Platform for Meta-Black-Box Optimization
abstract
Meta-Black-Box Optimization (MetaBBO) streamlines the automation of optimization algorithm design through meta-learning. It typically employs a bi-level structure: the meta-level policy undergoes meta-training to reduce the manual effort required in developing algorithms for low-level optimization tasks. The original MetaBox (2023) provided the first open-source framework for reinforcement learning-based single-objective MetaBBO. However, its relatively narrow scope no longer keep pace with the swift advancement in this field. In this paper, we introduce MetaBox-v2 (\url{https://github.com/MetaEvo/MetaBox}) as a milestone upgrade with four novel features: 1) a unified architecture supporting RL, evolutionary, and gradient-based approaches, by which we reproduce $23$ up-to-date baselines; 2) efficient parallelization schemes, which reduce the training/testing time by $10-40$x; 3) a comprehensive benchmark suite of $18$ synthetic/realistic tasks ($1900$+ instances) spanning single-objective, multi-objective, multi-model, and multi-task optimization scenarios; 4) plentiful and extensible interfaces for custom analysis/visualization and integrating to external optimization tools/benchmarks. To show the utility of MetaBox-v2, we carry out a systematic case study that evaluates the built-in baselines in terms of the optimization performance, generalization ability and learning efficiency. Valuable insights are concluded from thorough and detailed analysis for practitioners and those new to the field.
Zeyuan Ma, Yue-Jiao Gong, Hongshu Guo, Wenjie Qiu 0007, Sijie Ma, Hongqiao Lian, Jiajun Zhan, Kaixu Chen, Zhiyang Huang, Zechuan Huang, Guojun Peng, Yining Ma 0001
NeurIPS12
2025 MemoryTrap: Booby Trapping Memory to Counter Memory Disclosure Attacks with Hardware Support
Chenke Luo, Jiang Ming 0002, Dongpeng Xu 0001, Guojun Peng, Jianming Fu
USENIX ATC4
2025 VAPD: An Anomaly Detection Model for PDF Malware Forensics with Adversarial Robustness
Side Liu, Jiang Ming 0002, Jianming Fu, Guojun Peng
USENIX Security Symposium5
2025 The hidden complexities of Android TPL detection: An empirical analysis of techniques, challenges, and effectiveness
Lige Zhan, Jiang Ming 0002, Jianming Fu, Guojun Peng, Letian Sha, Lili Lan
Comput. Secur.4
2025 A survey on Android dynamic evasive malware: Taxonomy, countermeasures and open challenges
Dongni Zhang, Xiuzhang Yang, Side Liu, Jianming Fu, Guojun Peng
Comput. Secur.6
2025 XLM4Detector: Multistage Deobfuscation and Semantic-Driven Excel 4.0 Macro Malware Detection
abstract
Excel 4.0 Macro leverages XLM code to directly invoke system APIs and automate complex tasks, making it a widely used tool in phishing attacks, APT campaigns, and IoT intrusions in recent years. By constructing various obfuscated macro malware, attackers can easily evade firewalls and detection systems, thereby achieving persistent attacks. However, existing XLM malware defense mechanisms lack in-depth analysis of XLM malware families and behaviors, failing to integrate multi-dimensional features and semantic relationships effectively. As a result, detection systems struggle to accurately identify malicious operations in real-world attacks, leading to low robustness and accuracy. To this end, we propose XLM4Detector, a novel Excel 4.0 Macro malware detection framework based on multi-stage deobfuscation and multi-view semantic fusion. First, XLM4Detector integrates AST analysis, simulated execution, and regular expression matching to construct a multi-stage deobfuscation algorithm, enabling precise deobfuscation and XLM code extraction. Second, we introduce four feature extraction methods that capture fine-grained features at the word (string), token (function), abstract syntax tree, and semantic relationship levels. Then, we design four embedding representations (XlmWord2Vec, XlmToken2Vec, XlmAst2Vec, XlmRela2Vec) and employ a multi-view semantic fusion algorithm for feature alignment. Finally, we develop an MHSACNN-BiGRU model to capture hierarchical semantic relationships, effectively enabling XLM malware behavior detection and family classification. Experimental results demonstrate that XLM4Detector effectively reconstructs obfuscated XLM source code and accurately detects XLM malware families and behaviors. It outperforms state-of-the-art methods in detection accuracy, robustness, and generalization. Our framework provides critical technical support for IoT security defense, malicious document detection, and APT tracking.
Xiuzhang Yang, Yuling Chen 0002, Zhi Ouyang, Guojun Peng
IEEE Internet Things J.9
2025 MODFuzz: A Multiobjective Directed Fuzzer for USB Drivers
abstract
USB interfaces have become ubiquitous in various Internet of Things (IoT) devices, all adhering to the same universal serial bus (USB) protocol. While enhancing convenience, they also widen the potential attack surface. Fuzzing is a proactive way to identify potential security threats for USB drivers. However, existing USB driver fuzzers primarily prioritize the code coverage of USB drivers, leading to a significant waste of computational resources on irrelevant code segments. To this end, we combine directed fuzzing and USB driver fuzzing for the first time, and present multiobjective directed fuzzer (MODFuzz), a pioneering multiobjective directed fuzzing method for USB drivers. MODFuzz autonomously locates the most vulnerable parts within USB drivers, concentrating fuzzing efforts on these areas. Diverging from the existing directed fuzzers, MODFuzz employs a dynamic direction instead of predetermined addresses to guide the fuzzing campaign toward the triggered execution traces with a greater probability of containing vulnerabilities. MODFuzz outperforms the strong baseline in terms of execution speed (about 14% improvement) and crash generation capabilities (about 69% improvement). Meanwhile, we found six previously unknown bugs (all confirmed and assigned vulnerability IDs) in Linux kernel v6.4.10 and received acknowledgment from Red Hat.
Guojun Peng, Xingliang Wang, Zichuan Li, Side Liu, Xiuzhang Yang, Jianming Fu
IEEE Internet Things J.2
2025 Egalitarian Randomization for Multi-Language Applications on ARM64
abstract
Due to the inevitable information loss during IR lowering, compile-time metadata collection can provide more precise auxiliary information than binary analysis to achieve reliable fine-grained randomization. However, existing schemes build on deep modifications of compilers, making it challenging to provide consistent randomization protection for different high-level languages. Additionally, they are inadequate for securing widely used smartphones and embedded devices, since only ×86-64 applications are currently supported. In this paper, we present MLARandom, a compiler-assisted function-level randomization scheme designed for Multi-Language ARM64 applications. MLARandom employs a lightweight compilation standardization strategy that allows for uniform information collection at the assembly level, regardless of the high-level language or compiler used. Further, it combines ARM64 architecture specifications and collected relocation types to accurately repair all ARM64 pointers after randomization. Our experimental results show that MLARandom can equally randomize modules developed in different languages (e.g., C/C++, Rust, Fortran, Cangjie) with negligible runtime overhead (0.51%), to effectively counter against traditional Code Reuse Attacks as well as advanced Cross-Language Attacks. Although randomization approaches based on reassembly can achieve similar goals, our empirical evaluation highlights the imprecise pointer identification as a major obstacle to their practical deployment.
Mengfei Xie, Yan Lin 0003, Jianming Fu, Chenke Luo, Guojun Peng
IEEE Trans. Dependable Secur. Comput.5
2025 Efficient Intrusion Detection for In-Vehicle Networks Using Knowledge Distillation From BERT to CNN-BiLSTM
abstract
Under the development of intelligent transportation systems, In-Vehicle Networks (IVNs) serve as a critical channel for both internal and external communications. However, the inherent complexity and diversity of data traffic present significant challenges for the detection of IVN anomalous flows. Meanwhile, the introduction of various novel technologies has introduced new security vulnerabilities to IVNs. These vulnerabilities significantly impact the security of IVNs and the accuracy of in-vehicle Intrusion Detection Systems (IDS). To address these issues, this paper proposes a lightweight and efficient anomaly detection method based on knowledge distillation technology, termed Knowledge Distillation from BERT to CNN-BiLSTM (KDBC). Specifically, the KDBC distills the deep semantic knowledge from the BERT model into a more lightweight CNN-BiLSTM architecture, significantly reducing computational overhead and storage requirements without substantially compromising detection performance. Experimental results demonstrate that the KDBC model enhances both security and versatility, achieving superior detection accuracy in identifying abnormal attacks across diverse IVN data, including automotive Ethernet and CAN networks. Moreover, the KDBC model has been validated for its effectiveness and robustness in actual in-vehicle gateway environments, achieving an accuracy of over 0.98 and an F1 score greater than 0.98.
Yue Cao 0002, Guojun Peng, Meng Li 0006
IEEE Trans. Inf. Forensics Secur.3
2024 A survey on the evolution of fileless attacks and detection techniques
Side Liu, Guojun Peng, Haitao Zeng, Jianming Fu
Comput. Secur.2
2023 MetaBox: A Benchmark Platform for Meta-Black-Box Optimization with Reinforcement Learning
abstract
Recently, Meta-Black-Box Optimization with Reinforcement Learning (MetaBBO-RL) has showcased the power of leveraging RL at the meta-level to mitigate manual fine-tuning of low-level black-box optimizers. However, this field is hindered by the lack of a unified benchmark. To fill this gap, we introduce MetaBox, the first benchmark platform expressly tailored for developing and evaluating MetaBBO-RL methods. MetaBox offers a flexible algorithmic template that allows users to effortlessly implement their unique designs within the platform. Moreover, it provides a broad spectrum of over 300 problem instances, collected from synthetic to realistic scenarios, and an extensive library of 19 baseline methods, including both traditional black-box optimizers and recent MetaBBO-RL methods. Besides, MetaBox introduces three standardized performance metrics, enabling a more thorough assessment of the methods. In a bid to illustrate the utility of MetaBox for facilitating rigorous evaluation and in-depth analysis, we carry out a wide-ranging benchmarking study on existing MetaBBO-RL methods. Our MetaBox is open-source and accessible at: https://github.com/GMC-DRL/MetaBox.
Zeyuan Ma, Hongshu Guo, Zhenrui Li, Guojun Peng, Yue-Jiao Gong, Yining Ma 0001, Zhiguang Cao
NeurIPS5
2023 PointerScope: Understanding Pointer Patching for Code Randomization
abstract
Various fine-grained randomization schemes have been designed to increase the entropy of process space, while none of them can rise from an academic exercise to industrial deployment like Address Space Layout Randomization (ASLR). One of the critical reasons is the incorrectness of randomization caused by the mismatch between their pointer collection capabilities and the high accuracy requirements of the pointer patching task. In this article, we present PointerScope, an accurate compile-time pointer collection scheme deriving from a group of novel observations. The success of PointerScope relies on the complete tracing of the pointer generation process, including the compilation chain from compiler to static linker and the interface specification between them. From this view, PointerScope identifies four types of pointer-related static linker behaviors and clarifies five types of inherent addressing modes in the x86-64 architecture. The vague understanding of them causes the Compiler-assisted Code Randomization (CCR) to incorrectly collect pointers and patch them to the wrong values after randomization. Further, we measure the pointer collection capability of augmented binary analysis, the experimental results show that they can mitigate challenges from the traditional binary analysis by the given premises, but additional heuristics still need to be designed to support the fine-grained randomization.
Mengfei Xie, Yan Lin 0003, Chenke Luo, Guojun Peng, Jianming Fu
IEEE Trans. Dependable Secur. Comput.4
2023 Reverse Engineering of Obfuscated Lua Bytecode via Interpreter Semantics Testing
abstract
As an efficient and multi-platform scripting language, Lua is gaining increasing popularity in the industry. Unfortunately, Lua’s unique advantages also catch cybercriminals’ attention. A growing number of IoT malware authors switch to Lua for malicious payload development and then distribute malware in bytecode form. To impede malware code analysis, malware authors obfuscate standard Lua bytecode into a customized bytecode specification. Only the attached interpreter can execute that particular bytecode file. Rapid recovery of Lua obfuscated bytecode is essential for a swift response to new malware threats. However, existing generic code deobfuscation approaches cannot keep up with the pace of emerging threats. In this paper, we present a novel reverse engineering technique, calledinterpreter semantics testing. Given a customized interpreter used to execute obfuscated Lua bytecode, we construct a set ofLuaGadgetsthat can adapt to the customized interpreter. Each LuaGadget contains a carefully chosen opcode sequence to fulfill an observable calculation—it is designed to test one or two particular opcodes at a time. Next, we mutate unknown opcode values to generate a bunch of test cases and run them using the customized interpreter; we can observe the expected result only when the mutation hits the opcode’s right value. We perform test case prioritization to cost-effectively recover the semantics of all obfuscated opcodes. Our approach makes no assumptions about the interpreter’s structure and is free from analyzing the numerous execution traces of opcode handlers. We have evaluated our tool,LuaHunt, with Lua malware variants and real-world applications. LuaHunt is able to recover the obfuscated bytecode’s semantics within 90 seconds for each test case, and all of our deobfuscation results can pass the correctness testing. The encouraging results demonstrate that LuaHunt is a promising tool to lighten the burden of security analysts.
Chenke Luo, Jiang Ming 0002, Jianming Fu, Guojun Peng, Zhetao Li
IEEE Trans. Inf. Forensics Secur.4
2022 Ship Path Optimization That Accounts for Geographical Traffic Characteristics to Increase Maritime Port Safety
abstract
Maritime ports face challenges associated with navigation safety, operational efficiency, and management. With the development of the Internet of Things, artificial intelligence simulation technologies, geographical information systems, and cloud computing technologies as well as navigation aids and decision support systems in maritime transportation, ports have the potential to better manage traffic, loading, and unloading. Recently, there has been growing attention in unmanned shipping to support the maritime industry and the military. This paper aims to extend the application of geographical theory and methodology in unmanned ship path optimization. Automatic collision avoidance concerning maneuvering capabilities of ships as well as complying with maritime traffic rules remains a challenge. This study attempts to tackle development needs associated with path optimization in maritime travel. By integrating ship movement behavior, geographical features, and the International Regulations for Avoiding Collisions at Sea, the proposed methods seek to reduce the human error associated with maritime accidents. This paper proposes economic efficiency and safety-driven unmanned ship path planning that will promote the future growth of intelligent port development.
Hongchu Yu, Alan T. Murray, Zhixiang Fang, Jingxian Liu, Guojun Peng, Mohammad Solgi
IEEE Trans. Intell. Transp. Syst.5
2021 Towards Transparent and Stealthy Android OS Sandboxing via Customizable Container-Based Virtualization
abstract
A fast-growing demand from smartphone users is mobile virtualization.This technique supports running separate instances of virtual phone environments on the same device. In this way, users can run multiple copies of the same app simultaneously,and they can also run an untrusted app in an isolated virtual phone without causing damages to other apps. Traditional hypervisor-based virtualization is impractical to resource-constrained mobile devices.Recent app-level virtualization efforts suffer from the weak isolation mechanism. In contrast, container-based virtualization offers an isolated virtual environment with superior performance.However, existing Android containers do not meet the anti-evasion requirement for security applications: their designs are inherently incapable of providing transparency or stealthiness.
Wenna Song, Jiang Ming 0002, Xuanchen Pan, Jianming Fu, Guojun Peng
CCS7
2021 App's Auto-Login Function Security Testing via Android OS-Level Virtualization
abstract
Limited by the small keyboard, most mobile apps support the automatic login feature for better user experience. Therefore, users avoid the inconvenience of retyping their ID and password when an app runs in the foreground again. However, this auto-login function can be exploited to launch the so-called "data-clone attack": once the locally-stored, auto-login depended data are cloned by attackers and placed into their own smartphones, attackers can break through the login-device number limit and log in to the victim's account stealthily. A natural countermeasure is to check the consistency of device-specific attributes. As long as the new device shows different device fingerprints with the previous one, the app will disable the auto-login function and thus prevent data-clone attacks. In this paper, we develop VPDroid, a transparent Android OS-level virtualization platform tailored for security testing. With VPDroid, security analysts can customize different device artifacts, such as CPU model, Android ID, and phone number, in a virtual phone without user-level API hooking. VPDroid's isolation mechanism ensures that user-mode apps in the virtual phone cannot detect device-specific discrepancies. To assess Android apps' susceptibility to the data-clone attack, we use VPDroid to simulate data-clone attacks with 234 most-downloaded apps. Our experiments on five different virtual phone environments show that VPDroid's device attribute customization can deceive all tested apps that perform device-consistency checks, such as Twitter, WeChat, and PayPal. 19 vendors have confirmed our report as a zero-day vulnerability. Our findings paint a cautionary tale: only enforcing a device-consistency check at client side is still vulnerable to an advanced data-clone attack.
Wenna Song, Jiang Ming 0002, Han Yan 0013, Jianming Fu, Guojun Peng
ICSE8
2021 Obfuscation-Resilient Executable Payload Extraction From Packed Malware
Binlin Cheng, Jiang Ming 0002, Erika A. Leal, Haotian Zhang 0006, Jianming Fu, Guojun Peng, Jean-Yves Marion
USENIX Security Symposium6
2021 Context-Rich Privacy Leakage Analysis Through Inferring Apps in Smart Home IoT
abstract
Emerging Internet of Things (IoT) systems leverage connected devices to enable intelligent and automated functionalities. Despite the benefits, there exist privacy risks of network traffic, which have been studied by the previous research. However, with the current privacy inference remaining at the event-level, potential privacy risks are underestimated, which, as our study shows, can be much higher than previously reported through app-level traffic analysis. A key observation of our research is that IoT event-triggered traffic is generated by apps, which often adopt an if-trigger-then-action (trigger-action) programming paradigm. We utilize this feature to develop fingerprints to differentiate running apps and learn context-rich privacy-sensitive information from apps. In this article, we present a privacy leakage analysis called ALTA to infer running apps in smart home IoT environments. First, ALTA identifies app fingerprints through static analysis and extracts sensitive information from app descriptions and input prompts. Then, through dynamic traffic profiling, it learns traffic fingerprints of apps. Finally, ALTA matches the fingerprints of app and traffic, and thus is able to pinpoint which app is running from IoT traffic at runtime. To demonstrate the feasibility of our approach, we analyze 254 SmartThings applications via program and natural language processing (NLP) analysis. We also perform the app inference evaluation on 31 apps executed in a simulated smart home. The results suggest that ALTA can effectively infer running apps from IoT traffic and learn context-rich information (e.g., health conditions, daily routines, and user activities) from apps with high accuracy.
Long Cheng 0005, Hongxin Hu, Guojun Peng, Danfeng Yao
IEEE Internet Things J.4
2021 A Direction-Constrained Space-Time Prism-Based Approach for Quantifying Possible Multi-Ship Collision Risks
abstract
Maritime collision risk prediction is crucial for the safety management of ocean transportation. Previous studies have primarily focused on near-miss collision risk of ship pairs, yet the risk due to congestion caused by multiple ships is significant. This paper proposes a novel space-time geographical approach for addressing multi-ship near-miss collision risk based on vessel motion behavior. The direction-constrained space-time prism is used to characterize the interaction possibility of ships, enabling potential collision risk to be evaluated. The advantage of direction-constrained space-time prism in analyzing ship movements is that it accounts for direction limitation thereby eliminating unreasonable estimation associated with the assumption of arbitrary changes in sailing direction through the classical space-time prism. This paper uses the trajectories of ships traveling the southeast coast of China to investigate the viability of the proposed approach. In comparison to the fuzzy quaternion ship domain model and closest point of approach-based methods, the proposed approach is capable of identifying hierarchical near-miss collision risks for different ships to improve risk evaluation. This is essential for ship path optimization that accounts for changes in the speed and course of ships. This work supports maritime collision risk forecasting, but also provides detailed insights for actionable steps to reduce risks.
Hongchu Yu, Zhixiang Fang, Alan T. Murray, Guojun Peng
IEEE Trans. Intell. Transp. Syst.4
2020 JTaint: Finding Privacy-Leakage in Chrome Extensions
Mengfei Xie, Jianming Fu, Chenke Luo, Guojun Peng
ACISP5
2020 VAHunt: Warding Off New Repackaged Android Malware in App-Virtualization's Clothing
abstract
Repackaging popular benign apps with malicious payload used to be the most common way to spread Android malware. Nevertheless, since 2016, we have observed an alarming new trend to Android ecosystem: a growing number of Android malware samples abuse recent app-virtualization innovation as a new distribution channel. App-virtualization enables a user to run multiple copies of the same app on a single device, and tens of millions of users are enjoying this convenience. However, cybercriminals repackage various malicious APK files as plugins into an app-virtualization platform, which is flexible to launch arbitrary plugins without the hassle of installation. This new style of repackaging gains the ability to bypass anti-malware scanners by hiding the grafted malicious payload in plugins, and it also defies the basic premise embodied by existing repackaged app detection solutions.
Luman Shi, Jiang Ming 0002, Jianming Fu, Guojun Peng, Dongpeng Xu 0001, Xuanchen Pan
CCS4
2020 Detection of Repackaged Android Malware with Code-Heterogeneity Features
abstract
During repackaging, malware writers statically inject malcode and modify the control flow to ensure its execution. Repackaged malware is difficult to detect by existing classification techniques, partly because of their behavioral similarities to benign apps. By exploring the app's internal different behaviors, we propose a new Android repackaged malware detection technique based on code heterogeneity analysis. Our solution strategically partitions the code structure of an app into multiple dependence-based regions (subsets of the code). Each region is independently classified on its behavioral features. We point out the security challenges and design choices for partitioning code structures at the class and method level graphs, and present a solution based on multiple dependence relations. We have performed experimental evaluation with over 7,542 Android apps. For repackaged malware, our partition-based detection reduces false negatives (i.e., missed detection) by 30-fold, when compared to the non-partition-based approach. Overall, our approach achieves a false negative rate of 0.35 percent and a false positive rate of 2.97 percent.
Ke Tian, Danfeng Yao, Barbara G. Ryder, Gang Tan, Guojun Peng
IEEE Trans. Dependable Secur. Comput.5
2019 Automatic Identification System-Based Approach for Assessing the Near-Miss Collision Risk Dynamics of Ships in Ports
abstract
Vessel risk analysis is critical for safe ship navigation and maritime safety management. Near-miss collisions by ships comprise an significant risk, which may be complicated by factors, such as the ship conditions, waterway environment, and driving behavior of any ships encountered. Previous studies have rarely considered how to automatically and adaptively estimate the risk of near-miss collisions for different situations, particularly in port areas. In this paper, we propose an automatic identification system-based approach for adaptively calibrating near-miss collision risk model and assessing a ship's near-miss collision risk by using the vessel's speed and course patterns to obtain a robust estimate of the collision risk. Six measures are employed to determine the hierarchical geographical distribution of the near-miss collision risk for ships in port areas and to identify the high-risk areas. Some predicted high-risk areas were validated as official precautionary areas in the Xiamen Port area. All predicted areas may help the port administration to plan monitoring areas to ensure safe traffic flow in the port.
Zhixiang Fang, Hongchu Yu, Ran-Xuan Ke, Shih-Lung Shaw, Guojun Peng
IEEE Trans. Intell. Transp. Syst.5
2018 Towards Paving the Way for Large-Scale Windows Malware Analysis: Generic Binary Unpacking with Orders-of-Magnitude Performance Boost
abstract
Binary packing, encoding binary code prior to execution and decoding them at run time, is the most common obfuscation adopted by malware authors to camouflage malicious code. Especially, most packers recover the original code by going through a set of "written-then-executed" layers, which renders determining the end of the unpacking increasingly difficult. Many generic binary unpacking approaches have been proposed to extract packed binaries without the prior knowledge of packers. However, the high runtime overhead and lack of anti-analysis resistance have severely limited their adoptions. Over the past two decades, packed malware is always a veritable challenge to anti-malware landscape. This paper revisits the long-standing binary unpacking problem from a new angle: packers consistently obfuscate the standard use of API calls. Our in-depth study on an enormous variety of Windows malware packers at present leads to a common property: malware's Import Address Table (IAT), which acts as a lookup table for dynamically linked API calls, is typically erased by packers for further obfuscation; and then unpacking routine, like a custom dynamic loader, will reconstruct IAT before original code resumes execution. During a packed malware execution, if an API is invoked through looking up a rebuilt IAT, it indicates that the original payload has been restored. This insight motivates us to design an efficient unpacking approach, called BinUnpack. Compared to the previous methods that suffer from multiple "written-then-executed" unpacking layers, BinUnpack is free from tedious memory access monitoring, and therefore it introduces very small runtime overhead. To defeat a variety of ever-evolving evasion tricks, we design BinUnpack's API monitor module via a novel kernel-level DLL hijacking technique. We have evaluated BinUnpack's efficacy extensively with more than 238K packed malware and multiple Windows utilities. BinUnpack's success rate is significantly better than that of existing tools with several orders of magnitude performance boost. Our study demonstrates that BinUnpack can be applied to speeding up large-scale malware analysis.
Binlin Cheng, Jiang Ming 0002, Jianming Fu, Guojun Peng, Ting Chen 0002, Xiaosong Zhang 0001, Jean-Yves Marion
CCS4
2018 An Efficient SCA Leakage Model Construction Method Under Predictable Evaluation
abstract
Leakage models, regarded as a bridge between the physical signal and the sensitive operation, have a great influence on the effectiveness of the side channel analysis. The existing leakage models are usually divided into two categories, the non-profiled leakage models which have been chosen before sampling and analyzing, such as Hamming weight and Hamming distance, while the profiled leakage models, whose parameters have to be trained in the profiling phase, such as the Stochastic model of which both coefficient vector and pooled covariance matrix are required to be estimated based on the acquired samples. In general, a profiled leakage model is more accurate than a non-profiled one. However, it may lead to an inefficient attack if the leakage function is inaccurate, e.g., the over-fitting and under-fitting in the profiling phase. In this paper, we mathematically prove the relationship among different stochastic models, and propose a new method named ECM to solve the problem that much time is required to solve matrix in the profiling phase. Replacing the observations in the matrix solution with the average signals, the new method accelerates the construction of any stochastic model significantly, as long as the data-dependent signal has the property equal images under different subkeys. On the basis of theoretical results, we analyze the reasons why over-fitting and under-fitting happen, and quantify the condition when some of them occur. Finally, comparing with the existing construction method (HSS2012), we verify the effectiveness and efficiency of ECM with different metrics. Under the same accuracy, the ECM obviously has lower time complexity than HSS2012.
Ming Tang 0002, Xiaoqi Ma, Wenjie Chang, Huanguo Zhang, Guojun Peng, Jean-Luc Danger
IEEE Trans. Inf. Forensics Secur.6
2008 Further Research of RFID Applying on Exhibition Logistics
abstract
The world economy is becoming increasingly global and more technology-based. From 1999, RFID becomes popular and more and more retailers, banks, etc. apply the new technology to their products or services, operation and management; as well, some logistics providers, such as Schenker, test RFID for shuttle service. The biggest limitation of RFID's widespread is the cost of RFID, both tag and reader. However, RFID's advantages are also very obviously, which calls for the analysis of its application to Exhibition Logistics. The commodities or goods for an exhibition usually are high value and require strict supervision, safely and securely. With continuous improvement in technologies and transportation measurements, even the distribution of goods and services is taken on new meaning. This paper discusses how RFID employed in Exhibition Logistics and calls for applying RFID on worldwide Exhibition Logistics as soon as possible. The significant affect would help logistics enterprises to change accordingly, getting more effective and globally competitive.
Ran-Xuan Ke, Guojun Peng
CW2
2008 Research on 3D Simulation Technology Applied in Navigation-Aid Management
abstract
Application of 3D simulation technology in management of navigation-aid helps a lot for authorities to more scientific intuitively right placement, monitoring and maintaining the navigation-aid. Through deep research on the "3D navigation application" IALA recommendation and the analysis of international and domestic recent advances in high-tech information technology in the field of navigation applications, as well as combine navigation-aid custody of the day-to-day operational needs. This paper develop an exploration into management and monitoring of navigation-aid which based on 3D simulation technology, simultaneously, by the use of high-fidelity 3D visual simulation and ship handling simulator, the concrete simulation inspections applied research on the navigation-aid layout was underway. This paper briefly introduces the research of the application of 3D simulation technology in management of navigation-aid. And associated with the status of main channel of Xiamen, from two application layers--- one is inspection of navigation-aid interval, the other is verification of the most suitable position for navigation-aid in turning point of channel, this paper presents the way of navigation-aid layout by help from full mission ship handling simulator.
Guojun Peng, Ran-Xuan Ke, JinXing Shao
CW1
2008 Research on Navigation-Aids Information System
abstract
This thesis researches on the application of computer, modern communication, GIS, GPS, AIS and World-Wide-Web in the field of navigation-aids information system, and has realized an integrated system consisted of navigation-aids information GIS platform, navigation-aids monitoring system and navigation-aids information distribution system. This system has strong integration capability, and has realized navigation-aids information distribution based on WEBGIS at the first time. It strongly promotes navigation-aids daily management and maintenance, and this system provides technique guarantee for ships and marine departments to acquire navigation-aids information in time, by rule and line expediently.
Guojun Peng, XingGu Zhang, Ran-Xuan Ke, JinXing Shao
CW1
2006 Research on Sea Digital Map used for Ship Navigation
abstract
With rapid development of modern electronic information technique and navigation technique, the Sea Digital Map, presenting sea geographic information and maritime traffic information in digital style, has come into being. It brought birth to a technique revolution in the field of ship navigation, and accelerated the development of the informationalization, intelligentization, and information sharing of maritime ship traffic. Nowadays, Sea Digital Map and its application systems have been the key part in Ship Integrative Navigation and Maritime Traffic Control. This thesis has made a deeply research on the data structure, accessing control, standard map symbol display of the Sea Digital Map which is met with the criterion Lt Military Digital Sea Map for Ship Navigation Exchange Format Gt of Chinese Navy Navigation Guarantee Dept. Based on the data and display separate idea, and from the point of function encapsulated independently, the thesis developed the Sea Digital Map used for Ship Navigation. It is composed of basal platform of Sea Digital Map, GPS interface, AIS interface, B/S and C/S net communication interface and ARPA frame picture overlay. It's a Web Digital Navigation Map application platform made up with the functions of digital map display and edition, route designing, route monitoring, ship monitoring, maritime traffic data analyzing. Sea Digital Map used for Ship Navigation could be used widely in such fields as ship navigation, maritime transaction, maritime safety traffic management, sea function planning management, etc. This thesis gives a detailed introduction about Sea Digital Map used for Ship Navigation from the point of the composition and realizing technique.
XingGu Zhang, Guojun Peng, Tianhe Chi, Cuiling Ji
IGARSS2
2005 Research on Sea Digital Map based on WEBGIS
abstract
Sea digital map technique is one of the key developed techniques in forehand. The thesis develops WEB-ECDIS which is a Web Sea Digital Map System developed by WEBGIS technique and GeoBeans5.5, a WEBGIS tool software. WEB-ECDIS, working in the environment of Internet or intranet, is an Industrial WEBGIS system facing for the field of maritime traffic application which compatilizes, stores, processes, analyzes, displays and applies sea digital geographical map information. It can offer geographical query and obtaining the map service based on Sea Digital Map and related operational attribute data through Internet browser. At the point of view of realization technique and application, this thesis gives a detailed introduction.
Guojun Peng, Tianhe Chi, XingGu Zhang, Lu Xiang
IGARSS1
2005 Research on WEBGIS real-time distribution system of port navigation-supporting information
abstract
Port and its sostenuto stable development is the fundamental ensure of economic boom of littoral. Nowadays, the constructing scope of littoral cities is getting scale-up, and port traffic is getting crowded, so ensuring the safety of maritime traffic is one of the premises of port developing. Through deep research on navigation supporting data content and data-sharing distribution mode, this paper, using the technique of WEBGIS, constructed the instant WEBGIS port navigation supporting information distribution system, and this system is based on the map data standard of IHO-S57, and can be overlaid many kinds of real-time navigation-supporting information, and it works on the Web. This system makes it possible for the in-and-out port ships to get the latest navigation-supporting information, and ensure ship's navigation safety further. From the technique of realization, this paper gives a simple introduction.
Guojun Peng, Tianhe Chi, XingGu Zhang, Lu Xiang, Zongheng Chen
IGARSS1