EDBT 2026 Demo / reviewers in the wild / expert
Heyuan Shi
dblp:192/6867
· DBLP profile ↗
54ranked-venue papers
9as first author
44since 2021 · last 2026
0000-0002-9040-7247ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 4 first-author · 19 since 2021Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Computer networks · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Security and privacy · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective On-Hardware Fuzzing of Embedded Operating SystemsabstractFuzz testing embedded OSs is difficult because their implementations vary widely and rely on specialized hardware. These factors render many existing methods ineffective, since they prevent adapting established fuzzing routines, disrupt communication with the target OS, and impede observation of runtime behavior. This paper introduces EOF, a feedback-guided fuzzer designed to test embedded OSs running on actual hardware. Through the debug port, EOF communicates with the target embedded OS, executes test cases, and collect feedback data, with no dependence on OS services. Then, EOF deploys a cross-platform agent and executes API-aware input across diverse hardware. Last, EOF collects runtime coverage and critical execution events to identify interesting seeds and find potential bugs during fuzzing. We implemented EOF and evaluated its performance on four different embedded OSs, where EOF discovered 19 bugs and achieved a 50.84% coverage improvement on average compared with other comparable fuzzing methods. Yuheng Shen, Jianzhong Liu, Qiming Guo, Yifei Chu, Heyuan Shi, Yu Jiang 0001 |
EuroSys | 6 |
| 2026 | Scalable hierarchical protocol format inference via feature-heuristic message delimiter
Yanyang Zhao, Zhengxiong Luo 0002, Ronghua Shi, Yu Jiang 0001, Heyuan Shi |
Empir. Softw. Eng. | 8 |
| 2026 | Beyond sparse supervision: Diffusion-guided learning for few-shot graph fraud detection
Hu Chao, Mingfei Lu, Yiwei Ge, Xingle Li, Heyuan Shi |
Neurocomputing | 6 |
| 2026 | PowerEar: An Audio Eavesdropping Attack on Mobile Devices Through USB Power Side ChannelabstractWith the increasing popularity of voice-centric applications, acoustic eavesdropping attacks pose a significant threat to user privacy. Although smartphones require explicit user permission to access the microphone, such attacks can bypass this restriction by exploiting power consumption data through compromised power supplies, such as USB adapters, public charging stations, and power banks. However, previous attempts can only recognize a limited set of hotwords or digits. To address this limitation, we introduce PowerEar, an acoustic eavesdropping attack that leverages the power side channel to reconstruct any audio reproduced by the built-in loudspeaker of a mobile device with an unconstrained vocabulary. Our approach relies on a combination of signal processing and generative techniques to learn the mapping between power consumption and audio playback, enabling the reconstruction of such audio through spectrogram enhancement. To validate the effectiveness of PowerEar attack, we carry out a comprehensive set of experiments using audio samples from various public personalities. Our results obtained through objective and subjective evaluations clearly demonstrate that PowerEar can successfully recover user speeches from power consumption data in comprehensive realistic settings, including speech utterances of individuals and different devices, mobile operating systems, activities, charging technology, battery and volume levels. Riccardo Spolaor, Heyuan Shi, Zekun Miao, Yanni Yang 0003, Xiuzhen Cheng, Pengfei Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | DROIDFUZZ: Proprietary Driver Fuzzing for Embedded Android DevicesabstractEmbedded Android Devices have proliferated in many security-critical embedded scenarios, requiring sufficient testing to root out vulnerabilities. Due to Android’s architecture, which uses a Hardware ion Layer (HAL) for vendor-specific driver implementations, traditional kernel testing techniques cannot detect such bugs within the actual driver logic, which are commonly proprietary and vendorspecific. In this paper, we propose DroidFuzz, an embedded Android system fuzzer that targets such vendor-specific driver implementations to find such bugs. Through leveraging pre-testing HAL driver probing, kernel-user relational payload generation, and cross-boundary execution state feedback, we effectively test the proprietary drivers in both the kernel and the HAL layer. We implemented DroidFuzz and evaluated its effectiveness on 7 embedded Android devices, and found 12 security-critical previously unknown bugs, all of which have been confirmed by the respective vendors. Jianzhong Liu, Yuheng Shen, Yifei Chu, Heyuan Shi, Wanli Chang 0001, Yu Jiang 0001 |
DAC | 5 |
| 2025 | CMFuzz: Parallel Fuzzing of IoT Protocols by Configuration Model Identification and SchedulingabstractIoT protocols are essential for the communication among diverse devices. In real-world scenarios, IoT protocols utilize flexible configurations to meet various use cases. These configurations can significantly impact the protocols’ execution paths, with many bugs emerging only under specific configurations. Fuzzing has become a prominent technique for uncovering vulnerabilities in IoT protocol implementations. However, traditional fuzzing approaches are typically conducted using fixed or default configurations, overlooking potential issues that might arise in different settings. This limitation can lead to missing critical bugs that appear only under alternative configurations.In this paper, we propose CMFuzz, a parallel fuzzing framework designed to improve fuzzing effectiveness of IoT protocols through configuration identification and scheduling. CMFuzz first constructs a generalized protocol configuration model by systematically extracting configuration items from protocol implementations. Then, based on this model, CMFUZZ defines the relations among configuration items and introduces a relation-aware allocation mechanism to distribute them across parallel fuzzing instances. For evaluation, We implement CMFuzz on top of the widely-used protocol fuzzer Peach and conduct experiments on six popular IoT protocols. Compared to the original parallel mode of Peach and state-of-the-art parallel protocol fuzzer SPFuzz, CMFuzz covers an average of 34.4% and 28.5% more branches within 24 hours. Additionally, CMFuzz has detected 14 previously-unknown bugs in these real-world IoT protocols. Fuchen Ma, Yuanliang Chen, Feifan Wu, Yanyang Zhao, Heyuan Shi, Yu Jiang 0001 |
DAC | 7 |
| 2025 | Collaborative Countermeasure Against Network Traffic Analyses Based on Packet AggregationabstractNetwork traffic analysis attacks represent a significant threat to information security and user privacy. Leveraging machine learning, these attacks can infer sensitive information even when encryption and anonymization techniques are in place. This paper proposes TravelingTogether, a novel framework that leverages packet aggregation from multiple source hosts to defend users against network traffic analysis. TravelingTogether works at the network level, providing transparent and seamless protection to any hosts connected to the network. We evaluate our defense across different scenarios through a comprehensive set of experiments, demonstrating its effectiveness in thwarting website fingerprinting attacks as a representative use case, and its efficiency in terms of low time and bandwidth overhead. Riccardo Spolaor, Heyuan Shi, Fabio De Gaspari, Luigi V. Mancini, Dongxiao Yu, Xiuzhen Cheng |
ICC | 2 |
| 2025 | RPG: Linux Kernel Fuzzing Guided by Distribution-Specific Runtime Parameter InterfacesabstractThe Linux distribution kernel differs significantly from the mainline kernel, incorporating additional features and vendor-specific extensions. Among these additions, many runtime parameter interfaces are unique to distribution kernels, which expands the attack surface and increases the risk of potential vulnerabilities. Fuzzing has been used to assess Linux distributions, but existing tools cannot systematically test these distribution-specific interfaces due to two main challenges: (1) generating test cases for these runtime parameter interfaces, and (2) concentrating test resources on the distribution-specific interface code. To address these challenges, we propose RPG, a distribution-specific runtime parameter-guided kernel fuzzer. RPG operates in three phases: First, RPG extracts distribution-specific runtime parameter interfaces. Then, RPG uses LLM and tuning software databases to model each parameter range to generate meaningful interface test cases. Third, RPG utilizes the distribution kernel’s function control flow graph to guide the fuzzer to generate generic test cases that are more closely related to the distribution-specific interface code. We evaluated RPG on four Linux distribution kernels: Ubuntu 22.04, Fedora 42, OpenAnolis 8.8, and OpenAnolis 23.1. RPG detected 22 previously unknown bugs (13 distribution-specific), of which 15 were confirmed and 10 fixed by kernel maintainers. RPG also achieved 20.4% and 21.2% higher branch coverage than Syzkaller and Healer, respectively. Yuheng Shen, Guoyu Yin, Runzhe Wang, Tao Ma 0006, Xiaohai Shi, Heyuan Shi |
ASE | 10 |
| 2025 | Industry Practice of LLM-Assisted Protocol Fuzzing for Commercial Communication ModulesabstractFuzzing is widely used for software robustness testing. However, its application in commercial communication modules remains limited due to several key challenges, including labor-intensive template generation, lack of coverage collection support, limited testing performance, and inconsistencies between practical hardware and software CI/CD processes. In collaboration with China Mobile IoT, we present FuzzCM, a comprehensive protocol fuzzing framework tailored for commercial communication modules. FuzzCM employs a Retrieval-Augmented Generation (RAG)-enhanced large language model (LLM) to automate template generation and utilizes GPIO-based instrumentation for efficient runtime coverage data collection. Additionally, it leverages a knowledge base constructed from prior tests to guide hybrid mutation strategies and integrates CI/CD across both software and hardware layers, enabling continuous and environment-aware testing. We conducted an industrial practice with FuzzCM on five LTE Cat.1 bis modules, identifying 21 previously unknown bugs, 15 of which have been fixed. The results demonstrate that FuzzCM outperforms both manual methods and the Peach* approach, achieving average coverage improvements of 51% and 29%, respectively, with overall coverage reaching 85%. Yulai Fu, Ronghua Shi, Fuchen Ma, Heyuan Shi |
ASE | 10 |
| 2025 | LLM-assisted Industrial-Scale Differential Testing of Package Incompatibilities in Linux DistributionsabstractAn open source Linux distribution often undergoes version upgrades and migrations, which is prone to incompatibility issues especially when it comes to large-scale software changes. Although differential testing has been widely used in software testing, it is still challenging to apply it for detecting such incompatibilities in the context of industrial settings. In this paper, we report our experience in leveraging LLMs to address the challenges faced by the Linux distribution community. Specifically, we develop an LLM-based differential testing method called Versify to assist maintainers of Linux distributions in locating incompatibilities during version upgrades and migrations. Its trial operation period within the Linux distribution community shows that it uncovered 8,489 instances of differing behavior, of which 644 were prioritized for attention by developers. After deduplication and filtering, 39 unique compatibility reports were identified. Feedback from Linux distributions developers indicates that our reports have provided valuable recommendations for package selection in future OS releases. Chijin Zhou, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Tao Ma 0006, Zhe Wang 0015, Heyuan Shi |
ASE | 11 |
| 2025 | Tron: Fuzzing Linux Network Stack via Protocol-System Call Payload SynthesisabstractThe Linux kernel network stack is a critical component of modern operating systems, widely deployed across platforms and often exposed to untrusted inputs. Its complex and stateful nature makes it a frequent target of security vulnerabilities, particularly those triggered by subtle protocol interactions. While existing fuzzers like syzkaller have demonstrated strong capabilities in discovering kernel bugs, they face challenges in exercising deep protocol logic due to the lack of coordinated inputs and protocol awareness. In this paper, we present Tron, a tool designed for fuzzing the Linux kernel network stack. By synthesizing syscall–packet input sequences based on protocol structure and incorporating runtime feedback, Tron enables the exploration of protocol-dependent state transitions and deep execution paths. Our approach addresses the fundamental challenges in dual-input fuzzing by integrating protocol knowledge with execution feedback. We evaluate Tron on four recent Linux kernel versions and compare it against syzkaller and kernelGPT. The results show that Tron improves branch coverage by 22.9% and 12.1% over syzkaller and kernelGPT, respectively, and discovers 25 previously unknown bugs, 7 of which have been fixed. These results demonstrate the effectiveness of protocol–system call input synthesis in enhancing network stack fuzzing and uncovering hard-to-reach bugs in kernel protocol implementations. Yifei Chu, Yuheng Shen, Jianzhong Liu, Heyuan Shi, Yu Jiang 0001, Wanli Chang 0001 |
ASE | 5 |
| 2025 | DragonRadar: Fuzzing Linux Kernel Deployed in Cloud-Native EnvironmentabstractKata Containers is a secure container runtime with lightweight virtual machines and a customized Linux kernel optimized for cloud-native workloads, which is important for cloud-native systems. Fuzzing is a widely-used technique for detecting kernel vulnerability. However, current kernel fuzzers can't be simply applied to kernels in cloud-native environments because of the discrepancies between test and actual deployment scenarios. This paper introduces DragonRadar, a kernel fuzzing tool adapted for Kata Containers, which aligns the testing environment with cloud deployment realities. We extend to support kernel fuzzing in cloud-native environments by integrating Syzkaller's capability with a lightweight virtual machine manager called Dragonball. The evaluation shows that DragonRadar effectively identifies 25 kernel vulnerabilities in the mainline Linux kernel used in the Kata Containers environment, while maintaining code coverage similar to vanilla Syzkaller. DragonRadar is available at https://github.com/TOBESTONG//DragonRadar. Heyuan Shi, Weibo Zhang, Runzhe Wang, Xiaohai Shi, Guoyu Yin, Jianzhong Liu, Yuheng Shen |
SANER | 1 |
| 2025 | Protocol syntax recovery via knowledge transfer
Yanyang Zhao, Zhengxiong Luo 0002, Feifan Wu, Heyuan Shi, Yu Jiang 0001 |
Comput. Networks | 6 |
| 2025 | H3NI: Non-target-specific node injection attacks on hypergraph neural networks via genetic algorithm
Heyuan Shi, Binqi Zeng, Ruishi Yu, Zijian Zouxia, Ronghua Shi |
Neurocomputing | 1 |
| 2025 | CubeAgent: Efficient query-based video adversarial examples generation through deep reinforcement learning
Heyuan Shi, Binqi Zeng |
J. Syst. Softw. | 1 |
| 2025 | QuanTest: Entanglement-Guided Testing of Quantum Neural Network SystemsabstractQuantum Neural Network (QNN) combines the deep learning (DL) principle with the fundamental theory of quantum mechanics to achieve machine learning tasks with quantum acceleration. Recently, QNN systems have been found to manifest robustness issues similar to classical DL systems. There is an urgent need for ways to test their correctness and security. However, QNN systems differ significantly from traditional quantum software and classical DL systems, posing critical challenges for QNN testing. These challenges include the inapplicability of traditional quantum software testing methods to QNN systems due to differences in programming paradigms and decision logic representations, the dependence of quantum test sample generation on perturbation operators, and the absence of effective information in quantum neurons. In this article, we propose QuanTest, a quantum entanglement-guided adversarial testing framework to uncover potential erroneous behaviors in QNN systems. We design a quantum entanglement adequacy criterion to quantify the entanglement acquired by the input quantum states from the QNN system, along with two similarity metrics to measure the proximity of generated quantum adversarial examples to the original inputs. Subsequently, QuanTest formulates the problem of generating test inputs that maximize the quantum entanglement adequacy and capture incorrect behaviors of the QNN system as a joint optimization problem and solves it in a gradient-based manner to generate quantum adversarial examples. Experimental results demonstrate that QuanTest possesses the capability to capture erroneous behaviors in QNN systems (generating 67.48–96.05% more high-quality test samples than the random noise under the same perturbation size constraints). The entanglement-guided approach proves effective in adversarial testing, generating more adversarial examples (maximum increase reached 21.32%). Jinjing Shi, Zimeng Xiao, Heyuan Shi, Yu Jiang 0001, Xuelong Li 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | BiotaFormer: Detecting biota in microscopic activated sludge images for wastewater treatmentabstractRecognizing biota in microscopic activated sludge images is essential for effective wastewater treatment. General object detection methods face challenges in accurately recognizing the biota with various sizes and shapes, especially under constraints of computational efficiency in industry practice. This paper presents a framework called BiotaFormer to detect biota in microscopic activated sludge images for wastewater treatment. BiotaFormer recognizes and monitors biota in microscopic images to evaluate the condition of activated sludge, which optimizes the subsequent treatment process. To evaluate BiotaFormer, we introduce the Biota-12, the first dataset specifically for biota detection. Experimental results show that BiotaFormer has significant advantages in performance, demonstrating its overall effectiveness and efficiency in industry practice Huizhen Chen, Heyuan Shi, Ronghua Shi |
BIBM | 4 |
| 2024 | Effectively Sanitizing Embedded Operating SystemsabstractEmbedded operating systems, considering their widespread use in security-critical applications, are not effectively tested with sanitizers to effectively root out bugs. Sanitizers provide a means to detect bugs that are not visible directly through exceptional or erroneous behaviors, thus uncovering more potent bugs during testing. Jianzhong Liu, Yuheng Shen, Yiru Xu, Hao Sun 0021, Heyuan Shi, Yu Jiang 0001 |
DAC | 5 |
| 2024 | SPFuzz: Stateful Path based Parallel Fuzzing for Protocols in Autonomous VehiclesabstractProtocols in autonomous vehicles are essential for efficient in-vehicle network communication. To ensure their security, many research efforts have been paid to the fuzz testing of their implementations. However, those fuzzing optimizations often struggle to manage the protocols' complex state, resulting in low efficiency in branch covering and vulnerability detection. Junze Yu, Zhengxiong Luo 0002, Fangshangyuan Xia, Yanyang Zhao, Heyuan Shi, Yu Jiang 0001 |
DAC | 5 |
| 2024 | MDIplier: Protocol Format Recovery via Hierarchical InferenceabstractNetwork protocol reverse engineering is crucial for a wide range of security applications. Many existing techniques accomplish this task by analyzing network traces. However, these methods globally cluster messages and analyze each cluster separately, which causes the loss of valuable field information. To address this problem, we present MDIplier, a protocol reverse engineering tool that leverages the hierarchical structure of protocol messages and performs tailored analysis at each message layer. MDIplier performs an iterative inference process. During each iteration, it identifies the message delimiter for layer separation and infers the format for each layer separately, optimizing the use of available field information. Our evaluation of eight widely used protocols shows that MDIplier outperforms state-of-the-art methods. It identifies fields with a perfection score 4.6×, 1.4×, 5.8×, and 1.8× higher than that of Netzob, Netplier, FieldHunter, and BinaryInferno, respectively. Furthermore, the experiments on proprietary protocols used in three IoT devices demonstrate the effectiveness of MDIplier in real-world scenarios. Zhengxiong Luo 0002, Yanyang Zhao, Ronghua Shi, Yu Jiang 0001, Heyuan Shi |
ISSRE | 7 |
| 2024 | Towards More Complete Constraints for Deep Learning Library Testing via Complementary Set Guided RefinementabstractDeep learning library is important in AI systems. Recently, many works have been proposed to ensure its reliability. They often model inputs of tensor operations as constraints to guide the generation of test cases. However, these constraints may narrow the search space, resulting in incomplete testing. This paper introduces a complementary set-guided refinement that can enhance the completeness of constraints. The basic idea is to see if the complementary set of constraints yields valid test cases. If so, the original constraint is incomplete and needs refinement. Based on this idea, we design an automatic constraint refinement tool, DeepConstr, which adopts a genetic algorithm to refine constraints for better completeness. We evaluated it on two DL libraries, PyTorch and TensorFlow. DeepConstr discovered 84 unknown bugs, out of which 72 were confirmed, with 51 fixed. Compared to state-of-the-art fuzzers, DeepConstr increased coverage for 43.44% of operators supported by NNSmith, and 59.16% of operators supported by NeuRI. Gwihwan Go, Chijin Zhou, Quan Zhang 0003, Xiazijian Zou, Heyuan Shi, Yu Jiang 0001 |
ISSTA | 5 |
| 2024 | Enhancing ROS System Fuzzing through Callback TracingabstractThe Robot Operating System 2 (ROS) is the de-facto standard for robotic software development, with a wide application in diverse safety-critical domains. There are many efforts in testing that seek to deliver a more secure ROS codebase. However, existing testing methods are often inadequate to capture the complex and stateful behaviors inherent to ROS deployments, resulting in limited test- ing effectiveness. In this paper, we propose R2D2, a ROS system fuzzer that leverages ROS’s runtime states as guidance to increase fuzzing effectiveness and efficiency. Unlike traditional fuzzers, R2D2 employs a systematic instrumentation strategy that captures the system’s runtime behaviors and profiles the current system state in real-time. This approach provides a more in-depth understanding of system behaviors, thereby facilitating a more insightful explo- ration of ROS’s extensive state space. For evaluation, we applied it to four well-known ROS applications. Our evaluation shows that R2D2 achieves an improvement of 3.91× and 2.56× in code coverage compared to state-of-the-art ROS fuzzers, including Ros2Fuzz and RoboFuzz, while also uncovering 39 previously unknown vulnera- bilities, with 6 fixed in both ROS runtime and ROS applications. For its runtime overhead, R2D2 maintains an average execution and memory usage overhead with 10.4% and 1.0% in respect, making R2D2 effective in ROS testing. Yuheng Shen, Jianzhong Liu, Yiru Xu, Hao Sun 0021, Nan Guan, Heyuan Shi, Yu Jiang 0001 |
ISSTA | 7 |
| 2024 | Logos: Log Guided Fuzzing for Protocol ImplementationsabstractNetwork protocols are extensively used in a variety of network devices, making the security of their implementations crucial. Protocol fuzzing has shown promise in uncovering vulnerabilities in these implementations. However traditional methods often require instrumentation of the target implementation to provide guidance, which is intrusive, adds overhead, and can hinder black-box testing. This paper presents Logos, a protocol fuzzer that utilizes non-intrusive runtime log information for fuzzing guidance. Logos first standardizes the unstructured logs and embeds them into a high-dimensional vector space for semantic representation.Then, Logos filters the semantic representation and dynamically maintains a semantic coverage to chart the explored space for customized guidance.We evaluate Logos on eight widely used implementations of well-known protocols. Results show that, compared to existing intrusive or expert knowledge-driven protocol fuzzers, Logos achieves 26.75%-106.19% higher branch coverage within 24 hours. Furthermore, Logos exposed 12 security-critical vulnerabilities in these prominent protocol implementations, with 9 CVEs assigned. Feifan Wu, Zhengxiong Luo 0002, Yanyang Zhao, Qingpeng Du, Junze Yu, Ruikang Peng, Heyuan Shi, Yu Jiang 0001 |
ISSTA | 7 |
| 2024 | Industry Practice of Directed Kernel Fuzzing for Open-source Linux DistributionabstractDirected grey-box fuzzing is a widely used automatic testing technique that has helped developers test specific code space in the target program. Although many directed fuzzers are designed to test the Linux kernel, challenges still remain due to the complexity of industrial requirements and deployment environments. In this paper, we collaborate with developers from Alibaba and the OpenAnolis community to conduct an industry practice of directed kernel fuzzing for open-source Linux distribution. We highlight typical challenges in deploying directed kernel fuzzing, including target-related kernel configuration options being disabled, unrelated initial seeds limiting fuzzing startup performance, no support for kernel feature interface fuzzing, independent fuzzer execution limiting fuzzing effectiveness, much manual work to triage and analyze crashes, and hard to integrate into the existing fuzzing framework. We provide solutions to these challenges, which allowed us to discover 11 previously unknown kernel bugs related to cloud-native features, io_uring, and other components in the OpenAnolis Linux distribution. Heyuan Shi, Runzhe Wang, Weibo Zhang, Yuheng Shen, Xiaohai Shi, Yu Jiang 0001 |
ASE | 1 |
| 2024 | Imperceptible Content Poisoning in LLM-Powered ApplicationsabstractLarge Language Models (LLMs) have shown their superior capability in natural language processing, promoting extensive LLM-powered applications to be the new portals for people to access various content on the Internet. However, LLM-powered applications do not have sufficient security considerations on untrusted content, leading to potential threats. In this paper, we reveal content poisoning, where attackers can tailor attack content that appears benign to humans but causes LLM-powered applications to generate malicious responses. To highlight the impact of content poisoning and inspire the development of effective defenses, we systematically analyze the attack, focusing on the attack modes in various content, exploitable design features of LLM application frameworks, and the generation of attack content. We carry out a comprehensive evaluation on five LLMs, where content poisoning achieves an average attack success rate of 89.60%. Additionally, we assess content poisoning on four popular LLM-powered applications, achieving the attack on 72.00% of the content. Our experimental results also show that existing defenses are ineffective against content poisoning. Finally, we discuss potential mitigations for LLM application frameworks to counter content poisoning. Quan Zhang 0003, Chijin Zhou, Gwihwan Go, Binqi Zeng, Heyuan Shi, Zichen Xu 0001, Yu Jiang 0001 |
ASE | 5 |
| 2024 | DynPRE: Protocol Reverse Engineering via Dynamic Inference
Zhengxiong Luo 0002, Yanyang Zhao, Feifan Wu, Junze Yu, Heyuan Shi, Yu Jiang 0001 |
NDSS | 6 |
| 2024 | PatchBert: Continuous Stable Patch Identification for Linux Kernel via Pre-trained Model Fine-tuningabstractStable patch identification is crucial in merging patches into stable versions, which helps ensure the stability of the Linux kernel. Although many tools have been proposed to mitigate the manual effort of stable patch identification, challenges still arise because they neglect continuous stable patch tracking and advanced Natural Language Processing (NLP) pre-training techniques. In this paper, in collaboration with developers from the openAnolis Linux operating system distribution community, we present a stable patch identification model called PatchBERT. It utilizes BERT and CodeBERT to capture the semantic patch representation from the commit message and code changes in a patch. We then perform patch classification and output the probability that the patch should be merged into the stable versions. We perform experiments on the dataset used by the previous methods. The experimental results show the superior performance of PatchBERT over state-of-the-art baselines. Additionally, it is common practice to train the model using the latest Linux patches and implement it in a real-world industrial setting. In this exercise, we randomly select 10,000 patches for identification, accurately identifying 8,617 patches and incorrectly identifying 1,383 patches. This practical outcome further confirms the effectiveness and utility of PatchBERT in real-world scenarios. Heyuan Shi, Runzhe Wang, Yuheng Shen, Yuao Chen, Xiaohai Shi, Yu Jiang 0001 |
SANER | 2 |
| 2024 | EVMFuzz: Differential fuzz testing of Ethereum virtual machineabstractAbstract The vulnerabilities in Ethereum virtual machine (EVM) may lead to serious problems for the Ethereum ecosystem. With lots of techniques being developed for the validation of smart contracts, the testing of EVM has not been well‐studied. In this paper, we propose EVMFuzz, the first that uses the differential fuzzing technique to detect vulnerabilities in EVM. The core idea of EVMFuzz is to continuously generate seed contracts for different EVMs' execution, so as to find as many inconsistencies among execution results as possible, and eventually discover vulnerabilities with output cross‐referencing. First, we present the evaluation metric for the internal inconsistency indicator. Then, we construct seed contracts via predefined mutators and employ a dynamic priority scheduling algorithm to guide seed contract selection and maximize the inconsistency. Finally, we leverage different EVMs as cross‐referencing oracles avoiding manual checking. For evaluation, we selected four widely used EVMs for the test, conducted large‐scale mutation on 36,295 real‐world smart contracts, and generated 253,153 smart contracts as initial seeds. Accompanied by manual root cause analysis, we found five previously unknown security bugs and all had been included in the common vulnerabilities and exposures (CVE) database. Fuchen Ma, Heyuan Shi, Shanshan Li 0001, Xiangke Liao |
J. Softw. Evol. Process. | 5 |
| 2024 | Parallel Fuzzing of IoT Messaging Protocols Through Collaborative Packet GenerationabstractInternet of Things (IoT) messaging protocols play an important role in facilitating communications between users and IoT devices. Mainstream IoT platforms employ brokers, server-side implementations of IoT messaging protocols, to enable and mediate this user-device communication. Due to the complex nature of managing communications among devices with diverse roles and functionalities, comprehensive testing of the protocol brokers necessitates collaborative parallel fuzzing. However, being unaware of the relationship between test packets generated by different parties, existing parallel fuzzing methods fail to explore the brokers’ diverse processing logic effectively. This article introduces MPFuzz, a parallel fuzzing tool designed to secure IoT messaging protocols through collaborative packet generation. The approach leverages the critical role of certain fields within IoT messaging protocols that specify the logic for message forwarding and processing by protocol brokers. MPFuzzemploys an information synchronization mechanism to synchronize these key fields across different fuzzing instances and introduces a semantic-aware refinement module that optimizes generated test packets by utilizing the shared information and field semantics. This strategy facilitates a collaborative refinement of test packets across otherwise isolated fuzzing instances, thereby boosting the efficiency of parallel fuzzing. We evaluated MPFuzzon six widely used IoT messaging protocol implementations. Compared to two state-of-the-art protocol fuzzers with parallel capabilities, Peach and AFLNet, as well as two representative parallel fuzzers, SPFuzz and AFLTeam, MPFuzzachieves (6.1%,$174.5\times $), (20.2%,$607.2\times $), (1.9%,$4.1\times $), and (17.4%,$570.2\times $) higher branch coverage and fuzzing speed under the same computing resource. Furthermore, MPFuzzexposed seven previously unknown vulnerabilities in these extensively tested projects, all of which have been assigned with CVE identifiers. Zhengxiong Luo 0002, Junze Yu, Qingpeng Du, Yanyang Zhao, Feifan Wu, Heyuan Shi, Wanli Chang 0001, Yu Jiang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | ECG: Augmenting Embedded Operating System Fuzzing via LLM-Based Corpus GenerationabstractEmbedded operating systems (Embedded OSs) power much of our critical infrastructure but are, in general, much less tested for bugs than general-purpose operating systems. Fuzzing Embedded OSs encounter significant roadblocks due to much less documented specifications, an inherent ineffectiveness in generating high-quality payloads. In this article, we propose ECG, an Embedded OS fuzzer empowered by large language models (LLMs) to sufficiently mitigate the aforementioned issues. ECG approaches fuzzing Embedded OS by automatically generating input specifications based on readily available source code and documentation, instrumenting and intercepting execution behavior for directional guidance information, and generating inputs with payloads according to the pregenerated input specifications and directional hints provided from previous runs. These methods are empowered by using an interactive refinement method to extract the most from LLMs while using established parsing checkers to validate the outputs. Our evaluation results demonstrate that ECG uncovered 32 new vulnerabilities across three popular open-source Embedded OS (RT-Linux, RaspiOS, and OpenWrt) and detected ten bugs in a commercial Embedded OS running on an actual device. Moreover, compared to Syzkaller, Moonshine, KernelGPT, Rtkaller, and DRLF, ECG has achieved additional kernel code coverage improvements of 23.20%, 19.46%, 10.96%, 15.47%, and 11.05%, respectively, with an overall average improvement of 16.02%. These results underscore ECG’s enhanced capability in uncovering vulnerabilities, thus contributing to the overall robustness and security of the Embedded OS. Yuheng Shen, Jianzhong Liu, Yiru Xu, Heyuan Shi, Yu Jiang 0001, Wanli Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | A Robust and Lightweight Privacy-Preserving Data Aggregation Scheme for Smart GridabstractPrivacy-preserving data aggregation (PPDA) enables data availability and privacy preservation simultaneously in smart grid. However, existing methods, such as masking and homomorphic encryption, cannot simultaneously offer strong privacy preservation, fault tolerance for both smart meters and aggregators, verifiable aggregation, and lightweight encryption. To tackle these challenges, we design HTV-PRE, a homomorphic threshold proxy re-encryption scheme with re-encryption verifiability. HTV-PRE involves only linear operations and resists quantum attacks after being instanced by ideal lattices. By leveraging HTV-PRE, we propose a robust and lightweight data aggregation scheme with strong privacy preservation for smart grid. Robustness ensures fault tolerance and error detection. Even if some smart meters or aggregators are faulty, data aggregation can still work without imposing expensive computation on other smart meters or requiring additional trust assumptions. Additionally, to detect aggregators' errors, a proof for the aggregated result is presented so that anyone can verify whether the result has been correctly computed or not. The verifiable aggregation adds no computation/communication overhead on the user side. The performance evaluations demonstrate that our PPDA scheme significantly offloads computation overhead from smart meters and control center to the edge, and its user encryption is up to 4x faster than existing approaches. Liqiang Wu, Shaojing Fu, Yuchuan Luo, Hongyang Yan, Heyuan Shi, Ming Xu 0002 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | Exploiting Conversation-Branch-Tweet HyperGraph Structure to Detect Misinformation on Social MediaabstractThe spread of misinformation on social media is a serious issue that can have negative consequences for public health and political stability. While detecting and identifying misinformation can be challenging, many attempts have been made to address this problem. However, traditional models that focus on pairwise relationships on misinformation propagation paths may not be effective in capturing the underlying connections among multiple tweets. To address this limitation, the proposed “Conversation-Branch-Tweet” hypergraph convolutional network (CBT-HGCN) uses a hypergraph to represent the internal structure and content of tweet data, with tweets and their replies viewed as nodes and hyperedges, respectively. The model first pre-processes the tweets of a conversation and then uses a pre-trained model as an encoder to extract node information. Finally, a hypergraph convolution network is used as an information fuser for classification. Experimental results on three benchmark datasets (Twitter15, Twitter16, and Pheme) show that the proposed model outperforms several strong baseline models and achieves state-of-the-art performance. This indicates that the CBT-HGCN approach is effective in detecting and identifying misinformation on social media by capturing the underlying connections among multiple tweets. Fangfang Li 0004, Junwen Duan, Xingliang Mao, Heyuan Shi, Shichao Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Double-Layer Search and Adaptive Pooling Fusion for Reference-Based Image Super-ResolutionabstractReference-based image super-resolution (RefSR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) images by introducing HR reference images. The key step of RefSR is to transfer reference features to LR features. However, existing methods still lack an efficient transfer mechanism, resulting in blurry details in the generated image. In this article, we propose a double-layer search module and an adaptive pooling fusion module group for reference-based image super-resolution, called DLASR. Based on the re-search strategy, the double-layer search module can produce an accurate index map and score map. These two maps are used to filter out accurate reference features, which greatly increases the efficiency of feature transfer in the later stage. Through two continuous feature-enhancement steps, the adaptive pooling fusion module group can transfer more valuable reference features to the corresponding LR features. In addition, a structure reconstruction module is proposed to recover the geometric information of the images, which further improves the visual quality of the generated image. We conduct comparative experiments on a variety of datasets, and the results prove that DLASR achieves significant improvements over other state-of-the-art methods, in terms of quantitative accuracy and qualitative visual effect. The code is available at https://github.com/clttyou/DLASR. Kehua Guo, Xiangyuan Zhu, Xiaoyan Kui, Jian Zhang 0048, Heyuan Shi |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Cube-Evo: A Query-Efficient Black-Box Attack on Video Classification SystemabstractThe current progressive research in the domain of black-box adversarial attack enhances the reliability of deep neural network (DNN)-based video systems. Recent works mainly carry out black-box adversarial attacks on video systems by query-based parameter dimension reduction. However, the additional temporal dimension of video data leads to massive query consumption and low attack success rate. In this article, we embark on our efforts to design an effective adversarial attack on popular video classification systems. We deeply root the observations that the DNN-based systems are sensitive to adversarial perturbations with high frequency and reconstructed shape. Specifically, we propose a systematic attack pipeline Cube-Evo, aiming to reduce the search space dimension and obtain the effective adversarial perturbation via the optimal parameter group updating. We evaluate the proposed attack pipeline on two popular datasets: UCF101 and JESTER. Our attack pipeline reduces query consumption and achieves a high success rate on various DNN-based video classification systems. Compared with the state-of-the-art method Geo-Trap-Att, our pipeline averagely reduces 1.6× query consumption in untargeted attacks and 2.9× in targeted attacks. Besides, Cube-Evo improves 13% attack success rate on average, achieving new state-of-the-art results over diverse video classification systems. Jianmin Guo, Heyuan Shi, Houbing Song |
IEEE Trans. Reliab. | 5 |
| 2023 | KeenTune: Automated Tuning Tool for Cloud Application Performance Testing and OptimizationabstractThe performance testing and optimization of cloud applications is challenging, because manual tuning of cloud computing stacks is tedious and automated tuning tools are rare used for cloud services. To address this issue, we introduce KeenTune, an automated tuning tool designed to optimize application performance and facilitate performance testing. KeenTune is a lightweight and flexible tool that can be deployed with to-be-tuned applications with negligible impact on their performance. Specifically, KeenTune uses a surrogate model that can be implemented with machine learning models to filter out less relevant parameters for efficient tuning. Our empirical evaluation shows that KeenTune significantly enhances the throughput performance of Nginx web servers, resulting in performance improvements of up to 90.43% and 117.23% in certain cases. This study highlights the benefits of using KeenTune for achieving efficient and effective performance testing of cloud applications. The video and source code for KeenTune are provided as supplementary materials. Qinglong Wang 0003, Runzhe Wang, Xiaohai Shi, Zheng Liu 0022, Tao Ma 0006, Houbing Song, Heyuan Shi |
ISSTA | 8 |
| 2023 | Brief Industry Paper: Directed Kernel Fuzz Testing on Real-time LinuxabstractRt-Linux contains critical modifications that are much less tested than the vanilla kernel, thus placing many systems at risk. In this paper, we present DRLF, a directed fuzzer targeted towards fuzzing any code area in Rt- Linux, thus allowing for more efficient tests on Rt-Linux's unique code sections. DRLF performs directed fuzzing through a kernel-level weighted callgraph construction technique, and prioritizing input sequences that exhibit less distance to the target code. Evaluations show that DRLF delivers better cover speed while achieving a 24.70% coverage increase for the targeting code areas. DRLF also found 11 previously unknown bugs within Rt-Linux, and has been integrated into Alibaba's CI/CD pipeline. Yuheng Shen, Jianzhong Liu, Yiru Xu, Runzhe Wang, Heyuan Shi, Yu Jiang 0001 |
RTSS | 7 |
| 2023 | V-Gas: Generating High Gas Consumption Inputs to Avoid Out-of-Gas VulnerabilityabstractOut-of-gas errors occur when smart contract programs are provided with inputs that cause excessive gas consumption and which will be easily exploited to perform Denial-of-Service attacks. Various approaches have been proposed to estimate the gas limit of a function in smart contracts to avoid such error. However, underestimation often occurs when the contract is complex In this work, we propose V-Gas, which automatically generates inputs that maximize the gas cost and reduce underestimation. V-Gas is designed based on static analysis and feedback-directed mutational fuzz testing. First, V-Gas builds the gas weighted control flow graph of functions in smart contracts. Then, V-Gas develops gas consumption guided selection and mutation strategies to generate the input that maximize the gas consumption. For evaluation, we implement V-Gas based on js-evm, a widely used Ethereum virtual machine written in Javascript, and conduct experiments on 736 real-world transactions recorded on Ethereum. A total of 44.02% of the transactions would have out-of-gas errors based on the estimation results given by solc, meaning that the recorded real gas consumption for those transactions is larger than the gas limit estimated by solc. In comparison, V-Gas could reduce the underestimation ratio to 13.86%. To evaluate the performance of feedback-directed engine in V-Gas, we implemented other directed fuzzing engines and compared their performance with that of V-Gas. The results showed that V-Gas generates the same or higher gas estimation value on 97.8% of the transactions with less time, usually within 5 minutes. Furthermore, V-Gas has exposed 25 previously unknown out-of-gas vulnerabilities in widely used smart contracts, 6 of which have been assigned unique CVE identifiers in the U.S. National Vulnerability Database. Fuchen Ma, Houbing Song, Heyuan Shi, Yu Jiang 0001, Huizhong Li |
ACM Trans. Internet Techn. | 6 |
| 2022 | Industry practice of configuration auto-tuning for cloud applications and servicesabstractAuto-tuning attracts increasing attention in industry practice to optimize the performance of a system with many configurable parameters. It is particularly useful for cloud applications and services since they have complex system hierarchies and intricate knob correlations. However, existing tools and algorithms rarely consider practical problems such as workload pressure control, the support for distributed deployment, and expensive time costs, etc., which are utterly important for enterprise cloud applications and services. In this work, we significantly extend an open source tuning tool – KeenTune to optimize several typical enterprise cloud applications and services. Our practice is in collaboration with enterprise users and tuning tool developers to address the aforementioned problems. Specifically, we highlight five key challenges from our experiences and provide a set of solutions accordingly. Through applying the improved tuning tool to different application scenarios, we achieve 2%-14% improvements for the performance of MySQL, OceanBase, nginx, ingress-nginx, and 5%-70% improvements for the performance of ACK cloud container service. Runzhe Wang, Qinglong Wang 0003, Heyuan Shi, Yuheng Shen, Zheng Liu 0022, Xiaohai Shi, Yu Jiang 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Abaci-finder: Linux kernel crash classification through stack trace similarity learning
Heyuan Shi, Guyu Wang, Houbing Song, Jian Dong 0001 |
J. Parallel Distributed Comput. | 1 |
| 2022 | Tardis: Coverage-Guided Embedded Operating System FuzzingabstractEmbedded operating systems (Embedded OSs) are extensively deployed in many mission-critical industrial scenarios. Any defects within these systems may result in unacceptable losses. Therefore, it is imperative to develop tools to detect bugs within Embedded OSs, thus minimizing potential impacts on industrial infrastructures. Coverage-guided fuzzing is a vulnerability detection technique that has found numerous real-world vulnerabilities within both application programs as well as kernels. However, state-of-the-art kernel fuzzers, e.g., Syzkaller, mainly target general purpose-operating systems, such as Linux, macOS, and Windows, whereas Embedded OSs support is mostly lacking. In this article, we propose Tardis, the first Embedded OSs fuzzer capable of testing a wide selection of Embedded OSs while leveraging coverage feedback. Tardis conducts OS-agnostic code coverage collection and analysis, allowing developers and testers to test a wide range of Embedded OSs without significant manual efforts. We implemented and evaluated Tardis on several well-known Embedded OSs, such as UC/OS and FreeRTOS. Tardis can successfully perform fuzz testing on these kernels without significant manual effort for adaptation. By leveraging coverage feedback, Tardis can cover 51.32% more branches than black-box fuzzing on average on the respective Embedded OSs over 24 h. Tardis also found 17 previously unknown bugs among the target Embedded OSs. Yuheng Shen, Yiru Xu, Hao Sun 0021, Jianzhong Liu, Zichen Xu 0001, Aiguo Cui, Heyuan Shi, Yu Jiang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | RNN-Test: Towards Adversarial Testing for Recurrent Neural Network SystemsabstractWhile massive efforts have been investigated in adversarial testing of convolutional neural networks (CNN), testing for recurrent neural networks (RNN) is still limited and leaves threats for vast sequential application domains. In this paper, we propose an adversarial testing framework RNN-Test for RNN systems, focusing on sequence-to-sequence (seq2seq) tasks of widespread deployments, not only classification domains. First, we design a novel search methodology customized for RNN models by maximizing the inconsistency of RNN states against their inner dependencies to produce adversarial inputs. Next, we introduce two state-based coverage metrics according to the distinctive structure of RNNs to exercise more system behaviors. Finally, RNN-Test solves the joint optimization problem to maximize state inconsistency and state coverage, and crafts adversarial inputs for various tasks of different kinds of inputs. For evaluations, we apply RNN-Test on four RNN models of common structures. On the tested models, the RNN-Test approach is demonstrated to be competitive in generating adversarial inputs, outperforming FGSM-based and DLFuzz-based methods to reduce the model performance more sharply with 2.78% to 37.94% higher success (or generation) rate. RNN-Test could also achieve 52.65% to 66.45% higher adversary rate than testRNN on MNIST LSTM model, as well as 53.76% to 58.02% more perplexity with 16% higher generation rate than DeepStellar on PTB language model.Compared with the traditional neuron coverage, the proposed state coverage metrics as guidance excel with 4.17% to 97.22% higher success (or generation) rate. Jianmin Guo, Quan Zhang 0003, Yue Zhao 0040, Heyuan Shi, Yu Jiang 0001, Jia-Guang Sun 0001 |
IEEE Trans. Software Eng. | 4 |
| 2021 | Graph Attention Mechanism-based Deep Tensor Factorization for Predicting disease-associated miRNA-miRNA pairsabstractMicroRNAs (miRNAs) play a significant role in regulating gene transcription and tend to act in a combinatorial way, which provides great insights to explore disease-related miRNA pairs or modules for comprehending the synergistic roles of miRNAs in complex diseases. As wet experiments are often laborious and costly, computational methods offer great convenience for predicting potential associations between miRNAs and diseases. Existing methods focus on either the ‘one miRNA-one disease’ paradigm, or merely the synergetic miRNA network about specific diseases, which may lead to the incomplete understanding of the synergistic effect of miRNAs on the pathogenesis of complex diseases. In this work, we present a novel tensor-based framework, named GraphTF1, to predict disease-associated miRNA-miRNA pairs. GraphTF exploits graph attention network to effectively capture node features over multi-source biological network. Then, the learned miRNA and disease representations are used to reconstruct the association tensor for predicting potential disease-associated miRNA-miRNA pairs. Empirical results showed that the proposed method outperformed all other state-of-the-art methods under five-fold cross-validation. Robustness experiments also indicated the stability of GraphTF. Moreover, case studies for Breast Neoplasms and Lung Neoplasms further demonstrated the effectiveness of GraphTF in identifying potential disease-related miRNA-miRNA pairs. Jiawei Luo 0001, Zihan Lai, Cong Shen 0002, Heyuan Shi |
BIBM | 5 |
| 2021 | Detecting Malicious Gradients from Asynchronous SGD on Variational AutoencoderabstractIn asynchronous distributed learning, the parameter server updates the global model as soon as a new gradient is received from any device. The asynchronous systems are designed to address the existence of lagging devices which is inevitable due to device heterogeneity and network unreliability. However, the lack of synchrony incurs additional noise and makes detecting and defending against malicious model gradients a challenging task. Unlike existing works that struggle to design robust methods to tolerate untargeted model poisoning gradients, the paper considers detecting and removing targeted model poisoning gradients from the normal asynchronous training process. This paper proposes Asynvae, a robust distributed asynchronous learning framework where the parameter server uses variational autoencoder to detect and exclude malicious gradients. Since the reconstruction error of malicious updates is much larger than that of benign ones, it can be used as an anomaly score. We formulate a threshold of reconstruction error to differentiate malicious updates from normal ones based on this idea. Asynvae is tested with extensive experiments on distributed learning benchmarks, showing a competitive performance over existing distributed learning methods under untargeted model poisoning attack, targeted model poisoning attack and lagging attack. Zhipin Gu, Yuexiang Yang, Heyuan Shi |
SRDS | 3 |
| 2021 | Rtkaller: State-aware Task Generation for RTOS FuzzingabstractA real-time operating system (RTOS) is an operating system designed to meet certain real-time requirements. It is widely used in embedded applications, and its correctness is safety-critical. However, the validation of RTOS is challenging due to its complex real-time features and large code base. In this paper, we propose Rtkaller , a state-aware kernel fuzzer for the vulnerability detection in RTOS. First, Rtkaller implements an automatic task initialization to transform the syscall sequences into initial tasks with more real-time information. Then, a coverage-guided task mutation is designed to generate those tasks that explore more in-depth real-time related code for parallel execution. Moreover, Rtkaller realizes a task modification to correct those tasks that may hang during fuzzing. We evaluated it on recent versions of rt-Linux, which is one of the most widely used RTOS. Compared to the state-of-the-art kernel fuzzers Syzkaller and Moonshine, Rtkaller achieves the same code coverage at the speed of 1.7X and 1.6X, gains an increase of 26.1% and 22.0% branch coverage within 24 hours respectively. More importantly, Rtkaller has confirmed 28 previously unknown vulnerabilities that are missed by other fuzzers. Yuheng Shen, Hao Sun 0021, Yu Jiang 0001, Heyuan Shi, Yixiao Yang, Wanli Chang 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2019 | Research on treatment and medication rule of insomnia treated by TCM based on data miningabstractObjectives: Based on data mining technology, this research explores the rules of traditional Chinese medicine treatment of insomnia and obtains some significant medical implication through data mining. Technology or Method: Nearly 6,000 electronic medical records of Chinese medicine clinics are selected, and data pre-processing is performed. Then, Aripori's data mining algorithm is used to analyze prescription medication rule, and the association between symptoms and drugs. Finally, a network diagram is built through complex network analysis. Results: There are 631 insomnia cases that met the criteria, 22 typical symptoms, and 330 drugs. There are 6 symptom association items with a confidence level of more than 30%, 46 drug association items with a confidence level of more than 75%, and 40 symptom association rules with a confidence level of more than 40%. Conclusions: The results showed the rule of medication used in the diagnosis and treatment of insomnia in Chinese medicine outpatients, and found the association between prescription medication rules and symptomatic drugs. Clinical or Biological Impact: This paper provides a reference to the clinical diagnosis and treatment of young physicians, and provides a reference to the development of new drugs. And to promote the development of the theory and practice of traditional Chinese medicine prescriptions is conducive to inheriting and carrying forward the experience and advantages of Chinese medicine treatment of insomnia. Yi Ye, Li Ma 0013, Jiayan Zhu, Heyuan Shi, Xiaohong Cai |
BIBM | 5 |
| 2019 | EVMFuzzer: detect EVM vulnerabilities via fuzz testingabstractEthereum Virtual Machine (EVM) is the run-time environment for smart contracts and its vulnerabilities may lead to serious problems to the Ethereum ecology. With lots of techniques being continuously developed for the validation of smart contracts, the testing of EVM remains challenging because of the special test input format and the absence of oracles. In this paper, we propose EVMFuzzer, the first tool that uses differential fuzzing technique to detect vulnerabilities of EVM. The core idea is to continuously generate seed contracts and feed them to the target EVM and the benchmark EVMs, so as to find as many inconsistencies among execution results as possible, eventually discover vulnerabilities with output cross-referencing. Given a target EVM and its APIs, EVMFuzzer generates seed contracts via a set of predefined mutators, and then employs dynamic priority scheduling algorithm to guide seed contracts selection and maximize the inconsistency. Finally, EVMFuzzer leverages benchmark EVMs as cross-referencing oracles to avoid manual checking. With EVMFuzzer, we have found several previously unknown security bugs in four widely used EVMs, and 5 of which had been included in Common Vulnerabilities and Exposures (CVE) IDs in U.S. National Vulnerability Database. The video is presented at https://youtu.be/9Lejgf2GSOk. Fuchen Ma, Heyuan Shi, Yu Jiang 0001, Huizhong Li |
ESEC/SIGSOFT FSE | 4 |
| 2019 | Industry practice of coverage-guided enterprise Linux kernel fuzzingabstractCoverage-guided kernel fuzzing is a widely-used technique that has helped kernel developers and testers discover numerous vulnerabilities. However, due to the high complexity of application and hardware environment, there is little study on deploying fuzzing to the enterprise-level Linux kernel. In this paper, collaborating with the enterprise developers, we present the industry practice to deploy kernel fuzzing on four different enterprise Linux distributions that are responsible for internal business and external services of the company. We have addressed the following outstanding challenges when deploying a popular kernel fuzzer, syzkaller, to these enterprise Linux distributions: coverage support absence, kernel configuration inconsistency, bugs in shallow paths, and continuous fuzzing complexity. This leads to a vulnerability detection of 41 reproducible bugs which are previous unknown in these enterprise Linux kernel and 6 bugs with CVE IDs in U.S. National Vulnerability Database, including flaws that cause general protection fault, deadlock, and use-after-free. Heyuan Shi, Runzhe Wang, Xiaohai Shi, Xun Jiao 0002, Houbing Song, Yu Jiang 0001, Jia-Guang Sun 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2019 | Vulnerable Code Clone Detection for Operating System Through Correlation-Induced LearningabstractVulnerable code clones in the operating system (OS) threaten the safety of smart industrial environment, and most vulnerable OS code clone detection approaches neglect correlations between functions that limits the detection effectiveness. In this article, we propose a two-phase framework to find vulnerable OS code clones by learning on correlations between functions. On the training phase, functions as the training set are extracted from the latest code repository and function features are derived by their AST structure. Then, external and internal correlations are explored by graph modeling of functions. Finally, the graph convolutional network for code clone detection (GCN-CC) is trained using function features and correlations. On the detection phase, functions in the to-be-detected OS code repository are extracted and the vulnerable OS code clones are detected by the trained GCN-CC. We conduct experiments on five real OS code repositories, and experimental results show that our framework outperforms the state-of-the-art approaches. Heyuan Shi, Runzhe Wang, Yu Jiang 0001, Jian Dong 0001, Jia-Guang Sun 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Hypergraph-Induced Convolutional Networks for Visual ClassificationabstractAt present, convolutional neural networks (CNNs) have become popular in visual classification tasks because of their superior performance. However, CNN-based methods do not consider the correlation of visual data to be classified. Recently, graph convolutional networks (GCNs) have mitigated this problem by modeling the pairwise relationship in visual data. Real-world tasks of visual classification typically must address numerous complex relationships in the data, which are not fit for the modeling of the graph structure using GCNs. Therefore, it is vital to explore the underlying correlation of visual data. Regarding this issue, we propose a framework called the hypergraph-induced convolutional network to explore the high-order correlation in visual data during deep neural networks. First, a hypergraph structure is constructed to formulate the relationship in visual data. Then, the high-order correlation is optimized by a learning process based on the constructed hypergraph. The classification tasks are performed by considering the high-order correlation in the data. Thus, the convolution of the hypergraph-induced convolutional network is based on the corresponding high-order relationship, and the optimization on the network uses each data and considers the high-order correlation of the data. To evaluate the proposed hypergraph-induced convolutional network framework, we have conducted experiments on three visual data sets: the National Taiwan University 3-D model data set, Princeton Shape Benchmark, and multiview RGB-depth object data set. The experimental results and comparison in all data sets demonstrate the effectiveness of our proposed hypergraph-induced convolutional network compared with the state-of-the-art methods. Heyuan Shi, Yubo Zhang 0006, Zizhao Zhang 0003, Nan Ma 0012, Xibin Zhao, Yue Gao 0002, Jia-Guang Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Hypergraph Learning With Cost Interval OptimizationabstractIn many classification tasks, the misclassification costs of different categories usually vary significantly. Under such circumstances, it is essential to identify the importance of different categories and thus assign different misclassification losses in many applications, such as medical diagnosis, saliency detection and software defect prediction. However, we note that it is infeasible to determine the accurate cost value without great domain knowledge. In most common cases, we may just have the information that which category is more important than the other categories, i.e., the identification of defect-prone softwares is more important than that of defect-free. To tackle these issues, in this paper, we propose a hypergraph learning method with cost interval optimization, which is able to handle cost interval when data is formulated using the high-order relationships. In this way, data correlations are modeled by a hypergraph structure, which has the merit to exploit the underlying relationships behind the data. With a cost-sensitive hypergraph structure, in order to improve the performance of the classifier without precise cost value, we further introduce cost interval optimization to hypergraph learning. In this process, the optimization on cost interval achieves better performance instead of choosing uncertain fixed cost in the learning process. To evaluate the effectiveness of the proposed method, we have conducted experiments on two groups of dataset, i.e., the NASA Metrics Data Program (NASA) dataset and UCI Machine Learning Repository (UCI) dataset. Experimental results and comparisons with state-of-the-art methods have exhibited better performance of our proposed method. Xibin Zhao, Nan Wang 0015, Heyuan Shi, Hai Wan, Jin Huang 0002, Yue Gao 0002 |
AAAI | 3 |
| 2018 | VulSeeker-pro: enhanced semantic learning based binary vulnerability seeker with emulationabstractLearning-based clone detection is widely exploited for binary vulnerability search. Although they solve the problem of high time overhead of traditional dynamic and static search approaches to some extent, their accuracy is limited, and need to manually identify the true positive cases among the top-M search results during the industrial practice. This paper presents VulSeeker-Pro, an enhanced binary vulnerability seeker that integrates function semantic emulation at the back end of semantic learning, to release the engineers from the manual identification work. It first uses the semantic learning based predictor to quickly predict the top-M candidate functions which are the most similar to the vulnerability from the target binary. Then the top-M candidates are fed to the emulation engine to resort, and more accurate top-N candidate functions are obtained. With fast filtering of semantic learning and dynamic trace generation of function semantic emulation, VulSeeker-Pro can achieve higher search accuracy with little time overhead. The experimental results on 15 known CVE vulnerabilities involving 6 industry widely used programs show that VulSeeker-Pro significantly outperforms the state-of-the-art approaches in terms of accuracy. In a total of 45 searches, VulSeeker-Pro finds 40 and 43 real vulnerabilities in the top-1 and top-5 candidate functions, which are 12.33× and 2.58× more than the most recent and related work Gemini. In terms of efficiency, it takes 0.22 seconds on average to determine whether the target binary function contains a known vulnerability or not. Jian Gao 0008, Yu Jiang 0001, Heyuan Shi, Jia-Guang Sun 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2018 | Secure beamforming for cognitive cyber-physical systems based on cognitive radio with wireless energy harvesting
Ronghua Shi, Heyuan Shi, Md. Zakirul Alam Bhuiyan |
Ad Hoc Networks | 3 |
| 2018 | Multi-model induced network for participatory-sensing-based classification tasks in intelligent and connected transportation systems
Heyuan Shi, Xibin Zhao, Hai Wan, Huihui Wang 0001, Jian Dong 0001, Anfeng Liu |
Comput. Networks | 1 |
| 2018 | Cooperative spectrum sharing in cognitive radio networks with energy accumulation: design and analysisabstractThe authors propose an efficient spectrum sharing scheme in cooperative cognitive radio networks, where an energy‐constrained secondary transmitter (ST) first scavenges radio frequency (RF) energy from the received primary signals, and then the ST assists the primary transmission to obtain the opportunity of spectrum access. Specifically, the ST can forward the primary signal with its own signal by adopting both the Alamouti coding technique and superposition scheme only if the harvested energy is sufficient while the primary data is decoded correctly by the ST. Otherwise, the ST will continue to harvest RF energy. The authors use the discrete Markov chain to model the processes of charging and discharging of the battery. Moreover, two different joint decoding and interference cancellation schemes are employed at the receivers to restore the desired data. Closed‐form expressions of outage probabilities for both the primary and secondary systems are derived. Aiming to minimise the outage probability of the secondary system with guaranteeing the primary transmission, an optimal power allocation factor for the ST is determined by Monte‐Carlo simulation. Numerical results demonstrate that the proposed scheme can effectively improve the transfer performance of the secondary system while realising the transfer requirement of the primary system. Ronghua Shi, Jingchun Xi, Heyuan Shi, Wentai Lei |
IET Commun. | 4 |