Haichao Du

dblp:241/9769 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-2783-3232ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
abstract
Large language models (LLMs) increasingly support multilingual understanding and generation. Meanwhile, efforts to interpret their internal mechanisms have emerged, offering insights to enhance multilingual performance. While multi-head self-attention (MHA) has proven critical in many areas, its role in multilingual capabilities remains underexplored. In this work, we study the contribution of MHA in supporting multilingual processing in LLMs. We propose Language Attention Head Importance Scores (LAHIS), an effective and efficient method that identifies attention head importance for multilingual capabilities via a single forward and backward pass through the LLM. Applying LAHIS to Aya-23-8B, Llama-3.2-3B, and Mistral-7B-v0.1, we reveal the existence of both language-specific and language-general heads. Language-specific heads enable cross-lingual attention transfer to guide the model toward target language contexts and mitigate off-target language generation issue, contributing to addressing challenges in multilingual LLMs. We also introduce a lightweight adaptation that learns a soft head mask to modulate attention outputs over language heads, requiring only 20 tunable parameters to improve XQuAD accuracy. Overall, our work enhances both the interpretability and multilingual capabilities of LLMs from the perspective of MHA.
Qiyang Song, Qihang Zhou, Haichao Du, Shaowen Xu, Weijuan Zhang, Xiaoqi Jia
AAAI4
2026 FalconScope: Effective and Efficient Detection of Hidden Web Interfaces in IoT Devices
abstract
Hidden web interfaces in Internet of Things (IoT) devices pose significant security threats by unintentionally exposing inadequately protected functionalities, enabling attackers to bypass authentication, alter configurations, leak sensitive data, or execute arbitrary commands. Despite recent advancements, current detection approaches suffer from two critical challenges: 1) inadequately model the complex internal routing mechanisms of IoT firmware, leading to incomplete interface enumeration and substantial false negatives; and 2) inefficiently generate probing requests and verify unauthorized access due to limited semantic understanding of interface communication protocols. To overcome these challenges, we introduce FalconScope, a novel system combining precise firmware routing modeling and Large Language Model (LLM)-driven semantic analysis to detect hidden web interfaces effectively and efficiently. FalconScope achieves this through two key innovations: 1) a static analysis technique precisely reconstructs the device's internal routing mechanisms, enabling comprehensive enumeration of Routing Unique Identifiers and their corresponding backend handlers; and 2) an LLM-powered semantic engine automatically generates syntactically and semantically valid HTTP requests to efficiently trigger backend logic, coupled with semantic validation of device responses to accurately confirm unauthorized access. Evaluations on 11 real-world IoT devices from four major vendors demonstrate that FalconScope significantly surpasses existing state-of-the-art tools, detecting 620 hidden web interfaces—103 times more than IoTScope—while consuming only 3.6% of its analysis time. Following responsible disclosure, 50 issues have been assigned CVE IDs.
Jiaming Guo, Kuihao Yan, Jiekang Hu, Xiaoqi Jia, Haichao Du, Qihang Zhou
WWW6
2026 A hybrid framework of large language models and transformer networks for multimodal stock movement prediction
Haichao Du
Eng. Appl. Artif. Intell.1
2026 FlexClave: An Extensible and Secure Trusted Execution Environment Framework
abstract
As computer system software stacks become increasingly complex, the associated security risks also escalate. Trusted Execution Environments (TEEs) have emerged as a mainstream security solution to enhance system security. TEEs can be categorized into user-level TEEs, OS-level TEEs, and hybrid TEEs. However, these TEEs typically possess fixed security boundaries and isolation domains, limiting their adaptability to varying security requirements and dynamic scenarios. Moreover, the design of Trusted Computing Base (TCB) components in TEE frameworks often operates at the highest privilege levels of the architecture. This concentration of critical code at the highest privilege level increases the whole platform’s security risk due to the growing amount of code as more security functions are added. In this paper, we propose FlexClave, an extensible and secure TEE framework designed to address these issues. FlexClave leverages hardware primitives to create secure isolation boundaries tailored to different use cases. Additionally, our framework distributes TCB components across various privilege levels, reducing the concentration of security functions at the highest privilege levels and mitigating the risks associated with running extensive code in a single, highly privileged context. We implement two prototypes on ARMv9-A Fixed Virtual Platform and ARMv8 RK3399 SoC, each with two use cases (container and virtual machine), to evaluate the system’s security and performance.
Qihang Zhou, Wenzhuo Cao, Xiaoqi Jia, Shaowen Xu, Jiayun Chen, Haichao Du, Yamin Xie, Peijie Yin, Shengzhi Zhang, Peng Liu 0005
IEEE Trans. Computers9
2025 EMHunter: An Evasive Malware Detection Approach to Improve Dynamic Analysis Efficiency
abstract
Currently, a growing number of malware employ evasion techniques to hinder security analysts from analyzing their dynamic behavior, making them more likely to evade detection and pose a threat to users. The above malware is classified as Evasive Malware. To address this issue, we collect a dataset of labeled samples and propose a novel analysis method for evasive malware detection called EMHunter (Evasive Malware Hunter). After injecting samples into a specially designed software environment, EMHunter modifies the section table to launch from a designated location, and alters the export table to disable evasion-related behavior via identifying 48 commonly used APIs. When the malicious code attempts evasive actions, such as detecting if it's running in an analysis environment, the software cooperates with the dynamic analysis environment to determine the malware's key characteristics. It then selects appropriate countermeasures to lure the malicious code into continued execution, thereby exposing more malicious behavior and improving the accuracy of dynamic analysis. Our dataset consists of 12,543 samples, experiments show that this method successfully induced 4084 samples to exhibit their behavior. Furthermore, we integrated EMHunter into the open-source sandbox CAPE, enabling it to gather more behavioral information from the samples. Finally, we evaluated our approach using a LightGBM model, achieving the accuracy of 96.61%.
Yamin Xie, Zhengcai Chen, Haichao Du, Xiaoqi Jia, Jianwu Ni
CSCWD3
2025 Chameleon: Towards Building Least-privileged TEE via Functionality-based Resource Re-grouping
abstract
TrustZone-assisted Trusted Execution Environment (TEE) has been widely employed in mobile devices to protect sensitive applications. With increased customization demands, Trusted Applications (TAs) have become more flexible and complex, exposing numerous vulnerabilities within the TEE. Furthermore, due to the unrestricted Trusted Operating System (TOS) services provided to TA, an attacker can exploit vulnerabilities to compromise the whole TEE system. In this paper, we propose a novel customized TOS partition approach, called Chameleon, to enhance the security of the TrustZone-assisted TEE system. Inspired by the principle of least privilege and our TEE vulnerability analysis, we first categorize the TOS into TOS service modules and basic kernel modules. Then, we selectively encapsulate these modules into distinct Capsules based on the TA's functional requirements, providing each TA with a separate execution environment (TA-entity). To enforce access control and confine vulnerable modules within a TA-entity, we introduce T-Visor, which serves as our Trusted Computing Base. Our prototype implementation, built upon Linaro's OP-TEE, requires only 2.9K Lines of Code (LoC) modifications. Evaluation on a Hikey960 board demonstrates that Chameleon reduces the attack surface of TOS services to 51% and mitigates 122 out of 138 CVEs (88.41%) with negligible performance overhead.
Qihang Zhou, Feifan Qian, Jiayun Chen, Heqing Huang 0001, Xiaoqi Jia, Haichao Du
MobiSys7
2025 TGNS: A transformer-based graph neural network for stock trend forecasting
Haichao Du
Inf. Sci.1
2024 ConMonitor: Lightweight Container Protection with Virtualization and VM Functions
abstract
Containers are widely used in multi-tenant cloud computing for their ease of deployment, minimal overhead, and fast start-up. However, the intrinsic shared kernel model of containers poses significant security threats, risking confidentiality and integrity from co-located containers or compromised OS. Researchers have proposed various methods to protect containers from untrusted OS, but few consider both the universality and efficiency. In this paper, we present ConMonitor---a lightweight and efficient container protection architecture. ConMonitor protects the security of container application data by introducing a compact virtualization software, called ConVisor, as a trusted computing base. ConVisor enforces isolation of the physical memory between containers and the kernel, and monitors the sensitive operations performed by the OS. To ensure the security of ConMonitor, we implement a Container Guardian to serve as an intermediary for the kernel, managing sensitive operations. Moreover, we also leverage the VMFUNC feature to achieve fast context switching, thereby mitigating the performance penalty associated with frequent context switching. We have implemented ConMonitor on Intel CPU with Virtualization Technology, and the evaluation results show that ConMonitor can protect the security of container applications with a negligible performance overhead.
Shaowen Xu, Qihang Zhou, Xiaoqi Jia, Heqing Huang 0001, Haichao Du
SoCC7
2024 Structure-Sensitive Pointer Analysis for Multi-structure Objects
abstract
Static analysis is a method within software analysis, and pointer analysis is an important component of static analysis. An important dimension of pointer analysis is field-sensitivity, which has been proven to effectively enhance the accuracy of pointer analysis results. A crucial area of research within field-sensitivity is structure-sensitivity. Structure-sensitivity has been shown to further enhance the precision of pointer analysis. However, existing structure-sensitive methods cannot handle cases where an object possesses multiple structures.
Xun An, Xiaoqi Jia, Haichao Du, Yamin Xie
Internetware3
2024 LightArmor: A Lightweight Trusted Operating System Isolation Approach for Mobile Systems
Qihang Zhou, Xiaoqi Jia, Jiayun Chen, Qingjia Huang, Haichao Du
SEC6
2024 Malware Classification Method Based on Dynamic Features with Sensitive Behaviors
abstract
Traditional malware classification methods often just scratch the surface by analyzing the sequence of system commands (API calls) used by malware during its operation. These approaches miss out on deeper, complex behaviors that could significantly enhance accuracy in identifying different malware types. To address this, we introduce SenBeMC, a method that delves deeper into the behaviors exhibited by malware. SenBeMC combine API call information vectors with behavioral information to enhance the deep semantic information of input features, enriching the hierarchical structure of feature representation. SenBeMC stands out by employing soft thresholding and attention mechanisms to sift through the noise — extraneous information that can mask the malware's true nature, and a BiLSTM model that excels in understanding the sequence and timing of actions, crucial for spotting sophisticated threats. Experimental evaluations on real-world datasets affirm that SenBeMC effectively improves feature representation and accuracy of malware classification when compared to other contemporary state-of-the-art models.
Yamin Xie, Siyuan Li 0014, Zhengcai Chen, Haichao Du, Xiaoqi Jia, Yuejin Du
SMC4
2024 HClave: An isolated execution environment design for hypervisor runtime security
Qihang Zhou, Wenzhuo Cao, Xiaoqi Jia, Shengzhi Zhang, Jiayun Chen, Weijuan Zhang, Haichao Du, Qingjia Huang
Comput. Secur.8
2023 Log2Policy: An Approach to Generate Fine-Grained Access Control Rules for Microservices from Scratch
abstract
Microservice application architecture is one of the most widely used service architectures in the industry. To prevent a compromised microservice from abusing other microservices, authorization policy is applied to regulate the access among them. However, configuring access control policy manually is challenging due to the complexity and dynamic nature of microservice applications. In this paper, we present Log2Policy, a novel approach to generate microservice authorization policy based on access logs. Our approach consists of three fundamental techniques: (1) a log-based topological graph generation mechanism that automatically infers the invocation logic among microservices, (2) a machine learning based attributes mining method that extracts the relevant attributes of requests, and (3) a policy upgrade mechanism based on traffic management that can significantly reduce the upgrade time. We have implemented a prototype of Log2Policy on mainstream microservice infrastructures and have evaluated it with several microservice applications. The results show that Log2Policy can generate fine-grained and effective access control rules and upgrade them with negligible overhead.
Shaowen Xu, Qihang Zhou, Heqing Huang 0001, Xiaoqi Jia, Haichao Du, Yamin Xie
ACSAC5
2023 Refining Use-After-Free Defense: Eliminating Dangling Pointers in Registers and Memory
abstract
The prevalence of use-after-free (UAF) vulnerabilities poses a significant threat to software security, with dangling pointers identified as the primary cause. However, existing de-fense methods suffer from bypass attacks, high runtime overhead, or only address memory dangling pointers while neglecting register-based ones that also contribute to UAF vulnerabilities. To overcome these shortcomings, we introduce a novel approach, ISDE, that eliminates both register and memory dangling point-ers with minimal additional runtime overhead. ISDE leverages an inter-procedural static pointer analysis method to statically collect object pointers during compilation, and uses the call graph and data flow graph to identify and eliminate potential dangling pointers. Our implementation of ISDE demonstrated its effectiveness in defending against real-world UAF vulnerabilities while maintaining efficiency in the SPEC CPU2006 evaluation.
Xun An, Qihang Zhou, Haichao Du, Xiaoqi Jia
APSEC3
2023 Protecting Encrypted Virtual Machines from Nested Page Fault Controlled Channel
abstract
AMD Secure Encrypted Virtualization (SEV) assumes the hypervisor (HV) is untrusted and introduces hardware memory encryption support for virtual machines (VMs). Previous studies have proposed various attacks against encrypted VMs by exploiting SEV security flaws such as unencrypted VMCB and lack of memory integrity. Most of these flaws have been solved by the subsequent releases of SEV with Encrypted State (SEV-ES) and SEV with Secure Nested Paging (SEV-SNP). However, the latest SEV-SNP cannot stop the malicious HV tampering with critical flags in the nested page table (NPT). So SEV-SNP is still vulnerable to the nested page fault (NPF) controlled channel attack, which is a commonly shared step of most attacks against SEV. Existing works on SEV also cannot defend against NPF controlled channel. In this paper, we first analyze the root cause of NPF controlled channel. Then we propose a software-based approach to protect encrypted VMs from NPF controlled channel. We introduce a virtualization security module (VSM) as a software TCB to deprivilege the HV by modifing the HV to access critical resources indirectly through interfaces managed by VSM. To prevent the untrusted HV from compromising the VSM-based protection, we extend the nested kernel architecture to the virtualization layer to provide isolation for VSM at the same privilege level. A prototype of this approach is implemented based on KVM. The experiments show that the approach can protect encrypted VMs from NPF controlled channel with 1.21% average runtime overhead and 1.47% average I/O overhead.
Haoxiang Qin, Weijuan Zhang, Sicong Huang 0004, Xiaoqi Jia, Haichao Du
CODASPY8
2020 SEEF-ALDR: A Speaker Embedding Enhancement Framework via Adversarial Learning based Disentangled Representation
abstract
Speaker verification, as a biometric authentication mechanism, has been widely used due to the pervasiveness of voice control on smart devices. However, the task of “in-the-wild” speaker verification is still challenging, considering the speech samples may contain lots of identity-unrelated information, e.g., background noise, reverberation, emotion, etc. Previous works focus on optimizing the model to improve verification accuracy, without taking into account the elimination of the impact from the identity-unrelated information. To solve the above problem, we propose SEEF-ALDR, a novel Speaker Embedding Enhancement Framework via Adversarial Learning based Disentangled Representation, to reinforce the performance of existing models on speaker verification. The key idea is to retrieve as much speaker identity information as possible from the original speech, thus minimizing the impact of identity-unrelated information on the speaker verification task by using adversarial learning. Experimental results demonstrate that the proposed framework can significantly improve the performance of speaker verification by 20.3% and 23.8% on average over 13 tested baselines on dataset Voxceleb1 and 8 tested baselines on dataset Voxceleb2 respectively, without adjusting the structure or hyper-parameters of them. Furthermore, the ablation study was conducted to evaluate the contribution of each module in SEEF-ALDR. Finally, porting an existing model into the proposed framework is straightforward and cost-efficient, with very little effort from the model owners due to the modular design of the framework.
Jianwei Tai, Xiaoqi Jia, Qingjia Huang, Weijuan Zhang, Haichao Du, Shengzhi Zhang
ACSAC5
2020 ET-GAN: Cross-Language Emotion Transfer Based on Cycle-Consistent Generative Adversarial Networks
abstract
Despite the remarkable progress made in synthesizing emotional speech from text, it is still challenging to provide emotion information to existing speech segments. Previous methods mainly rely on parallel data, and few works have studied the generalization ability for one model to transfer emotion information across different languages. To cope with such problems, we propose an emotion transfer system named ET-GAN, for learning language-independent emotion transfer from one emotion to another without parallel training samples. Based on cycle-consistent generative adversarial network, our method ensures the transfer of only emotion information across speeches with simple loss designs. Besides, we introduce an approach for migrating emotion information across different languages by using transfer learning. The experiment results show that our method can efficiently generate high-quality emotional speech for any given emotion category, without aligned speech pairs.
Xiaoqi Jia, Jianwei Tai, Yakai Li, Weijuan Zhang, Haichao Du, Qingjia Huang
ECAI6