Bei Guan

dblp:117/5499 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0001-5071-1546ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Security and privacy · 4Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Private-library-oriented code generation with large language models
Daoguang Zan, Bei Chen 0008, Yongshun Gong, Junzhi Cao, Fengji Zhang, Bingchao Wu, Bei Guan, Yilong Yin, Yongji Wang 0002
Knowl. Based Syst.7
2024 A GAN-Based Data Poisoning Framework Against Anomaly Detection in Vertical Federated Learning
abstract
In vertical federated learning (VFL), commercial entities collaboratively train a model while preserving data privacy. However, a malicious participant's poisoning attack may degrade the performance of this collaborative model. The main challenge in achieving the poisoning attack is the absence of access to the server-side top model, leaving the malicious par-ticipant without a clear target model. To address this challenge, we introduce an innovative end-to-end poisoning framework P-GAN. Specifically, the malicious participant initially employs semi-supervised learning to train a surrogate target model. Subsequently, this participant employs a GAN-based method to produce adversarial perturbations to degrade the surrogate target model's performance. Finally, the generator is obtained and tailored for VFL poisoning. Besides, we develop an anomaly detection algorithm based on a deep auto-encoder (DAE), offering a robust defense mechanism to VFL scenarios. Through extensive experiments, we evaluate the efficacy of P-GAN and DAE, and further analyze the factors that influence their performance.
Daoguang Zan, Wei Li 0326, Bei Guan, Yongji Wang 0002
ICC4
2024 FIA-TE: Feature Inference Attack on Decision Tree Ensembles in Vertical Federated Learning
abstract
Vertical federated learning (VFL) enables multiple parties to collaboratively train a model while preserving privacy. However, recent studies have raised concerns about the susceptibility of VFL models, including those using logistic regression and neural networks, to feature inference attacks. Meanwhile, the non-differentiable characteristics of decision tree ensembles make conducting such attacks impractical. To address this challenge, we introduce a feature inference attack framework FIA-TE tailored for decision tree ensembles, including gradient boosted decision trees (GBDT) and random forest. Specifically, we distill the knowledge from trees into neural networks by leaf embedding and structure distillation to create a targeted model for the inference attack. We then employ a generative model based on the deconvolutional network for capturing correlation features and reconstructing the target features. Through extensive experiments on table and image data, we evaluate the effectiveness of our framework and provide an analysis of potential influencing factors.
Daoguang Zan, Wei Li 0326, Bei Guan, Yongji Wang 0002
ICME4
2023 Large Language Models Meet NL2Code: A Survey
abstract
Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, Jian-Guang Lou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Daoguang Zan, Bei Chen 0008, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang 0002, Jian-Guang Lou
ACL (1)6
2023 Hierarchical and Contrastive Representation Learning for Knowledge-Aware Recommendation
abstract
Incorporating knowledge graph into recommendation is an effective way to alleviate data sparsity. Most existing knowledge-aware methods usually perform recursive embedding propagation by enumerating graph neighbors. However, the number of nodes’ neighbors grows exponentially as the hop number increases, forcing the nodes to be aware of vast neighbors under this recursive propagation for distilling the high-order semantic relatedness. This may induce more harmful noise than useful information into recommendation, leading the learned node representations to be indistinguishable from each other, that is, the well-known over-smoothing issue. To relieve this issue, we propose a Hierarchical and CONtrastive representation learning framework for knowledge-aware recommendation named HiCON. Specifically, for avoiding the exponential expansion of neighbors, we propose a hierarchical message aggregation mechanism to interact separately with low-order neighbors and meta-path-constrained high-order neighbors. Moreover, we also perform cross-order contrastive learning to enforce the representations to be more discriminative. Extensive experiments on three datasets show the remarkable superiority of HiCON over state-of-the-art approaches. The code is available now1.
Bingchao Wu, Yangyuxuan Kang, Daoguang Zan, Bei Guan, Yongji Wang 0002
ICME4
2023 We Are Not So Similar: Alleviating User Representation Collapse in Social Recommendation
abstract
Integrating social relations into recommendation is an effective way to mitigate data sparsity. Most social recommendation methods encode user representations from a unified graph that includes user-user and user-item relations. Due to the enriched relations on this graph, a large fraction of users are aware of each other within only a few hops, and the user representations generated by existing methods may encode the information received from a large number of neighbors. Thus, many user representations are enforced to be too similar, which hinders modeling fine-grained user interest. Here, we name this phenomenon as user representation collapse. To address this problem, in this paper we propose a robust user representation learning method named RobustSR with social regularization and multi-view contrastive learning, which aim to enhance the model’s awareness of relation informativeness and the discriminativeness of user representations, respectively. Concretely, the social regularization mechanism encourages the model to learn from the relation importance weights derived from graph topologies, which helps recognize important observed relations meanwhile mining potential useful relations. To enhance the discriminativeness of user representations, we further perform multi-view contrastive learning between collaborative and social-enhanced user representations. Extensive experiments on four benchmark datasets show that RobustSR effectively alleviates user representation collapse and improves recommendation performance. Our code is deposited at https://github.com/paulpig/RobustSR.
Bingchao Wu, Yangyuxuan Kang, Bei Guan, Yongji Wang 0002
ICMR3
2022 Enhancing Sequential Recommendation via Decoupled Knowledge Graphs
Bingchao Wu, Chenglong Deng, Bei Guan, Yongji Wang 0002, Yuxuan Kangyang
ESWC3
2022 CERT: Continual Pre-training on Sketches for Library-oriented Code Generation
abstract
Code generation is a longstanding challenge, aiming to generate a code snippet based on a natural language description. Usually, expensive text-code paired data is essential for training a code generation model. Recently, thanks to the success of pre-training techniques, large language models are trained on large unlabelled code corpora and perform well in generating code. In this paper, we investigate how to leverage an unlabelled code corpus to train a model for library-oriented code generation. Since it is a common practice for programmers to reuse third-party libraries, in which case the text-code paired data are harder to obtain due to the huge number of libraries. We observe that library-oriented code snippets are more likely to share similar code sketches. Hence, we present CERT with two steps: a sketcher generates the sketch, then a generator fills the details in the sketch. Both the sketcher and generator are continually pre-trained upon a base model using unlabelled data. Also, we carefully craft two benchmarks to evaluate library-oriented code generation named PandasEval and NumpyEval. Experimental results have shown the impressive performance of CERT. For example, it surpasses the base model by an absolute 15.67% improvement in terms of pass@1 on PandasEval. Our work is available at https://github.com/microsoft/PyCodeGPT.
Daoguang Zan, Bei Chen 0008, Dejian Yang, Zeqi Lin, Bei Guan, Yongji Wang 0002, Weizhu Chen, Jian-Guang Lou
IJCAI6
2022 Complex Question Answering over Incomplete Knowledge Graph as N-ary Link Prediction
abstract
The Question Answering over Knowledge Graph (KGQA) task seeks entities (answers) from the Knowledge Graph (KG) in order to answer natural language questions. In practice, KG is often incomplete, with numerous missing links and nodes. With such an incomplete KG, it is tricky to use the semantics inside the KG to get the golden answers, particularly for complex questions. Some current efforts concentrate on using external corpora to overcome KG sparsity; however, identifying and obtaining the corpora is challenging. Other types of work aim to leverage the pre-trained embeddings to resolve the issue but perform slightly worse on complex questions involving numerous triple facts in KG. To address the aforementioned problems, we present a framework CAPKGQA, which transforms Complex KGQA into an n-Ary link Prediction task capable of explicitly modeling complex questions. Furthermore, previous methods also suffer from incomplete KG throughout the candidate answer generation phase. Therefore, we devise an embedding-based retrieval strategy to extract more reliable candidate answers from incomplete KG. Extensive experiments reveal that our approach beats the state-of-the-art models on incomplete and complex KGQA tasks by a significant margin.
Daoguang Zan, Kun Zhou 0002, Wei Wu 0014, Wayne Xin Zhao, Bingchao Wu, Bei Guan, Yongji Wang 0002
IJCNN8
2022 S2QL: Retrieval Augmented Zero-Shot Question Answering over Knowledge Graph
Daoguang Zan, Yuanmeng Yan, Wei Wu 0014, Bei Guan, Yongji Wang 0002
PAKDD (3)6
2021 Fed-EINI: An Efficient and Interpretable Inference Framework for Decision Tree Ensembles in Vertical Federated Learning
abstract
Vertical federated learning has a great potential of driving a great variety of business cooperation among enterprises in many fields. In machine learning, decision tree ensembles such as gradient boosting decision trees (GBDT) and random forest are widely applied powerful models with high interpretability and modeling efficiency. However, state-of-art framework for decision tree ensembles in vertical federated learning frameworks adapt anonymous features to avoid possible data breaches, makes the interpretability of the model compromised.To address this issue in the inference process, in this paper, we firstly make a problem analysis about the necessity of disclosure meanings of feature to Guest Party in vertical federated learning. We protect data privacy and allow the disclosure of feature meaning by concealing decision paths and adapt a communication-efficient secure computation method for inference outputs. The advantages of Fed-EINI will be demonstrated through both theoretical analysis and extensive numerical results. We improve the interpretability of the model by disclosing the meaning of features while ensuring efficiency and accuracy.
Shuai Zhou 0001, Bei Guan, Hao Fao, Yongji Wang 0002
IEEE BigData3
2020 Following Passive DNS Traces to Detect Stealthy Malicious Domains Via Graph Inference
abstract
Malicious domains, including phishing websites, spam servers, and command and control servers, are the reason for many of the cyber attacks nowadays. Thus, detecting them in a timely manner is important to not only identify cyber attacks but also take preventive measures. There has been a plethora of techniques proposed to detect malicious domains by analyzing Domain Name System (DNS) traffic data. Traditionally, DNS acts as an Internet miscreant’s best friend, but we observe that the subtle traces in DNS logs left by such miscreants can be used against them to detect malicious domains. Our approach is to build a set of domain graphs by connecting “related” domains together and injecting known malicious and benign domains into these graphs so that we can make inferences about the other domains in the domain graphs. A key challenge in building these graphs is how to accurately identify related domains so that incorrect associations are minimized and the number of domains connected from the dataset is maximized. Based on our observations, we first train two classifiers and then devise a set of association rules that assist in linking domains together. We perform an in-depth empirical analysis of the graphs built using these association rules on passive DNS data and show that our techniques can detect many more malicious domains than the state-of-the-art.
Mohamed Nabeel, Issa M. Khalil, Bei Guan, Ting Yu 0001
ACM Trans. Priv. Secur.3
2019 An Efficient Approach for Mitigating Covert Storage Channel Attacks in Virtual Machines by the Anti-Detection Criterion
Nasro Min-Allah, Bei Guan, Yuqi Lin, JingZheng Wu, Yongji Wang 0002
J. Comput. Sci. Technol.3
2018 A Domain is only as Good as its Buddies: Detecting Stealthy Malicious Domains via Graph Inference
abstract
Inference based techniques are one of the major approaches to analyze DNS data and detect malicious domains. The key idea of inference techniques is to first define associations between domains based on features extracted from DNS data. Then, an inference algorithm is deployed to infer potential malicious domains based on their direct/indirect associations with known malicious ones. The way associations are defined is key to the effectiveness of an inference technique. It is desirable to be both accurate (i.e., avoid falsely associating domains with no meaningful connections) and with good coverage (i.e., identify all associations between domains with meaningful connections). Due to the limited scope of information provided by DNS data, it becomes a challenge to design an association scheme that achieves both high accuracy and good coverage.
Issa M. Khalil, Bei Guan, Mohamed Nabeel, Ting Yu 0001
CODASPY2
2016 Discovering Malicious Domains through Passive DNS Data Graph Analysis
abstract
Malicious domains are key components to a variety of cyber attacks. Several recent techniques are proposed to identify malicious domains through analysis of DNS data. The general approach is to build classifiers based on DNS-related local domain features. One potential problem is that many local features, e.g., domain name patterns and temporal patterns, tend to be not robust. Attackers could easily alter these features to evade detection without affecting much their attack capabilities. In this paper, we take a complementary approach. Instead of focusing on local features, we propose to discover and analyze global associations among domains. The key challenges are (1) to build meaningful associations among domains; and (2) to use these associations to reason about the potential maliciousness of domains. For the first challenge, we take advantage of the modus operandi of attackers. To avoid detection, malicious domains exhibit dynamic behavior by, for example, frequently changing the malicious domain-IP resolutions and creating new domains. This makes it very likely for attackers to reuse resources. It is indeed commonly observed that over a period of time multiple malicious domains are hosted on the same IPs and multiple IPs host the same malicious domains, which creates intrinsic association among them. For the second challenge, we develop a graph-based inference technique over associated domains. Our approach is based on the intuition that a domain having strong associations with known malicious domains is likely to be malicious. Carefully established associations enable the discovery of a large set of new malicious domains using a very small set of previously known malicious ones. Our experiments over a public passive DNS database show that the proposed technique can achieve high true positive rates (over 95%) while maintaining low false positive rates (less than 0.5%). Further, even with a small set of known malicious domains (a couple of hundreds), our technique can discover a large set of potential malicious domains (in the scale of up to tens of thousands).
Issa M. Khalil, Ting Yu 0001, Bei Guan
AsiaCCS3
2014 CIVSched: A Communication-Aware Inter-VM Scheduling Technique for Decreased Network Latency between Co-Located VMs
abstract
Server consolidation in cloud computing environments makes it possible for multiple servers or desktops to run on a single physical server for high resource utilization, low cost, and reduced energy consumption. However, the scheduler in the virtual machine monitor (VMM), such as Xen credit scheduler, is agnostic about the communication behavior between the guest operating systems (OS). The aforementioned behavior leads to increased network communication latency in consolidated environments. In particular, the CPU resources management has a critical impact on the network latency between co-located virtual machines (VMs) when there are CPUand I/O-intensive workloads running simultaneously. This paper presents the design and implementation of a communication-aware inter-VM scheduling (CIVSched) technique that takes into account the communication behavior between inter-VMs running on the same virtualization platform. The CIVSched technique inspects the network packets transmitted between local co-resident domains to identify the target VM and process that will receive the packets. Thereafter, the target VM and process are preferentially scheduled by the VMM and the guest OS. The cooperation of these two schedulers makes the network packets to be timely received by the target application. Experimental results on the Xen virtualization platform depict that the CIVSched technique can reduce the average response time of network traffic by approximately 19 percent for the highly consolidated environment, while keeping the inherent fairness of the VMM scheduler.
Bei Guan, JingZheng Wu, Yongji Wang 0002, Samee Ullah Khan
IEEE Trans. Cloud Comput.1
2013 CIVSched: Communication-aware Inter-VM Scheduling in Virtual Machine Monitor Based on the Process
abstract
Server consolidation in Cloud Computing makes it possible for multiple servers or desktops to run on one physical server to get high resource utilization, low cost and less energy consumption. However, the scheduler in virtual machine monitor (VMM) is agnostic about the communication behavior between the guest operating systems. It leads to inefficient network communication in consolidated environment. In particular, the CPU resource management has a critical impact on the network latency between co-resident virtual machines (VMs) when there are CPU-bound and I/O-bound workloads existing simultaneously. It brings a negative impact on latency-sensitive VMs. In this paper, we present the design and implementation of CIVSched scheduling to make the VMM aware of the communication behavior between two inter-VMs running on the same virtual platform. CIVSched inspects the network packets transmitted between local domains and find the destination VM and the target process inside that will receive the packets. Then, the destination VM and the target process are preferentially scheduled by VMM scheduler and guest OS scheduler respectively. The cooperation of these two schedulers makes the network packets received by the target application timely. Experimental results show that the CIVSched scheduling can reduce the average response time of network traffic by up to 18% for the highly consolidated environment while keeping the fairness of the VMM scheduler.
Bei Guan, Liping Ding, Yongji Wang 0002
CCGRID1
2012 Return-Oriented Programming Attack on the Xen Hypervisor
abstract
In this paper, we present an approach to attack on the Xen hypervisor utilizing return-oriented programming (ROP). It modifies the data in the hypervisor that controls whether a VM is privileged or not and thus can escalate the privilege of an unprivileged domain (domU) at run time. As ROP technique makes use of existed code to implement attack, not modifying or injecting any code, it can bypass the integrity protections that base on code measurement. By constructing such kind of attack at the virtualization layer, it can motivate further research work towards preventing or detecting ROP attack on the hypervisor.
Baozeng Ding, Yeping He, Shuo Tian, Bei Guan
ARES5