VLDB 2026 Research / reviewers in the wild / expert
Yuhan Chai
dblp:232/4944
· DBLP profile ↗
8ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-4332-0234ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dual-Surrogate Competition-Assisted Evolutionary Algorithm With Triggered Constraint First Search for Expensive Constrained OptimizationabstractExpensive constrained optimization problems (ECOPs), which frequently arise in real-world engineering optimization, are often limited by the number of evaluations. Using surrogate-assisted evolutionary algorithms (EAs) to reduce computational costs is a common approach. However, surrogate model errors are inevitable and often mislead the search direction of EAs. Most existing algorithms overlook the inevitability of such errors and attempt to minimize them through techniques like data selection, which might be ineffective for problems with highly complex constraint and objective functions. Therefore, we propose a dual-surrogate competition assisted EA with triggered constraint-first search (DC-TCFS) for ECOPs, aiming to reduce the misleading effects of surrogate models on the evolutionary process. In this study, two surrogate models are used to assist local searches through competition, effectively mitigating the impact of errors from a single surrogate model. A triggered constraint-first search method is proposed to quickly identify a feasible solution for problems with complex constraints. Additionally, an adaptive sampling criterion is designed to guide the algorithm toward solutions that are more beneficial to the evolutionary process. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art methods across various benchmark problems, real-world problems and a motor design optimization problem, highlighting its effectiveness in solving ECOPs. Kunjie Yu, Yuhan Chai, Fan Chen 0011, Ke Chen 0022, Rui Nie 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | MalFSCIL: A Few-Shot Class-Incremental Learning Approach for Malware DetectionabstractThe continuous evolution of malware is posing a serious threat to personal privacy, enterprise data security, and global network infrastructure. For example, attackers can use phishing emails, botnets, etc. to induce victims to execute malware for nefarious purposes such as stealing sensitive information. Therefore, it is significant to develop effective and efficient methods to detect malware. Towards this, most state-of-the-art methods are focused on learning-based method. In order to adapt to the characteristics of sample scarcity and dynamic evolution of malware detection tasks, few-shot class incremental learning has been proposed as an efficient pairwise solution. Nevertheless, they still face two major challenges: 1) Catastrophic Forgetting: the erosion of existing knowledge by newly acquired knowledge during incremental learning. 2) Decision boundary confusion: after continuous multiple incremental sessions, the discriminative ability of the classification model is weakened. To address the above challenges, we propose a new Malware detection framework based on Few-Shot Class Incremental Learning, MalFSCIL, which utilizes a decoupled training strategy combined with a variational autocoder to mitigate catastrophic forgetting, and designs a dynamic boundary delineation method based on class prototyping to achieve accurate delineation of incremental decision boundaries. Extensive experimental results show that the proposed method outperforms the state-of-the-art techniques in malware detection and classification with high classification accuracy with open-source dataset and Internal enterprise dataset. Yuhan Chai, Ximing Chen 0004, Jing Qiu 0002, Yanjun Xiao 0001, Qiying Feng, Shouling Ji, Zhihong Tian 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Toward Open-World Network Intrusion Detection via Open Recognition and InspectionabstractDeep learning is promising in open-world network intrusion detection, but current deep learning-based methods mainly focus on open recognition with properties that may not always hold and significantly neglect the inspection of unknown samples, increasing open space risks and manual inspection overhead for deployed models. To address these challenges in real-world environments, we propose a novel system, ORI, designed to tackle two critical tasks: 1) open recognition, including classifying known class samples while recognizing unknown ones, and 2) inspection, involving further inspecting samples recognized as unknown. Specifically, we reformulate open recognition as a binary classification task and propose a density-based method to recognize low-density samples as unknown while classifying known class samples with a closed-world classifier, thereby minimizing the risk associated with open spaces. To reduce the inspection overhead of samples recognized as unknown, we treat unknown sample inspection as a constrained clustering task, using a few manually inspected samples as constraints, and then assign labels to the remaining unknown samples via clustering. We evaluate our system against established open recognition and unknown sample inspection baselines through extensive experiments on three public datasets. Additionally, we simulated a security analyst inspecting unknown samples labeled by ORI. The experimental results demonstrate that ORI accurately classifies known class samples, recognizes unknown samples, and effectively labels samples recognized as unknown, enhancing both open recognition and inspection capabilities. Yuhan Chai, Yan Jia 0001, Binxing Fang, Hao Li 0027, Zhaoquan Gu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Dynamic Prototype Network Based on Sample Adaptation for Few-Shot Malware DetectionabstractThe continuous increase and spread of malware have caused immeasurable losses to social enterprises and even the country, especially unknown malware. Most existing methods use predefined class samples to train models, which cannot handle unknown malware detection. In this paper, we formalize unknown malware detection as a Few-Shot Learning problem. However, the existing model cannot dynamically adjust the model parameters according to the samples and does not deeply consider the influence of the correlation between samples, so it achieves sub-optimal performance. We propose a Dynamic Prototype Network based on Sample Adaptation for few-shot malware detection (DPNSA). Specifically, we use dynamic convolution to realize dynamic feature extraction based on sample adaptation. Secondly, we define the class feature (prototype) as the mean of the dynamic embedding of all malware samples of each class in the support set. Then, a dual-sample dynamic activation function is proposed, which uses the correlation of the dual-sample to reduce the impact of unrelated features between samples on the metric. Finally, we use the metric-based method to calculate the distance between the query sample and the prototype to realize malware detection. Experiments show that our method outperforms the existing few-shot malware detection models and achieves significant improvement. Yuhan Chai, Jing Qiu 0002, Lihua Yin, Zhihong Tian 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | TPRPF: a preserving framework of privacy relations based on adversarial training for texts in big data
Yuhan Chai, Zhe Sun 0005, Jing Qiu 0002, Lihua Yin, Zhihong Tian 0001 |
Frontiers Comput. Sci. | 1 |
| 2022 | From Data and Model Levels: Improve the Performance of Few-Shot Malware ClassificationabstractExisting malware classification methods cannot handle the open-ended growth of new or unknown malware well because it only focuses on pre-defined malware classes with sufficient training data. Due to the superiority of the visualization method, some researchers use it for solving few-shot malware classification. However, the malware images generated by existing visualization methods contain insufficient semantic information. At the same time, existing few-shot models tend to converge to sharp minima resulting in poor generalization performance. By synthesizing the observations, we think that accurate and effective few-shot malware classification methods are affected by generated malware images and classification models, which can be called data and model levels, respectively. To solve the above problems, we propose a novel method from the Data and Model levels, which is used to classify new or unknown malware well, called DMMal. More specifically, we propose a multi-channel malware image generation method based on multi-view so that malware images can contain more prosperous information at the data level. In addition, we investigated adaptive sharpness-aware minimization in a few-shot scenario from the perspective of model optimization at the model level to minimize the loss value and sharpness simultaneously. This enhances the generalization ability of the model and improves the ability of the model to classify new or unknown classes. Experiments on two few-shot malware classification datasets show that the method proposed can improve the performance of few-shot malware classification from the data and model levels. Yuhan Chai, Jing Qiu 0002, Lihua Yin, Lejun Zhang, Brij B. Gupta, Zhihong Tian 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2020 | LGMal: A Joint Framework Based on Local and Global Features for Malware DetectionabstractWith the gradual advancement of smart city construction, various information systems have been widely used in smart cities. In order to obtain huge economic benefits, criminals frequently invade the information system, which leads to the increase of malware. Malware attacks not only seriously infringe on the legitimate rights and interests of users, but also cause huge economic losses. Signature-based malware detection algorithms can only detect known malware, and are susceptible to evasion techniques such as binary obfuscation. Behavior-based malware detection methods can solve this problem well. Although there are some malware behavior analysis works, they may ignore semantic information in the malware API call sequence. In this paper, we design a joint framework based on local and global features for malware detection to solve the problem of network security of smart cities, called LGMal, which combines the stacked convolutional neural network and graph convolutional networks. Specially, the stacked convolutional neural network is used to learn API call sequence information to capture local semantic features and the graph convolutional networks is used to learn API call semantic graph structure information to capture global semantic features. Experiments on Alibaba Cloud Security Malware Detection datasets show that the joint framework gets better results. The experimental results show that the precision is 87.76%, the recall is 88.08%, and the F1-measure is 87.79%. We hope this paper can provide a useful way for malware detection and protect the network security of smart city. Yuhan Chai, Jing Qiu 0002, Shen Su, Chunsheng Zhu, Lihua Yin, Zhihong Tian 0001 |
IWCMC | 1 |
| 2020 | Automatic Concept Extraction Based on Semantic Graphs From Big Data in Smart CityabstractWith the rapid development of smart cities, various types of sensors can rapidly collect a large amount of data, and it becomes increasingly important to discover effective knowledge and process information from massive amounts of data. Currently, in the field of knowledge engineering, knowledge graphs, especially domain knowledge graphs, play important roles and become the infrastructure of Internet knowledge-driven intelligent applications. Domain concept extraction is critical to the construction of domain knowledge graphs. Although there have been some works that have extracted concepts, semantic information has not been fully used. However, the excellent concept extraction results can be obtained by making full use of semantic information. In this article, a novel concept extraction method, Semantic Graph-Based Concept Extraction (SGCCE), is proposed. First, the similarities between terms are calculated using the word co-occurrence, the LDA topic model and Word2Vec. Then, a semantic graph of terms is constructed based on the similarities between the terms. Finally, according to the semantic graph of the terms, community detection algorithms are used to divide the terms into different communities where each community acts as a concept. In the experiments, we compare the concept extraction results that are obtained by different community detection algorithms to analyze the different semantic graphs. The experimental results show the effectiveness of our proposed method. This method can effectively use semantic information, and the results of the concept extraction are better from domain big data in smart cities. Jing Qiu 0002, Yuhan Chai, Zhihong Tian 0001, Xiaojiang Du, Mohsen Guizani |
IEEE Trans. Comput. Soc. Syst. | 2 |