VLDB 2026 Research / reviewers in the wild / expert
Donghong Sun
dblp:51/7665
· DBLP profile ↗
14ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wolf in Sheep's Clothing: Understanding and Detecting Mobile Cloaking in Blackhat SEO
Yan Li 0195, Zhenrui Zhang, Zhiyu Lang, Xiang Li 0041, Donghong Sun |
ACNS (2) | 6 |
| 2025 | CodeSAGE: A multi-feature fusion vulnerability detection approach using code attribute graphs and attention mechanisms
Tianyu Yao, Jiawei Qin, Qiao Ma, Donghong Sun |
J. Inf. Secur. Appl. | 6 |
| 2023 | Semi-supervised named entity recognition in multi-level contexts
Yubo Chen 0002, Chuhan Wu, Tao Qi 0001, Zhigang Yuan, Yuesong Zhang, Jian Guan 0009, Donghong Sun, Yongfeng Huang 0001 |
Neurocomputing | 8 |
| 2023 | Automatic Generation of Adversarial Readable Chinese TextsabstractNatural language processing (NLP) models are known vulnerable to adversarial examples, similar to image processing models. Studying adversarial texts is an essential step to improve the robustness of NLP models. However, existing studies mainly focus on generating adversarial texts for English, with no prior knowledge that whether those attacks could be applied to Chinese. After analyzing the differences between Chinese and English, we propose a novel adversarial Chinese text generation solution Argot, by utilizing the method for adversarial English examples and several novel methods developed on Chinese characteristics. Argot could effectively and efficiently generate adversarial Chinese texts with good readability in both white-box and black-box settings. Argot could also automatically generatetargetedChinese adversarial texts, achieving a high success rate and ensuring the readability of the generated texts. Furthermore, we apply Argot to the spam detection task in both local detection models and a public toxic content detection system from a well-known security company. Argot achieves a relatively high bypass success rate with fluent readability, which proves that the real-world toxic content detection system is vulnerable to adversarial example attacks. We also evaluate some available defense strategies, and the results indicate that Argot can still achieve high attack success rates. Mingxuan Liu 0006, Yiming Zhang 0009, Chao Zhang 0008, Zhou Li 0001, Qi Li 0002, Hai-Xin Duan, Donghong Sun |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2022 | ValCAT: Variable-Length Contextualized Adversarial Transformations Using Encoder-Decoder Language ModelabstractChuyun Deng, Mingxuan Liu, Yue Qin, Jia Zhang, Hai-Xin Duan, Donghong Sun. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Chuyun Deng, Mingxuan Liu 0006, Jia Zhang 0004, Hai-Xin Duan, Donghong Sun |
NAACL-HLT | 6 |
| 2021 | Detecting and Characterizing SMS Spearphishing AttacksabstractAlthough spearphishing is a well-known security issue and has been widely researched, it is still an evolving threat with emerging forms. In recent years, Short Message Service (SMS) has been revealed as a new distribution channel for spearphishing messages, which already has caused a serious impact in the real world, but has not yet attracted enough attention from the academic community. In this paper, we report the first systemic study to spotlight this emerging threat, SMS spearphishing attack. Through cooperating with a leading security vendor, we obtain 31.96M real-world spam messages that span three months. We design and implement a novel NLP-based detection algorithm, and uncover 90,801 spearphishing messages on the entire dataset. And then, a large-scale measurement was performed on the detected messages to reveal and understand the characteristics of SMS spearphishing attack. Our findings are multi-fold. We discover that SMS spearphishing has a significant negative impact on the real-world, and a large number of victims have been affected. And the distribution of active illicit types between spearphishing message and common spam is quite inconsistent. At the micro-level, to evade detection and increase the probability of success, adversary campaigns have evolved a set of sophisticated strategies. Our research highlights the impact of SMS spearphishing attack is prominent. We call on different communities to work together to mitigate this emerging security threat. Mingxuan Liu 0006, Yiming Zhang 0009, Baojun Liu 0002, Zhou Li 0001, Hai-Xin Duan, Donghong Sun |
ACSAC | 6 |
| 2020 | Argot: Generating Adversarial Readable Chinese TextsabstractNatural language processing (NLP) models are known vulnerable to adversarial examples, similar to image processing models. Studying adversarial texts is an essential step to improve the robustness of NLP models. However, existing studies mainly focus on analyzing English texts and generating adversarial examples for English texts. There is no work studying the possibility and effect of the transformation to another language, e.g, Chinese. In this paper, we analyze the differences between Chinese and English, and explore the methodology to transform the existing English adversarial generation method to Chinese. We propose a novel black-box adversarial Chinese texts generation solution Argot, by utilizing the method for adversarial English samples and several novel methods developed on Chinese characteristics. Argot could effectively and efficiently generate adversarial Chinese texts with good readability. Furthermore, Argot could also automatically generate targeted Chinese adversarial text, achieving a high success rate and ensuring readability of the Chinese. Mingxuan Liu 0006, Chao Zhang 0008, Yiming Zhang 0009, Zhou Li 0001, Qi Li 0002, Hai-Xin Duan, Donghong Sun |
IJCAI | 8 |
| 2020 | How to Keep an Online Learning Chatbot From Being CorruptedabstractOnline learning can improve chatbots' conversational abilities. Although the online learning method has enhanced the diversity of chatbots' statements, it also brings opportunities for corruption. The chatbot may be corrupted to generate offensive responses such as racist and hate speech. The key component to keeping chatbots from being corrupted is offensive-response detection. Until now, the training datasets for offensive detection have focused only on individual response sentences, disregarding user input sentences. In this paper, we introduce a dialogue-based offensive-response dataset, which consists of 110K input-response chat records. The dataset fills the gap in response detection for chatbots. Then, we build two challenging tasks based on the dataset: an offensive-response detection task and a corrupted chatbot purification task. In addition, we propose a strong benchmark method for the tasks: an encoder-classifier model to detect input-response pairs and a one-shot reinforcement learning (RL) method to reduce rapidly the probability of generating offensive responses. Yixuan Chai, Ziwei Jin, Donghong Sun |
IJCNN | 4 |
| 2019 | Understanding Distributed Poisoning Attack in Federated LearningabstractFederated learning is inherently vulnerable to poisoning attacks, since no training samples will be released to and checked by trustworthy authority. Poisoning attacks are widely investigated in centralized learning paradigm, however distributed poisoning attacks, in which more than one attacker colludes with each other, and injects malicious training samples into local models of their own, may result in a greater catastrophe in federated learning intuitively. In this paper, through real implementation of a federated learning system and distributed poisoning attacks, we obtain several observations about the relations between the number of poisoned training samples, attackers, and attack success rate. Moreover, we propose a scheme, Sniper, to eliminate poisoned local models from malicious participants during training. Sniper identifies benign local models by solving a maximum clique problem, and suspected (poisoned) local models will be ignored during global model updating. Experimental results demonstrate the efficacy of Sniper. The attack success rates are reduced to around 2% even a third of participants are attackers. Shan Chang, Zhijian Lin, Donghong Sun |
ICPADS | 5 |
| 2019 | Utility-Aware Participant Selection with Budget Constraints for Mobile Crowd Sensing
Shanila Azhar, Shan Chang, Yuting Tao, Donghong Sun |
QSHINE | 6 |
| 2016 | Analysis and forensics for Behavior Characteristics of Malware in InternetabstractAll sorts of Malwares severely threaten users in Internet. These malwares do share some common characteristics, despite malware and its variants may vary a lot from content signatures. The common characteristics they shared can be used to reveal the real intent of malware. In this paper, we study on the behavior characteristics of malwares in Internet, and based on which we present the method to extract the Formal Malware Behavior Characteristics (MBC), and finally, we design and implement the formal Malware Behavior Characteristics (MBC) extraction method, and propose the malicious behavior characteristics based malware detection algorithm. Finally we designed and implemented the malware analysis and forensics system. And the experimental results show that it can detect newly appeared unknown malwares. Donghong Sun |
PST | 3 |
| 2016 | MrBayes tgMC3++: A High Performance and Resource-Efficient GPU-Oriented Phylogenetic Analysis MethodabstractMrBayes is a widespread phylogenetic inference tool harnessing empirical evolutionary models and Bayesian statistics. However, the computational cost on the likelihood estimation is very expensive, resulting in undesirably long execution time. Although a number of multi-threaded optimizations have been proposed to speed up MrBayes, there are bottlenecks that severely limit the GPU thread-level parallelism of likelihood estimations. This study proposes a high performance and resource-efficient method for GPU-oriented parallelization of likelihood estimations. Instead of having to rely on empirical programming, the proposed novel decomposition storage model implements high performance data transfers implicitly. In terms of performance improvement, a speedup factor of up to 178 can be achieved on the analysis of simulated datasets by four Tesla K40 cards. In comparison to the other publicly available GPU-oriented MrBayes, the tgMC3++ method (proposed herein) outperforms the tgMC3(v1.0), nMC3(v2.1.1) and oMC3(v1.00) methods by speedup factors of up to 1.6, 1.9 and 2.9, respectively. Moreover, tgMC3++ supports more evolutionary models and gamma categories, which previous GPU-oriented methods fail to take into analysis. Cheng Ling, Tsuyoshi Hamada, Jingyang Gao, Guoguang Zhao, Donghong Sun, Weifeng Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2015 | SparkSW: Scalable Distributed Computing System for Large-Scale Biological Sequence AlignmentabstractThe Smith-Waterman (SW) algorithm is universally used for a database search owing to its high sensitively. The widespread impact of the algorithm is reflected in over 8000 citations that the algorithm has received in the past decades. However, the algorithm is prohibitively high in terms of time and space complexity, and so poses significant computational challenges. Apache Spark is an increasingly popular fast big data analytics engine, which has been highly successful in implementing large-scale data-intensive applications on commercial hardware. This paper presents the first ever reported system that implements the SW algorithm on Apache Spark based distributed computing framework, with a couple of off-the-shelf workstations, which is named as SparkSW. The scalability and load-balancing efficiency of the system are investigated by realistic ultra-large database from the state-of-the-art UniRef100. The experimental results indicate that 1) SparkSW is load-balancing for parallel adaptive on workloads and scales extremely well with the increases of computing resource, 2) SparkSW provides a fast and universal option high sensitively biological sequence alignments. The success of SparkSW also reveals that Apache Spark framework provides an efficient solution to facilitate coping with ever increasing sizes of biological sequence databases, especially generated by second-generation sequencing technologies. Guoguang Zhao, Cheng Ling, Donghong Sun |
CCGRID | 3 |
| 2014 | Designing a new TCP based on FAST TCP for datacenterabstractTCP incast problem has become a severe problem in datacenters due to the catastrophic collapse of goodput at the receiver side. Many solutions have been proposed, however, none of them solved it fundamentally. In this paper, we take a deep dive into the causes of incast problem, and find out that the root cause of TCP incast is the droptails induced by TCP congestion control algorithm based on packet losses. According to the root cause, we design a new TCP based on FAST TCP, a delay based congestion control algorithm, to radically solve TCP incast problem. The new TCP aims at maintaining a relatively small queue length that does not exceed the switch buffer and meanwhile fully occupy the bottleneck link. The simulation results demonstrate that our method cuts off most of the timeouts and attains high goodput under various conditions. Yongfeng Huang 0001, Donghong Sun |
ICC | 3 |