VLDB 2026 Research / reviewers in the wild / expert
Nan Sun 0002
dblp:17/5023-2
· DBLP profile ↗
19ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0001-9123-9022ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ST-Attention-XAI: Intrinsic Spatio-Temporal Explainability for IoT Intrusion Detection via Attention Analysis
Nimesha Dilini, Nan Sun 0002, Yuantian Miao, Nour Moustafa |
ACISP (2) | 2 |
| 2026 | Harnessing Large Language Models for Real-Time Cyber Threat Detection and Response: A Comprehensive Survey
Xiaoyu Li 0001, Nan Sun 0002, Jiaojiao Jiang 0001 |
ACISP (2) | 2 |
| 2026 | PufferDoS: Efficient and Effective Attack String Generation for Regular Expression Denial of Service Vulnerabilities
Shangzhi Xu, Yuekang Li, Nan Sun 0002, Benjamin Turnbull, Shuangxiang Kan, Siqi Ma 0001 |
SP | 5 |
| 2026 | DagFC: Dependency-Aware Fact-Checking via Claim-Constructed Knowledge Graphs and Large Language ModelsabstractFact-checking, also referred to as fact verification, is essential for evaluating the accuracy of claims and curbing the dissemination and influence of misinformation. Recent advancements in Large Language Models (LLMs) have enabled their use in automated fact-checking systems. These approaches frequently adopt prompting techniques within a ''divide-and-conquer'' framework, where complex claims are broken down into simpler sub-claims that are individually verified to reduce the overall complexity of the task. These existing works often neglect the dependency between sub-claims and verify them in isolation. For complex claims, particularly those requiring multi-hop reasoning, the interconnections between sub-claims are crucial, as verifying each one independently often fails to capture the full context and reasoning needed for accurate verification. To address this, we propose DagFC, a novel LLM-based framework that performs Dependency-Aware Task Generation, Scheduling and Processing for Fact-Checking. DagFC constructs Knowledge Graphs (KGs) from claims to guide the decomposition of fact-checking problems and build dependent verification sub-tasks that capture the interrelations between sub-claims. This dependency-aware approach ensures more coherent and accurate verification by integrating intermediate results. Additionally, DagFC leverages LLMs throughout the verification process, from KG construction to final veracity prediction, enhancing reasoning and generation capabilities. Extensive experiments on two benchmark datasets, FEVEROUS and HoVer, demonstrate that DagFC outperforms state-of-the-art methods in both accuracy and Macro-F1 score. Furthermore, we present a user-friendly fact-checking prototype based on our framework, offering practical value for both research and public use. Zhouhui Wu, Zhuohua Yang, Jiaojiao Jiang 0001, Shuiqiao Yang, Nan Sun 0002 |
WSDM | 5 |
| 2025 | What Lies Beneath: An Empirical Study of Silent Vulnerability Fixes in Open-Source SoftwareabstractUnlike standard vulnerability disclosure, fixing vulnerabilities "silently" is another common approach in software development. While silent fixes can prevent potential targeted attacks without disclosing any details, they only offer short-term protection. Users often rely on publicly disclosed vulnerabilities to identify and eliminate vulnerabilities in their own software, particularly for open-source software (OSS). We conduct the first comprehensive empirical study in OSS to investigate potential security threats brought by silent fixes. By examining disclosed vulnerabilities and their corresponding patches in real-world OSSes, we investigated the prevalence of silent vulnerabilities and assessed the potential impacts they might cause. After analyzing 3,515 vulnerabilities, we observed that nearly 50% of the vulnerabilities, with half of them classified as high-severity, were exposed for over 30 days. Over 10% of them published exploits during the "silent" period, enabling adversaries to replicate the exploits and attack other OSS users. Due to delayed vulnerability disclosure, we even found one HIGH-severity vulnerability that was silently fixed still exists in downstream software, indicating that silent fixes may increase the risks for other OSSes, amplifying their impact throughout software supply chains. Jialiang Dong, Xinzhang Chen, Willy Susilo, Nan Sun 0002, Arash Shaghaghi, Siqi Ma 0001 |
DSN | 4 |
| 2025 | Arms Race in Deep Learning: A Survey of Backdoor Defenses and Adaptive Attacks
Xiaoxing Mo, Nan Sun 0002, Leo Yu Zhang, Wei Luo 0001, Shang Gao 0003, Yong Xiang 0001 |
PAKDD (4) | 2 |
| 2025 | Large Language Models for Cybersecurity Education: A Survey of Current Practices and Future Directions
Nan Sun 0002, Yuantian Miao, Xiaoxing Mo, Jun Zhang 0010 |
PAKDD (6) | 1 |
| 2025 | Systematic Approaches to Fact Verification: Evidence Retrieval, Veracity Prediction, and Beyond
Zhouhui Wu, Nan Sun 0002, Jiaojiao Jiang 0001, Shuiqiao Yang |
PAKDD (4) | 2 |
| 2025 | AI-Powered Pavement Moduli Prediction: Advancing Smart City Infrastructure with Scalable and Interpretable Solutions*abstractWith the growth of smart cities, data-driven approaches are transforming infrastructure management by enabling efficient and accurate decision-making. Pavement condition evaluation, a critical aspect of transportation infrastructure, plays a vital role in ensuring mobility, safety, and economic productivity. This study compares traditional machine learning (ML) and deep learning (DL) models for predicting pavement layer moduli, which are key indicators of structural health, using Falling Weight Deflectometer (FWD) data from the Long-Term Pavement Performance (LTPP) database. Ensemble machine learning models, such as Random Forest and XGBoost, offer high predictive accuracy with minimal training time, making them suitable for scenarios with limited data. In contrast, advanced DL models—particularly hybrid ResRNN architectures with Wide & Deep (W&D), LSTM, and GRU components—demonstrate superior scalability and predictive power, which aligns with the growing availability of data in smart infrastructure systems. To enhance interpretability, SHapley Additive exPlanations (SHAP) analysis is employed, uncovering feature importance patterns that align with pavement engineering principles and informing future sensor placement strategies. The findings support the selection of AI models that balance accuracy, efficiency, and transparency in smart infrastructure systems. To facilitate reproducibility and future research, we make our code publicly available on GitHub: https://github.com/Gxy9527/LTPP. Nan Sun 0002, Jianfeng Xue |
SMC | 2 |
| 2025 | Dual-View Evidence Learning and Cross-View Fusion for Enhanced Text-Table Fact VerificationabstractFact verification involves assessing the factual-ity of claims to detect false information. This work focuses on a specific fact verification subtask: verifying claims using retrieved textual and tabular evidence. Existing approaches often overlook the distinct features and interactions of table and text evidence, which are essential for accurate claim verification by providing a comprehensive understanding. Moreover, current evidence fusion strategies used by existing work fail to model complex distinctions, leading to ineffective integration. This work introduces a novel veracity prediction model that leverages dual-view evidence learning and graph-based evidence fusion to address these limitations. Our model incorporates a local view, capturing the unique information within each sentence and table, and a global view, modeling the interactions between these evidence pieces. We further employ graph networks to fuse information within each view and across views, generating richer evidence representations for improved claim verification. Extensive experiments demonstrate the effectiveness of our method. Zhouhui Wu, Jiaojiao Jiang 0001, Shuiqiao Yang, Nan Sun 0002 |
SMC | 4 |
| 2025 | Securing AI Code Generation - A Prompt Rectification Approach for Mitigating Cyber RisksabstractThe past decade has witnessed the wide adoption of AI code generators, such as GitHub Copilot, AskCodi, and OpenAI Codex. They offer intelligent solution code for code completion to achieve faster development, cleaner code, and a significant boost in overall productivity. However, such significant productivity advantages also inadvertently lead to the generation of insecure solution code because most AI code generators derive their knowledge from existing projects, which typically prioritize functionality over security. Although numerous tools have been developed to integrate with the code generators for identifying vulnerabilities, the inconsistency in syntactic features and variability in coding rules make the detection task challenging across different programming languages. To address the challenges, we devise a prompt-enhancing approach, PECKER. It examines textual prompts provided by users to identify risky prompts that could lead to insecure code generation. Given the risky prompts, PECKER conducts security-centric rewriting to strengthen the "potentially insecure" descriptions, thereby guiding AI code generators in mitigating vulnerabilities during code generation. We integrated PECKER with one of the most prevalent AI code generators, GitHub Copilot for evaluation. Among 509 risky prompts, PECKER successfully identified and rectified 471 risky prompts. Jialiang Dong, Zihan Ni, Nan Sun 0002, Sanjay K. Jha, Yiwei Zhang 0008, Elisa Bertino, Surya Nepal, Siqi Ma 0001 |
TrustCom | 3 |
| 2025 | Can LLM-generated misinformation be detected: A study on Cyber Threat IntelligenceabstractGiven the increasing number and severity of cyber attacks, there has been a surge in cybersecurity information across various mediums such as posts, news articles, reports, and other resources. Cyber Threat Intelligence (CTI) involves processing data from these cybersecurity sources, enabling professionals and organizations to gain valuable insights. However, with the rapid dissemination of cybersecurity information, the inclusion of fake CTI can lead to severe consequences, including data poisoning attacks. To address this challenge, we have implemented a three-step strategy: generating synthetic CTI, evaluating the quality of the generated CTI, and detecting fake CTI. Unlike other subdomains, such as fake COVID news detection, there is currently no publicly available dataset specifically tailored for fake CTI detection research. To address this gap, we first establish a reliable groundtruth dataset by utilizing domain-specific cybersecurity data to fine-tune a Large Language Model (LLM) for synthetic CTI generation. We then employ crowdsourcing techniques and advanced synthetic data verification methods to evaluate the quality of the generated dataset, introducing a novel evaluation methodology that combines quantitative and qualitative approaches. Our comprehensive evaluation reveals that the generated CTI cannot be distinguished from genuine CTI by human annotators, regardless of their computer science background, demonstrating the effectiveness of our generation approach. We benchmark various misinformation detection techniques against our groundtruth dataset to establish baseline performance metrics for identifying fake CTI. By leveraging existing techniques and adapting them to the context of fake CTI detection, we provide a foundation for future research in this critical field. To facilitate further research, we make our code, dataset, and experimental results publicly available on GitHub . Nan Sun 0002, Massimiliano Tani, Yu Zhang 0217, Jiaojiao Jiang 0001, Sanjay K. Jha |
Future Gener. Comput. Syst. | 2 |
| 2024 | DiHAN: A Novel Dynamic Hierarchical Graph Attention Network for Fake News DetectionabstractThe rapid spread of fake news on social media has caused great harm to society in recent years, which raises the detection of fake news as an urgent task. Recent methods utilize the interactions among different entities such as authors, subjects, and news articles to model news propagation as a static heterogeneous information network (HIN). However, this is suboptimal since fake news emerges dynamically, and the latent chronological interactions between news in HIN are essential signals for fake news detection. To this end, we model the dynamics of news and associated entities as a News-Driven Dynamic Heterogeneous Information Network (News-DyHIN), where the temporal relationships among news articles are well captured with meta-path based temporal neighbors. With the support of News-DyHIN, we propose a novel fake news detection framework, named D ynam i c H ierarchical A ttention N etwork (DiHAN), which learns news representations via a hierarchical attention mechanism to fuse temporal interactions among news articles. In particular, DiHAN first employs a temporal node level attention to learn the temporal information from meta-path based news neighbors through the modeled News-DyHIN. Then, a semantic attention layer is adopted to fuse different types of meta-path based temporal information for news representation learning. Extensive evaluations conducted on two public real-world datasets demonstrate that our proposed DiHAN achieves significant improvements over established baseline models. Ya-Ting Chang, Zhibo Hu, Xiaoyu Li 0001, Shuiqiao Yang, Jiaojiao Jiang 0001, Nan Sun 0002 |
CIKM | 6 |
| 2024 | Robust Backdoor Detection for Deep Learning via Topological Evolution DynamicsabstractA backdoor attack in deep learning inserts a hidden backdoor in the model to trigger malicious behavior upon specific input patterns. Existing detection approaches assume a metric space (for either the original inputs or their latent representations) in which normal samples and malicious samples are separable. We show that this assumption has a severe limitation by introducing a novel SSDT (Source-Specific and Dynamic-Triggers) backdoor, which obscures the difference between normal samples and malicious samples.To overcome this limitation, we move beyond looking for a perfect metric space that would work for different deep-learning models, and instead resort to more robust topological constructs. We propose TED (Topological Evolution Dynamics) as a model-agnostic basis for robust backdoor detection. The main idea of TED is to view a deep-learning model as a dynamical system that evolves inputs to outputs. In such a dynamical system, a benign input follows a natural evolution trajectory similar to other benign inputs. In contrast, a malicious sample displays a distinct trajectory, since it starts close to benign samples but eventually shifts towards the neighborhood of attacker-specified target samples to activate the backdoor.Extensive evaluations are conducted on vision and natural language datasets across different network architectures. The results demonstrate that TED not only achieves a high detection rate, but also significantly outperforms existing state-of-the-art detection approaches, particularly in addressing the sophisticated SSDT attack. The code to reproduce the results is made public on GitHub. Xiaoxing Mo, Yechao Zhang, Leo Yu Zhang, Wei Luo 0001, Nan Sun 0002, Shengshan Hu, Shang Gao 0003, Yang Xiang 0001 |
SP | 5 |
| 2023 | Backdoor Attack on Deep Neural Networks in Perception DomainabstractAs deep neural networks (DNNs) are widely deployed in various applications, the security of pretrained DNNs is crucial since backdoors can be introduced through poisoned training. A backdoored DNN model works properly when benign inputs are provided, but it produces targeted misclassification on the inputs with an intended pattern known as a trojan trigger. Current technologies for trigger generation mainly focus on the physical and model domains. In this work, we investigate trojan triggers from the perception domain, especially the physical process of collecting light rays when they pass through the lens and hit the optical sensors. A new type of backdoor attack, Lens Flare attack, is introduced. It concentrates on the perception domain and is more physically plausible and stealthy. Experiments show that the DNNs with Lens Flare backdoor can achieve accuracy comparable to their original counterpart on benign input while misclassifying the input with high certainty if the Lens Flare trigger is present. It is also demonstrated that the Lens Flare backdoor is resistant to state-of-the-art backdoor defenses. Xiaoxing Mo, Leo Yu Zhang, Nan Sun 0002, Wei Luo 0001, Shang Gao 0003 |
IJCNN | 3 |
| 2023 | Cyber Information Retrieval Through Pragmatics Understanding and VisualizationabstractThe amount of cybersecurity-related information is extraordinarily increasing, given the fast-growing number of cybersecurity attacks and the significant influence brought by them. How to efficiently obtain and precisely understand the relevant knowledge in the sea of information on cybersecurity becomes a challenge. In this article, we propose an innovative cybersecurity retrieval scheme that supports automatic indexing and searching of cybersecurity information based on semantic contents and hidden metadata. The proposed scheme leverages a customized neural model that incorporates new linguistic features and word embedding by identifying the entities related to cybersecurity incidents from the text. We implement a novel cybersecurity search engine to demonstrate effective, understandable and pragmatic cybersecurity information retrieval based on the proposed schema. Comprehensive performance evaluation over real-world datasets has been conducted to validate the new algorithms and techniques developed for cybersecurity information retrieval. The new engine makes it possible to conduct augmented search, cybersecurity analytics, and visualization, with the ultimate goal of providing direct and efficient results to help people obtain and truly understand cybersecurity information. Nan Sun 0002, Jun Zhang 0010, Shang Gao 0003, Leo Yu Zhang, Seyit Ahmet Çamtepe, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Blockwise Spectral Analysis for Deepfake Detection in High-fidelity VideosabstractDeepfakes have gained widespread attention as they may give rise to a series of risks ranging from personal reputation damages to national security breaches. A mainstream approach to generating deepfakes is based on Generative Adversarial Networks (GAN). Various methods have been proposed to detect GAN-generated fake content. However, most of them only target a specific GAN and do not generalize well to other unseen GAN architectures. Moreover, many existing methods show poor performance on deepfakes that are imperceptible to the human eye as they heavily rely on visual artifacts, such as unblinking eyes and asymmetric faces. In this work, we exploit the spectral artifacts left by up-sampling operations that are universally used in GAN architectures for detecting high-fidelity deepfake videos. We first divide video frames into blocks containing the most informative areas, e.g., face, eyes, and mouth areas, and use their spectrum to train a ResNet-based classifier to detect GAN-generated images. Experimental results on public datasets show that our method is effective in detecting high-fidelity deepfakes and generalizes well across GANs with the same or similar up-sampling operations. Nan Sun 0002, Xufeng Lin |
DSAA | 2 |
| 2022 | Advanced Face Anti-Spoofing with Depth SegmentationabstractFace anti-spoofing (FAS) plays a vital role in securing face recognition systems. In state-of-the-art FAS methods, face depth is determined for every position in a facial image. However, face depth varies at different positions, which leads to low accuracy when predicting face depth. As we observe, spoof faces have depth values that are 0, while live faces have depths that are equal to or greater than 0. As a result, for faces that have a depth greater than 0, if they are estimated as merely a positive value, instead of an accurate value, the prediction of whether they are real or not will not change. Further, if a range of depths are considered as one category, then there are more samples per category for the network training. Based on the above observation, in this paper, we propose to aggregate simple depth values to the same category and perform classification to optimize the FAS network. To evaluate the performance of the proposed approach, we perform extensive experiments on four benchmark databases, respectively, OULU-NPU, SiW, CASIA-FASD, and Replay-Attack. The results demonstrate that the proposed approach outperforms state-of-the-art methods on intra-database testing. Furthermore, our proposed approach shows advanced performance on cross-database testing. Nan Sun 0002, Xihong Wu, Dingsheng Luo |
IJCNN | 2 |
| 2020 | Data Analytics of Crowdsourced Resources for Cybersecurity Intelligence
Nan Sun 0002, Jun Zhang 0010, Shang Gao 0003, Leo Yu Zhang, Seyit Ahmet Çamtepe, Yang Xiang 0001 |
NSS | 1 |